-
Notifications
You must be signed in to change notification settings - Fork 2.7k
Pull requests: NVIDIA/TensorRT-LLM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[TRTLLM-14660][feat] Early reused cache transfer in Transceiver V2
#18688
opened Sep 3, 2026 by
athena-nv
Collaborator
Loading…
1 task done
[https://nvbugs/6720250][fix] Use adjusted clock for VisualGen timing
VisualGen
#18686
opened Sep 3, 2026 by
chienchunhung
Collaborator
Loading…
[TRTLLM-16022][feat] Add sub-agent conversation affinity for disaggregated serving
#18684
opened Sep 3, 2026 by
xwang233
Collaborator
Loading…
[None][fix] Self-sampling top-k host: physical row-width envelope and exact-row warmup population
#18683
opened Sep 3, 2026 by
longcheng-nv
Collaborator
Loading…
1 task done
[https://nvbugs/6707518][fix] Fix Kimi K3 spec dec test
#18682
opened Sep 3, 2026 by
mikeiovine
Collaborator
Loading…
1 task done
[None][refactor] Table-drive the lazy-safetensors model-type check in HfWeightLoader
#18681
opened Sep 3, 2026 by
moraxu
Collaborator
Loading…
1 task done
[None][fix] Use the single custom-tokenizer alias table in llm_args
#18680
opened Sep 3, 2026 by
moraxu
Collaborator
Loading…
1 task done
[None][feat] Add VisualGen usage telemetry
api-compatible
Accepted LLM API contract change that is backwards-compatible
VisualGen
#18679
opened Sep 3, 2026 by
Mgluhovskoi
Collaborator
•
Draft
[https://nvbugs/6695563][fix] Densify the bmm LHS with
a.contiguous() gated on get_sm_version() in (120…
#18678
opened Sep 3, 2026 by
trtllm-agent
Collaborator
Loading…
2 tasks done
[TRTLLMINF-336][infra] Enable BOLT premerge consume
#18677
opened Sep 3, 2026 by
mlefeb01
Collaborator
Loading…
1 task done
[https://nvbugs/6193837][fix] Include FINALIZE-fusion workspace for SM>=90 in the MoE autotuner
#18675
opened Sep 3, 2026 by
farazkh80
Collaborator
Loading…
4 tasks done
[https://nvbugs/6480110][fix] Fall back to FP8 KV cache when NVFP4 KV cache is requested on SM107
#18674
opened Sep 3, 2026 by
farazkh80
Collaborator
Loading…
4 tasks done
[None][fix] Guard MiniMax-M3 FP8 indexer against padded -1 cache slots
#18673
opened Sep 3, 2026 by
brb-nv
Collaborator
Loading…
1 task done
[None][docs] fix legacy benchmark KV cache config
#18671
opened Sep 3, 2026 by
imitater-dou
Loading…
2 tasks done
[None][feat] perf-optimize: optimize one half of a disaggregated deployment
#18670
opened Sep 3, 2026 by
hyukn
Collaborator
Loading…
5 tasks done
[https://nvbugs/6701493][fix] Latch warmup peak before KV cache size estimation resets it
#18669
opened Sep 3, 2026 by
trtllm-agent
Collaborator
Loading…
2 tasks done
[None][fix] Derive KV capacity once and select windowed blocks by ordinal
#18668
opened Sep 3, 2026 by
Shixiaowei02
Collaborator
•
Draft
1 task
[None][fix] improve BCG + ADP trigger time in agg mode
#18667
opened Sep 3, 2026 by
GuanhuaWang2001
Collaborator
Loading…
4 tasks done
[#18659][fix] Apply the DeepSeek kv_b_proj default exclusion on the MIXED_PRECISION quant-config path
#18666
opened Sep 3, 2026 by
PierreLeGuen
Loading…
1 task done
[#18658][fix] Dequantize FP8 block-scaled weights for unquantized DeepSeek-V3.2/GLM indexer projections
#18665
opened Sep 3, 2026 by
PierreLeGuen
Loading…
1 task done
[TRTLLMINF-396][ci] Automate full pre-merge approval
#18656
opened Sep 3, 2026 by
ZhanruiSunCh
Collaborator
Loading…
1 task
[None][test] Add func and perf cases for Qwen3.6-35B-A3B and gemma4 on Spark
#18655
opened Sep 3, 2026 by
JennyLiu-nv
Collaborator
Loading…
1 task done
Previous Next
ProTip!
no:milestone will show everything without a milestone.