Yuewei Na
8483e4a07f
chore: Upgrade to tensorrt-llm==1.3.0rc5 ( #6579 )
...
Signed-off-by: Yuewei Na <nv-yna@users.noreply.github.com>
Co-authored-by: Yuewei Na <nv-yna@users.noreply.github.com>
2026-02-25 12:22:58 -08:00
Keiven C
ff06b17e7f
fix: guarantee RouterRequestMetrics availability & documentation updates ( #6558 )
...
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-25 10:18:01 -08:00
Qi Wang
5a4c96dbaf
test: introduce multimodal benchmark toolkit ( #6330 )
2026-02-24 16:55:04 -08:00
Jason Zhou
4ea80e9585
chore: use aic release/0.7.0 ( #6494 )
2026-02-24 16:27:19 -06:00
Alec
7893f2684e
feat: add --disaggregation-mode enum to vLLM backend ( #6483 )
...
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 19:34:30 -07:00
Yan Ru Pei
0cb1d73383
fix: fold scheduling into queue so backpressure actually works ( #6470 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Signed-off-by: Yan Ru Pei <yanrpei@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 18:33:03 -08:00
Alec
7bbacce196
feat: default kv-events-config to empty (align with vLLM defaults) ( #6404 )
...
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-20 06:44:51 +00:00
GuanLuo
a2a6917fa6
feat: use embedding transfer classes for EPD ( #6223 )
...
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
2026-02-19 12:27:01 -08:00
Yuewei Na
9a15730a7f
chore: Upgrade to tensorrt-llm==1.3.0rc3 ( #6402 )
...
Signed-off-by: Yuewei Na <nv-yna@users.noreply.github.com>
Co-authored-by: Yuewei Na <nv-yna@users.noreply.github.com>
Co-authored-by: dagil-nvidia <dagil@nvidia.com>
2026-02-19 14:12:32 -05:00
Yan Ru Pei
4ede59a269
feat: speculative prefill ( #6230 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: Janelle Cai <jcai18@mit.edu>
2026-02-15 09:27:34 +00:00
Yan Ru Pei
bc514fbee6
feat: router priority queue ( #6010 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
2026-02-13 02:43:17 +00:00
Hongkuan Zhou
a04b56310c
feat: support AIC DGD gen call (WILL BREAK DGDR) ( #6216 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-02-12 08:47:55 -08:00
Jonathan Tong
39d645e586
docs: migrate Fern docs from fern/ into docs/ ( #6206 )
...
Signed-off-by: Jont828 <jt572@cornell.edu>
2026-02-11 16:22:27 -08:00
Karen Chung
8707dc2cf5
fix: scale synthesized data length correctly for expected cache hit stats ( #6117 )
2026-02-10 12:05:42 -08:00
hhzhang16
1aab7f6b6d
fix: use actual service names for profiler logs and handle FileNotFoundError correctly ( #6112 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-02-10 10:02:03 -08:00
MatejKosec
67329d1023
fix: profiler deployment timeout handling for MoE models ( #6086 )
...
Wrap wait_for_deployment_ready() in try/except TimeoutError for both prefill and decode profiling sweeps
On timeout: log error, record via add_profiling_error(), clean up the timed-out deployment, and continue to the next parallelization mapping
Previously, a single deployment timeout would crash the entire profiler job
2026-02-09 15:57:59 -08:00
Yan Ru Pei
5035447f2f
feat: NAT telemetry conversion ( #6022 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-02-05 20:52:43 -08:00
dagil-nvidia
b19de4ed77
docs: cleanup of docs refactor for components, integrations, and features ( #6019 )
...
Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-05 19:50:17 -08:00
akshatha-k
80e7bafd37
docs: Migrate router documentation to three-tier structure ( #5979 )
...
Signed-off-by: akshatha-k <akshutk@gmail.com>
Signed-off-by: dagil-nvidia <dagil@nvidia.com>
Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: dagil-nvidia <dagil@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-06 01:55:50 +00:00
Yan Ru Pei
50e17783d9
fix: ignore benchmarks pytest for now ( #5983 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-02-05 08:26:25 +00:00
Yan Ru Pei
88efa735a5
fix: e2e aiperf profiling on NAT dataset ( #5990 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-02-04 21:50:30 -08:00
hhzhang16
7e4bc71671
feat: remove default model name in Profiler; validate for one of served model name and model path in Profiler ( #5950 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-02-04 16:21:31 -08:00
Yan Ru Pei
4fcee92f60
chore: use aiperf utils in prefix synthesizer ( #5906 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-02-04 12:57:05 -08:00
hhzhang16
0268aea4e3
fix: Add status file to prevent output-copier hang on failures ( #5898 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-02-03 17:25:08 -08:00
Yan Ru Pei
cad453f280
chore: applying rolling hasher in prefix synthesizer ( #5903 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-02-02 20:20:55 -08:00
hhzhang16
2b19954060
fix: check --served-model-name first before --model/--model-path ( #5881 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-02-02 16:29:42 -08:00
Tanmay Verma
ba711cc1ac
chore: Upgrade to Tensorrt-LLM 1.3.0rc1 ( #5700 )
...
Co-authored-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com>
2026-01-29 15:20:58 -08:00
GuanLuo
903f818498
chore: add multimodal image benchmark scripts for performance evaluation ( #5509 )
...
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
2026-01-27 17:39:58 -08:00
Yan Ru Pei
e00ada2071
feat: convert NAT trace to mooncake-style trace ( #5675 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-01-27 07:49:36 +00:00
Dmitry Tokarev
546f1bbc0d
fix: sglang version in build.sh and docs ( #5665 )
...
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com>
2026-01-26 19:53:55 -05:00
Jason Zhou
aecf005136
chore: use aic release/0.6.0 ( #5600 )
2026-01-26 13:34:10 -08:00
jthomson04
64ba7dd068
chore: Bump TRTLLM to 1.2.0rc6.post2 ( #5580 )
...
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
2026-01-22 20:23:40 -08:00
Hongkuan Zhou
b0959cfdbc
feat: warmup dataset for planner load predictor ( #5529 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-01-21 10:55:03 -08:00
Yan Ru Pei
050906b593
feat: track output tokens / blocks in the Router (optional) ( #5452 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-01-18 09:30:25 +08:00
hhzhang16
a2c7d0f95c
fix: make it clear what features are only available 0.8.1 and on ( #5492 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-01-16 14:06:33 -08:00
Yan Ru Pei
abe9d127df
fix: include missing reqs in benchmark ( #5477 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-01-16 14:29:32 +08:00
Yan Ru Pei
d83ab662f7
chore: make precentile plots in router benchmark ( #5476 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-01-16 14:25:17 +08:00
dagil-nvidia
ef09f718e3
fix: use correct argument name max_isl in synthesize_requests ( #5441 )
...
Signed-off-by: Dan Gil <dagil@nvidia.com>
2026-01-15 19:01:12 +00:00
MatejKosec
7cfd966e51
feat: SLA Planner configuration support for DEP and TEP VLLM ( #4783 )
...
Extends the MOE planner profiler to support TEP (tensor expert parallel) and DEP (data expert parallel) configs with the vllm backend
2026-01-14 13:13:44 -08:00
Elias Bermudez
c3dc3de48d
chore: Update aiperf pinned version ( #5331 )
...
Signed-off-by: Anant Sharma <anants@nvidia.com>
Co-authored-by: Anant Sharma <anants@nvidia.com>
2026-01-13 12:34:21 -05:00
hhzhang16
c8770464ab
feat: normalize dynamo namespace computation ( #5231 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-01-12 19:09:43 +00:00
Tanmay Verma
abd4b5d9a0
chore: Upgrade to tensorrt_llm==1.2.0rc6.post1 ( #5356 )
2026-01-12 18:58:00 +00:00
hhzhang16
ec5630ead9
feat: mount model path to Profiler if specified ( #5212 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-01-09 17:13:53 +00:00
ishandhanani
92748c937f
feat: sglang update to 0.5.7 ( #5148 )
2026-01-08 17:00:03 +08:00
Ryan McCormick
6306afa6f6
fix: Remove asymmetric --request-plane nats from run_engines.sh script ( #5245 )
2026-01-07 10:50:20 -08:00
Hongkuan Zhou
a82137877c
fix: add sweep range for moe dgdr example ( #5225 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-01-06 14:28:54 -08:00
Alec
19b5917c2b
fix: assign unique ports per vLLM worker to avoid ZMQ bind conflicts ( #5224 )
...
Signed-off-by: alec-flowers <aflowers@nvidia.com>
2026-01-06 13:45:28 -08:00
Hongkuan Zhou
fbe6bb0a9e
feat: support PVC model cache in profiler ( #5124 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-01-05 08:52:34 -08:00
Tushar Sharma
cf433e6825
chore: update all copyright headers in repo to 2026 ( #5130 )
...
Signed-off-by: Tushar Sharma <tusharma@nvidia.com>
2026-01-02 22:08:23 +00:00
Nate Mailhot
3c7ed61d05
chore: TRTLLM 1.2.0rc6 ( #5017 )
...
Signed-off-by: Nate Mailhot <nmailhot@nvidia.com>
2026-01-02 10:29:04 -08:00