Commit Graph

177 Commits

Author SHA1 Message Date
Yuewei Na 8483e4a07f
chore: Upgrade to tensorrt-llm==1.3.0rc5 (#6579)
Signed-off-by: Yuewei Na <nv-yna@users.noreply.github.com>
Co-authored-by: Yuewei Na <nv-yna@users.noreply.github.com>
2026-02-25 12:22:58 -08:00
Keiven C ff06b17e7f
fix: guarantee RouterRequestMetrics availability & documentation updates (#6558)
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-25 10:18:01 -08:00
Qi Wang 5a4c96dbaf
test: introduce multimodal benchmark toolkit (#6330) 2026-02-24 16:55:04 -08:00
Jason Zhou 4ea80e9585
chore: use aic release/0.7.0 (#6494) 2026-02-24 16:27:19 -06:00
Alec 7893f2684e
feat: add --disaggregation-mode enum to vLLM backend (#6483)
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 19:34:30 -07:00
Yan Ru Pei 0cb1d73383
fix: fold scheduling into queue so backpressure actually works (#6470)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Signed-off-by: Yan Ru Pei <yanrpei@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 18:33:03 -08:00
Alec 7bbacce196
feat: default kv-events-config to empty (align with vLLM defaults) (#6404)
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-20 06:44:51 +00:00
GuanLuo a2a6917fa6
feat: use embedding transfer classes for EPD (#6223)
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
2026-02-19 12:27:01 -08:00
Yuewei Na 9a15730a7f
chore: Upgrade to tensorrt-llm==1.3.0rc3 (#6402)
Signed-off-by: Yuewei Na <nv-yna@users.noreply.github.com>
Co-authored-by: Yuewei Na <nv-yna@users.noreply.github.com>
Co-authored-by: dagil-nvidia <dagil@nvidia.com>
2026-02-19 14:12:32 -05:00
Yan Ru Pei 4ede59a269
feat: speculative prefill (#6230)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: Janelle Cai <jcai18@mit.edu>
2026-02-15 09:27:34 +00:00
Yan Ru Pei bc514fbee6
feat: router priority queue (#6010)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
2026-02-13 02:43:17 +00:00
Hongkuan Zhou a04b56310c
feat: support AIC DGD gen call (WILL BREAK DGDR) (#6216)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-02-12 08:47:55 -08:00
Jonathan Tong 39d645e586
docs: migrate Fern docs from fern/ into docs/ (#6206)
Signed-off-by: Jont828 <jt572@cornell.edu>
2026-02-11 16:22:27 -08:00
Karen Chung 8707dc2cf5
fix: scale synthesized data length correctly for expected cache hit stats (#6117) 2026-02-10 12:05:42 -08:00
hhzhang16 1aab7f6b6d
fix: use actual service names for profiler logs and handle FileNotFoundError correctly (#6112)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-02-10 10:02:03 -08:00
MatejKosec 67329d1023
fix: profiler deployment timeout handling for MoE models (#6086)
Wrap wait_for_deployment_ready() in try/except TimeoutError for both prefill and decode profiling sweeps
On timeout: log error, record via add_profiling_error(), clean up the timed-out deployment, and continue to the next parallelization mapping
Previously, a single deployment timeout would crash the entire profiler job
2026-02-09 15:57:59 -08:00
Yan Ru Pei 5035447f2f
feat: NAT telemetry conversion (#6022)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-02-05 20:52:43 -08:00
dagil-nvidia b19de4ed77
docs: cleanup of docs refactor for components, integrations, and features (#6019)
Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-05 19:50:17 -08:00
akshatha-k 80e7bafd37
docs: Migrate router documentation to three-tier structure (#5979)
Signed-off-by: akshatha-k <akshutk@gmail.com>
Signed-off-by: dagil-nvidia <dagil@nvidia.com>
Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: dagil-nvidia <dagil@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-06 01:55:50 +00:00
Yan Ru Pei 50e17783d9
fix: ignore benchmarks pytest for now (#5983)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-02-05 08:26:25 +00:00
Yan Ru Pei 88efa735a5
fix: e2e aiperf profiling on NAT dataset (#5990)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-02-04 21:50:30 -08:00
hhzhang16 7e4bc71671
feat: remove default model name in Profiler; validate for one of served model name and model path in Profiler (#5950)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-02-04 16:21:31 -08:00
Yan Ru Pei 4fcee92f60
chore: use aiperf utils in prefix synthesizer (#5906)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-02-04 12:57:05 -08:00
hhzhang16 0268aea4e3
fix: Add status file to prevent output-copier hang on failures (#5898)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-02-03 17:25:08 -08:00
Yan Ru Pei cad453f280
chore: applying rolling hasher in prefix synthesizer (#5903)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-02-02 20:20:55 -08:00
hhzhang16 2b19954060
fix: check --served-model-name first before --model/--model-path (#5881)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-02-02 16:29:42 -08:00
Tanmay Verma ba711cc1ac
chore: Upgrade to Tensorrt-LLM 1.3.0rc1 (#5700)
Co-authored-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com>
2026-01-29 15:20:58 -08:00
GuanLuo 903f818498
chore: add multimodal image benchmark scripts for performance evaluation (#5509)
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
2026-01-27 17:39:58 -08:00
Yan Ru Pei e00ada2071
feat: convert NAT trace to mooncake-style trace (#5675)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-01-27 07:49:36 +00:00
Dmitry Tokarev 546f1bbc0d
fix: sglang version in build.sh and docs (#5665)
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com>
2026-01-26 19:53:55 -05:00
Jason Zhou aecf005136
chore: use aic release/0.6.0 (#5600) 2026-01-26 13:34:10 -08:00
jthomson04 64ba7dd068
chore: Bump TRTLLM to 1.2.0rc6.post2 (#5580)
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
2026-01-22 20:23:40 -08:00
Hongkuan Zhou b0959cfdbc
feat: warmup dataset for planner load predictor (#5529)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-01-21 10:55:03 -08:00
Yan Ru Pei 050906b593
feat: track output tokens / blocks in the Router (optional) (#5452)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-01-18 09:30:25 +08:00
hhzhang16 a2c7d0f95c
fix: make it clear what features are only available 0.8.1 and on (#5492)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-01-16 14:06:33 -08:00
Yan Ru Pei abe9d127df
fix: include missing reqs in benchmark (#5477)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-01-16 14:29:32 +08:00
Yan Ru Pei d83ab662f7
chore: make precentile plots in router benchmark (#5476)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-01-16 14:25:17 +08:00
dagil-nvidia ef09f718e3
fix: use correct argument name max_isl in synthesize_requests (#5441)
Signed-off-by: Dan Gil <dagil@nvidia.com>
2026-01-15 19:01:12 +00:00
MatejKosec 7cfd966e51
feat: SLA Planner configuration support for DEP and TEP VLLM (#4783)
Extends the MOE planner profiler to support  TEP (tensor expert parallel) and DEP (data expert parallel) configs with the vllm backend
2026-01-14 13:13:44 -08:00
Elias Bermudez c3dc3de48d
chore: Update aiperf pinned version (#5331)
Signed-off-by: Anant Sharma <anants@nvidia.com>
Co-authored-by: Anant Sharma <anants@nvidia.com>
2026-01-13 12:34:21 -05:00
hhzhang16 c8770464ab
feat: normalize dynamo namespace computation (#5231)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-01-12 19:09:43 +00:00
Tanmay Verma abd4b5d9a0
chore: Upgrade to tensorrt_llm==1.2.0rc6.post1 (#5356) 2026-01-12 18:58:00 +00:00
hhzhang16 ec5630ead9
feat: mount model path to Profiler if specified (#5212)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-01-09 17:13:53 +00:00
ishandhanani 92748c937f
feat: sglang update to 0.5.7 (#5148) 2026-01-08 17:00:03 +08:00
Ryan McCormick 6306afa6f6
fix: Remove asymmetric --request-plane nats from run_engines.sh script (#5245) 2026-01-07 10:50:20 -08:00
Hongkuan Zhou a82137877c
fix: add sweep range for moe dgdr example (#5225)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-01-06 14:28:54 -08:00
Alec 19b5917c2b
fix: assign unique ports per vLLM worker to avoid ZMQ bind conflicts (#5224)
Signed-off-by: alec-flowers <aflowers@nvidia.com>
2026-01-06 13:45:28 -08:00
Hongkuan Zhou fbe6bb0a9e
feat: support PVC model cache in profiler (#5124)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-01-05 08:52:34 -08:00
Tushar Sharma cf433e6825
chore: update all copyright headers in repo to 2026 (#5130)
Signed-off-by: Tushar Sharma <tusharma@nvidia.com>
2026-01-02 22:08:23 +00:00
Nate Mailhot 3c7ed61d05
chore: TRTLLM 1.2.0rc6 (#5017)
Signed-off-by: Nate Mailhot <nmailhot@nvidia.com>
2026-01-02 10:29:04 -08:00