Commit Graph

766 Commits

Author SHA1 Message Date
jh-nv c263a99ed2
feat: propagate OTEL trace context across E/P/D multimodal workers (#7239) 2026-03-14 01:20:41 +00:00
Biswa Panda 947939c727
fix: populate logprobs bytes and token fields in OpenAI-compatible responses (#6953) 2026-03-13 22:59:29 +00:00
KrishnanPrash 4105de62ca
fix: reject multimodal requests when worker lacks --modality multimodal (#7065)
Signed-off-by: Krishnan Prashanth <kprashanth@nvidia.com>
2026-03-13 15:20:11 -07:00
Yan Ru Pei 7e07495f6c
feat(mocker): add --decode-speedup-ratio for speculative decoding simulation (#7349)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-03-13 12:35:40 -07:00
hhzhang16 20f1c5a3c0
fix: inject tolerations for interpolation (profiling) (#7344)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-03-13 18:44:50 +00:00
Wang, Yi f29753dc89
feat: Add NVTX markers for sglang EPD (#7079)
Signed-off-by: Wang, Yi <yi.a.wang@intel.com>
2026-03-13 10:47:05 -07:00
Yan Ru Pei bddaaa2658
feat(kv-router): pluggable scheduling policy for router queue [DYN-2454] (#7260)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-03-13 04:16:49 +00:00
Tzu-Ling Kan dba69e0f2f
chore(deps): bump vLLM 0.16.0 → 0.17.1 (#7170)
Signed-off-by: Tzu-Ling <tzulingk@nvidia.com>
2026-03-12 20:44:11 -07:00
Hongkuan Zhou cd4773fb31
feat: ForwardPassMetrics dynamo event plane integration (#7250)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-03-12 20:58:01 +00:00
Yan Ru Pei bb07b2f455
feat(kv-router): ZMQ gap detection + replay for standalone indexer [LLM-126] (#7209)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-03-12 10:46:01 -07:00
Ayush Agarwal 1182e2071d
feat: vllm omni image to video support (#6530)
Signed-off-by: ayushag <ayushag@nvidia.com>
2026-03-12 17:30:46 +00:00
hhzhang16 5c7e66ece1
docs: add docs for DGDR usage -- golden path (#6946)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-03-12 17:18:09 +00:00
Daniel Socek 387100c8d5
feat: replaces PersistentConnector monkey-patch with proper nixl_conn… (#6913)
Signed-off-by: Daniel Socek <daniel.socek@intel.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>
2026-03-12 05:50:17 +00:00
Tushar Sharma cdb7218a65
fix: address sglang failures on pre-merge/post-merge workflows (#7255)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-11 22:03:42 -04:00
Daniel Socek 6a46e3a0a1
chore(vllm): expand qwen3 vl multimodal support list (#6163)
Signed-off-by: Daniel Socek <daniel.socek@intel.com>
2026-03-12 01:10:22 +00:00
Daniel Socek f01a5c7182
fix: vision model loader fixes (#6952)
Signed-off-by: Daniel Socek <daniel.socek@intel.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>
2026-03-11 17:33:22 -07:00
MatejKosec 5178a4a485
feat: streaming tool call and reasoning dispatch SSE events (#7114)
Signed-off-by: Matej Kosec <mkosec@nvidia.com>
2026-03-11 15:24:39 -07:00
GuanLuo d16862ada6
chore: add context manager based timer (#7007)
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
2026-03-11 15:15:06 -07:00
daiyaanarfeen 5d5fd243da
feat: GlobalPlanner --max-total-gpus for cluster-wide GPU budget (#7103)
Signed-off-by: Daiyaan <darfeen@nvidia.com>
Signed-off-by: athreesh <anish.maddipoti@utexas.edu>
Signed-off-by: Anish <80174047+athreesh@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: athreesh <anish.maddipoti@utexas.edu>
Co-authored-by: Anish <80174047+athreesh@users.noreply.github.com>
2026-03-11 21:54:55 +00:00
Vladislav Nosivskoy cf5f65f7df
feat: add generate health check support for PD SGLang (#6004)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
2026-03-11 14:24:23 -07:00
hhzhang16 5fd39ade53
feat: apply DGD overrides before running interpolation (#7226)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-03-11 20:49:45 +00:00
Graham King fda022b1c3
test(frontend): Minimal integration test for vllm processor (#7173)
Signed-off-by: Graham King <grahamk@nvidia.com>
2026-03-11 15:47:49 -04:00
Hongkuan Zhou af30b779d5
feat: forward pass metric via ZMQ in vllm (#7200)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-03-11 11:07:58 -07:00
Yuewei Na e930526b14
fix: serialize disagg first_gen_log_probs int keys for Rust transport (#7145)
Signed-off-by: Yuewei Na <nv-yna@users.noreply.github.com>
Co-authored-by: Yuewei Na <nv-yna@users.noreply.github.com>
2026-03-11 06:24:37 +00:00
hhzhang16 611e856d87
feat: hide optimizationType (#7160)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-03-11 00:36:18 +00:00
MatejKosec 012236ee4e
feat(anthropic): add thinking block support and preamble stripping to /v1/messages (#7137)
Signed-off-by: Matej Kosec <mkosec@nvidia.com>
2026-03-11 00:07:09 +00:00
Hongkuan Zhou f435da1470
fix: set correct component type in agg planner (#7176)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-03-10 15:25:35 -07:00
hhzhang16 0c6a802487
fix: resolve 'auto' backend to concrete value in every situation (#7158)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-03-10 13:35:41 -07:00
Hongkuan Zhou 14d928cbf4
fix: only add profiling data to mocker workers (#7164)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-03-10 12:58:33 -07:00
zhongdaor-nv 6634f33f27
fix(perf): Skip duplicate image downloads and unnecessary image processing in MM Router (vLLM) (#7080)
Signed-off-by: zhongdaor <zhongdaor@nvidia.com>
Signed-off-by: zhongdaor-nv <zhongdaor@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-10 12:19:36 -07:00
Yan Ru Pei 236cb17d00
fix(mocker): align vLLM scheduler with v1 — drop watermark, LIFO preemption, retry loop (#7020)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-10 09:53:50 +00:00
hhzhang16 50818575d1
fix: propagate resolved backend and skip interpolation for aggregated configs (#7106)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-03-09 19:48:48 -07:00
ishandhanani 51dfd76045
feat: add SGLang chat processor for frontend pre/post processing (#6834)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 23:27:34 +00:00
GuanLuo 2cc92bfa9c
fix: restrict dummy embedding value range for bypassing vLLM check in E/P/D (#7117)
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
2026-03-09 23:12:40 +00:00
Hongkuan Zhou 41b357d872
fix: use profiler DGD gen route in naive fallback mode (#7099)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-03-09 15:28:32 -07:00
hhzhang16 175d9196d0
fix: strip apiVersion/kind/metadata from overrides.dgd before merging (#7109)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-03-09 14:36:03 -07:00
Hongkuan Zhou 8a0657cbe3
fix: correct planner entrypoint in profiler's planner config gen (#7095)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-03-09 13:36:12 -07:00
Anant Sharma 809c04c985
revert: "fix: strip apiVersion/kind/metadata from overrides.dgd before merging" (#7100) 2026-03-09 14:44:28 -04:00
hhzhang16 440d72eeb5
fix: strip apiVersion/kind/metadata from overrides.dgd before merging (#7025)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-03-09 12:59:01 -05:00
hhzhang16 71f9e7a9de
fix: normalize GPUSKU to AIC system identifier (#6984)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-03-09 12:58:39 -05:00
Schwinn Saereesitthipitak 6831020f35
chore: rename chrek to Dynamo Snapshot (#7028)
Signed-off-by: Schwinn Saereesitthipitak <17022745+galletas1712@users.noreply.github.com>
2026-03-08 15:13:36 -04:00
Graham King 76c96c5d6c
chore(frontend): Remove the debug_perf flag used for perf work (#7024)
Signed-off-by: Graham King <grahamk@nvidia.com>
2026-03-06 22:47:04 +00:00
Yifan Jiang 100819299f
feat(trtllm): add additional metrics for dynamo-trtllm (#6668)
Signed-off-by: Yifan Jiang <19356972+yifjiang@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 14:34:38 -08:00
jh-nv 2831bfecde
chore: add mypy typing to vllm (#6858) 2026-03-06 21:45:43 +00:00
Graham King 35b0ce6296
chore(frontend): Remove the multi-processing vllm processor path (#7005)
Signed-off-by: Graham King <grahamk@nvidia.com>
2026-03-06 20:57:45 +00:00
Schwinn Saereesitthipitak c09a9aad10
fix: guard SGLang/vLLM memory occupation control endpoints (#6967) 2026-03-06 10:09:25 -08:00
Graham King abc02c689f
fix: llm/mocker: Remove the llm -> mocker crate dependency, move config (#6998)
Signed-off-by: Graham King <grahamk@nvidia.com>
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-03-06 09:41:39 -08:00
Yan Ru Pei f934744fb9
fix: restore --enforce-disagg to reject requests before prefill router activates (#6957)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-03-06 04:39:28 +00:00
Michal Guzek e6ddf0eaf9
fix: TRT-LLM multimodal preprocessor - remove default_multimodal_input_loader from the embedding paths (#6924)
Signed-off-by: Michal Guzek <mguzek@nvidia.com>
2026-03-06 01:03:10 +00:00
hhzhang16 b97fde10f9
fix: propagate tolerations and cap auto-discovered GPUs (#6947)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-03-06 00:33:19 +00:00