jh-nv
|
c263a99ed2
|
feat: propagate OTEL trace context across E/P/D multimodal workers (#7239)
|
2026-03-14 01:20:41 +00:00 |
Biswa Panda
|
947939c727
|
fix: populate logprobs bytes and token fields in OpenAI-compatible responses (#6953)
|
2026-03-13 22:59:29 +00:00 |
KrishnanPrash
|
4105de62ca
|
fix: reject multimodal requests when worker lacks --modality multimodal (#7065)
Signed-off-by: Krishnan Prashanth <kprashanth@nvidia.com>
|
2026-03-13 15:20:11 -07:00 |
Yan Ru Pei
|
7e07495f6c
|
feat(mocker): add --decode-speedup-ratio for speculative decoding simulation (#7349)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
|
2026-03-13 12:35:40 -07:00 |
hhzhang16
|
20f1c5a3c0
|
fix: inject tolerations for interpolation (profiling) (#7344)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
|
2026-03-13 18:44:50 +00:00 |
Wang, Yi
|
f29753dc89
|
feat: Add NVTX markers for sglang EPD (#7079)
Signed-off-by: Wang, Yi <yi.a.wang@intel.com>
|
2026-03-13 10:47:05 -07:00 |
Yan Ru Pei
|
bddaaa2658
|
feat(kv-router): pluggable scheduling policy for router queue [DYN-2454] (#7260)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
|
2026-03-13 04:16:49 +00:00 |
Tzu-Ling Kan
|
dba69e0f2f
|
chore(deps): bump vLLM 0.16.0 → 0.17.1 (#7170)
Signed-off-by: Tzu-Ling <tzulingk@nvidia.com>
|
2026-03-12 20:44:11 -07:00 |
Hongkuan Zhou
|
cd4773fb31
|
feat: ForwardPassMetrics dynamo event plane integration (#7250)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
|
2026-03-12 20:58:01 +00:00 |
Yan Ru Pei
|
bb07b2f455
|
feat(kv-router): ZMQ gap detection + replay for standalone indexer [LLM-126] (#7209)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
|
2026-03-12 10:46:01 -07:00 |
Ayush Agarwal
|
1182e2071d
|
feat: vllm omni image to video support (#6530)
Signed-off-by: ayushag <ayushag@nvidia.com>
|
2026-03-12 17:30:46 +00:00 |
hhzhang16
|
5c7e66ece1
|
docs: add docs for DGDR usage -- golden path (#6946)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
|
2026-03-12 17:18:09 +00:00 |
Daniel Socek
|
387100c8d5
|
feat: replaces PersistentConnector monkey-patch with proper nixl_conn… (#6913)
Signed-off-by: Daniel Socek <daniel.socek@intel.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>
|
2026-03-12 05:50:17 +00:00 |
Tushar Sharma
|
cdb7218a65
|
fix: address sglang failures on pre-merge/post-merge workflows (#7255)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-11 22:03:42 -04:00 |
Daniel Socek
|
6a46e3a0a1
|
chore(vllm): expand qwen3 vl multimodal support list (#6163)
Signed-off-by: Daniel Socek <daniel.socek@intel.com>
|
2026-03-12 01:10:22 +00:00 |
Daniel Socek
|
f01a5c7182
|
fix: vision model loader fixes (#6952)
Signed-off-by: Daniel Socek <daniel.socek@intel.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>
|
2026-03-11 17:33:22 -07:00 |
MatejKosec
|
5178a4a485
|
feat: streaming tool call and reasoning dispatch SSE events (#7114)
Signed-off-by: Matej Kosec <mkosec@nvidia.com>
|
2026-03-11 15:24:39 -07:00 |
GuanLuo
|
d16862ada6
|
chore: add context manager based timer (#7007)
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
|
2026-03-11 15:15:06 -07:00 |
daiyaanarfeen
|
5d5fd243da
|
feat: GlobalPlanner --max-total-gpus for cluster-wide GPU budget (#7103)
Signed-off-by: Daiyaan <darfeen@nvidia.com>
Signed-off-by: athreesh <anish.maddipoti@utexas.edu>
Signed-off-by: Anish <80174047+athreesh@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: athreesh <anish.maddipoti@utexas.edu>
Co-authored-by: Anish <80174047+athreesh@users.noreply.github.com>
|
2026-03-11 21:54:55 +00:00 |
Vladislav Nosivskoy
|
cf5f65f7df
|
feat: add generate health check support for PD SGLang (#6004)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
|
2026-03-11 14:24:23 -07:00 |
hhzhang16
|
5fd39ade53
|
feat: apply DGD overrides before running interpolation (#7226)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
|
2026-03-11 20:49:45 +00:00 |
Graham King
|
fda022b1c3
|
test(frontend): Minimal integration test for vllm processor (#7173)
Signed-off-by: Graham King <grahamk@nvidia.com>
|
2026-03-11 15:47:49 -04:00 |
Hongkuan Zhou
|
af30b779d5
|
feat: forward pass metric via ZMQ in vllm (#7200)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
|
2026-03-11 11:07:58 -07:00 |
Yuewei Na
|
e930526b14
|
fix: serialize disagg first_gen_log_probs int keys for Rust transport (#7145)
Signed-off-by: Yuewei Na <nv-yna@users.noreply.github.com>
Co-authored-by: Yuewei Na <nv-yna@users.noreply.github.com>
|
2026-03-11 06:24:37 +00:00 |
hhzhang16
|
611e856d87
|
feat: hide optimizationType (#7160)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
|
2026-03-11 00:36:18 +00:00 |
MatejKosec
|
012236ee4e
|
feat(anthropic): add thinking block support and preamble stripping to /v1/messages (#7137)
Signed-off-by: Matej Kosec <mkosec@nvidia.com>
|
2026-03-11 00:07:09 +00:00 |
Hongkuan Zhou
|
f435da1470
|
fix: set correct component type in agg planner (#7176)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
|
2026-03-10 15:25:35 -07:00 |
hhzhang16
|
0c6a802487
|
fix: resolve 'auto' backend to concrete value in every situation (#7158)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
|
2026-03-10 13:35:41 -07:00 |
Hongkuan Zhou
|
14d928cbf4
|
fix: only add profiling data to mocker workers (#7164)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
|
2026-03-10 12:58:33 -07:00 |
zhongdaor-nv
|
6634f33f27
|
fix(perf): Skip duplicate image downloads and unnecessary image processing in MM Router (vLLM) (#7080)
Signed-off-by: zhongdaor <zhongdaor@nvidia.com>
Signed-off-by: zhongdaor-nv <zhongdaor@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-10 12:19:36 -07:00 |
Yan Ru Pei
|
236cb17d00
|
fix(mocker): align vLLM scheduler with v1 — drop watermark, LIFO preemption, retry loop (#7020)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-10 09:53:50 +00:00 |
hhzhang16
|
50818575d1
|
fix: propagate resolved backend and skip interpolation for aggregated configs (#7106)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
|
2026-03-09 19:48:48 -07:00 |
ishandhanani
|
51dfd76045
|
feat: add SGLang chat processor for frontend pre/post processing (#6834)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-03-09 23:27:34 +00:00 |
GuanLuo
|
2cc92bfa9c
|
fix: restrict dummy embedding value range for bypassing vLLM check in E/P/D (#7117)
Signed-off-by: Guan Luo <41310872+GuanLuo@users.noreply.github.com>
|
2026-03-09 23:12:40 +00:00 |
Hongkuan Zhou
|
41b357d872
|
fix: use profiler DGD gen route in naive fallback mode (#7099)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
|
2026-03-09 15:28:32 -07:00 |
hhzhang16
|
175d9196d0
|
fix: strip apiVersion/kind/metadata from overrides.dgd before merging (#7109)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
|
2026-03-09 14:36:03 -07:00 |
Hongkuan Zhou
|
8a0657cbe3
|
fix: correct planner entrypoint in profiler's planner config gen (#7095)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
|
2026-03-09 13:36:12 -07:00 |
Anant Sharma
|
809c04c985
|
revert: "fix: strip apiVersion/kind/metadata from overrides.dgd before merging" (#7100)
|
2026-03-09 14:44:28 -04:00 |
hhzhang16
|
440d72eeb5
|
fix: strip apiVersion/kind/metadata from overrides.dgd before merging (#7025)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
|
2026-03-09 12:59:01 -05:00 |
hhzhang16
|
71f9e7a9de
|
fix: normalize GPUSKU to AIC system identifier (#6984)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
|
2026-03-09 12:58:39 -05:00 |
Schwinn Saereesitthipitak
|
6831020f35
|
chore: rename chrek to Dynamo Snapshot (#7028)
Signed-off-by: Schwinn Saereesitthipitak <17022745+galletas1712@users.noreply.github.com>
|
2026-03-08 15:13:36 -04:00 |
Graham King
|
76c96c5d6c
|
chore(frontend): Remove the debug_perf flag used for perf work (#7024)
Signed-off-by: Graham King <grahamk@nvidia.com>
|
2026-03-06 22:47:04 +00:00 |
Yifan Jiang
|
100819299f
|
feat(trtllm): add additional metrics for dynamo-trtllm (#6668)
Signed-off-by: Yifan Jiang <19356972+yifjiang@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-03-06 14:34:38 -08:00 |
jh-nv
|
2831bfecde
|
chore: add mypy typing to vllm (#6858)
|
2026-03-06 21:45:43 +00:00 |
Graham King
|
35b0ce6296
|
chore(frontend): Remove the multi-processing vllm processor path (#7005)
Signed-off-by: Graham King <grahamk@nvidia.com>
|
2026-03-06 20:57:45 +00:00 |
Schwinn Saereesitthipitak
|
c09a9aad10
|
fix: guard SGLang/vLLM memory occupation control endpoints (#6967)
|
2026-03-06 10:09:25 -08:00 |
Graham King
|
abc02c689f
|
fix: llm/mocker: Remove the llm -> mocker crate dependency, move config (#6998)
Signed-off-by: Graham King <grahamk@nvidia.com>
Signed-off-by: PeaBrane <yanrpei@gmail.com>
|
2026-03-06 09:41:39 -08:00 |
Yan Ru Pei
|
f934744fb9
|
fix: restore --enforce-disagg to reject requests before prefill router activates (#6957)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
|
2026-03-06 04:39:28 +00:00 |
Michal Guzek
|
e6ddf0eaf9
|
fix: TRT-LLM multimodal preprocessor - remove default_multimodal_input_loader from the embedding paths (#6924)
Signed-off-by: Michal Guzek <mguzek@nvidia.com>
|
2026-03-06 01:03:10 +00:00 |
hhzhang16
|
b97fde10f9
|
fix: propagate tolerations and cap auto-discovered GPUs (#6947)
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
|
2026-03-06 00:33:19 +00:00 |