Schwinn Saereesitthipitak
a6d970e961
feat: refactor GMS client memory manager with tiered API ( #6549 )
2026-02-25 23:24:17 +00:00
daiyaanarfeen
651ef5b506
feat: throughput-metrics-source for SLA planner + GlobalPlanner disagg scaling ( #6500 )
...
Signed-off-by: Daiyaan <darfeen@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-25 22:58:24 +00:00
Tzu-Ling Kan
c18b475851
feat: Use --request-rate and --request-rate-mode for aiper client ( #6585 )
...
Signed-off-by: Tzu-Ling <tzulingk@nvidia.com>
2026-02-25 14:20:44 -08:00
Tzu-Ling Kan
539117546a
fix: Parse the latest successful attempt ( #6565 )
...
Signed-off-by: Tzu-Ling <tzulingk@nvidia.com>
2026-02-25 10:55:25 -07:00
Yan Ru Pei
65dc451db9
feat: add ZMQ KV event publishing to mocker [DYN-2221] ( #6528 )
...
Signed-off-by: Ru Pei <rupei@nvidia.com>
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-24 18:57:28 -08:00
KrishnanPrash
8e2363758e
fix: restore E/P/D multimodal disagg serving and add Qwen3-VL-30B-A3B support ( #6533 )
...
Signed-off-by: Krishnan Prashanth <kprashanth@nvidia.com>
Signed-off-by: KrishnanPrash <140860868+KrishnanPrash@users.noreply.github.com>
2026-02-24 14:13:10 -08:00
Biswa Panda
75fea78797
feat: make KV cache events and routing LoRA-aware ( #6517 )
2026-02-24 18:35:08 +00:00
Yan Ru Pei
08db2844b8
test: more stringent test for routing decisions ( #6531 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-02-24 10:22:38 -08:00
Alec
6d3b92f04e
feat: remove --connector flag for vLLM backend (LLM-90) ( #6450 )
...
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-24 17:49:00 +00:00
Yan Ru Pei
c9ff623538
chore: replace --enforce-disagg with --decode-fallback, default to enforcing disagg ( #6515 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-02-24 09:24:11 -08:00
Alec
7893f2684e
feat: add --disaggregation-mode enum to vLLM backend ( #6483 )
...
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 19:34:30 -07:00
Richard Huo
15d217606a
chore: revert the kvbm workaround since trtllm v1.3.0rc3 is upgraded ( #6495 )
2026-02-23 15:05:04 -08:00
Tzu-Ling Kan
80cac7c14b
feat: Remove Component from public ( #6403 )
...
Signed-off-by: tzulingk@nvidia.com <tzulingk@nvidia.com>
2026-02-23 22:58:38 +00:00
Yan Ru Pei
038b50d2b6
feat: add standalone KV indexer with query endpoint [DYN-2164] ( #6446 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-23 12:12:25 -08:00
jh-nv
ea86df2985
feat: Backend accept new requests during shutdown grace period ( #6093 )
...
Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com>
Co-authored-by: Jacky <18255193+kthui@users.noreply.github.com>
2026-02-23 03:18:46 +00:00
Alec
4ebb244b2a
feat: add --headless mode for multi-node TP/PP in dynamo.vllm ( #6204 )
...
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-22 22:38:17 +00:00
knarangN
da7d3e9e2a
test: add TensorRT-LLM multimodal EPD test for nightly CI ( #6193 )
...
Signed-off-by: Kavita Narang <knarang@nvidia.com>
2026-02-20 22:38:34 +00:00
Jacky
41d7d5490f
feat: Configurable Request Cancellation abort passage to TRT-LLM ( #6445 )
...
Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com>
2026-02-20 10:36:18 -08:00
Qi Wang
c82fe88847
feat: add embedding cache to pd worker ( #6061 )
2026-02-20 08:27:44 -08:00
Alec
7bbacce196
feat: default kv-events-config to empty (align with vLLM defaults) ( #6404 )
...
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-20 06:44:51 +00:00
zhongdaor-nv
23de4e86ae
feat: e2e mm aware kv cache routing support for trtllm backend ( #5480 )
...
Signed-off-by: zhongdaor <zhongdaor@nvidia.com>
Signed-off-by: zhongdaor-nv <zhongdaor@nvidia.com>
Signed-off-by: Zhongdao Ren <zhongdaor@zhongdaor-mlt.client.nvidia.com>
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
Co-authored-by: Zhongdao Ren <zhongdaor@zhongdaor-mlt.client.nvidia.com>
2026-02-19 20:52:00 -07:00
Dmitry Tokarev
121d805020
fix: Fix async pytests - added missing marker ( #6439 )
...
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com>
2026-02-19 19:06:13 -05:00
Tzu-Ling Kan
0ce3461a9e
feat: Add runtime.endpoint() method to eliminate namespace chaining ( #6386 )
...
Signed-off-by: tzulingk@nvidia.com <tzulingk@nvidia.com>
2026-02-19 16:02:37 -05:00
knarangN
6f4b33f7b5
test: add vllm audio tests to nightly ci pipeline ( #6392 )
...
Signed-off-by: Kavita Narang <knarang@nvidia.com>
2026-02-19 12:55:07 -08:00
jh-nv
44a76f96b3
refactor: update frontend kv-router flags to be consistent with router ( #6361 )
2026-02-19 00:27:15 +00:00
Hongkuan Zhou
b2075619c9
feat: planner argparse CLI -> config file ( #6356 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-02-18 23:17:36 +00:00
knarangN
638d8e68f3
test: Add multimodal video tests to nightly CI pipeline ( #6023 )
...
Signed-off-by: Kavita Narang <knarang@nvidia.com>
2026-02-18 18:54:09 +00:00
Alec
c02cefb5cb
fix: reduce pytest-marker-report output noise and move to tests/ ( #6359 )
...
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 12:39:04 -05:00
ishandhanani
b6603d90d6
feat: add Anthropic Messages API endpoint (/v1/messages) ( #6231 )
...
Signed-off-by: Matej Kosec <mkosec@nvidia.com>
Signed-off-by: Marko Kosec <mkosec@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Matej Kosec <mkosec@nvidia.com>
2026-02-18 06:11:52 +00:00
Alec
d86937f955
test: relocate test output to /tmp to keep git working tree clean ( #6289 )
...
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
2026-02-17 21:21:01 -08:00
Qi Wang
d2a5783938
refactor: delete handlers and disagg EC producer/consumer ( #6051 )
2026-02-17 18:05:54 -08:00
daiyaanarfeen
0a26665303
feat: add GlobalPlanner component for centralized scaling ( #5702 )
...
Signed-off-by: Daiyaan <darfeen@nvidia.com>
Signed-off-by: daiyaanarfeen <darfeen@nvidia.com>
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
2026-02-17 23:12:10 +00:00
Hongkuan Zhou
359765d354
feat: load-based scaling in SLA Planner ( #6145 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-02-14 09:43:02 -08:00
zhongdaor-nv
815b129126
feat: mm aware routing for vllm ( #6235 )
...
Signed-off-by: zhongdaor <zhongdaor@nvidia.com>
Signed-off-by: zhongdaor-nv <zhongdaor@nvidia.com>
2026-02-13 21:16:06 -08:00
hhzhang16
d56439ec28
feat: migrate GPU discovery from Dynamo Profiler to Dynamo Operator with automatic injection ( #6224 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-02-14 00:34:21 +00:00
Tzu-Ling Kan
5624d14481
Rename fetch_llm to fetch_model ( #6268 )
...
Signed-off-by: tzulingk@nvidia.com <tzulingk@nvidia.com>
2026-02-13 23:22:35 +00:00
mohammedabdulwahhab
a289695c37
fix: consolidate dyn_discovery_backend and dyn_kv_store ( #6167 )
...
Signed-off-by: mohammedabdulwahhab <furkhan324@berkeley.edu>
2026-02-13 19:07:11 +00:00
Yan Ru Pei
14eceb43df
chore: rename KvPushRouter to KvRouter in python + more bindings removal ( #6238 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2026-02-12 23:44:52 -08:00
Keiven C
166e1f4d94
feat: use dynamic port allocation for DYN_SYSTEM_PORT in e2e router t… ( #6262 )
...
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
2026-02-12 20:37:50 -08:00
Graham King
bbe82f182a
chore: Remove dynamo-run and mistral-rs engine ( #6203 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2026-02-12 21:50:58 +00:00
Hongkuan Zhou
a04b56310c
feat: support AIC DGD gen call (WILL BREAK DGDR) ( #6216 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2026-02-12 08:47:55 -08:00
Yan Ru Pei
937398cf61
feat: Flash Indexer ( #5785 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Signed-off-by: jthomson04 <jothomson@nvidia.com>
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Signed-off-by: Janelle Cai <jcai18@mit.edu>
Co-authored-by: jthomson04 <jwillthomson19@gmail.com>
Co-authored-by: Janelle Cai <jcai18@mit.edu>
2026-02-11 22:26:56 -08:00
MatejKosec
8cb47d04d2
feat: responses API compliance with upstream type alignment ( #6089 )
...
Signed-off-by: Matej Kosec <mkosec@nvidia.com>
Co-authored-by: Ishan Dhanani <ishandhanani@gmail.com>
2026-02-12 01:18:28 +00:00
Alec
f8d0a9f9c6
test: change KVBM vLLM interface tests to gpu_0 marker ( #6205 )
...
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
2026-02-11 16:31:21 -08:00
Jonathan Tong
39d645e586
docs: migrate Fern docs from fern/ into docs/ ( #6206 )
...
Signed-off-by: Jont828 <jt572@cornell.edu>
2026-02-11 16:22:27 -08:00
MatejKosec
45bc1b798c
fix: use tempfile.TemporaryDirectory in load generator test ( #6196 )
...
Signed-off-by: Matej Kosec <mkosec@nvidia.com>
2026-02-11 16:15:03 -05:00
Dmitry Tokarev
4220771fba
fix: Cleanup pytest markers, enable gpu_0 tests on trtllm arm, reduce log noise ( #6124 )
...
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com>
2026-02-11 16:00:58 -05:00
Olga Andreeva
f9d20c1097
test: adding interface contract tests for KVBM vLLM integration ( #5847 )
2026-02-11 10:42:05 -08:00
Dillon Cullinan
3188c70a3b
chore: Templating Feedback Followup ( #6125 )
...
Signed-off-by: Dillon Cullinan <dcullinan@nvidia.com>
2026-02-11 13:09:27 -05:00
Keiven C
e18840cef3
feat: add Prometheus auto and custom label injection for engine metrics ( #5989 )
...
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
2026-02-10 21:22:53 -08:00