Alec
|
e10787894d
|
fix: add periodic vLLM engine stats logging and fix log routing (#6566)
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-02-25 23:45:34 +00:00 |
daiyaanarfeen
|
651ef5b506
|
feat: throughput-metrics-source for SLA planner + GlobalPlanner disagg scaling (#6500)
Signed-off-by: Daiyaan <darfeen@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
2026-02-25 22:58:24 +00:00 |
Nikita
|
21fce9ba0d
|
feat: Tiktoken support (#6460)
Signed-off-by: Nikita Sukharev <kaonael@gmail.com>
|
2026-02-25 12:46:47 -08:00 |
Yan Ru Pei
|
ebaf048d16
|
chore: plumb allowed_worker_ids through RoutingHints (#6580)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
|
2026-02-25 12:12:53 -08:00 |
Keiven C
|
ff06b17e7f
|
fix: guarantee RouterRequestMetrics availability & documentation updates (#6558)
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-02-25 10:18:01 -08:00 |
atchernych
|
c916cd42ff
|
feat: Support epp's "pods" interface in Dynamo fixes [DEP-424] (#6302)
Signed-off-by: Anna Tchernych <atchernych@nvidia.com>
|
2026-02-25 01:54:38 +00:00 |
Keiven C
|
1df620b4d0
|
feat: router metrics with dynamo_router_* {worker_id=...}. Update docs (#6227)
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-02-24 21:34:14 +00:00 |
Yan Ru Pei
|
bc00ef3879
|
chore: split JetStream subscriber into dedicated module and deprecate durable_kv_events [DYN-2203] (#6477)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
|
2026-02-24 11:30:38 -08:00 |
Biswa Panda
|
75fea78797
|
feat: make KV cache events and routing LoRA-aware (#6517)
|
2026-02-24 18:35:08 +00:00 |
Yan Ru Pei
|
c9ff623538
|
chore: replace --enforce-disagg with --decode-fallback, default to enforcing disagg (#6515)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
|
2026-02-24 09:24:11 -08:00 |
Jacky
|
c8276cd28b
|
feat: Standardized Dynamo Error Type (#6303)
Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com>
|
2026-02-23 20:23:58 -08:00 |
Tzu-Ling Kan
|
80cac7c14b
|
feat: Remove Component from public (#6403)
Signed-off-by: tzulingk@nvidia.com <tzulingk@nvidia.com>
|
2026-02-23 22:58:38 +00:00 |
Thomas Montfort
|
cb55766c47
|
feat(runtime): add hierarchical Model/WorkerSet architecture for multi-namespace support (#6054)
Signed-off-by: tmontfort <tmontfort@nvidia.com>
|
2026-02-23 22:44:38 +00:00 |
Yan Ru Pei
|
038b50d2b6
|
feat: add standalone KV indexer with query endpoint [DYN-2164] (#6446)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-02-23 12:12:25 -08:00 |
Ryan Olson
|
61c6780469
|
feat: add kvbm-kernels crate and upgrade cudarc to 0.19 (#6309)
Signed-off-by: Ryan Olson <rolson@nvidia.com>
|
2026-02-23 16:15:28 +00:00 |
Tzu-Ling Kan
|
42d6980546
|
feat: Remove public uses of CancellationToken (#6405)
Signed-off-by: tzulingk@nvidia.com <tzulingk@nvidia.com>
|
2026-02-20 21:01:13 -08:00 |
Ryan Olson
|
7a6283e927
|
chore: cargo machete - remove unused dependencies (#6453)
Signed-off-by: Ryan Olson <rolson@nvidia.com>
|
2026-02-20 19:11:52 +00:00 |
zhongdaor-nv
|
23de4e86ae
|
feat: e2e mm aware kv cache routing support for trtllm backend (#5480)
Signed-off-by: zhongdaor <zhongdaor@nvidia.com>
Signed-off-by: zhongdaor-nv <zhongdaor@nvidia.com>
Signed-off-by: Zhongdao Ren <zhongdaor@zhongdaor-mlt.client.nvidia.com>
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
Co-authored-by: Zhongdao Ren <zhongdaor@zhongdaor-mlt.client.nvidia.com>
|
2026-02-19 20:52:00 -07:00 |
Tzu-Ling Kan
|
0ce3461a9e
|
feat: Add runtime.endpoint() method to eliminate namespace chaining (#6386)
Signed-off-by: tzulingk@nvidia.com <tzulingk@nvidia.com>
|
2026-02-19 16:02:37 -05:00 |
Yan Ru Pei
|
fc229004b5
|
chore: Remove ZmqKvEventListener binding and rework standalone TRT-LLM example to use native Python ZMQ (#6164)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
|
2026-02-19 04:15:13 +00:00 |
Tzu-Ling Kan
|
bc8f1170ce
|
feat: Move ModelDeploymentCard to _internal (#6378)
Signed-off-by: tzulingk@nvidia.com <tzulingk@nvidia.com>
|
2026-02-18 23:12:18 +00:00 |
Yan Ru Pei
|
cf51a0c4fd
|
chore: gate plotters dependency behind bench feature (#6380)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-02-18 15:27:19 -05:00 |
Anant Sharma
|
aa16ccf545
|
fix: remove default-members in workspace (#6279)
Signed-off-by: Anant Sharma <anants@nvidia.com>
|
2026-02-18 12:39:17 -05:00 |
Yan Ru Pei
|
c5c6a55190
|
feat: more flash indexer optimizations (#6305)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-02-18 09:13:37 -08:00 |
Ayush Agarwal
|
2ace5a4a3d
|
chore: multimodal endpoint registration via cli modality (#6270)
Signed-off-by: ayushag <ayushag@nvidia.com>
|
2026-02-18 07:59:39 +00:00 |
zhongdaor-nv
|
815b129126
|
feat: mm aware routing for vllm (#6235)
Signed-off-by: zhongdaor <zhongdaor@nvidia.com>
Signed-off-by: zhongdaor-nv <zhongdaor@nvidia.com>
|
2026-02-13 21:16:06 -08:00 |
Tzu-Ling Kan
|
5624d14481
|
Rename fetch_llm to fetch_model (#6268)
Signed-off-by: tzulingk@nvidia.com <tzulingk@nvidia.com>
|
2026-02-13 23:22:35 +00:00 |
Yan Ru Pei
|
e30696e5aa
|
chore: allow frontend + mockers to run on macos (#6282)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
|
2026-02-13 19:33:45 +00:00 |
mohammedabdulwahhab
|
a289695c37
|
fix: consolidate dyn_discovery_backend and dyn_kv_store (#6167)
Signed-off-by: mohammedabdulwahhab <furkhan324@berkeley.edu>
|
2026-02-13 19:07:11 +00:00 |
ishandhanani
|
2be83be2d8
|
feat: add video generation support (T2V) (#5793)
|
2026-02-13 09:09:56 +00:00 |
Yan Ru Pei
|
14eceb43df
|
chore: rename KvPushRouter to KvRouter in python + more bindings removal (#6238)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
|
2026-02-12 23:44:52 -08:00 |
atchernych
|
5227176057
|
feat: Decomposed pipeline for EPP integration [DEP-730] (#5446)
Signed-off-by: Anna Tchernych <atchernych@nvidia.com>
|
2026-02-13 06:18:29 +00:00 |
Yan Ru Pei
|
bc514fbee6
|
feat: router priority queue (#6010)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
|
2026-02-13 02:43:17 +00:00 |
Yongming Ding
|
2d517e7780
|
feat(mocker): improve mocker's perf timing accuracy (#6100)
Signed-off-by: Yongming Ding <yongmingd@nvidia.com>
|
2026-02-12 17:33:43 -08:00 |
Karen Chung
|
cd6984b951
|
feat: use RNG when dp routing targets are tied; override no-assume-kv-reuse for decode requests (#6253)
Signed-off-by: Karen Chung <karenc@nvidia.com>
|
2026-02-12 15:59:37 -08:00 |
Graham King
|
bbe82f182a
|
chore: Remove dynamo-run and mistral-rs engine (#6203)
Signed-off-by: Graham King <grahamk@nvidia.com>
|
2026-02-12 21:50:58 +00:00 |
Yan Ru Pei
|
937398cf61
|
feat: Flash Indexer (#5785)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Signed-off-by: jthomson04 <jothomson@nvidia.com>
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Signed-off-by: Janelle Cai <jcai18@mit.edu>
Co-authored-by: jthomson04 <jwillthomson19@gmail.com>
Co-authored-by: Janelle Cai <jcai18@mit.edu>
|
2026-02-11 22:26:56 -08:00 |
Jonathan Tong
|
39d645e586
|
docs: migrate Fern docs from fern/ into docs/ (#6206)
Signed-off-by: Jont828 <jt572@cornell.edu>
|
2026-02-11 16:22:27 -08:00 |
Yan Ru Pei
|
6a728d1044
|
chore: remove kv indexers bindings (#6159)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
|
2026-02-11 20:23:48 +00:00 |
Graham King
|
73dc3be81a
|
chore: Merge bindings client and client2 functions (#6158)
Signed-off-by: Graham King <grahamk@nvidia.com>
|
2026-02-11 11:51:01 -05:00 |
Yan Ru Pei
|
f46720c996
|
feat: add router-level Prometheus metrics and centralize request tracking (#6146)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-02-11 07:36:08 +00:00 |
Keiven C
|
e18840cef3
|
feat: add Prometheus auto and custom label injection for engine metrics (#5989)
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
|
2026-02-10 21:22:53 -08:00 |
Graham King
|
4f99451bb0
|
feat(frontend): Use vllm for pre and post processing (#5544)
Signed-off-by: Graham King <grahamk@nvidia.com>
|
2026-02-10 22:39:11 +00:00 |
Graham King
|
f1bcb17542
|
feat: Add metric tokenizer_latency_ms (#6092)
Signed-off-by: Graham King <grahamk@nvidia.com>
|
2026-02-10 08:52:31 -08:00 |
Keiven C
|
027d2653a5
|
feat: expose Python Prometheus metric via DynamoComponentMetrics (#5817)
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
|
2026-02-09 17:10:37 -08:00 |
Ayush Agarwal
|
9f76d0606c
|
feat: text to image vLLM Omni (#5912)
Signed-off-by: ayushag <ayushag@nvidia.com>
|
2026-02-09 18:19:27 -05:00 |
Yan Ru Pei
|
6783bdcaa9
|
chore: enable local indexers by default, and use normal event plane by default (not jetstream) (#5941)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
|
2026-02-08 20:48:02 +00:00 |
Yongming Ding
|
7c25f70291
|
feat(mocker): add optional KV cache allocation/eviction trace (#6052)
Signed-off-by: Yongming Ding <yongmingd@nvidia.com>
Co-authored-by: Yan Ru Pei <yanrpei@gmail.com>
|
2026-02-07 03:02:28 +00:00 |
Jacky
|
1ffa489ea1
|
refactor: Move --migration-limit flag from backend to frontend (#5918)
Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com>
|
2026-02-06 20:50:55 +00:00 |
Anant Sharma
|
3842b24479
|
fix: update bytes crate version to latest (#6041)
Signed-off-by: Anant Sharma <anants@nvidia.com>
|
2026-02-06 15:11:04 -05:00 |