Olga Andreeva
27fad26faf
refactor: Split ModelType to ModelInput for request and response type; ModelType for the supported workloads ( #2714 )
...
Signed-off-by: Guan Luo <gluo@nvidia.com>
Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
Co-authored-by: Guan Luo <gluo@nvidia.com>
Co-authored-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
2025-09-03 22:22:37 +00:00
KrishnanPrash
c920cbd9dc
feat: Add --custom-jinja-template argument to pass a custom chat template for vLLM ( #2829 )
...
Signed-off-by: Krishnan Prashanth <kprashanth@nvidia.com>
2025-09-03 14:22:15 -07:00
Biswa Panda
c6becbc859
feat: dynamo namespace isolation ( #2394 )
...
Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
2025-09-03 15:47:06 +00:00
Yan Ru Pei
383e3b3a52
feat: don't modify kv scheduler states on query + more python binding ( #2798 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-09-02 17:09:43 -07:00
Ayush Agarwal
87a721a83e
feat: added parser name bindings ( #2808 )
...
Signed-off-by: Ayush Agarwal <ayushag@nvidia.com>
2025-09-02 20:37:29 +00:00
Harrison Saturley-Hall
561ecb98a2
chore: bump version numbers ahead of 0.5.0 release ( #2812 )
...
Signed-off-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com>
2025-09-02 16:13:28 -04:00
Jacky
6c539fbdac
feat: FT Request Cancellation feature and test for 0.5.0 ( #2500 )
2025-09-02 08:26:21 -07:00
Yan Ru Pei
7fabe7bfe2
fix: do not delete KV events jetstream ( #2800 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-09-01 21:25:49 +00:00
Yan Ru Pei
488c87095c
feat: Router warm restarts via durable KV event consumers and radix snapshotting ( #2756 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-08-30 23:42:57 +00:00
Richard Huo
a68c2f8f12
feat: DIS-373 dynamo KVBM connector API integration with TRTLLM ( #2544 )
...
Signed-off-by: richardhuo-nv <rihuo@nvidia.com>
2025-08-29 19:27:15 -07:00
Bhuvan Agrawal
43a26958d9
feat: add logits processor support for trtllm backend ( #2702 )
...
Signed-off-by: Bhuvan Agrawal <11240550+bhuvan002@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2025-08-30 01:06:03 +00:00
Keiven C
15539fd093
feat: add Prometheus metrics integration for KvStats ( #2704 )
...
Signed-off-by: Keiven C <213854356+keivenchang@users.noreply.github.com>
2025-08-28 22:12:22 -07:00
KavinKrishnan
95ce83d59a
feat: Integrate Model Express Client into Dynamo Model Downloads ( #2574 )
...
Signed-off-by: Kavin Krishnan <kavink@nvidia.com>
Co-authored-by: KavinKrishnan <kavin.krishnan@nvidia.com>
2025-08-28 16:09:04 -07:00
GuanLuo
91a459c038
feat: KServe gRPC support ( #2638 )
2025-08-26 22:57:31 -07:00
Tzu-Ling Kan
ef535edb98
feat: Trtllm metric_labels. ( #2666 )
2025-08-27 05:50:05 +00:00
Yan Ru Pei
f08729ae15
feat: python bindings for the entire KvPushRouter + per-request router configs ( #2658 )
2025-08-25 22:27:10 +00:00
nachiketb-nvidia
3036e60b1e
feat: add gpt oss reasoning parser through harmony ( #2656 )
...
- couple of refactors
- added a new dependency, openai-harmony
- implemented the gpt oss parser
2025-08-25 17:13:38 +00:00
Ziqi Fan
b39382ba68
feat: add initial batch of KVBM metrics on match, offload and onboard ( #2673 )
2025-08-25 09:28:25 -07:00
Alec
9b9f2ce4f6
fix: pytest robustness and parsing error ( #2676 )
2025-08-24 10:13:41 -07:00
Ayush Agarwal
cbe854fc5f
feat: [vLLM] implement cli args for tool and reasoning parsers ( #2619 )
2025-08-22 20:29:18 +00:00
Ziqi Fan
b658ba6139
feat: enable dynamo metrics on KVBM ( #2626 )
2025-08-22 19:58:05 +00:00
Bhuvan Agrawal
b92a805edb
feat: add BaseLogitsProcessor core interface ( #2613 )
...
Signed-off-by: Bhuvan Agrawal <11240550+bhuvan002@users.noreply.github.com>
2025-08-22 13:41:30 -04:00
Graham King
6a358f7c8c
chore(llm): Rename protocols::Endpoint to EndpointId ( #2615 )
2025-08-22 15:07:34 +00:00
Michael Feil
174389e6d3
fix: Httpengine sync-enable-endpoint ( #2591 )
2025-08-21 18:29:52 -04:00
Tzu-Ling Kan
57728909cf
feat: Add model label for vllm backend metrics ( #2474 )
...
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>
2025-08-21 17:32:34 +00:00
Michael Feil
626d7e182d
feat(request cancellation): pycontext, propagating the `is_stopped` into python land. ( #2158 )
2025-08-20 12:06:13 -07:00
Graham King
49958435eb
chore: Remove async-openai-macros ( #2554 )
2025-08-20 11:33:10 -07:00
Yan Ru Pei
d319abf3a3
feat: upload/download rust structs directly through NATs object store ( #2540 )
2025-08-20 17:17:21 +00:00
Dmitry Tokarev
9a02188531
chore: Bumped Dynamo version to 0.4.1 ( #2545 )
2025-08-19 21:55:37 -04:00
Dmitry Tokarev
177d662f86
fix: Dockerfile.sglang - Fixed sglang and dynamo wheels installaiton in ru… ( #2537 )
2025-08-19 20:56:12 -04:00
nachiketb-nvidia
199b9a30f4
chore: Bring async-openai into repo as request starter ( #2520 )
...
Co-authored-by: Graham King <grahamk@nvidia.com>
2025-08-19 17:47:01 -04:00
Ryan Olson
07cfc3a11b
feat: kvbm + connector ( #2258 )
...
Signed-off-by: Ryan Olson <rolson@nvidia.com>
Co-authored-by: Olga Andreeva <oandreeva@nvidia.com>
Co-authored-by: Ziqi Fan <ziqif@nvidia.com>
Co-authored-by: John Thompson <jothomson@nvidia.com>
Co-authored-by: Richard Huo <rihuo@nvidia.com>
Co-authored-by: Zicheng Ma <zichengm@nvidia.com>
2025-08-19 12:36:53 -07:00
Ryan Olson
a33033b7f6
feat: task scheduler ( #2406 )
...
Signed-off-by: Ryan Olson <ryanolson@users.noreply.github.com>
2025-08-19 12:35:31 -06:00
suzu
c5d9d26703
feat(frontend): support setting HTTP host via CLI (--http-host) ( #2523 )
2025-08-19 09:01:04 -04:00
Yan Ru Pei
85d8310806
feat: router-level request rejection ( #2465 )
2025-08-19 01:51:00 -07:00
Graham King
a4bbe49228
feat(http): TLS support ( #2492 )
2025-08-18 16:06:29 -04:00
Keiven C
0444217339
fix: replace metrics callback with background scraping to prevent tim… ( #2480 )
...
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>
2025-08-18 12:38:50 -07:00
Harrison Saturley-Hall
ffae72b7d2
fix: remove kvmanager feature from python 3.12 ai-dynamo-runtime wheel ( #2456 )
2025-08-15 16:43:55 -04:00
Keiven C
acbdabc464
feat(metrics): add NATS client metrics to prometheus_metrics_fmt ( #2292 )
...
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>
2025-08-14 19:33:53 -07:00
Tzu-Ling Kan
3a3f5bf275
feat: Add a "model" label to Component metrics ( #2389 )
2025-08-14 13:48:25 -05:00
Jorge António
d0a6363584
feat: add RuntimeConfig to ModelEntry ( #2311 )
...
Co-authored-by: Yan Ru Pei <yanrpei@gmail.com>
2025-08-14 10:19:53 -07:00
Dan Aloni
c12c25787f
fix: upgrade cudarc to 0.17.1 ( #2341 )
...
Signed-off-by: Dan Aloni <dan.aloni@vastdata.com>
Co-authored-by: Tushar Sharma <tusharma@nvidia.com>
2025-08-13 15:45:08 -04:00
Graham King
72ec5f5c7b
feat: Allow an endpoint to serve multiple models ( #2418 )
2025-08-13 10:48:42 -04:00
GuanLuo
9b87c89c41
feat: multi-modal example with vLLM v1 and UX v2 ( #2040 )
...
Co-authored-by: krishung5 <krish@nvidia.com>
2025-08-12 23:55:31 +00:00
Yan Ru Pei
5166a3dd44
feat: Router replicas with state-sharing ( #2264 )
2025-08-07 22:57:57 +00:00
Neelay Shah
bd4fe1a7d1
feat: cross process instrumentation ( #2243 )
...
Signed-off-by: Neelay Shah <neelays@nvidia.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: Olga Andreeva <124622579+oandreeva-nv@users.noreply.github.com>
2025-08-07 15:32:48 +00:00
Graham King
1954fcfa09
chore: Remove service_name from ModelDeploymentCard ( #2349 )
2025-08-07 11:02:41 -04:00
Graham King
dbe48a1d2b
chore: Bump mistral.rs, llama.cpp and tokenizers deps ( #2338 )
2025-08-06 16:23:54 -04:00
Dan Aloni
b2aa504b47
fix: upgrade axum to 0.8 and etcd-client to 0.16 ( #2317 )
...
Signed-off-by: Dan Aloni <dan.aloni@vastdata.com>
2025-08-06 18:25:49 +00:00
Graham King
6a1a801c2d
feat: Support static workers, run without etcd. ( #2281 )
2025-08-06 09:51:43 -04:00