Graham King
bdad6f1a50
fix: Make planner VirtualConnectorClient also use v1/ prefix. ( #3468 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-10-07 22:34:30 +00:00
Graham King
a5371bfc50
feat(etcd): Version the etcd keys ( #3458 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-10-07 16:48:26 +00:00
Graham King
81162dfeb9
chore(discovery): Watch/publish ModelDeploymentCard instead of ModelEntry ( #3350 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-10-07 10:02:02 -04:00
Yan Ru Pei
30610e7371
feat: use KvPushRouter for prefill router ( #3401 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-10-03 18:21:48 -07:00
Olga Andreeva
d2e3b66e65
feat: Transition to FullyContiguous Host and Disk layouts ( #3090 )
...
Signed-off-by: Olga Andreeva <oandreeva@nvidia.com>
Signed-off-by: Olga Andreeva <124622579+oandreeva-nv@users.noreply.github.com>
Co-authored-by: oandreeva-nv <oandreeva-nv@nvidia.com>
2025-10-01 16:26:39 -07:00
Richard Huo
713e9e481b
fix: DIS-706 skip offloading the G1 matched blocks during offloading ( #3299 )
...
Signed-off-by: richardhuo-nv <rihuo@nvidia.com>
2025-10-01 10:28:06 -07:00
Yan Ru Pei
9b9536d0d7
feat: make prefill router general ( #3329 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-09-30 20:38:47 -07:00
Keiven C
f4a3a6b66a
refactor: standardize Prometheus metric naming conventions (part 1) ( #3035 )
...
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
2025-09-30 19:01:39 -07:00
Michael Feil
5b457b70a6
feat: python add abi compatability for cross-platform builds + add a unit test to HttpServer ( #3044 )
...
Signed-off-by: michaelfeil <me@michaelfeil.eu>
Signed-off-by: Michael Feil <63565275+michaelfeil@users.noreply.github.com>
Signed-off-by: root <root@michaelfeil2-dev-pod-b200-0.michaelfeil2-dev-pod-b200.baseten.svc.cluster.local>
Signed-off-by: root <root@michaelfeildns-dev-pod-h100-0.michaelfeildns-dev-pod-h100.baseten.svc.cluster.local>
Co-authored-by: root <root@michaelfeil2-dev-pod-b200-0.michaelfeil2-dev-pod-b200.baseten.svc.cluster.local>
Co-authored-by: root <root@michaelfeildns-dev-pod-h100-0.michaelfeildns-dev-pod-h100.baseten.svc.cluster.local>
2025-09-30 14:42:42 -07:00
Yan Ru Pei
d354763c40
fix: python bindings for router should register to etcd as well ( #3302 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-09-30 12:38:56 -07:00
Keiven C
cacac9b9f4
feat: add Python const for Prometheus metric names ( #3244 )
...
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
2025-09-30 10:28:47 -07:00
Yan Ru Pei
3aa3077808
fix: more fixes for stable router benchmarking ( #3264 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-09-29 23:08:56 +00:00
Ziqi Fan
e21dcf6cab
feat: enable KVBM emit metrics in Dynamo TRTLLM ( #3254 )
...
Signed-off-by: Ziqi Fan <ziqif@nvidia.com>
2025-09-29 08:48:51 -07:00
Graham King
7ebbd001c0
chore: Remove etcd from python bindings ( #3238 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-09-26 12:58:30 -04:00
Alec
5bb7490448
chore: bump vllm version to 0.10.2 ( #3180 )
...
Signed-off-by: Alec <aflowers@nvidia.com>
Signed-off-by: Alec <35311602+alec-flowers@users.noreply.github.com>
Signed-off-by: krishung5 <krish@nvidia.com>
Co-authored-by: Kris Hung <krish@nvidia.com>
2025-09-26 04:01:57 +00:00
Neelay Shah
f2e2e935b2
feat: Add distributed tracing context support to Python bindings ( #3160 )
...
Signed-off-by: nnshah1 <neelays@nvidia.com>
Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com>
Co-authored-by: Jacky <18255193+kthui@users.noreply.github.com>
2025-09-25 17:26:35 -07:00
Graham King
c03e2f6bbe
chore: Migrate planner virtual_connector internals into bindings ( #3205 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-09-25 18:04:58 -04:00
GuanLuo
6ba64c31f5
feat: tensor type for generic inference. ( #2746 )
...
Signed-off-by: Guan Luo <gluo@nvidia.com>
Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
Co-authored-by: Olga Andreeva <124622579+oandreeva-nv@users.noreply.github.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>
2025-09-24 07:31:09 +00:00
Yan Ru Pei
031590fc14
feat: vllm prefill router ( #3155 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-09-22 16:38:56 -07:00
Graham King
7a5a0bd6cd
chore: Upgrade Rust to 1.90 ( #3147 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-09-19 18:54:29 -04:00
Yan Ru Pei
5b19a39bab
feat: allow router to not track active blocks (prefill), and to not track cached blocks (decode) ( #3135 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-09-19 20:33:07 +00:00
Graham King
3865a94148
feat: Port vllm port allocator to Rust in bindings ( #3125 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-09-19 18:34:52 +00:00
Olga Andreeva
2d39f1b1bb
feat: KVBM connector : enabling vectorized copy from pinned memory to device memory and vice versa ( #2989 )
...
Signed-off-by: Olga Andreeva <oandreeva@nvidia.com>
Signed-off-by: oandreeva-nv <oandreeva-nv@nvidia.com>
Co-authored-by: Ziqi Fan <ziqif@nvidia.com>
Co-authored-by: oandreeva-nv <oandreeva-nv@nvidia.com>
2025-09-19 09:35:07 -07:00
Graham King
b6595e2484
chore(bindings): Provide a binding to clear etcd namespace ( #3094 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-09-18 11:49:08 -04:00
Yan Ru Pei
78a3fedab9
fix: hook up worker removals for indexer ( #3095 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-09-17 21:33:25 +00:00
Graham King
f88d7dc74b
chore(bindings): Remove NatsQueue ( #3086 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-09-17 12:46:54 -04:00
Graham King
9060ce12ce
feat: Make part of discovery re-usable ( #3073 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-09-17 10:43:24 -04:00
Tzu-Ling Kan
08cb08c1bc
feat: Canary Health Check. ( #2903 )
...
Signed-off-by: tzulingk@nvidia.com <tzulingk@nvidia.com>
2025-09-17 03:31:56 +00:00
Graham King
723f2da74b
chore: Remove more extended Apache headers ( #3063 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-09-16 15:48:51 -04:00
Graham King
87e6e0529d
fix: Interactive inputs actually stops, does not ignore stop token ( #3057 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-09-16 14:39:05 -04:00
Ziqi Fan
55659eae70
fix: early stop if CPU or disk space not set when using KVBM ( #2997 )
...
Signed-off-by: Ziqi Fan <ziqif@nvidia.com>
2025-09-15 10:14:01 -07:00
blarson-b10
4000097653
feat: adds kv indexer metrics ( #2905 )
...
Signed-off-by: Brian Larson <brian.larson@baseten.co>
2025-09-10 21:33:15 +00:00
Graham King
cb5a657a6a
fix: Load the tokenizer JSON once for chat and completions. ( #2910 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-09-05 16:48:16 -04:00
Olga Andreeva
27fad26faf
refactor: Split ModelType to ModelInput for request and response type; ModelType for the supported workloads ( #2714 )
...
Signed-off-by: Guan Luo <gluo@nvidia.com>
Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
Co-authored-by: Guan Luo <gluo@nvidia.com>
Co-authored-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
2025-09-03 22:22:37 +00:00
KrishnanPrash
c920cbd9dc
feat: Add --custom-jinja-template argument to pass a custom chat template for vLLM ( #2829 )
...
Signed-off-by: Krishnan Prashanth <kprashanth@nvidia.com>
2025-09-03 14:22:15 -07:00
Biswa Panda
c6becbc859
feat: dynamo namespace isolation ( #2394 )
...
Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
2025-09-03 15:47:06 +00:00
Yan Ru Pei
383e3b3a52
feat: don't modify kv scheduler states on query + more python binding ( #2798 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-09-02 17:09:43 -07:00
Ayush Agarwal
87a721a83e
feat: added parser name bindings ( #2808 )
...
Signed-off-by: Ayush Agarwal <ayushag@nvidia.com>
2025-09-02 20:37:29 +00:00
Jacky
6c539fbdac
feat: FT Request Cancellation feature and test for 0.5.0 ( #2500 )
2025-09-02 08:26:21 -07:00
Yan Ru Pei
7fabe7bfe2
fix: do not delete KV events jetstream ( #2800 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-09-01 21:25:49 +00:00
Yan Ru Pei
488c87095c
feat: Router warm restarts via durable KV event consumers and radix snapshotting ( #2756 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-08-30 23:42:57 +00:00
Richard Huo
a68c2f8f12
feat: DIS-373 dynamo KVBM connector API integration with TRTLLM ( #2544 )
...
Signed-off-by: richardhuo-nv <rihuo@nvidia.com>
2025-08-29 19:27:15 -07:00
Keiven C
15539fd093
feat: add Prometheus metrics integration for KvStats ( #2704 )
...
Signed-off-by: Keiven C <213854356+keivenchang@users.noreply.github.com>
2025-08-28 22:12:22 -07:00
Yan Ru Pei
f08729ae15
feat: python bindings for the entire KvPushRouter + per-request router configs ( #2658 )
2025-08-25 22:27:10 +00:00
Ziqi Fan
b39382ba68
feat: add initial batch of KVBM metrics on match, offload and onboard ( #2673 )
2025-08-25 09:28:25 -07:00
Ayush Agarwal
cbe854fc5f
feat: [vLLM] implement cli args for tool and reasoning parsers ( #2619 )
2025-08-22 20:29:18 +00:00
Ziqi Fan
b658ba6139
feat: enable dynamo metrics on KVBM ( #2626 )
2025-08-22 19:58:05 +00:00
Graham King
6a358f7c8c
chore(llm): Rename protocols::Endpoint to EndpointId ( #2615 )
2025-08-22 15:07:34 +00:00
Michael Feil
174389e6d3
fix: Httpengine sync-enable-endpoint ( #2591 )
2025-08-21 18:29:52 -04:00
Tzu-Ling Kan
57728909cf
feat: Add model label for vllm backend metrics ( #2474 )
...
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>
2025-08-21 17:32:34 +00:00
Michael Feil
626d7e182d
feat(request cancellation): pycontext, propagating the `is_stopped` into python land. ( #2158 )
2025-08-20 12:06:13 -07:00
Ryan Olson
07cfc3a11b
feat: kvbm + connector ( #2258 )
...
Signed-off-by: Ryan Olson <rolson@nvidia.com>
Co-authored-by: Olga Andreeva <oandreeva@nvidia.com>
Co-authored-by: Ziqi Fan <ziqif@nvidia.com>
Co-authored-by: John Thompson <jothomson@nvidia.com>
Co-authored-by: Richard Huo <rihuo@nvidia.com>
Co-authored-by: Zicheng Ma <zichengm@nvidia.com>
2025-08-19 12:36:53 -07:00
suzu
c5d9d26703
feat(frontend): support setting HTTP host via CLI (--http-host) ( #2523 )
2025-08-19 09:01:04 -04:00
Yan Ru Pei
85d8310806
feat: router-level request rejection ( #2465 )
2025-08-19 01:51:00 -07:00
Graham King
a4bbe49228
feat(http): TLS support ( #2492 )
2025-08-18 16:06:29 -04:00
Keiven C
acbdabc464
feat(metrics): add NATS client metrics to prometheus_metrics_fmt ( #2292 )
...
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>
2025-08-14 19:33:53 -07:00
Tzu-Ling Kan
3a3f5bf275
feat: Add a "model" label to Component metrics ( #2389 )
2025-08-14 13:48:25 -05:00
Jorge António
d0a6363584
feat: add RuntimeConfig to ModelEntry ( #2311 )
...
Co-authored-by: Yan Ru Pei <yanrpei@gmail.com>
2025-08-14 10:19:53 -07:00
Graham King
72ec5f5c7b
feat: Allow an endpoint to serve multiple models ( #2418 )
2025-08-13 10:48:42 -04:00
Yan Ru Pei
5166a3dd44
feat: Router replicas with state-sharing ( #2264 )
2025-08-07 22:57:57 +00:00
Graham King
1954fcfa09
chore: Remove service_name from ModelDeploymentCard ( #2349 )
2025-08-07 11:02:41 -04:00
Graham King
6a1a801c2d
feat: Support static workers, run without etcd. ( #2281 )
2025-08-06 09:51:43 -04:00
Hongkuan Zhou
36c4ef5eb2
feat: migrate requests when planner shutdown decode engine (vllm) ( #2280 )
...
Signed-off-by: Hongkuan Zhou <tedzhouhk@gmail.com>
Co-authored-by: Jacky <18255193+kthui@users.noreply.github.com>
Co-authored-by: hhzhang16 <54051230+hhzhang16@users.noreply.github.com>
2025-08-05 12:24:07 -07:00
Jacky
347620a1ff
feat: Allow Python Engine to end stream before final ( #2270 )
2025-08-05 10:44:18 -07:00
Chi
433f60121a
feat: Pass user_data to register_llm for LoRA support ( #2286 )
2025-08-05 10:16:43 -04:00
Yan Ru Pei
bae25dc6d4
feat: skip downloading model weights if using mocker (only tokenizer) ( #2213 )
2025-07-31 18:15:57 +00:00
Jacky
1f07dab7bd
feat: Add migration to LLM requests ( #1930 )
2025-07-18 20:04:20 +00:00
Graham King
fc12436048
feat(frontend): router-mode settings ( #2001 )
2025-07-18 18:52:57 +00:00
Graham King
182d3b5dc7
chore(bindings): Remove mistralrs / llama.cpp ( #1970 )
2025-07-16 16:12:40 -04:00
Yan Ru Pei
f31732a22d
feat: integrate mocker with dynamo-run and python cli ( #1927 )
2025-07-16 18:22:15 +00:00
Graham King
2bf27924a1
feat(python): Python bindings for the Dynamo CLI tools ( #1799 )
2025-07-08 21:40:36 +00:00
Yan Ru Pei
84e71e27d3
feat: predictive active blocks for routing without load metrics ( #1731 )
...
Signed-off-by: Yan Ru Pei <yanrpei@gmail.com>
Co-authored-by: Alec <35311602+alec-flowers@users.noreply.github.com>
2025-07-08 00:18:22 -07:00
jain-ria
439e977d9c
feat: vllm speculative decoding metrics ( #1549 )
...
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: Alec <35311602+alec-flowers@users.noreply.github.com>
2025-07-07 23:40:43 +00:00
Jacky
b4ddca99a0
feat: Failure Detection while Responses are returning ( #1671 )
2025-07-07 14:00:17 -07:00
Alec
0a32b3443f
fix: default to None initialization of routing config ( #1713 )
2025-07-01 10:55:05 -07:00
jthomson04
6365a015b3
fix: Fix main ( #1712 )
2025-06-30 22:08:06 -07:00
jthomson04
aaf283bbb8
feat: Approximate KV Routing ( #1636 )
2025-06-30 20:34:08 -07:00
Graham King
92f06b0e7f
chore(dynamo-run): Refactor to library ( #1687 )
...
Move much of what was in the `dynamo-run` crate into `dynamo-llm` so that everyone can use it.
Example usage:
1. Create a `LocalModel`:
```
let local_model = LocalModelBuilder::default()
.model_path("Qwen/Qwen3-0.6B")
.http_port(8080)
.build().await?;
```
2. Make an engine:
```
let engine_config = EngineConfig::StaticFull {
engine: dynamo_engine_mistralrs::make_engine(&local_model).await?,
model: Box::new(local_model),
};
```
3. Connect it to an input and run it
```
dynamo_llm::entrypoint::input::run_input(Input::Http, runtime, engine_config).await?;
```
For https://github.com/ai-dynamo/dynamo/issues/1647
Code Rabbit summary, thanks:
* Introduced a flexible builder pattern for local model configuration, allowing advanced customization and easier initialization.
* Added new input modes and unified input handling, supporting interactive chat, HTTP server, batch file, and distributed endpoint modes.
* Centralized engine configuration and routing, enabling more extensible and maintainable engine management.
* Simplified and modularized the codebase by moving input and engine logic into dedicated modules.
* Replaced direct model construction with an asynchronous builder for improved clarity and extensibility.
* Streamlined configuration and validation for flags and router settings.
* Added validation to prevent incompatible input and output combinations in endpoint and dynamic modes.
2025-06-30 21:06:24 +00:00
Yan Ru Pei
8392e7a190
feat: Unnormalize waiting requests + predictive load updates for Python router (mirroring Rust) + softmax sampling to reduce thrashing ( #1638 )
2025-06-27 09:01:59 +00:00
Yan Ru Pei
13a99b7f76
feat: Standalone Router ( #1409 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Signed-off-by: Yan Ru Pei <yanrpei@gmail.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
2025-06-14 19:02:31 +00:00
Graham King
3f6a74723f
chore: Remove PreprocessedRequest alias BackendInput ( #1307 )
...
It was confusing to have two names for one type.
This tidy up started in #1064 , is now complete.
2025-06-02 11:22:27 -04:00
Alec
2f8da9ad1d
refactor: rename KvMetricsPublisher to WorkerMetricsPublisher ( #1284 )
2025-05-30 17:23:32 +00:00
jthomson04
9210a26d90
refactor: Refactor kv event publishers ( #1287 )
2025-05-30 09:04:28 -07:00
Alec
f67dc38b28
fix: Renamed event publisher classes and configuration ( #1273 )
2025-05-29 13:00:06 -07:00
Jacky
7677f74f37
feat: KVBM async Python bindings and Layer class ( #1141 )
2025-05-29 17:49:05 +00:00
Alec
0df6d462d7
feat: add KV Event Publishing to vLLM v1 ( #1181 )
2025-05-29 07:56:49 -07:00
Graham King
0a1d1fbe3d
feat(dynamo-llm): Remove bring-your-own-engine ( #1216 )
...
It was removed from the docs in 0.2.1 and replaced with writing a [standalone Python engine](https://github.com/ai-dynamo/dynamo/blob/main/docs/guides/dynamo_run.md#writing-your-own-engine-in-python ).
Also remove the associated `dynamo-run` feature `python`.
Releasing this in 0.3.0 will resolve #784 and #1109 .
2025-05-28 18:52:27 +00:00
Graham King
183f2b3286
feat(dynamo-run): Allow setting KV cache block size ( #1175 )
...
Example:
```
dynamo-run out=<engine> <model> --kv-cache-block-size 64
```
In a distributed system this goes on the worker node and is propagated to ingress via the model deployment card.
Previously hard coded to 16, which is now the default.
- Load context_length from model. Closes #1172
- Store context length and KV cache block size in Model Deployment Card #1170
2025-05-22 14:55:36 -07:00
Graham King
3e8e38a942
fix(llmctl): Use ModelWatcher instead of direct etcd operations ( #1150 )
2025-05-21 19:06:56 +00:00
Ryan Olson
80256acf15
feat: adding outer dimension to isolate k/v blocks ( #1126 )
2025-05-20 08:51:41 -04:00
Graham King
aeb79e6277
feat: Support multiple models on single ingress node ( #1127 )
...
We can now do this:
- Node 1:
```
dynamo-run in=http out=dyn
```
- Node 2 and 3, two instances of component 'backend' in the nemotron_ultra pipeline:
```
dynamo-run in=dyn://nemotron_ultra.backend.generate out=vllm /data/models/NemotronUltra
```
- Node 4 and 5, two instances of the 'backend' component in nemotron_super pipeline:
```
dynamo-run in=dyn://nemotron_super.backend.generate out=vllm /data/models/NemotronSuper
```
The ingress node will discover all four instances and route correctly. We have been planning for this for a long time now.
As part of this auto-discovery is now always `out=dyn`, with no extra URL parts. Previously it could only route to a single pipeline.
Also:
- Refactor endpoint / instance naming now that I understand them
- Fix removing models when their instance stops.
2025-05-19 18:41:07 -04:00
Jacky
437cae0ad0
feat: KV Block Manager Python bindings ( #1022 )
2025-05-19 12:33:41 -04:00
Tom O'Brien
73fdfb8ab8
feat: Add OpenAI Embeddings interface in rust lib ( #1110 )
...
Implements OpenAI embeddings (interface only).
- Adds ModelType::Embedding
- Adds OpenAI embedding request/response structs
- Adds support for embedding model discovery
2025-05-19 11:59:46 -04:00
Graham King
2981350880
feat(dynamo-run): KV-aware routing ( #1064 )
...
Router:
```
dynamo-run in=http out=dyn://dynamo.endpoint.generate --router-mode kv
```
Worker (* N):
```
dynamo-run in=dyn://dynamo.endpoint.generate out=vllm /data/llms/Qwen/Qwen3-4B
```
You need patched vllm and the C bindings `.so`. Full docs in the updated guide: `docs/guides/dynamo_run.md`.
This gives us a pure-Rust ingress node: OpenAI compliant HTTP server + Pre-processor + KV-aware router.
2025-05-14 15:21:46 -04:00
Hongkuan Zhou
466b8e5fb7
refactor: use primary lease + self-contained graceful shutdown trigged by SIGINT/SIGTERM ( #1001 )
2025-05-08 16:53:39 -07:00
Hongkuan Zhou
a590d10317
feat: cleanup EtcdKvCache and PrefillQueue before and after launch ( #925 )
2025-05-07 16:35:55 -07:00
Graham King
92bbbc3923
fix: Fix vllm/sglang engine model name if using HF repo ( #986 )
...
Signed-off-by: Graham King <graham@gkgk.org>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>
2025-05-07 23:20:51 +00:00
jthomson04
c4213899fe
feat: Migrate NATS Queue to Rust ( #669 ) ( #961 )
2025-05-06 15:36:55 -07:00
Graham King
99cd9d85a9
feat: dynamo-run <-> python interop ( #934 )
...
Adding this to a Python script makes it register on the network so that `dynamo-run` can discover it and send it requests:
```
from dynamo.llm import register_llm
MODEL = "Qwen/Qwen2.5-0.5B-Instruct"
await register_llm(endpoint, MODEL, 3)
```
Full vllm example, with pre-processing in dynamo:
- `dynamo-run in=text out=dyn://dynamo.backend.generate`
- `cd lib/bindings/python/examples/hello_world`
- `python server_vllm.py`
This builds on top of the work to move pre-processor to ingress side. It means we can decouple Rust and Python using NATS as the bus.
The `register_llm` call does this:
- Download the model from HF if necessary
- Load the model deployment card from the HF folder or extract from GGUF
- Push the tokenizer config etc into NATS object store so ingress can access it from a different machine
- Publish the model deployment card to ETCD
2025-05-05 20:49:19 -04:00
Hongkuan Zhou
9d643f1eed
fix: use primary lease for NixlMetadataStore ( #928 )
2025-05-05 08:38:52 -07:00
Graham King
a1a10365a7
chore: Split PushRouter from Client ( #817 )
...
In a distributed system we don't know if the remote workers need pre-processing done ingress-side or not. Previously Client required us to decide this before discovering the remote endpoints, which was fine because pre-processing was worker-side.
As part of moving pre-processing back to ingress-side we need to split this into two steps:
- Client discovers the endpoints, and (later PR) will fetch their Model Deployment Card.
- PushRouter will use the Model Deployment Card to decide if they need pre-processing or not, which affects the types of the generic parameters.
Part of #743
2025-04-29 14:39:56 +00:00
Hongkuan Zhou
7d5d6f8c08
feat: local planner for 0.2.0 release ( #398 )
...
Signed-off-by: Hongkuan Zhou <tedzhouhk@gmail.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: ishandhanani <ishandhanani@gmail.com>
Co-authored-by: Ubuntu <ubuntu@dev-inst-2w1vokvyuts83rzn4n1k7mnzew9.us-central1-a.c.brevdevprod.internal>
Co-authored-by: Biswa Panda <biswa.panda@gmail.com>
Co-authored-by: Anant Sharma <anants@nvidia.com>
2025-04-26 00:05:54 +00:00
Pankaj Gupta
420b7a82be
fix: Fix cancellation flow in python component graph ( #765 )
2025-04-21 16:07:49 -04:00
ishandhanani
c392c3419b
feat: add custom lease to worker components ( #748 )
2025-04-21 17:43:02 +00:00
Hongkuan Zhou
4c38680eac
feat: gracefully shutdown endpoint by revoking etcd lease + python binding ( #730 )
...
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
2025-04-18 18:36:50 +00:00
Hongkuan Zhou
08fd28978c
feat: ETCD prefix watcher + python binding + runtime reconfiguration for router and disagg router ( #581 )
...
Signed-off-by: Hongkuan Zhou <tedzhouhk@gmail.com>
Co-authored-by: Neelay Shah <neelays@nvidia.com>
2025-04-11 18:15:42 -07:00
Cole
447840c223
docs: add docstring for llm.rs ( #267 )
2025-04-10 22:40:09 -04:00
Yan Ru Pei
4b6cfc1be0
feat: KV recorder for dumping router events into a jsonl ( #505 )
2025-04-04 15:31:32 -07:00
Graham King
88ad3425c4
feat: Python decorator dynamo_worker takes optional `static` parameter without etcd ( #494 )
...
Adds `@dynamo_worker(static = True)` to create a static worker which has a predictable name and hence does not require discovery or `etcd` to be running. There can only be a single static worker per namespace / component / endpoint trio.
This contrasts with the default dynamic `dynamo_worker` endpoints we have now, which get a unique random name (based on namespace/component/endpoint), and are discovered by ingress components using etcd.
Also change the hello_world example to use `dynamo_worker(static = True)` so that it is exercised and demonstrated somewhere.
For NIM.
2025-04-04 09:26:52 -04:00
Ryan Olson
84985d3f1d
refactor: migrate engines to standalone crates ( #453 )
...
Moved all of `lib/llm/src/engines` to their own crates as e.g. `lib/engines/mistralrs`. This will allow publishing of the `dynamo-llm` crate as it won't have any github dependencies.
The only engines in dynamo-llm will be the demo `echo` ones.
Co-authored-by: Graham King <grahamk@nvidia.com>
2025-04-03 21:53:54 +00:00
Ryan Olson
c4106e6a27
feat: kv aware router executable ( #399 )
2025-04-02 17:58:21 +02:00
Ryan Olson
5b682f4839
feat: unified logging ( #472 )
2025-04-01 19:49:20 +00:00
GuanLuo
6e09681e0e
feat: expose Python binding for KVEventPublisher. Use event pub/sub trait for KV events ( #169 )
2025-03-17 00:17:33 -07:00
Alec
3f84cdadfa
feat: add new metrics and simple router cost fn ( #88 )
2025-03-11 11:53:21 -07:00
Biswa Panda
dd6208254d
feat: add openai http service ( #82 )
2025-03-10 17:27:22 -07:00
Alec
989bb3d59c
feat: make block_size input for indexer, router, publisher ( #66 )
2025-03-09 15:58:27 -07:00
Hongkuan Zhou
19844fc07e
feat: kv aware router + disagg router + prefill queue ( #11 )
...
Signed-off-by: Hongkuan Zhou <tedzhouhk@gmail.com>
Co-authored-by: hongkuan <hongkuanz@nvidia.com>
Co-authored-by: Piotr Tarasiewicz <ptarasiewicz@nvidia.com>
Co-authored-by: Piotr Tarasiewicz Nvidia <ptarasiewicznv@Piotrs-MacBook-Pro.local>
Co-authored-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: Neelay Shah <neelays@nvidia.com>
2025-03-08 17:09:11 -08:00
Neelay Shah
602352ce19
chore: rename dynamo ( #44 )
...
Co-authored-by: Biswa Panda <biswa.panda@gmail.com>
2025-03-08 01:42:40 -08:00
Graham King
12714d9080
feat: Python bring-your-own-engine with our tokenizer ( #47 )
...
Instead of using `out=pystr:<my.py>` we can now do this:
```
dynemo-run out=pytok:/home/graham/my_python_engine.py --model-path <hf-repo-checkout>
```
That engine will receive and respond with tokens. Here's an example engine file:
```
import asyncio
async def generate(request):
yield {"token_ids":[791]}
await asyncio.sleep(0.1)
yield {"token_ids":[6864]}
await asyncio.sleep(0.1)
yield {"token_ids":[315]}
await asyncio.sleep(0.1)
yield {"token_ids":[9822]}
await asyncio.sleep(0.1)
yield {"token_ids":[374]}
await asyncio.sleep(0.1)
yield {"token_ids":[12366]}
await asyncio.sleep(0.1)
yield {"token_ids":[13]}
```
Also reduce duplication by making the bindings engine use the llm lib engine.
2025-03-07 10:48:37 -05:00
GuanLuo
e159e53fe6
feat: expose KV routing components for easier router customization ( #15 )
2025-03-05 16:06:19 -08:00
Neelay Shah
1af7433bff
refactor: rename triton_distributed to dynemo ( #22 )
...
Co-authored-by: Graham King <grahamk@nvidia.com>
2025-03-05 09:15:46 -08:00
Biswa Panda
a32cdad622
feat: add python binding for rust llm modules ( #13 )
2025-03-04 14:52:09 -08:00
Neelay Shah
3a5fe17db9
feat: nixl metadata store and retrieved from etcd ( #6 )
...
Co-authored-by: hongkuanz <hongkuanz@nvidia.com>
Co-authored-by: Piotr Tarasiewicz <ptarasiewicz@nvidia.com>
Co-authored-by: Piotr Tarasiewicz Nvidia <ptarasiewicznv@Piotrs-MacBook-Pro.local>
Co-authored-by: Neelay Shah <neelays@ipp2-0493.ipp2u1.colossus.nvidia.com>
Co-authored-by: Neelay Shah <neelays@ipp1-1941.ipp1a1.colossus.nvidia.com>
Co-authored-by: ishandhanani <ishandhanani@gmail.com>
Co-authored-by: Neelay Shah <neelays@4u8g-gen-0078.ipp3a2.colossus.nvidia.com>
Co-authored-by: ptarasiewiczNV <104908264+ptarasiewiczNV@users.noreply.github.com>
2025-03-04 13:10:16 -08:00
Alec
11a36651a6
[fix] KV Router Example fixes ( #314 )
...
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>
2025-02-28 12:20:47 -08:00
Ryan Olson
85cc7b67c6
refactor: service/endpoint stats_handler ( #282 )
2025-02-27 11:30:18 -07:00
Alec
b760c5694d
feat: Add completion endpoint to http server and llmctl ( #230 )
...
Co-authored-by: aflowers <aflowers@nvidia.com>
2025-02-25 12:32:09 -08:00
GuanLuo
861c50982b
feat: enable metrics polling
...
Signed-off-by: Meenakshi Sharma <163925564+nvda-mesharma@users.noreply.github.com>
Signed-off-by: Neelay Shah <neelays@nvidia.com>
Co-authored-by: Neelay Shah <neelays@nvidia.com>
Co-authored-by: Ryan Olson <ryanolson@users.noreply.github.com>
Co-authored-by: Meenakshi Sharma <163925564+nvda-mesharma@users.noreply.github.com>
Co-authored-by: Biswa Panda <biswapanda@users.noreply.github.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>
2025-02-25 11:12:54 -08:00
Ryan McCormick
c06b95ffdb
ci: Add rust checks to missing directories ( #239 )
...
Signed-off-by: Ryan McCormick <rmccormick@nvidia.com>
2025-02-25 07:42:33 -08:00
Neelay Shah
08fcd7e93b
refactor: move libs to lib dir
...
Signed-off-by: Neelay Shah <neelays@nvidia.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>
2025-02-24 18:21:02 -08:00