Harrison Saturley-Hall
cd2389baef
chore: pre-0.6.0 activities ( #3592 )
...
Signed-off-by: Harrison Saturley-Hall <hsaturleyhal@nvidia.com>
2025-10-13 17:45:19 -04:00
Michael Feil
5b457b70a6
feat: python add abi compatability for cross-platform builds + add a unit test to HttpServer ( #3044 )
...
Signed-off-by: michaelfeil <me@michaelfeil.eu>
Signed-off-by: Michael Feil <63565275+michaelfeil@users.noreply.github.com>
Signed-off-by: root <root@michaelfeil2-dev-pod-b200-0.michaelfeil2-dev-pod-b200.baseten.svc.cluster.local>
Signed-off-by: root <root@michaelfeildns-dev-pod-h100-0.michaelfeildns-dev-pod-h100.baseten.svc.cluster.local>
Co-authored-by: root <root@michaelfeil2-dev-pod-b200-0.michaelfeil2-dev-pod-b200.baseten.svc.cluster.local>
Co-authored-by: root <root@michaelfeildns-dev-pod-h100-0.michaelfeildns-dev-pod-h100.baseten.svc.cluster.local>
2025-09-30 14:42:42 -07:00
Graham King
c03e2f6bbe
chore: Migrate planner virtual_connector internals into bindings ( #3205 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-09-25 18:04:58 -04:00
Harrison Saturley-Hall
980727bba1
chore: bump versions ahead of 0.5.1 release ( #3209 )
...
Signed-off-by: Harrison Saturley-Hall <hsaturleyhal@nvidia.com>
2025-09-24 18:26:38 -04:00
Graham King
7a5a0bd6cd
chore: Upgrade Rust to 1.90 ( #3147 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-09-19 18:54:29 -04:00
Graham King
3865a94148
feat: Port vllm port allocator to Rust in bindings ( #3125 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-09-19 18:34:52 +00:00
Ayush Agarwal
87a721a83e
feat: added parser name bindings ( #2808 )
...
Signed-off-by: Ayush Agarwal <ayushag@nvidia.com>
2025-09-02 20:37:29 +00:00
Harrison Saturley-Hall
561ecb98a2
chore: bump version numbers ahead of 0.5.0 release ( #2812 )
...
Signed-off-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com>
2025-09-02 16:13:28 -04:00
Ziqi Fan
b658ba6139
feat: enable dynamo metrics on KVBM ( #2626 )
2025-08-22 19:58:05 +00:00
Dmitry Tokarev
9a02188531
chore: Bumped Dynamo version to 0.4.1 ( #2545 )
2025-08-19 21:55:37 -04:00
nachiketb-nvidia
199b9a30f4
chore: Bring async-openai into repo as request starter ( #2520 )
...
Co-authored-by: Graham King <grahamk@nvidia.com>
2025-08-19 17:47:01 -04:00
Ryan Olson
07cfc3a11b
feat: kvbm + connector ( #2258 )
...
Signed-off-by: Ryan Olson <rolson@nvidia.com>
Co-authored-by: Olga Andreeva <oandreeva@nvidia.com>
Co-authored-by: Ziqi Fan <ziqif@nvidia.com>
Co-authored-by: John Thompson <jothomson@nvidia.com>
Co-authored-by: Richard Huo <rihuo@nvidia.com>
Co-authored-by: Zicheng Ma <zichengm@nvidia.com>
2025-08-19 12:36:53 -07:00
Harrison Saturley-Hall
ffae72b7d2
fix: remove kvmanager feature from python 3.12 ai-dynamo-runtime wheel ( #2456 )
2025-08-15 16:43:55 -04:00
Dmitry Tokarev
4c90b1b924
chore: Version bump to 0.4.0 ( #2179 )
2025-07-30 02:07:31 +00:00
Graham King
182d3b5dc7
chore(bindings): Remove mistralrs / llama.cpp ( #1970 )
2025-07-16 16:12:40 -04:00
Graham King
2bf27924a1
feat(python): Python bindings for the Dynamo CLI tools ( #1799 )
2025-07-08 21:40:36 +00:00
Anant Sharma
c4935b3497
chore: update versions for 0.3.2 release ( #1793 )
2025-07-07 17:21:02 -04:00
Paul Hendricks
dfbd741de2
feat: Support for Responses API ( #1694 )
2025-07-01 12:08:15 -04:00
Paul Hendricks
82eae1fdf5
refactor: Upgrade async-openai ( #1693 )
2025-06-30 13:35:10 -04:00
Yan Ru Pei
13a99b7f76
feat: Standalone Router ( #1409 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Signed-off-by: Yan Ru Pei <yanrpei@gmail.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: jain-ria <riajain@NVIDIA.com>
2025-06-14 19:02:31 +00:00
Anant Sharma
99e67e607a
chore: update dynamo and nixl versions for 0.3.1 ( #1517 )
2025-06-13 21:13:43 +00:00
Anant Sharma
9d9a1d9b74
chore: update dynamo and nixl versions for 0.3.0 ( #1240 )
2025-05-29 17:30:17 +00:00
Graham King
0a1d1fbe3d
feat(dynamo-llm): Remove bring-your-own-engine ( #1216 )
...
It was removed from the docs in 0.2.1 and replaced with writing a [standalone Python engine](https://github.com/ai-dynamo/dynamo/blob/main/docs/guides/dynamo_run.md#writing-your-own-engine-in-python ).
Also remove the associated `dynamo-run` feature `python`.
Releasing this in 0.3.0 will resolve #784 and #1109 .
2025-05-28 18:52:27 +00:00
Jacky
437cae0ad0
feat: KV Block Manager Python bindings ( #1022 )
2025-05-19 12:33:41 -04:00
Ryan McCormick
34f3fc6d12
test: Add doc tests to Rust CI ( #1102 )
2025-05-16 07:56:52 -07:00
Harrison Saturley-Hall
e9cb035ac7
chore: bump versions and NIXL dependencies for 0.2.1 ( #1012 )
2025-05-09 15:35:01 +00:00
Harrison Saturley-Hall
0715d4691f
chore: bump NIXL version and package versions ( #836 )
...
Signed-off-by: Harrison Saturley-Hall <454891+saturley-hall@users.noreply.github.com>
2025-04-25 19:24:38 -04:00
Anant Sharma
fa7ee14c2a
chore: update versions to 0.1.1 ( #552 )
2025-04-09 09:50:30 -04:00
Ryan Olson
84985d3f1d
refactor: migrate engines to standalone crates ( #453 )
...
Moved all of `lib/llm/src/engines` to their own crates as e.g. `lib/engines/mistralrs`. This will allow publishing of the `dynamo-llm` crate as it won't have any github dependencies.
The only engines in dynamo-llm will be the demo `echo` ones.
Co-authored-by: Graham King <grahamk@nvidia.com>
2025-04-03 21:53:54 +00:00
Anant Sharma
3d2928510c
build: add top level rust workspace ( #137 )
2025-03-13 17:22:19 -04:00
Anant Sharma
fc4da34502
chore: update wheel name and reset versions ( #73 )
2025-03-10 17:41:13 -04:00
Neelay Shah
678cffb4e5
chore: left over renaming ( #67 )
...
Co-authored-by: Harrison Saturley-Hall <454891+saturley-hall@users.noreply.github.com>
Co-authored-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com>
2025-03-09 14:09:04 -04:00
Neelay Shah
602352ce19
chore: rename dynamo ( #44 )
...
Co-authored-by: Biswa Panda <biswa.panda@gmail.com>
2025-03-08 01:42:40 -08:00
Graham King
12714d9080
feat: Python bring-your-own-engine with our tokenizer ( #47 )
...
Instead of using `out=pystr:<my.py>` we can now do this:
```
dynemo-run out=pytok:/home/graham/my_python_engine.py --model-path <hf-repo-checkout>
```
That engine will receive and respond with tokens. Here's an example engine file:
```
import asyncio
async def generate(request):
yield {"token_ids":[791]}
await asyncio.sleep(0.1)
yield {"token_ids":[6864]}
await asyncio.sleep(0.1)
yield {"token_ids":[315]}
await asyncio.sleep(0.1)
yield {"token_ids":[9822]}
await asyncio.sleep(0.1)
yield {"token_ids":[374]}
await asyncio.sleep(0.1)
yield {"token_ids":[12366]}
await asyncio.sleep(0.1)
yield {"token_ids":[13]}
```
Also reduce duplication by making the bindings engine use the llm lib engine.
2025-03-07 10:48:37 -05:00
Neelay Shah
1af7433bff
refactor: rename triton_distributed to dynemo ( #22 )
...
Co-authored-by: Graham King <grahamk@nvidia.com>
2025-03-05 09:15:46 -08:00
Anant Sharma
ea401e3bc0
ci: build wheel from root directory ( #274 )
2025-02-27 11:56:51 -05:00
Neelay Shah
08fcd7e93b
refactor: move libs to lib dir
...
Signed-off-by: Neelay Shah <neelays@nvidia.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>
2025-02-24 18:21:02 -08:00