dagil-nvidia
46dad85b03
docs: update KVBM diagram and bump container image tags to 1.0.0 ( #7365 )
...
Signed-off-by: Dan Gil <dagil@nvidia.com>
2026-03-13 18:08:10 -05:00
Ayush Agarwal
1182e2071d
feat: vllm omni image to video support ( #6530 )
...
Signed-off-by: ayushag <ayushag@nvidia.com>
2026-03-12 17:30:46 +00:00
dagil-nvidia
96f3bdccef
docs: 1.0 documentation improvements ( #7168 )
...
Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: athreesh <anish.maddipoti@utexas.edu>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 17:08:31 -05:00
ishandhanani
51dfd76045
feat: add SGLang chat processor for frontend pre/post processing ( #6834 )
...
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-09 23:27:34 +00:00
Yifan Jiang
100819299f
feat(trtllm): add additional metrics for dynamo-trtllm ( #6668 )
...
Signed-off-by: Yifan Jiang <19356972+yifjiang@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 14:34:38 -08:00
Yan Ru Pei
9f3b7b330d
chore(kv-router): deduplicate KvRouterConfig args into shared ArgGroup ( #6805 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 19:53:03 +00:00
Alec
9fe03dd864
docs: vLLM README container instructions and KV offloading page ( #6793 )
...
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 05:07:41 +00:00
Yue Yu
4d7c98458b
feat: FlexKV integration in Dynamo ( #5858 )
...
Signed-off-by: zhuofan1123 <zhuofanl@nvidia.com>
Co-authored-by: zhuofan1123 <zhuofanl@nvidia.com>
2026-03-02 20:18:39 -08:00
Tanmay Verma
1bb28d6ff3
docs: Refactor TensorRT-LLM backend docs ( #6782 )
2026-03-03 01:28:45 +00:00
Alec
75bf1e0993
docs: restructure vLLM docs and add startup banners to launch scripts ( #6698 )
...
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: ishandhanani <ishandhanani@gmail.com>
2026-03-03 00:37:55 +00:00
Neal Vaidya
ece08dc926
docs: restructure docs directory and move fern config to fern/ ( #6700 )
...
Signed-off-by: Neal Vaidya <nealv@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 19:58:22 -08:00
Jonathan Tong
39d645e586
docs: migrate Fern docs from fern/ into docs/ ( #6206 )
...
Signed-off-by: Jont828 <jt572@cornell.edu>
2026-02-11 16:22:27 -08:00
Dillon Cullinan
3188c70a3b
chore: Templating Feedback Followup ( #6125 )
...
Signed-off-by: Dillon Cullinan <dcullinan@nvidia.com>
2026-02-11 13:09:27 -05:00
Keiven C
e18840cef3
feat: add Prometheus auto and custom label injection for engine metrics ( #5989 )
...
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
2026-02-10 21:22:53 -08:00
Yuewei Na
df2daaddab
feat: add video diffusion support to TRTLLM backend (wan_t2v only) ( #5926 )
...
Signed-off-by: Yuewei Na <nv-yna@users.noreply.github.com>
Signed-off-by: Yuewei Na <248773860+nv-yna@users.noreply.github.com>
Co-authored-by: Yuewei Na <nv-yna@users.noreply.github.com>
Co-authored-by: Tanmay Verma <tanmayv@nvidia.com>
2026-02-10 12:06:25 -08:00
Dillon Cullinan
ac0206293c
feat: Dockerfile templating ( #5633 )
...
Signed-off-by: Dillon Cullinan <dcullinan@nvidia.com>
2026-02-10 08:56:17 -05:00
Ayush Agarwal
9f76d0606c
feat: text to image vLLM Omni ( #5912 )
...
Signed-off-by: ayushag <ayushag@nvidia.com>
2026-02-09 18:19:27 -05:00
Jacky
1ffa489ea1
refactor: Move --migration-limit flag from backend to frontend ( #5918 )
...
Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com>
2026-02-06 20:50:55 +00:00
dagil-nvidia
b19de4ed77
docs: cleanup of docs refactor for components, integrations, and features ( #6019 )
...
Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-05 19:50:17 -08:00
akshatha-k
80e7bafd37
docs: Migrate router documentation to three-tier structure ( #5979 )
...
Signed-off-by: akshatha-k <akshutk@gmail.com>
Signed-off-by: dagil-nvidia <dagil@nvidia.com>
Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: dagil-nvidia <dagil@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-06 01:55:50 +00:00
Karen Chung
b5c0db63c3
feat: TRTLLM DP Rank Routing ( #5936 )
2026-02-05 17:43:04 -08:00
dagil-nvidia
3023c6258a
docs: migrate Profiler docs to three-tier structure ( #6003 )
...
Signed-off-by: Dan Gil <dagil@nvidia.com>
Signed-off-by: dagil-nvidia <dagil@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Jonathan Tong <jt572@cornell.edu>
2026-02-05 18:05:49 -06:00
akshatha-k
4c3eba2aeb
docs: restructure KVBM documentation into three-tier format ( #5905 )
...
Signed-off-by: akshatha-k <akshutk@gmail.com>
Signed-off-by: Dan Gil <dagil@nvidia.com>
Signed-off-by: dagil-nvidia <dagil@nvidia.com>
Co-authored-by: dagil-nvidia <dagil@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Kyle McGill <101670481+nv-kmcgill53@users.noreply.github.com>
2026-02-05 18:02:05 -06:00
dagil-nvidia
8aa7335e35
docs: migrate Speculative Decoding docs to three-tier structure ( #6001 )
...
Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-05 17:45:10 -06:00
dagil-nvidia
6720dfb658
docs: fix SGLang docs links (docs.sglang.ai → docs.sglang.io) ( #5894 )
...
Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-03 04:29:57 +00:00
Yuewei Na
e2982e1f5a
fix(docs): remove trailing transition causing Sphinx build failure ( #5831 )
...
Signed-off-by: Yuewei Na <nv-yna@users.noreply.github.com>
Co-authored-by: Yuewei Na <nv-yna@users.noreply.github.com>
2026-01-30 12:51:26 -08:00
Yuewei Na
284f772b25
docs: add known issues section for TRT-LLM backend ( #5801 )
...
Signed-off-by: Yuewei Na <nv-yna@users.noreply.github.com>
Signed-off-by: Yuewei Na <248773860+nv-yna@users.noreply.github.com>
Co-authored-by: Yuewei Na <nv-yna@users.noreply.github.com>
2026-01-30 11:21:28 -08:00
Richard Huo
8cdfe68242
docs: fix the statements for setting up the trtllm NIXL backends ( #5763 )
2026-01-29 08:15:45 -08:00
ishandhanani
ed5e4fd021
docs: remove incorrect HighConfidence dLLM algorithm reference ( #5723 )
...
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-28 08:16:08 -08:00
Abhishek Gupta
912a4d4b7e
docs: add NIXL backend configuration and fix multiple typos ( #5564 )
...
Signed-off-by: Abhishek Kumar Gupta (AbhiOnGithub)
Signed-off-by: Abhishek Kumar Gupta <mail2abhishekgupta@gmail.com>
Co-authored-by: dagil-nvidia <dagil@nvidia.com>
2026-01-23 16:43:46 -08:00
Ayush Agarwal
91ddf4189b
feat: v0 diffusion handler support ( #5533 )
...
Signed-off-by: ayushag <ayushag@nvidia.com>
2026-01-21 22:25:54 +00:00
Anish
0d597e7c55
docs: Documentation audit and updates for 0.8 ( #5380 )
...
Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
Co-authored-by: Dan Gil <dagil@nvidia.com>
2026-01-15 00:23:09 +00:00
Neelay Shah
52ce68e60b
docs: Fix KV cache transfer UCX configuration instructions ( #5247 )
...
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-12 13:06:04 -08:00
dagil-nvidia
4f6996c733
docs: fix broken documentation links ( #5330 )
2026-01-09 17:34:41 -06:00
Wenqi Glantz
6f9619a210
feat(vllm): Add prompt embeds support for pre-computed inference inputs ( #4739 )
...
Signed-off-by: Wenqi Glantz <wglantz@nvidia.com>
2026-01-08 20:03:10 +00:00
ishandhanani
5a158552db
fix: remove old docs and unify model paths ( #5179 )
2026-01-05 19:27:44 +00:00
Tushar Sharma
cf433e6825
chore: update all copyright headers in repo to 2026 ( #5130 )
...
Signed-off-by: Tushar Sharma <tusharma@nvidia.com>
2026-01-02 22:08:23 +00:00
Yan Ru Pei
0055f2a471
docs: remove stale router diagram in trtllm docs ( #5034 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-12-19 18:20:55 +00:00
Keiven C
4d0b1a119e
docs: document LMCache Prometheus metrics harmless error for vLLM 0.12.0 ( #5026 )
...
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>
2025-12-19 08:57:40 -08:00
Tanmay Verma
e1c685b105
fix: Remove debug logging from aiperf benchmarking ( #5010 )
2025-12-18 02:34:41 +00:00
Neal Vaidya
3c4b3069d2
docs: consolidate multimodal docs ( #4842 )
...
Signed-off-by: Neal Vaidya <nealv@nvidia.com>
2025-12-16 02:08:58 +00:00
Yan Ru Pei
8711237357
docs: router cost function mermaid ( #4943 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-12-13 07:13:58 +00:00
Dmitry Tokarev
525030324e
chore: TRTLLM 1.2.0rc4 ( #4836 )
...
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com>
2025-12-11 03:02:24 +00:00
J Wyman
96fe63fe2e
feat: nixl_connect: Improve Concurrency Support ( #4433 )
...
Signed-off-by: J Wyman <jwyman@nvidia.com>
2025-12-09 17:29:22 -05:00
Indrajit Bhosale
94d145a9b4
docs: Add multimodal documentation vllm, sglang, and trtllm backends ( #4510 )
...
Signed-off-by: Indrajit Bhosale <iamindrajitb@gmail.com>
Co-authored-by: krishung5 <krish@nvidia.com>
2025-12-09 23:38:31 +08:00
Anant Sharma
8b197406b4
fix: move new launch script to examples ( #4800 )
...
Signed-off-by: Anant Sharma <anants@nvidia.com>
2025-12-08 20:27:54 +00:00
Dilreet Raju
f315374fdc
docs: Guide for Speculative Decoding in VLLM using Eagle3 and Meta-Llama-3.1-8B-Instruct ( #3895 )
...
Signed-off-by: DilreetRaju <dilreetraju@gmail.com>
2025-12-05 21:09:38 +00:00
William Arnold
3bc129dc37
feat: Add python-configurable engine routes for sglang profiling ( #4617 )
...
Signed-off-by: William Arnold <7565007+Aphoh@users.noreply.github.com>
2025-12-05 10:28:36 -08:00
dagil-nvidia
d5a75dd8c5
docs: update TensorRT-LLM commit for multimodal EPD to v1.2.0rc3 ( #4713 )
...
Signed-off-by: Dan Gil <dagil@nvidia.com>
2025-12-03 18:18:28 +00:00
Keiven C
c6555852ee
fix: enable LMCache metrics visibility with PROMETHEUS_MULTIPROC_DIR ( #4654 )
...
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com>
2025-12-02 18:47:30 -08:00