Tzu-Ling Kan
0ce3461a9e
feat: Add runtime.endpoint() method to eliminate namespace chaining ( #6386 )
...
Signed-off-by: tzulingk@nvidia.com <tzulingk@nvidia.com>
2026-02-19 16:02:37 -05:00
Tzu-Ling Kan
5624d14481
Rename fetch_llm to fetch_model ( #6268 )
...
Signed-off-by: tzulingk@nvidia.com <tzulingk@nvidia.com>
2026-02-13 23:22:35 +00:00
Tushar Sharma
cf433e6825
chore: update all copyright headers in repo to 2026 ( #5130 )
...
Signed-off-by: Tushar Sharma <tusharma@nvidia.com>
2026-01-02 22:08:23 +00:00
Neal Vaidya
118323f26e
fix: skip HuggingFace download for non-llms ( #4686 )
...
Signed-off-by: Neal Vaidya <nealv@nvidia.com>
2025-12-04 10:23:24 -08:00
Yan Ru Pei
d821a8b9f7
chore: parallelize planner profile tests + bindings test cleanup ( #4532 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com>
2025-11-21 09:21:43 +00:00
Graham King
69797b5ab3
feat: Only monitor NATS metrics if using NATS request plane ( #4442 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-11-19 17:48:49 +00:00
Graham King
e1af3af6ee
chore: Remove static mode ( #4235 )
...
Signed-off-by: Graham King <grahamk@nvidia.com>
2025-11-11 19:25:12 +00:00
GuanLuo
6ba64c31f5
feat: tensor type for generic inference. ( #2746 )
...
Signed-off-by: Guan Luo <gluo@nvidia.com>
Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
Co-authored-by: Olga Andreeva <124622579+oandreeva-nv@users.noreply.github.com>
Co-authored-by: Ryan McCormick <rmccormick@nvidia.com>
2025-09-24 07:31:09 +00:00