dagil-nvidia
b19de4ed77
docs: cleanup of docs refactor for components, integrations, and features ( #6019 )
...
Signed-off-by: Dan Gil <dagil@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-05 19:50:17 -08:00
dagil-nvidia
3023c6258a
docs: migrate Profiler docs to three-tier structure ( #6003 )
...
Signed-off-by: Dan Gil <dagil@nvidia.com>
Signed-off-by: dagil-nvidia <dagil@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Jonathan Tong <jt572@cornell.edu>
2026-02-05 18:05:49 -06:00
hhzhang16
ec5630ead9
feat: mount model path to Profiler if specified ( #5212 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2026-01-09 17:13:53 +00:00
Janelle Cai
d0cfc40a7d
docs: broken links in benchmarking documentation ( #5258 )
2026-01-08 16:08:07 -08:00
Tushar Sharma
cf433e6825
chore: update all copyright headers in repo to 2026 ( #5130 )
...
Signed-off-by: Tushar Sharma <tusharma@nvidia.com>
2026-01-02 22:08:23 +00:00
Thomas Montfort
3057af00b6
feat(dep-715): refactor helm directory and remove reference to cloud ( #5042 )
2025-12-23 12:08:19 -05:00
hhzhang16
6b5842eebe
feat: Profiler WebUI improvements -- error handling, GPU hours, style fixes, preview configs ( #4968 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2025-12-16 23:05:38 +00:00
Hongkuan Zhou
5652d670c7
feat: support MQA + MoE (Qwen3 MoE) TEP/DEP in Planner Profiler ( #4612 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2025-11-26 01:41:46 +00:00
Harrison Saturley-Hall
a4eb4e8a8d
fix: misnamed tensorrtllm-runtime image and incorrect tag ( #4289 )
...
Signed-off-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com>
Signed-off-by: Harrison Saturley-Hall <harrison.saturley.hall@gmail.com>
Co-authored-by: hhzhang16 <54051230+hhzhang16@users.noreply.github.com>
2025-11-14 15:16:34 -05:00
Jason Zhou
43986372f5
feat: DynamoPlanner profiler to use hf_id for AIConfigurator 0.4.0 ( #4167 )
...
Signed-off-by: Jason Zhou <jasonzho@jasonzho-mlt.client.nvidia.com>
Signed-off-by: Jason Zhou <jasonzho@nvidia.com>
Co-authored-by: Jason Zhou <jasonzho@jasonzho-mlt.client.nvidia.com>
2025-11-10 23:17:42 +00:00
Hongkuan Zhou
5e4a339a1f
feat: support TEP/DEP for DSR1 in SGLang MoE Planner ( #4203 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
Signed-off-by: Hongkuan Zhou <tedzhouhk@gmail.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2025-11-10 09:56:26 -08:00
hhzhang16
0e623146a8
feat: remove scripts for manipulating pvcs ( #4206 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2025-11-10 11:07:58 +00:00
Anant Sharma
8bd37c96d6
refactor: move backend deploy, launch and slurm files from components to examples ( #3849 )
...
Signed-off-by: Anant Sharma <anants@nvidia.com>
2025-10-31 13:09:46 -04:00
hhzhang16
6fc4c59599
feat: remove cluster wide logic from namespace restricted operator ( #3934 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2025-10-30 23:08:54 +00:00
hhzhang16
a7b703bdc6
fix: profiler sidecar and other SLA-driven autodeployment fixes ( #3932 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2025-10-29 00:02:43 +00:00
Anant Sharma
e3346dabb0
fix: use image relative path for sphinx build ( #3933 )
...
Signed-off-by: Anant Sharma <anants@nvidia.com>
2025-10-28 16:16:36 -04:00
hhzhang16
6a84ffd347
feat: turn profiling k8s jobs into sample DGDR requests ( #3864 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
Signed-off-by: Hongkuan Zhou <tedzhouhk@gmail.com>
Co-authored-by: hongkuanz <hongkuanz@nvidia.com>
Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com>
2025-10-27 13:36:31 -07:00
Hongkuan Zhou
a1b38af25e
feat: automatic profiling config generation ( #3787 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
Signed-off-by: Hongkuan Zhou <tedzhouhk@gmail.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2025-10-23 08:07:39 -07:00
Anish
c6b5904579
docs: address Harry/VDR feedback + fixing broken links across repository ( #3802 )
...
Signed-off-by: Harry Kim <harry_kim@live.com>
Signed-off-by: athreesh <anish.maddipoti@utexas.edu>
Signed-off-by: akshatha-k <33278067+akshatha-k@users.noreply.github.com>
Signed-off-by: Harrison Saturley-Hall <hsaturleyhal@nvidia.com>
Signed-off-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com>
Co-authored-by: Harry Kim <harry_kim@live.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: akshatha-k <33278067+akshatha-k@users.noreply.github.com>
Co-authored-by: Harrison Saturley-Hall <hsaturleyhal@nvidia.com>
2025-10-22 19:56:24 -04:00
hhzhang16
eb73c2b09f
feat: remove deploy/utils rbac ( #3771 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2025-10-21 16:27:29 -07:00
Ben Hamm
a665886a42
docs: Add KV Smart Router A/B Benchmarking Guide ( #3696 )
...
Signed-off-by: Ben Hamm <bhamm@nvidia.com>
Co-authored-by: Ben Hamm <bhamm@bhamm-mlt.client.nvidia.com>
Co-authored-by: Ben Hamm <bhamm@nvidia.com>
2025-10-20 20:47:27 +02:00
William Arnold
ee56782b98
fix: standardize all planner ttft/itl units to float ms and fix docs ( #3673 )
...
Signed-off-by: William Arnold <7565007+Aphoh@users.noreply.github.com>
2025-10-17 18:33:41 +00:00
Anish
598cbbb73b
docs: reorganizing documentation to make things clearer ( #3658 )
...
Signed-off-by: athreesh <anish.maddipoti@utexas.edu>
Co-authored-by: Claude <noreply@anthropic.com>
2025-10-16 23:59:59 +00:00
Harshini Komali
9f310225eb
feat: Replace genai-perf with aiperf ( #3533 )
...
Signed-off-by: lkomali <lkomali@nvidia.com>
2025-10-16 09:58:36 -07:00
Hongkuan Zhou
400dceae6a
feat: support yaml config input for pre-deployment sweep script ( #3622 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2025-10-14 17:10:55 -07:00
Biswa Panda
de3ca70b43
feat: update benchmarking script to use aiperf ( #3306 )
...
Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
Signed-off-by: lkomali <lkomali@nvidia.com>
Co-authored-by: lkomali <lkomali@nvidia.com>
Co-authored-by: Harshini Komali <157742537+lkomali@users.noreply.github.com>
2025-10-14 05:26:29 +00:00
Alec
179f993a54
fix: profilier bug fixes and doc improvements ( #3530 )
...
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
Co-authored-by: hongkuanz <hongkuanz@nvidia.com>
2025-10-09 14:47:26 -07:00
Hongkuan Zhou
83e259a7cc
feat: pre-deployment profiling automatically generates and deploys optimized DGD with planner ( #3441 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2025-10-07 16:52:51 -07:00
hhzhang16
89cf9107db
docs: add Planner Quickstart doc ( #3358 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
Signed-off-by: hhzhang16 <54051230+hhzhang16@users.noreply.github.com>
Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com>
2025-10-03 18:49:28 +00:00
Hongkuan Zhou
cfe74445fc
fix: correctly parse subcomponent in sla-planner + deploy/doc update ( #3363 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
2025-10-02 16:06:47 -07:00
Ilya Sherstyuk
9d73be1255
docs: Fix profile_sla command in docs ( #3258 )
...
Signed-off-by: Ilya Sherstyuk <isherstyuk@nvidia.com>
2025-09-26 23:01:01 +00:00
hhzhang16
fb12b67ff7
feat: update how inputs are input into the benchmark script ( #3187 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2025-09-24 22:27:11 +00:00
Harrison Saturley-Hall
9e8f67ed0f
fix: update the tags for consistency and remove 0.4.1 refs ( #3058 )
...
Signed-off-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com>
Signed-off-by: Harrison Saturley-Hall <hsaturleyhal@nvidia.com>
2025-09-24 14:53:43 -04:00
hhzhang16
4800d0be7f
feat: allow users to only plot certain subdirs ( #3190 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2025-09-24 11:39:07 -07:00
Hongkuan Zhou
6243bcbe04
feat: support MoE model in SLA Planner Sglang ( #3185 )
...
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
Signed-off-by: Hongkuan Zhou <tedzhouhk@gmail.com>
Co-authored-by: hhzhang16 <54051230+hhzhang16@users.noreply.github.com>
2025-09-23 19:42:43 -07:00
Julien Mancuso
c1907d129d
fix: fix broken links ( #3186 )
2025-09-23 15:17:18 -06:00
Julien Mancuso
4a71802842
feat: revamp kubernetes doc ( #3173 )
...
Signed-off-by: Julien Mancuso <161955438+julienmancuso@users.noreply.github.com>
Co-authored-by: hhzhang16 <54051230+hhzhang16@users.noreply.github.com>
2025-09-23 14:09:29 -06:00
hhzhang16
ce36d9f4ce
feat: allow in-cluster perf benchmarks with a kubectl one-liner ( #3144 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
Signed-off-by: zhongdaor <zhongdaor@nvidia.com>
Signed-off-by: tmontfort <tmontfort@nvidia.com>
Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com>
Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com>
Signed-off-by: richardhuo-nv <rihuo@nvidia.com>
Signed-off-by: Tushar Sharma <tusharma@nvidia.com>
Signed-off-by: tzulingk@nvidia.com <tzulingk@nvidia.com>
Signed-off-by: Krishnan Prashanth <kprashanth@nvidia.com>
Signed-off-by: Olga Andreeva <oandreeva@nvidia.com>
Signed-off-by: oandreeva-nv <oandreeva-nv@nvidia.com>
Co-authored-by: zhongdaor-nv <zhongdaor@nvidia.com>
Co-authored-by: Thomas Montfort <61255722+tmonty12@users.noreply.github.com>
Co-authored-by: Keiven C <213854356+keivenchang@users.noreply.github.com>
Co-authored-by: Jacky <18255193+kthui@users.noreply.github.com>
Co-authored-by: Richard Huo <rihuo@nvidia.com>
Co-authored-by: Tushar Sharma <tusharma@nvidia.com>
Co-authored-by: Tzu-Ling Kan <tzulingk@nvidia.com>
Co-authored-by: KrishnanPrash <140860868+KrishnanPrash@users.noreply.github.com>
Co-authored-by: Olga Andreeva <124622579+oandreeva-nv@users.noreply.github.com>
Co-authored-by: Ziqi Fan <ziqif@nvidia.com>
Co-authored-by: oandreeva-nv <oandreeva-nv@nvidia.com>
2025-09-23 10:03:09 -07:00
nv-nmailhot
406c4d4e87
ci: add broken links check ( #2927 )
...
Signed-off-by: Nate Mailhot <nmailhot@nvidia.com>
Signed-off-by: nv-nmailhot <nmailhot@nvidia.com>
2025-09-19 21:42:05 +00:00
Thomas Montfort
7d2fc13edb
fix: small planner manifest/doc fixes ( #3129 )
...
Signed-off-by: tmontfort <tmontfort@nvidia.com>
2025-09-19 12:46:02 -07:00
Ilya Sherstyuk
19948b7f5a
feat: Add --use-ai-configurator to profile_sla.py ( #3079 )
...
Signed-off-by: Ilya Sherstyuk <isherstyuk@nvidia.com>
Signed-off-by: Ilya Sherstyuk <46343317+ilyasher@users.noreply.github.com>
Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com>
2025-09-19 11:23:24 -07:00
hhzhang16
20b7a8ae2f
feat: remove kubectl dependencies from benchmarking ( #3098 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
Signed-off-by: hhzhang16 <54051230+hhzhang16@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2025-09-18 15:52:57 -07:00
hhzhang16
26889b099a
feat: decouple existing benchmarking within docs ( #3072 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
Signed-off-by: hhzhang16 <54051230+hhzhang16@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2025-09-17 11:16:13 -07:00
hhzhang16
f77511ff2c
chore: a few small benchmarking and profiling improvements ( #3069 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2025-09-16 14:43:47 -07:00
hhzhang16
c82f477dcc
feat: decouple dynamo k8s setup from additional benchmarking requirements ( #2973 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2025-09-15 15:04:23 -07:00
Hongkuan Zhou
dcee4dbd17
feat: support trtllm in sla-planner ( #2980 )
...
Signed-off-by: Biswa Panda <biswa.panda@gmail.com>
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
Signed-off-by: ayushag <ayushag@nvidia.com>
Signed-off-by: Dillon Cullinan <dcullinan@nvidia.com>
Signed-off-by: Harrison Saturley-Hall <hsaturleyhal@nvidia.com>
Signed-off-by: alec-flowers <aflowers@nvidia.com>
Signed-off-by: Julien Mancuso <jmancuso@nvidia.com>
Signed-off-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com>
Signed-off-by: Harry Kim <harry_kim@live.com>
Signed-off-by: Tushar Sharma <tusharma@nvidia.com>
Signed-off-by: GuanLuo <gluo@nvidia.com>
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
Signed-off-by: Graham King <grahamk@nvidia.com>
Signed-off-by: Anant Sharma <anants@nvidia.com>
Signed-off-by: Greg Clark <grclark@nvidia.com>
Signed-off-by: Neal Vaidya <nealv@nvidia.com>
Co-authored-by: Biswa Panda <biswa.panda@gmail.com>
Co-authored-by: Ayush Agarwal <ayushag@nvidia.com>
Co-authored-by: Dillon Cullinan <dcullinan92@gmail.com>
Co-authored-by: Harrison Saturley-Hall <hsaturleyhal@nvidia.com>
Co-authored-by: alec-flowers <aflowers@nvidia.com>
Co-authored-by: julienmancuso <161955438+julienmancuso@users.noreply.github.com>
Co-authored-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com>
Co-authored-by: Harry Kim <harry_kim@live.com>
Co-authored-by: Tushar Sharma <tusharma@nvidia.com>
Co-authored-by: Alec <35311602+alec-flowers@users.noreply.github.com>
Co-authored-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
Co-authored-by: hhzhang16 <54051230+hhzhang16@users.noreply.github.com>
Co-authored-by: Graham King <grahamk@nvidia.com>
Co-authored-by: Anant Sharma <anants@nvidia.com>
Co-authored-by: Greg Clark <grclark@nvidia.com>
Co-authored-by: Neal Vaidya <nealv@nvidia.com>
Co-authored-by: nv-nmailhot <nmailhot@nvidia.com>
2025-09-10 14:41:29 -07:00
hhzhang16
241bd01495
docs: update how deploy util helper files are mentioned ( #2957 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
Signed-off-by: hhzhang16 <54051230+hhzhang16@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2025-09-09 13:50:29 -07:00
hhzhang16
09c7b73caf
feat: update benchmarking and deploy utils ( #2933 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
2025-09-08 22:53:55 -04:00
hhzhang16
7eef9ac332
docs: update profiling-related docs ( #2816 )
...
Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
Signed-off-by: hhzhang16 <54051230+hhzhang16@users.noreply.github.com>
Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com>
2025-09-02 22:00:12 +00:00
Jason Zhou
3f09c39559
fix: fix #2653 : links of h100_prefill_performance.png and h100_decode_performance.png ( #2650 )
2025-08-30 22:26:49 +00:00