Signed-off-by: Hannah Zhang <hannahz@nvidia.com> Signed-off-by: hongkuanz <hongkuanz@nvidia.com> Signed-off-by: Hongkuan Zhou <tedzhouhk@gmail.com> Co-authored-by: hongkuanz <hongkuanz@nvidia.com> Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com> |
||
|---|---|---|
| .. | ||
| manifests | ||
| README.md | ||
| __init__.py | ||
| download_pvc_results.py | ||
| dynamo_deployment.py | ||
| gpu_inventory.py | ||
| inject_manifest.py | ||
| kubernetes.py | ||
| requirements.txt | ||
| setup_benchmarking_resources.sh | ||
README.md
Kubernetes utilities for Dynamo Benchmarking and Profiling
This directory contains utilities and manifests for Dynamo benchmarking and profiling workflows.
Prerequisites
Before using these utilities, you must first set up Dynamo Cloud following the main installation guide:
👉 Follow the Dynamo Cloud installation guide to install the Dynamo Kubernetes Platform first.
This includes:
- Installing the Dynamo CRDs
- Installing the Dynamo Platform (operator, etcd, NATS)
- Setting up your target namespace
Contents
setup_benchmarking_resources.sh— Sets up benchmarking and profiling resources in your existing Dynamo namespacemanifests/pvc.yaml— PVCdynamo-pvcfor storing profiler results and configurationspvc-access-pod.yaml— short‑lived pod for copying profiler results from the PVC
kubernetes.py— helper used by tooling to apply/read resources (e.g., access pod for PVC downloads)inject_manifest.py— utility for injecting deployment configurations into the PVC for profilingdownload_pvc_results.py— utility for downloading benchmark/profiling results from the PVCdynamo_deployment.py— utilities for working with DynamoGraphDeployment resourcesrequirements.txt— Python dependencies for benchmarking utilities
Quick start
Benchmarking Resource Setup
After setting up Dynamo Cloud, use this script to prepare your namespace with the additional resources needed for benchmarking and profiling workflows:
The setup script creates a dynamo-pvc with ReadWriteMany (RWX). If your cluster's default storageClassName does not support RWX, set storageClassName in deploy/utils/manifests/pvc.yaml to an RWX-capable class before running the script.
Example (add under spec in deploy/utils/manifests/pvc.yaml):
...
spec:
accessModes:
- ReadWriteMany
storageClassName: <your-rwx-storageclass>
...
[!TIP] Check your clusters storage classes
- List storage classes and provisioners:
kubectl get sc -o wide
export NAMESPACE=your-dynamo-namespace
export HF_TOKEN=<HF_TOKEN> # Optional: for HuggingFace model access
deploy/utils/setup_benchmarking_resources.sh
This script applies the following manifests to your existing Dynamo namespace:
deploy/utils/manifests/pvc.yaml- PVCdynamo-pvc
If HF_TOKEN is provided, it also creates a secret for HuggingFace model access.
After running the setup script, verify the resources by checking:
kubectl get pvc dynamo-pvc -n $NAMESPACE
PVC Manipulation Scripts
These scripts interact with the Persistent Volume Claim (PVC) that stores configuration files and benchmark/profiling results. They're essential for the Dynamo benchmarking and profiling workflows.
Why These Scripts Are Needed
- For Pre-Deployment Profiling: The profiling job needs access to your Dynamo deployment configurations (DGD manifests) to test different parallelization strategies
- For Retrieving Results: Both benchmarking and profiling jobs write their results to the PVC, which you need to download for analysis
Script Usage
Inject deployment configurations for profiling:
# The profiling job reads your DGD config from the PVC
# IMPORTANT: All paths must start with /data/ for security reasons
python3 -m deploy.utils.inject_manifest \
--namespace $NAMESPACE \
--src ./my-disagg.yaml \
--dest /data/configs/disagg.yaml
Download benchmark/profiling results:
# After benchmarking or profiling completes, download results
python3 -m deploy.utils.download_pvc_results \
--namespace $NAMESPACE \
--output-dir ./pvc_files \
--folder /data/results \
--no-config # optional: skip *.yaml/*.yml in the download
Path Requirements
Important: The PVC is mounted at /data in the access pod for security reasons. All destination paths must start with /data/.
Common path patterns:
/data/configs/- Configuration files (DGD manifests)/data/results/- Benchmark results/data/profiling_results/- Profiling data/data/benchmarking/- Benchmarking artifacts
User-friendly error messages: If you forget the /data/ prefix, the script will show a helpful error message with the correct path and example commands.
Next Steps
For complete benchmarking and profiling workflows:
- Benchmarking Guide: See docs/benchmarks/benchmarking.md for comparing DynamoGraphDeployments and external endpoints
- Pre-Deployment Profiling: See docs/benchmarks/sla_driven_profiling.md for optimizing configurations before deployment
Notes
- This setup is focused on benchmarking and profiling resources only - the main Dynamo platform must be installed separately.