2467 lines
112 KiB
Markdown
2467 lines
112 KiB
Markdown
---
|
||
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||
# SPDX-License-Identifier: Apache-2.0
|
||
---
|
||
|
||
> **⚠️ Important**: This documentation is automatically generated from source code.
|
||
> Do not edit this file directly.
|
||
|
||
# API Reference
|
||
|
||
## Packages
|
||
- [nvidia.com/v1alpha1](#nvidiacomv1alpha1)
|
||
- [nvidia.com/v1beta1](#nvidiacomv1beta1)
|
||
- [operator.config.dynamo.nvidia.com/v1alpha1](#operatorconfigdynamonvidiacomv1alpha1)
|
||
|
||
|
||
## nvidia.com/v1alpha1
|
||
|
||
Package v1alpha1 contains API Schema definitions for the nvidia.com v1alpha1 API group.
|
||
|
||
This package defines the DynamoGraphDeploymentRequest (DGDR) custom resource, which provides
|
||
a high-level, SLA-driven interface for deploying machine learning models on Dynamo.
|
||
|
||
Package v1alpha1 contains API Schema definitions for the nvidia.com v1alpha1 API group.
|
||
|
||
### Resource Types
|
||
- [DynamoCheckpoint](#dynamocheckpoint)
|
||
- [DynamoComponentDeployment](#dynamocomponentdeployment)
|
||
- [DynamoGraphDeployment](#dynamographdeployment)
|
||
- [DynamoGraphDeploymentRequest](#dynamographdeploymentrequest)
|
||
- [DynamoGraphDeploymentScalingAdapter](#dynamographdeploymentscalingadapter)
|
||
- [DynamoModel](#dynamomodel)
|
||
|
||
|
||
|
||
#### Autoscaling
|
||
|
||
|
||
|
||
Deprecated: This field is deprecated and ignored. Use DynamoGraphDeploymentScalingAdapter
|
||
with HPA, KEDA, or Planner for autoscaling instead. See docs/kubernetes/autoscaling.md
|
||
for migration guidance. This field will be removed in a future API version.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec)
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `enabled` _boolean_ | Deprecated: This field is ignored. | | |
|
||
| `minReplicas` _integer_ | Deprecated: This field is ignored. | | |
|
||
| `maxReplicas` _integer_ | Deprecated: This field is ignored. | | |
|
||
| `behavior` _[HorizontalPodAutoscalerBehavior](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#horizontalpodautoscalerbehavior-v2-autoscaling)_ | Deprecated: This field is ignored. | | |
|
||
| `metrics` _[MetricSpec](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#metricspec-v2-autoscaling) array_ | Deprecated: This field is ignored. | | |
|
||
|
||
|
||
|
||
|
||
#### CheckpointMode
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
CheckpointMode defines how checkpoint creation is handled
|
||
|
||
_Validation:_
|
||
- Enum: [Auto Manual]
|
||
|
||
_Appears in:_
|
||
- [ServiceCheckpointConfig](#servicecheckpointconfig)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `Auto` | CheckpointModeAuto means the DGD controller will automatically create a Checkpoint CR<br /> |
|
||
| `Manual` | CheckpointModeManual means the user must create the Checkpoint CR themselves<br /> |
|
||
|
||
|
||
#### ComponentKind
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
ComponentKind represents the type of underlying Kubernetes resource.
|
||
|
||
_Validation:_
|
||
- Enum: [PodClique PodCliqueScalingGroup Deployment LeaderWorkerSet]
|
||
|
||
_Appears in:_
|
||
- [ServiceReplicaStatus](#servicereplicastatus)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `PodClique` | ComponentKindPodClique represents a PodClique resource.<br /> |
|
||
| `PodCliqueScalingGroup` | ComponentKindPodCliqueScalingGroup represents a PodCliqueScalingGroup resource.<br /> |
|
||
| `Deployment` | ComponentKindDeployment represents a Deployment resource.<br /> |
|
||
| `LeaderWorkerSet` | ComponentKindLeaderWorkerSet represents a LeaderWorkerSet resource.<br /> |
|
||
|
||
|
||
#### ConfigMapKeySelector
|
||
|
||
|
||
|
||
ConfigMapKeySelector selects a specific key from a ConfigMap.
|
||
Used to reference external configuration data stored in ConfigMaps.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [ProfilingConfigSpec](#profilingconfigspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `name` _string_ | Name of the ConfigMap containing the desired data. | | Required: \{\} <br /> |
|
||
| `key` _string_ | Key in the ConfigMap to select. If not specified, defaults to "disagg.yaml". | disagg.yaml | |
|
||
|
||
|
||
#### DGDRState
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
|
||
|
||
_Validation:_
|
||
- Enum: [Initializing Pending Profiling Deploying Ready DeploymentDeleted Failed]
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestStatus](#dynamographdeploymentrequeststatus)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `Initializing` | |
|
||
| `Pending` | |
|
||
| `Profiling` | |
|
||
| `Deploying` | |
|
||
| `Ready` | |
|
||
| `DeploymentDeleted` | |
|
||
| `Failed` | |
|
||
|
||
|
||
#### DGDState
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
|
||
|
||
_Validation:_
|
||
- Enum: [initializing pending successful failed]
|
||
|
||
_Appears in:_
|
||
- [DeploymentStatus](#deploymentstatus)
|
||
- [DynamoGraphDeploymentStatus](#dynamographdeploymentstatus)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `initializing` | |
|
||
| `pending` | |
|
||
| `successful` | |
|
||
| `failed` | |
|
||
|
||
|
||
#### DeploymentOverridesSpec
|
||
|
||
|
||
|
||
DeploymentOverridesSpec allows users to customize metadata for auto-created DynamoGraphDeployments.
|
||
When autoApply is enabled, these overrides are applied to the generated DGD resource.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestSpec](#dynamographdeploymentrequestspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `name` _string_ | Name is the desired name for the created DynamoGraphDeployment.<br />If not specified, defaults to the DGDR name. | | Optional: \{\} <br /> |
|
||
| `namespace` _string_ | Namespace is the desired namespace for the created DynamoGraphDeployment.<br />If not specified, defaults to the DGDR namespace. | | Optional: \{\} <br /> |
|
||
| `labels` _object (keys:string, values:string)_ | Labels are additional labels to add to the DynamoGraphDeployment metadata.<br />These are merged with auto-generated labels from the profiling process. | | Optional: \{\} <br /> |
|
||
| `annotations` _object (keys:string, values:string)_ | Annotations are additional annotations to add to the DynamoGraphDeployment metadata. | | Optional: \{\} <br /> |
|
||
| `workersImage` _string_ | WorkersImage specifies the container image to use for DynamoGraphDeployment worker components.<br />This image is used for both temporary DGDs created during online profiling and the final DGD.<br />If omitted, the image from the base config file (e.g., disagg.yaml) is used.<br />Example: "nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.0.0" | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### DeploymentStatus
|
||
|
||
|
||
|
||
DeploymentStatus tracks the state of an auto-created DynamoGraphDeployment.
|
||
This status is populated when autoApply is enabled and a DGD is created.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestStatus](#dynamographdeploymentrequeststatus)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `name` _string_ | Name is the name of the created DynamoGraphDeployment. | | |
|
||
| `namespace` _string_ | Namespace is the namespace of the created DynamoGraphDeployment. | | |
|
||
| `state` _[DGDState](#dgdstate)_ | State is the current state of the DynamoGraphDeployment.<br />This value is mirrored from the DGD's status.state field. | initializing | Enum: [initializing pending successful failed] <br /> |
|
||
| `created` _boolean_ | Created indicates whether the DGD has been successfully created.<br />Used to prevent recreation if the DGD is manually deleted by users. | | |
|
||
|
||
|
||
|
||
|
||
#### DynamoCheckpoint
|
||
|
||
|
||
|
||
DynamoCheckpoint is the Schema for the dynamocheckpoints API
|
||
It represents a container checkpoint that can be used to restore pods to a warm state
|
||
|
||
|
||
|
||
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `apiVersion` _string_ | `nvidia.com/v1alpha1` | | |
|
||
| `kind` _string_ | `DynamoCheckpoint` | | |
|
||
| `metadata` _[ObjectMeta](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#objectmeta-v1-meta)_ | Refer to Kubernetes API documentation for fields of `metadata`. | | |
|
||
| `spec` _[DynamoCheckpointSpec](#dynamocheckpointspec)_ | | | |
|
||
| `status` _[DynamoCheckpointStatus](#dynamocheckpointstatus)_ | | | |
|
||
|
||
|
||
|
||
|
||
#### DynamoCheckpointIdentity
|
||
|
||
|
||
|
||
DynamoCheckpointIdentity defines the inputs that determine checkpoint equivalence
|
||
Two checkpoints with the same identity hash are considered equivalent
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoCheckpointSpec](#dynamocheckpointspec)
|
||
- [ServiceCheckpointConfig](#servicecheckpointconfig)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `model` _string_ | Model is the model identifier (e.g., "meta-llama/Llama-3-70B") | | Required: \{\} <br /> |
|
||
| `backendFramework` _string_ | BackendFramework is the runtime framework (vllm, sglang, trtllm) | | Enum: [vllm sglang trtllm] <br />Required: \{\} <br /> |
|
||
| `dynamoVersion` _string_ | DynamoVersion is the Dynamo platform version (optional)<br />If not specified, version is not included in identity hash<br />This ensures checkpoint compatibility across Dynamo releases | | Optional: \{\} <br /> |
|
||
| `tensorParallelSize` _integer_ | TensorParallelSize is the tensor parallel configuration | 1 | Minimum: 1 <br />Optional: \{\} <br /> |
|
||
| `pipelineParallelSize` _integer_ | PipelineParallelSize is the pipeline parallel configuration | 1 | Minimum: 1 <br />Optional: \{\} <br /> |
|
||
| `dtype` _string_ | Dtype is the data type (fp16, bf16, fp8, etc.) | | Optional: \{\} <br /> |
|
||
| `maxModelLen` _integer_ | MaxModelLen is the maximum sequence length | | Minimum: 1 <br />Optional: \{\} <br /> |
|
||
| `extraParameters` _object (keys:string, values:string)_ | ExtraParameters are additional parameters that affect the checkpoint hash<br />Use for any framework-specific or custom parameters not covered above | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### DynamoCheckpointJobConfig
|
||
|
||
|
||
|
||
DynamoCheckpointJobConfig defines the configuration for the checkpoint creation Job
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoCheckpointSpec](#dynamocheckpointspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `podTemplateSpec` _[PodTemplateSpec](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#podtemplatespec-v1-core)_ | PodTemplateSpec allows customizing the checkpoint Job pod<br />This should include the container that runs the workload to be checkpointed | | Required: \{\} <br /> |
|
||
| `activeDeadlineSeconds` _integer_ | ActiveDeadlineSeconds specifies the maximum time the Job can run | 3600 | Optional: \{\} <br /> |
|
||
| `backoffLimit` _integer_ | BackoffLimit specifies the number of retries before marking the Job failed | 3 | Optional: \{\} <br /> |
|
||
| `ttlSecondsAfterFinished` _integer_ | TTLSecondsAfterFinished specifies how long to keep the Job after completion | 300 | Optional: \{\} <br /> |
|
||
|
||
|
||
#### DynamoCheckpointPhase
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
DynamoCheckpointPhase represents the current phase of the checkpoint lifecycle
|
||
|
||
_Validation:_
|
||
- Enum: [Pending Creating Ready Failed]
|
||
|
||
_Appears in:_
|
||
- [DynamoCheckpointStatus](#dynamocheckpointstatus)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `Pending` | DynamoCheckpointPhasePending indicates the checkpoint CR has been created but the Job has not started<br /> |
|
||
| `Creating` | DynamoCheckpointPhaseCreating indicates the checkpoint Job is running<br /> |
|
||
| `Ready` | DynamoCheckpointPhaseReady indicates the checkpoint tar file is available on the PVC<br /> |
|
||
| `Failed` | DynamoCheckpointPhaseFailed indicates the checkpoint creation failed<br /> |
|
||
|
||
|
||
#### DynamoCheckpointSpec
|
||
|
||
|
||
|
||
DynamoCheckpointSpec defines the desired state of DynamoCheckpoint
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoCheckpoint](#dynamocheckpoint)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `identity` _[DynamoCheckpointIdentity](#dynamocheckpointidentity)_ | Identity defines the inputs that determine checkpoint equivalence | | Required: \{\} <br /> |
|
||
| `job` _[DynamoCheckpointJobConfig](#dynamocheckpointjobconfig)_ | Job defines the configuration for the checkpoint creation Job | | Required: \{\} <br /> |
|
||
|
||
|
||
#### DynamoCheckpointStatus
|
||
|
||
|
||
|
||
DynamoCheckpointStatus defines the observed state of DynamoCheckpoint
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoCheckpoint](#dynamocheckpoint)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `phase` _[DynamoCheckpointPhase](#dynamocheckpointphase)_ | Phase represents the current phase of the checkpoint lifecycle | | Enum: [Pending Creating Ready Failed] <br />Optional: \{\} <br /> |
|
||
| `identityHash` _string_ | IdentityHash is the computed hash of the checkpoint identity<br />This hash is used to identify equivalent checkpoints | | Optional: \{\} <br /> |
|
||
| `location` _string_ | Location is the full URI/path to the checkpoint in the storage backend<br />For PVC: same as TarPath (e.g., /checkpoints/\{hash\}.tar)<br />For S3: s3://bucket/prefix/\{hash\}.tar<br />For OCI: oci://registry/repo:\{hash\} | | Optional: \{\} <br /> |
|
||
| `storageType` _[DynamoCheckpointStorageType](#dynamocheckpointstoragetype)_ | StorageType indicates the storage backend type used for this checkpoint | | Enum: [pvc s3 oci] <br />Optional: \{\} <br /> |
|
||
| `jobName` _string_ | JobName is the name of the checkpoint creation Job | | Optional: \{\} <br /> |
|
||
| `createdAt` _[Time](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#time-v1-meta)_ | CreatedAt is the timestamp when the checkpoint tar was created | | Optional: \{\} <br /> |
|
||
| `message` _string_ | Message provides additional information about the current state | | Optional: \{\} <br /> |
|
||
| `conditions` _[Condition](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#condition-v1-meta) array_ | Conditions represent the latest available observations of the checkpoint's state | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### DynamoCheckpointStorageType
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
DynamoCheckpointStorageType defines the supported storage backends for checkpoints
|
||
|
||
_Validation:_
|
||
- Enum: [pvc s3 oci]
|
||
|
||
_Appears in:_
|
||
- [DynamoCheckpointStatus](#dynamocheckpointstatus)
|
||
|
||
|
||
|
||
#### DynamoComponentDeployment
|
||
|
||
|
||
|
||
DynamoComponentDeployment is the Schema for the dynamocomponentdeployments API
|
||
|
||
|
||
|
||
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `apiVersion` _string_ | `nvidia.com/v1alpha1` | | |
|
||
| `kind` _string_ | `DynamoComponentDeployment` | | |
|
||
| `metadata` _[ObjectMeta](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#objectmeta-v1-meta)_ | Refer to Kubernetes API documentation for fields of `metadata`. | | |
|
||
| `spec` _[DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)_ | Spec defines the desired state for this Dynamo component deployment. | | |
|
||
|
||
|
||
#### DynamoComponentDeploymentSharedSpec
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
- [DynamoGraphDeploymentSpec](#dynamographdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `annotations` _object (keys:string, values:string)_ | Annotations to add to generated Kubernetes resources for this component<br />(such as Pod, Service, and Ingress when applicable). | | |
|
||
| `labels` _object (keys:string, values:string)_ | Labels to add to generated Kubernetes resources for this component. | | |
|
||
| `serviceName` _string_ | The name of the component | | |
|
||
| `componentType` _string_ | ComponentType indicates the role of this component (for example, "main"). | | |
|
||
| `subComponentType` _string_ | SubComponentType indicates the sub-role of this component (for example, "prefill"). | | |
|
||
| `dynamoNamespace` _string_ | DynamoNamespace is deprecated and will be removed in a future version.<br />The DGD Kubernetes namespace and DynamoGraphDeployment name are used to construct the Dynamo namespace for each component | | Optional: \{\} <br /> |
|
||
| `globalDynamoNamespace` _boolean_ | GlobalDynamoNamespace indicates that the Component will be placed in the global Dynamo namespace | | |
|
||
| `resources` _[Resources](#resources)_ | Resources requested and limits for this component, including CPU, memory,<br />GPUs/devices, and any runtime-specific resources. | | |
|
||
| `autoscaling` _[Autoscaling](#autoscaling)_ | Deprecated: This field is deprecated and ignored. Use DynamoGraphDeploymentScalingAdapter<br />with HPA, KEDA, or Planner for autoscaling instead. See docs/kubernetes/autoscaling.md<br />for migration guidance. This field will be removed in a future API version. | | |
|
||
| `envs` _[EnvVar](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#envvar-v1-core) array_ | Envs defines additional environment variables to inject into the component containers. | | |
|
||
| `envFromSecret` _string_ | EnvFromSecret references a Secret whose key/value pairs will be exposed as<br />environment variables in the component containers. | | |
|
||
| `volumeMounts` _[VolumeMount](#volumemount) array_ | VolumeMounts references PVCs defined at the top level for volumes to be mounted by the component. | | |
|
||
| `ingress` _[IngressSpec](#ingressspec)_ | Ingress config to expose the component outside the cluster (or through a service mesh). | | |
|
||
| `modelRef` _[ModelReference](#modelreference)_ | ModelRef references a model that this component serves<br />When specified, a headless service will be created for endpoint discovery | | Optional: \{\} <br /> |
|
||
| `sharedMemory` _[SharedMemorySpec](#sharedmemoryspec)_ | SharedMemory controls the tmpfs mounted at /dev/shm (enable/disable and size). | | |
|
||
| `extraPodMetadata` _[ExtraPodMetadata](#extrapodmetadata)_ | ExtraPodMetadata adds labels/annotations to the created Pods. | | Optional: \{\} <br /> |
|
||
| `extraPodSpec` _[ExtraPodSpec](#extrapodspec)_ | ExtraPodSpec allows to override the main pod spec configuration.<br />It is a k8s standard PodSpec. It also contains a MainContainer (standard k8s Container) field<br />that allows overriding the main container configuration. | | Optional: \{\} <br /> |
|
||
| `livenessProbe` _[Probe](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#probe-v1-core)_ | LivenessProbe to detect and restart unhealthy containers. | | |
|
||
| `readinessProbe` _[Probe](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#probe-v1-core)_ | ReadinessProbe to signal when the container is ready to receive traffic. | | |
|
||
| `replicas` _integer_ | Replicas is the desired number of Pods for this component.<br />When scalingAdapter is enabled, this field is managed by the<br />DynamoGraphDeploymentScalingAdapter and should not be modified directly. | | Minimum: 0 <br /> |
|
||
| `multinode` _[MultinodeSpec](#multinodespec)_ | Multinode is the configuration for multinode components. | | |
|
||
| `scalingAdapter` _[ScalingAdapter](#scalingadapter)_ | ScalingAdapter configures whether this service uses the DynamoGraphDeploymentScalingAdapter.<br />When enabled, replicas are managed via DGDSA and external autoscalers can scale<br />the service using the Scale subresource. When disabled, replicas can be modified directly. | | Optional: \{\} <br /> |
|
||
| `eppConfig` _[EPPConfig](#eppconfig)_ | EPPConfig defines EPP-specific configuration options for Endpoint Picker Plugin components.<br />Only applicable when ComponentType is "epp". | | Optional: \{\} <br /> |
|
||
| `frontendSidecar` _[FrontendSidecarSpec](#frontendsidecarspec)_ | FrontendSidecar configures an auto-generated frontend sidecar container.<br />When specified, the operator injects a fully configured frontend container<br />with all standard Dynamo environment variables, health probes, and ports.<br />This eliminates the need to manually specify these in extraPodSpec.containers. (GAIE) | | Optional: \{\} <br /> |
|
||
| `checkpoint` _[ServiceCheckpointConfig](#servicecheckpointconfig)_ | Checkpoint configures container checkpointing for this service.<br />When enabled, pods can be restored from a checkpoint files for faster cold start. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### DynamoComponentDeploymentSpec
|
||
|
||
|
||
|
||
DynamoComponentDeploymentSpec defines the desired state of DynamoComponentDeployment
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeployment](#dynamocomponentdeployment)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `backendFramework` _string_ | BackendFramework specifies the backend framework (e.g., "sglang", "vllm", "trtllm") | | Enum: [sglang vllm trtllm] <br /> |
|
||
| `annotations` _object (keys:string, values:string)_ | Annotations to add to generated Kubernetes resources for this component<br />(such as Pod, Service, and Ingress when applicable). | | |
|
||
| `labels` _object (keys:string, values:string)_ | Labels to add to generated Kubernetes resources for this component. | | |
|
||
| `serviceName` _string_ | The name of the component | | |
|
||
| `componentType` _string_ | ComponentType indicates the role of this component (for example, "main"). | | |
|
||
| `subComponentType` _string_ | SubComponentType indicates the sub-role of this component (for example, "prefill"). | | |
|
||
| `dynamoNamespace` _string_ | DynamoNamespace is deprecated and will be removed in a future version.<br />The DGD Kubernetes namespace and DynamoGraphDeployment name are used to construct the Dynamo namespace for each component | | Optional: \{\} <br /> |
|
||
| `globalDynamoNamespace` _boolean_ | GlobalDynamoNamespace indicates that the Component will be placed in the global Dynamo namespace | | |
|
||
| `resources` _[Resources](#resources)_ | Resources requested and limits for this component, including CPU, memory,<br />GPUs/devices, and any runtime-specific resources. | | |
|
||
| `autoscaling` _[Autoscaling](#autoscaling)_ | Deprecated: This field is deprecated and ignored. Use DynamoGraphDeploymentScalingAdapter<br />with HPA, KEDA, or Planner for autoscaling instead. See docs/kubernetes/autoscaling.md<br />for migration guidance. This field will be removed in a future API version. | | |
|
||
| `envs` _[EnvVar](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#envvar-v1-core) array_ | Envs defines additional environment variables to inject into the component containers. | | |
|
||
| `envFromSecret` _string_ | EnvFromSecret references a Secret whose key/value pairs will be exposed as<br />environment variables in the component containers. | | |
|
||
| `volumeMounts` _[VolumeMount](#volumemount) array_ | VolumeMounts references PVCs defined at the top level for volumes to be mounted by the component. | | |
|
||
| `ingress` _[IngressSpec](#ingressspec)_ | Ingress config to expose the component outside the cluster (or through a service mesh). | | |
|
||
| `modelRef` _[ModelReference](#modelreference)_ | ModelRef references a model that this component serves<br />When specified, a headless service will be created for endpoint discovery | | Optional: \{\} <br /> |
|
||
| `sharedMemory` _[SharedMemorySpec](#sharedmemoryspec)_ | SharedMemory controls the tmpfs mounted at /dev/shm (enable/disable and size). | | |
|
||
| `extraPodMetadata` _[ExtraPodMetadata](#extrapodmetadata)_ | ExtraPodMetadata adds labels/annotations to the created Pods. | | Optional: \{\} <br /> |
|
||
| `extraPodSpec` _[ExtraPodSpec](#extrapodspec)_ | ExtraPodSpec allows to override the main pod spec configuration.<br />It is a k8s standard PodSpec. It also contains a MainContainer (standard k8s Container) field<br />that allows overriding the main container configuration. | | Optional: \{\} <br /> |
|
||
| `livenessProbe` _[Probe](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#probe-v1-core)_ | LivenessProbe to detect and restart unhealthy containers. | | |
|
||
| `readinessProbe` _[Probe](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#probe-v1-core)_ | ReadinessProbe to signal when the container is ready to receive traffic. | | |
|
||
| `replicas` _integer_ | Replicas is the desired number of Pods for this component.<br />When scalingAdapter is enabled, this field is managed by the<br />DynamoGraphDeploymentScalingAdapter and should not be modified directly. | | Minimum: 0 <br /> |
|
||
| `multinode` _[MultinodeSpec](#multinodespec)_ | Multinode is the configuration for multinode components. | | |
|
||
| `scalingAdapter` _[ScalingAdapter](#scalingadapter)_ | ScalingAdapter configures whether this service uses the DynamoGraphDeploymentScalingAdapter.<br />When enabled, replicas are managed via DGDSA and external autoscalers can scale<br />the service using the Scale subresource. When disabled, replicas can be modified directly. | | Optional: \{\} <br /> |
|
||
| `eppConfig` _[EPPConfig](#eppconfig)_ | EPPConfig defines EPP-specific configuration options for Endpoint Picker Plugin components.<br />Only applicable when ComponentType is "epp". | | Optional: \{\} <br /> |
|
||
| `frontendSidecar` _[FrontendSidecarSpec](#frontendsidecarspec)_ | FrontendSidecar configures an auto-generated frontend sidecar container.<br />When specified, the operator injects a fully configured frontend container<br />with all standard Dynamo environment variables, health probes, and ports.<br />This eliminates the need to manually specify these in extraPodSpec.containers. (GAIE) | | Optional: \{\} <br /> |
|
||
| `checkpoint` _[ServiceCheckpointConfig](#servicecheckpointconfig)_ | Checkpoint configures container checkpointing for this service.<br />When enabled, pods can be restored from a checkpoint files for faster cold start. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### DynamoGraphDeployment
|
||
|
||
|
||
|
||
DynamoGraphDeployment is the Schema for the dynamographdeployments API.
|
||
|
||
|
||
|
||
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `apiVersion` _string_ | `nvidia.com/v1alpha1` | | |
|
||
| `kind` _string_ | `DynamoGraphDeployment` | | |
|
||
| `metadata` _[ObjectMeta](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#objectmeta-v1-meta)_ | Refer to Kubernetes API documentation for fields of `metadata`. | | |
|
||
| `spec` _[DynamoGraphDeploymentSpec](#dynamographdeploymentspec)_ | Spec defines the desired state for this graph deployment. | | |
|
||
| `status` _[DynamoGraphDeploymentStatus](#dynamographdeploymentstatus)_ | Status reflects the current observed state of this graph deployment. | | |
|
||
|
||
|
||
#### DynamoGraphDeploymentRequest
|
||
|
||
|
||
|
||
DynamoGraphDeploymentRequest is the Schema for the dynamographdeploymentrequests API.
|
||
It serves as the primary interface for users to request model deployments with
|
||
specific performance and resource constraints, enabling SLA-driven deployments.
|
||
|
||
Lifecycle:
|
||
1. Initializing → Pending: Validates spec and prepares for profiling
|
||
2. Pending → Profiling: Creates and runs profiling job (online or AIC)
|
||
3. Profiling → Ready/Deploying: Generates DGD spec after profiling completes
|
||
4. Deploying → Ready: When autoApply=true, monitors DGD until Ready
|
||
5. Ready: Terminal state when DGD is operational or spec is available
|
||
6. DeploymentDeleted: Terminal state when auto-created DGD is manually deleted
|
||
|
||
The spec becomes immutable once profiling starts. Users must delete and recreate
|
||
the DGDR to modify configuration after this point.
|
||
|
||
DEPRECATION NOTICE: v1alpha1 DynamoGraphDeploymentRequest is deprecated.
|
||
Please migrate to nvidia.com/v1beta1 DynamoGraphDeploymentRequest.
|
||
v1alpha1 will be removed in a future release.
|
||
|
||
|
||
|
||
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `apiVersion` _string_ | `nvidia.com/v1alpha1` | | |
|
||
| `kind` _string_ | `DynamoGraphDeploymentRequest` | | |
|
||
| `metadata` _[ObjectMeta](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#objectmeta-v1-meta)_ | Refer to Kubernetes API documentation for fields of `metadata`. | | |
|
||
| `spec` _[DynamoGraphDeploymentRequestSpec](#dynamographdeploymentrequestspec)_ | Spec defines the desired state for this deployment request. | | |
|
||
| `status` _[DynamoGraphDeploymentRequestStatus](#dynamographdeploymentrequeststatus)_ | Status reflects the current observed state of this deployment request. | | |
|
||
|
||
|
||
#### DynamoGraphDeploymentRequestSpec
|
||
|
||
|
||
|
||
DynamoGraphDeploymentRequestSpec defines the desired state of a DynamoGraphDeploymentRequest.
|
||
This CRD serves as the primary interface for users to request model deployments with
|
||
specific performance constraints and resource requirements, enabling SLA-driven deployments.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequest](#dynamographdeploymentrequest)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `model` _string_ | Model specifies the model to deploy (e.g., "Qwen/Qwen3-0.6B", "meta-llama/Llama-3-70b").<br />This is a high-level identifier for easy reference in kubectl output and logs.<br />The controller automatically sets this value in profilingConfig.config.deployment.model. | | Required: \{\} <br /> |
|
||
| `backend` _string_ | Backend specifies the inference backend for profiling.<br />The controller automatically sets this value in profilingConfig.config.engine.backend.<br />Profiling runs on real GPUs or via AIC simulation to collect performance data. | | Enum: [auto vllm sglang trtllm] <br />Required: \{\} <br /> |
|
||
| `useMocker` _boolean_ | UseMocker indicates whether to deploy a mocker DynamoGraphDeployment instead of<br />a real backend deployment. When true, the deployment uses simulated engines that<br />don't require GPUs, using the profiling data to simulate realistic timing behavior.<br />Mocker is available in all backend images and useful for large-scale experiments.<br />Profiling still runs against the real backend (specified above) to collect performance data. | false | |
|
||
| `profilingConfig` _[ProfilingConfigSpec](#profilingconfigspec)_ | ProfilingConfig provides the complete configuration for the profiling job.<br />Note: GPU discovery is automatically attempted to detect GPU resources from Kubernetes<br />cluster nodes. If the operator has node read permissions (cluster-wide or explicitly granted),<br />discovered GPU configuration is used as defaults when hardware configuration is not manually<br />specified (minNumGpusPerEngine, maxNumGpusPerEngine, numGpusPerNode). User-specified values<br />always take precedence over auto-discovered values. If GPU discovery fails (e.g.,<br />namespace-restricted operator without node permissions), manual hardware config is required.<br />This configuration is passed directly to the profiler.<br />The structure matches the profile_sla config format exactly (see ProfilingConfigSpec for schema).<br />Note: deployment.model and engine.backend are automatically set from the high-level<br />modelName and backend fields and should not be specified in this config. | | Required: \{\} <br /> |
|
||
| `enableGpuDiscovery` _boolean_ | EnableGPUDiscovery controls whether the operator attempts to discover GPU hardware from cluster nodes.<br />DEPRECATED: This field is deprecated and will be removed in v1beta1. GPU discovery is now always<br />attempted automatically. Setting this field has no effect - the operator will always try to discover<br />GPU hardware when node read permissions are available. If discovery is unavailable (e.g., namespace-scoped<br />operator without permissions), manual hardware configuration is required regardless of this setting. | true | Optional: \{\} <br /> |
|
||
| `autoApply` _boolean_ | AutoApply indicates whether to automatically create a DynamoGraphDeployment<br />after profiling completes. If false, only the spec is generated and stored in status.<br />Users can then manually create a DGD using the generated spec. | false | |
|
||
| `deploymentOverrides` _[DeploymentOverridesSpec](#deploymentoverridesspec)_ | DeploymentOverrides allows customizing metadata for the auto-created DGD.<br />Only applicable when AutoApply is true. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### DynamoGraphDeploymentRequestStatus
|
||
|
||
|
||
|
||
DynamoGraphDeploymentRequestStatus represents the observed state of a DynamoGraphDeploymentRequest.
|
||
The controller updates this status as the DGDR progresses through its lifecycle.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequest](#dynamographdeploymentrequest)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `state` _[DGDRState](#dgdrstate)_ | State is a high-level textual status of the deployment request lifecycle. | Initializing | Enum: [Initializing Pending Profiling Deploying Ready DeploymentDeleted Failed] <br /> |
|
||
| `backend` _string_ | Backend is extracted from profilingConfig.config.engine.backend for display purposes.<br />This field is populated by the controller and shown in kubectl output. | | Optional: \{\} <br /> |
|
||
| `observedGeneration` _integer_ | ObservedGeneration reflects the generation of the most recently observed spec.<br />Used to detect spec changes and enforce immutability after profiling starts. | | |
|
||
| `conditions` _[Condition](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#condition-v1-meta) array_ | Conditions contains the latest observed conditions of the deployment request.<br />Standard condition types include: Validation, Profiling, SpecGenerated, DeploymentReady.<br />Conditions are merged by type on patch updates. | | |
|
||
| `profilingResults` _string_ | ProfilingResults contains a reference to the ConfigMap holding profiling data.<br />Format: "configmap/\<name\>" | | Optional: \{\} <br /> |
|
||
| `generatedDeployment` _[RawExtension](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#rawextension-runtime-pkg)_ | GeneratedDeployment contains the full generated DynamoGraphDeployment specification<br />including metadata, based on profiling results. Users can extract this to create<br />a DGD manually, or it's used automatically when autoApply is true.<br />Stored as RawExtension to preserve all fields including metadata.<br />For mocker backends, this contains the mocker DGD spec. | | EmbeddedResource: \{\} <br />Optional: \{\} <br /> |
|
||
| `deployment` _[DeploymentStatus](#deploymentstatus)_ | Deployment tracks the auto-created DGD when AutoApply is true.<br />Contains name, namespace, state, and creation status of the managed DGD. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### DynamoGraphDeploymentScalingAdapter
|
||
|
||
|
||
|
||
DynamoGraphDeploymentScalingAdapter provides a scaling interface for individual services
|
||
within a DynamoGraphDeployment. It implements the Kubernetes scale
|
||
subresource, enabling integration with HPA, KEDA, and custom autoscalers.
|
||
|
||
The adapter acts as an intermediary between autoscalers and the DGD,
|
||
ensuring that only the adapter controller modifies the DGD's service replicas.
|
||
This prevents conflicts when multiple autoscaling mechanisms are in play.
|
||
|
||
|
||
|
||
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `apiVersion` _string_ | `nvidia.com/v1alpha1` | | |
|
||
| `kind` _string_ | `DynamoGraphDeploymentScalingAdapter` | | |
|
||
| `metadata` _[ObjectMeta](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#objectmeta-v1-meta)_ | Refer to Kubernetes API documentation for fields of `metadata`. | | |
|
||
| `spec` _[DynamoGraphDeploymentScalingAdapterSpec](#dynamographdeploymentscalingadapterspec)_ | | | |
|
||
| `status` _[DynamoGraphDeploymentScalingAdapterStatus](#dynamographdeploymentscalingadapterstatus)_ | | | |
|
||
|
||
|
||
#### DynamoGraphDeploymentScalingAdapterSpec
|
||
|
||
|
||
|
||
DynamoGraphDeploymentScalingAdapterSpec defines the desired state of DynamoGraphDeploymentScalingAdapter
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentScalingAdapter](#dynamographdeploymentscalingadapter)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `replicas` _integer_ | Replicas is the desired number of replicas for the target service.<br />This field is modified by external autoscalers (HPA/KEDA/Planner) or manually by users. | | Minimum: 0 <br />Required: \{\} <br /> |
|
||
| `dgdRef` _[DynamoGraphDeploymentServiceRef](#dynamographdeploymentserviceref)_ | DGDRef references the DynamoGraphDeployment and the specific service to scale. | | Required: \{\} <br /> |
|
||
|
||
|
||
#### DynamoGraphDeploymentScalingAdapterStatus
|
||
|
||
|
||
|
||
DynamoGraphDeploymentScalingAdapterStatus defines the observed state of DynamoGraphDeploymentScalingAdapter
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentScalingAdapter](#dynamographdeploymentscalingadapter)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `replicas` _integer_ | Replicas is the current number of replicas for the target service.<br />This is synced from the DGD's service replicas and is required for the scale subresource. | | Optional: \{\} <br /> |
|
||
| `selector` _string_ | Selector is a label selector string for the pods managed by this adapter.<br />Required for HPA compatibility via the scale subresource. | | Optional: \{\} <br /> |
|
||
| `lastScaleTime` _[Time](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#time-v1-meta)_ | LastScaleTime is the last time the adapter scaled the target service. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### DynamoGraphDeploymentServiceRef
|
||
|
||
|
||
|
||
DynamoGraphDeploymentServiceRef identifies a specific service within a DynamoGraphDeployment
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentScalingAdapterSpec](#dynamographdeploymentscalingadapterspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `name` _string_ | Name of the DynamoGraphDeployment | | MinLength: 1 <br />Required: \{\} <br /> |
|
||
| `serviceName` _string_ | ServiceName is the key name of the service within the DGD's spec.services map to scale | | MinLength: 1 <br />Required: \{\} <br /> |
|
||
|
||
|
||
#### DynamoGraphDeploymentSpec
|
||
|
||
|
||
|
||
DynamoGraphDeploymentSpec defines the desired state of DynamoGraphDeployment.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeployment](#dynamographdeployment)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `pvcs` _[PVC](#pvc) array_ | PVCs defines a list of persistent volume claims that can be referenced by components.<br />Each PVC must have a unique name that can be referenced in component specifications. | | MaxItems: 100 <br />Optional: \{\} <br /> |
|
||
| `services` _object (keys:string, values:[DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec))_ | Services are the services to deploy as part of this deployment. | | MaxProperties: 25 <br />Optional: \{\} <br /> |
|
||
| `envs` _[EnvVar](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#envvar-v1-core) array_ | Envs are environment variables applied to all services in the deployment unless<br />overridden by service-specific configuration. | | Optional: \{\} <br /> |
|
||
| `backendFramework` _string_ | BackendFramework specifies the backend framework (e.g., "sglang", "vllm", "trtllm"). | | Enum: [sglang vllm trtllm] <br /> |
|
||
| `restart` _[Restart](#restart)_ | Restart specifies the restart policy for the graph deployment. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### DynamoGraphDeploymentStatus
|
||
|
||
|
||
|
||
DynamoGraphDeploymentStatus defines the observed state of DynamoGraphDeployment.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeployment](#dynamographdeployment)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `observedGeneration` _integer_ | ObservedGeneration is the most recent generation observed by the controller. | | Optional: \{\} <br /> |
|
||
| `state` _[DGDState](#dgdstate)_ | State is a high-level textual status of the graph deployment lifecycle. | initializing | Enum: [initializing pending successful failed] <br /> |
|
||
| `conditions` _[Condition](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#condition-v1-meta) array_ | Conditions contains the latest observed conditions of the graph deployment.<br />The slice is merged by type on patch updates. | | |
|
||
| `services` _object (keys:string, values:[ServiceReplicaStatus](#servicereplicastatus))_ | Services contains per-service replica status information.<br />The map key is the service name from spec.services. | | Optional: \{\} <br /> |
|
||
| `restart` _[RestartStatus](#restartstatus)_ | Restart contains the status of the restart of the graph deployment. | | Optional: \{\} <br /> |
|
||
| `checkpoints` _object (keys:string, values:[ServiceCheckpointStatus](#servicecheckpointstatus))_ | Checkpoints contains per-service checkpoint status information.<br />The map key is the service name from spec.services. | | Optional: \{\} <br /> |
|
||
| `rollingUpdate` _[RollingUpdateStatus](#rollingupdatestatus)_ | RollingUpdate tracks the progress of operator manged rolling updates.<br />Currently only supported for singl-node, non-Grove deployments (DCD/Deployment). | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### DynamoModel
|
||
|
||
|
||
|
||
DynamoModel is the Schema for the dynamo models API
|
||
|
||
|
||
|
||
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `apiVersion` _string_ | `nvidia.com/v1alpha1` | | |
|
||
| `kind` _string_ | `DynamoModel` | | |
|
||
| `metadata` _[ObjectMeta](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#objectmeta-v1-meta)_ | Refer to Kubernetes API documentation for fields of `metadata`. | | |
|
||
| `spec` _[DynamoModelSpec](#dynamomodelspec)_ | | | |
|
||
| `status` _[DynamoModelStatus](#dynamomodelstatus)_ | | | |
|
||
|
||
|
||
#### DynamoModelSpec
|
||
|
||
|
||
|
||
DynamoModelSpec defines the desired state of DynamoModel
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoModel](#dynamomodel)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `modelName` _string_ | ModelName is the full model identifier (e.g., "meta-llama/Llama-3.3-70B-Instruct-lora") | | Required: \{\} <br /> |
|
||
| `baseModelName` _string_ | BaseModelName is the base model identifier that matches the service label<br />This is used to discover endpoints via headless services | | Required: \{\} <br /> |
|
||
| `modelType` _string_ | ModelType specifies the type of model (e.g., "base", "lora", "adapter") | base | Enum: [base lora adapter] <br />Optional: \{\} <br /> |
|
||
| `source` _[ModelSource](#modelsource)_ | Source specifies the model source location (only applicable for lora model type) | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### DynamoModelStatus
|
||
|
||
|
||
|
||
DynamoModelStatus defines the observed state of DynamoModel
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoModel](#dynamomodel)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `endpoints` _[EndpointInfo](#endpointinfo) array_ | Endpoints is the current list of all endpoints for this model | | Optional: \{\} <br /> |
|
||
| `readyEndpoints` _integer_ | ReadyEndpoints is the count of endpoints that are ready | | |
|
||
| `totalEndpoints` _integer_ | TotalEndpoints is the total count of endpoints | | |
|
||
| `conditions` _[Condition](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#condition-v1-meta) array_ | Conditions represents the latest available observations of the model's state | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### EPPConfig
|
||
|
||
|
||
|
||
EPPConfig contains configuration for EPP (Endpoint Picker Plugin) components.
|
||
EPP is responsible for intelligent endpoint selection and KV-aware routing.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec)
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `configMapRef` _[ConfigMapKeySelector](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#configmapkeyselector-v1-core)_ | ConfigMapRef references a user-provided ConfigMap containing EPP configuration.<br />The ConfigMap should contain EndpointPickerConfig YAML.<br />Mutually exclusive with Config. | | Optional: \{\} <br /> |
|
||
| `config` _[EndpointPickerConfig](#endpointpickerconfig)_ | Config allows specifying EPP EndpointPickerConfig directly as a structured object.<br />The operator will marshal this to YAML and create a ConfigMap automatically.<br />Mutually exclusive with ConfigMapRef.<br />One of ConfigMapRef or Config must be specified (no default configuration).<br />Uses the upstream type from github.com/kubernetes-sigs/gateway-api-inference-extension | | Type: object <br />Optional: \{\} <br /> |
|
||
|
||
|
||
#### EndpointInfo
|
||
|
||
|
||
|
||
EndpointInfo represents a single endpoint (pod) serving the model
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoModelStatus](#dynamomodelstatus)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `address` _string_ | Address is the full address of the endpoint (e.g., "http://10.0.1.5:9090") | | |
|
||
| `podName` _string_ | PodName is the name of the pod serving this endpoint | | Optional: \{\} <br /> |
|
||
| `ready` _boolean_ | Ready indicates whether the endpoint is ready to serve traffic<br />For LoRA models: true if the POST /loras request succeeded with a 2xx status code<br />For base models: always false (no probing performed) | | |
|
||
|
||
|
||
#### ExtraPodMetadata
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec)
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `annotations` _object (keys:string, values:string)_ | | | |
|
||
| `labels` _object (keys:string, values:string)_ | | | |
|
||
|
||
|
||
#### ExtraPodSpec
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec)
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `mainContainer` _[Container](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#container-v1-core)_ | | | |
|
||
|
||
|
||
#### FrontendSidecarSpec
|
||
|
||
|
||
|
||
FrontendSidecarSpec configures the auto-generated frontend sidecar container.
|
||
The operator uses these fields together with built-in frontend defaults (command, probes, ports,
|
||
and Dynamo env vars) to produce a fully configured sidecar container.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec)
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `image` _string_ | Image is the container image for the frontend sidecar. | | Required: \{\} <br /> |
|
||
| `args` _string array_ | Args overrides the default frontend arguments. When specified, these replace<br />the default ["-m", "dynamo.frontend"] entirely.<br />For example, ["-m", "dynamo.frontend", "--router-mode", "direct"] for GAIE deployments. | | Optional: \{\} <br /> |
|
||
| `envFromSecret` _string_ | EnvFromSecret references a Secret whose key/value pairs will be exposed as<br />environment variables in the frontend sidecar container. | | Optional: \{\} <br /> |
|
||
| `envs` _[EnvVar](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#envvar-v1-core) array_ | Envs defines additional environment variables for the frontend sidecar.<br />These are merged with (and can override) the auto-generated Dynamo env vars. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### IngressSpec
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec)
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `enabled` _boolean_ | Enabled exposes the component through an ingress or virtual service when true. | | |
|
||
| `host` _string_ | Host is the base host name to route external traffic to this component. | | |
|
||
| `useVirtualService` _boolean_ | UseVirtualService indicates whether to configure a service-mesh VirtualService instead of a standard Ingress. | | |
|
||
| `virtualServiceGateway` _string_ | VirtualServiceGateway optionally specifies the gateway name to attach the VirtualService to. | | |
|
||
| `hostPrefix` _string_ | HostPrefix is an optional prefix added before the host. | | |
|
||
| `annotations` _object (keys:string, values:string)_ | Annotations to set on the generated Ingress/VirtualService resources. | | |
|
||
| `labels` _object (keys:string, values:string)_ | Labels to set on the generated Ingress/VirtualService resources. | | |
|
||
| `tls` _[IngressTLSSpec](#ingresstlsspec)_ | TLS holds the TLS configuration used by the Ingress/VirtualService. | | |
|
||
| `hostSuffix` _string_ | HostSuffix is an optional suffix appended after the host. | | |
|
||
| `ingressControllerClassName` _string_ | IngressControllerClassName selects the ingress controller class (e.g., "nginx"). | | |
|
||
|
||
|
||
#### IngressTLSSpec
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [IngressSpec](#ingressspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `secretName` _string_ | SecretName is the name of a Kubernetes Secret containing the TLS certificate and key. | | |
|
||
|
||
|
||
|
||
|
||
#### ModelReference
|
||
|
||
|
||
|
||
ModelReference identifies a model served by this component
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec)
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `name` _string_ | Name is the base model identifier (e.g., "llama-3-70b-instruct-v1") | | Required: \{\} <br /> |
|
||
| `revision` _string_ | Revision is the model revision/version (optional) | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### ModelSource
|
||
|
||
|
||
|
||
ModelSource defines the source location of a model
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoModelSpec](#dynamomodelspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `uri` _string_ | URI is the model source URI<br />Supported formats:<br />- S3: s3://bucket/path/to/model<br />- HuggingFace: hf://org/model@revision_sha | | Required: \{\} <br /> |
|
||
|
||
|
||
#### MultinodeSpec
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec)
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `nodeCount` _integer_ | Indicates the number of nodes to deploy for multinode components.<br />Total number of GPUs is NumberOfNodes * GPU limit.<br />Must be greater than 1. | 2 | Minimum: 2 <br /> |
|
||
|
||
|
||
#### PVC
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentSpec](#dynamographdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `create` _boolean_ | Create indicates to create a new PVC | | |
|
||
| `name` _string_ | Name is the name of the PVC | | Required: \{\} <br /> |
|
||
| `storageClass` _string_ | StorageClass to be used for PVC creation. Required when create is true. | | |
|
||
| `size` _[Quantity](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#quantity-resource-api)_ | Size of the volume in Gi, used during PVC creation. Required when create is true. | | |
|
||
| `volumeAccessMode` _[PersistentVolumeAccessMode](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#persistentvolumeaccessmode-v1-core)_ | VolumeAccessMode is the volume access mode of the PVC. Required when create is true. | | |
|
||
|
||
|
||
#### ProfilingConfigSpec
|
||
|
||
|
||
|
||
ProfilingConfigSpec defines configuration for the profiling process.
|
||
This structure maps directly to the profile_sla.py config format.
|
||
See dynamo/profiler/utils/profiler_argparse.py for the complete schema.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestSpec](#dynamographdeploymentrequestspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `config` _[JSON](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#json-v1-apiextensions-k8s-io)_ | Config is the profiling configuration as arbitrary JSON/YAML. This will be passed directly to the profiler.<br />The profiler will validate the configuration and report any errors. | | Optional: \{\} <br />Type: object <br /> |
|
||
| `configMapRef` _[ConfigMapKeySelector](#configmapkeyselector)_ | ConfigMapRef is an optional reference to a ConfigMap containing the DynamoGraphDeployment<br />base config file (disagg.yaml). This is separate from the profiling config above.<br />The path to this config will be set as engine.config in the profiling config. | | Optional: \{\} <br /> |
|
||
| `profilerImage` _string_ | ProfilerImage specifies the container image to use for profiling jobs.<br />This image contains the profiler code and dependencies needed for SLA-based profiling.<br />Example: "nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.0.0" | | Required: \{\} <br /> |
|
||
| `outputPVC` _string_ | OutputPVC is an optional PersistentVolumeClaim name for storing profiling output.<br />If specified, all profiling artifacts (logs, plots, configs, raw data) will be written<br />to this PVC instead of an ephemeral emptyDir volume. This allows users to access<br />complete profiling results after the job completes by mounting the PVC.<br />The PVC must exist in the same namespace as the DGDR.<br />If not specified, profiling uses emptyDir and only essential data is saved to ConfigMaps.<br />Note: ConfigMaps are still created regardless of this setting for planner integration. | | Optional: \{\} <br /> |
|
||
| `resources` _[ResourceRequirements](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#resourcerequirements-v1-core)_ | Resources specifies the compute resource requirements for the profiling job container.<br />If not specified, no resource requests or limits are set. | | Optional: \{\} <br /> |
|
||
| `tolerations` _[Toleration](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#toleration-v1-core) array_ | Tolerations allows the profiling job to be scheduled on nodes with matching taints.<br />For example, to schedule on GPU nodes, add a toleration for the nvidia.com/gpu taint. | | Optional: \{\} <br /> |
|
||
| `nodeSelector` _object (keys:string, values:string)_ | NodeSelector is a selector which must match a node's labels for the profiling pod to be scheduled on that node.<br />For example, to schedule on ARM64 nodes, use \{"kubernetes.io/arch": "arm64"\}. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### ResourceItem
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [Resources](#resources)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `cpu` _string_ | CPU specifies the CPU resource request/limit (e.g., "1000m", "2") | | |
|
||
| `memory` _string_ | Memory specifies the memory resource request/limit (e.g., "4Gi", "8Gi") | | |
|
||
| `gpu` _string_ | GPU indicates the number of GPUs to request.<br />Total number of GPUs is NumberOfNodes * GPU in case of multinode deployment. | | |
|
||
| `gpuType` _string_ | GPUType can specify a custom GPU type, e.g. "gpu.intel.com/xe"<br />By default if not specified, the GPU type is "nvidia.com/gpu" | | |
|
||
| `custom` _object (keys:string, values:string)_ | Custom specifies additional custom resource requests/limits | | |
|
||
|
||
|
||
#### Resources
|
||
|
||
|
||
|
||
Resources defines requested and limits for a component, including CPU, memory,
|
||
GPUs/devices, and any runtime-specific resources.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec)
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `requests` _[ResourceItem](#resourceitem)_ | Requests specifies the minimum resources required by the component | | |
|
||
| `limits` _[ResourceItem](#resourceitem)_ | Limits specifies the maximum resources allowed for the component | | |
|
||
| `claims` _[ResourceClaim](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#resourceclaim-v1-core) array_ | Claims specifies resource claims for dynamic resource allocation | | |
|
||
|
||
|
||
#### Restart
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentSpec](#dynamographdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `id` _string_ | ID is an arbitrary string that triggers a restart when changed.<br />Any modification to this value will initiate a restart of the graph deployment according to the strategy. | | MinLength: 1 <br />Required: \{\} <br /> |
|
||
| `strategy` _[RestartStrategy](#restartstrategy)_ | Strategy specifies the restart strategy for the graph deployment. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### RestartPhase
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [RestartStatus](#restartstatus)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `Pending` | |
|
||
| `Restarting` | |
|
||
| `Completed` | |
|
||
| `Failed` | |
|
||
| `Superseded` | |
|
||
|
||
|
||
#### RestartStatus
|
||
|
||
|
||
|
||
RestartStatus contains the status of the restart of the graph deployment.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentStatus](#dynamographdeploymentstatus)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `observedID` _string_ | ObservedID is the restart ID that has been observed and is being processed.<br />Matches the Restart.ID field in the spec. | | |
|
||
| `phase` _[RestartPhase](#restartphase)_ | Phase is the phase of the restart. | | |
|
||
| `inProgress` _string array_ | InProgress contains the names of the services that are currently being restarted. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### RestartStrategy
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [Restart](#restart)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `type` _[RestartStrategyType](#restartstrategytype)_ | Type specifies the restart strategy type. | Sequential | Enum: [Sequential Parallel] <br /> |
|
||
| `order` _string array_ | Order specifies the order in which the services should be restarted. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### RestartStrategyType
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [RestartStrategy](#restartstrategy)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `Sequential` | |
|
||
| `Parallel` | |
|
||
|
||
|
||
#### RollingUpdatePhase
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
RollingUpdatePhase represents the current phase of a rolling update.
|
||
|
||
_Validation:_
|
||
- Enum: [Pending InProgress Completed Failed ]
|
||
|
||
_Appears in:_
|
||
- [RollingUpdateStatus](#rollingupdatestatus)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `Pending` | |
|
||
| `InProgress` | |
|
||
| `Completed` | |
|
||
| `` | |
|
||
|
||
|
||
#### RollingUpdateStatus
|
||
|
||
|
||
|
||
RollingUpdateStatus tracks the progress of a rolling update.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentStatus](#dynamographdeploymentstatus)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `phase` _[RollingUpdatePhase](#rollingupdatephase)_ | Phase indicates the current phase of the rolling update. | | Enum: [Pending InProgress Completed Failed ] <br />Optional: \{\} <br /> |
|
||
| `startTime` _[Time](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#time-v1-meta)_ | StartTime is when the rolling update began. | | Optional: \{\} <br /> |
|
||
| `endTime` _[Time](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#time-v1-meta)_ | EndTime is when the rolling update completed (successfully or failed). | | Optional: \{\} <br /> |
|
||
| `updatedServices` _string array_ | UpdatedServices is the list of services that have completed the rolling update.<br />A service is considered updated when its new replicas are all ready and old replicas are fully scaled down.<br />Only services of componentType Worker (or Prefill/Decode) are considered. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### ScalingAdapter
|
||
|
||
|
||
|
||
ScalingAdapter configures whether a service uses the DynamoGraphDeploymentScalingAdapter
|
||
for replica management. When enabled, the DGDSA owns the replicas field and
|
||
external autoscalers (HPA, KEDA, Planner) can control scaling via the Scale subresource.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec)
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `enabled` _boolean_ | Enabled indicates whether the ScalingAdapter should be enabled for this service.<br />When true, a DGDSA is created and owns the replicas field.<br />When false (default), no DGDSA is created and replicas can be modified directly in the DGD. | false | Optional: \{\} <br /> |
|
||
|
||
|
||
#### ServiceCheckpointConfig
|
||
|
||
|
||
|
||
ServiceCheckpointConfig configures checkpointing for a DGD service
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec)
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `enabled` _boolean_ | Enabled indicates whether checkpointing is enabled for this service | false | Optional: \{\} <br /> |
|
||
| `mode` _[CheckpointMode](#checkpointmode)_ | Mode defines how checkpoint creation is handled<br />- Auto: DGD controller creates Checkpoint CR automatically<br />- Manual: User must create Checkpoint CR | Auto | Enum: [Auto Manual] <br />Optional: \{\} <br /> |
|
||
| `checkpointRef` _string_ | CheckpointRef references an existing Checkpoint CR to use<br />If specified, Identity is ignored and this checkpoint is used directly | | Optional: \{\} <br /> |
|
||
| `identity` _[DynamoCheckpointIdentity](#dynamocheckpointidentity)_ | Identity defines the checkpoint identity for hash computation<br />Used when Mode is Auto or when looking up existing checkpoints<br />Required when checkpointRef is not specified | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### ServiceCheckpointStatus
|
||
|
||
|
||
|
||
ServiceCheckpointStatus contains checkpoint information for a single service.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentStatus](#dynamographdeploymentstatus)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `checkpointName` _string_ | CheckpointName is the name of the associated Checkpoint CR | | Optional: \{\} <br /> |
|
||
| `identityHash` _string_ | IdentityHash is the computed hash of the checkpoint identity | | Optional: \{\} <br /> |
|
||
| `ready` _boolean_ | Ready indicates if the checkpoint is ready for use | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### ServiceReplicaStatus
|
||
|
||
|
||
|
||
ServiceReplicaStatus contains replica information for a single service.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentStatus](#dynamographdeploymentstatus)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `componentKind` _[ComponentKind](#componentkind)_ | ComponentKind is the underlying resource kind (e.g., "PodClique", "PodCliqueScalingGroup", "Deployment", "LeaderWorkerSet"). | | Enum: [PodClique PodCliqueScalingGroup Deployment LeaderWorkerSet] <br /> |
|
||
| `componentName` _string_ | ComponentName is the name of the primary underlying resource.<br />DEPRECATED: Use ComponentNames instead. This field will be removed in a future release.<br />During rolling updates, this reflects the new (target) component name. | | |
|
||
| `componentNames` _string array_ | ComponentNames is the list of underlying resource names for this service.<br />During normal operation, this contains a single name.<br />During rolling updates, this contains both old and new component names. | | Optional: \{\} <br /> |
|
||
| `replicas` _integer_ | Replicas is the total number of non-terminated replicas.<br />Required for all component kinds. | | Minimum: 0 <br /> |
|
||
| `updatedReplicas` _integer_ | UpdatedReplicas is the number of replicas at the current/desired revision.<br />Required for all component kinds. | | Minimum: 0 <br /> |
|
||
| `readyReplicas` _integer_ | ReadyReplicas is the number of ready replicas.<br />Populated for PodClique, Deployment, and LeaderWorkerSet.<br />Not available for PodCliqueScalingGroup.<br />When nil, the field is omitted from the API response. | | Minimum: 0 <br />Optional: \{\} <br /> |
|
||
| `availableReplicas` _integer_ | AvailableReplicas is the number of available replicas.<br />For Deployment: replicas ready for >= minReadySeconds.<br />For PodCliqueScalingGroup: replicas where all constituent PodCliques have >= MinAvailable ready pods.<br />Not available for PodClique or LeaderWorkerSet.<br />When nil, the field is omitted from the API response. | | Minimum: 0 <br />Optional: \{\} <br /> |
|
||
|
||
|
||
#### SharedMemorySpec
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec)
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `disabled` _boolean_ | | | |
|
||
| `size` _[Quantity](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#quantity-resource-api)_ | | | |
|
||
|
||
|
||
#### VolumeMount
|
||
|
||
|
||
|
||
VolumeMount references a PVC defined at the top level for volumes to be mounted by the component
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoComponentDeploymentSharedSpec](#dynamocomponentdeploymentsharedspec)
|
||
- [DynamoComponentDeploymentSpec](#dynamocomponentdeploymentspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `name` _string_ | Name references a PVC name defined in the top-level PVCs map | | Required: \{\} <br /> |
|
||
| `mountPoint` _string_ | MountPoint specifies where to mount the volume.<br />If useAsCompilationCache is true and mountPoint is not specified,<br />a backend-specific default will be used. | | |
|
||
| `useAsCompilationCache` _boolean_ | UseAsCompilationCache indicates this volume should be used as a compilation cache.<br />When true, backend-specific environment variables will be set and default mount points may be used. | false | |
|
||
|
||
|
||
|
||
## nvidia.com/v1beta1
|
||
|
||
Package v1beta1 contains API Schema definitions for the nvidia.com v1beta1 API group.
|
||
|
||
### Resource Types
|
||
- [DynamoGraphDeploymentRequest](#v1beta1-dynamographdeploymentrequest)
|
||
|
||
|
||
|
||
#### BackendType
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
BackendType specifies the inference backend.
|
||
|
||
_Validation:_
|
||
- Enum: [auto sglang trtllm vllm]
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestSpec](#v1beta1-dynamographdeploymentrequestspec)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `auto` | |
|
||
| `sglang` | |
|
||
| `trtllm` | |
|
||
| `vllm` | |
|
||
|
||
|
||
#### DGDRPhase
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
DGDRPhase represents the lifecycle phase of a DynamoGraphDeploymentRequest.
|
||
|
||
_Validation:_
|
||
- Enum: [Pending Profiling Ready Deploying Deployed Failed]
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestStatus](#v1beta1-dynamographdeploymentrequeststatus)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `Pending` | |
|
||
| `Profiling` | |
|
||
| `Ready` | |
|
||
| `Deploying` | |
|
||
| `Deployed` | |
|
||
| `Failed` | |
|
||
|
||
|
||
#### DeploymentInfoStatus
|
||
|
||
|
||
|
||
DeploymentInfoStatus tracks the state of the deployed DynamoGraphDeployment.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestStatus](#v1beta1-dynamographdeploymentrequeststatus)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `replicas` _integer_ | Replicas is the desired number of replicas. | | Optional: \{\} <br /> |
|
||
| `availableReplicas` _integer_ | AvailableReplicas is the number of replicas that are available and ready. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### v1beta1 DynamoGraphDeploymentRequest
|
||
|
||
|
||
|
||
DynamoGraphDeploymentRequest is the Schema for the dynamographdeploymentrequests API.
|
||
It provides a simplified, SLA-driven interface for deploying inference models on Dynamo.
|
||
Users specify a model and optional performance targets; the controller handles profiling,
|
||
configuration selection, and deployment.
|
||
|
||
Lifecycle:
|
||
1. Pending: Spec validated, preparing for profiling
|
||
2. Profiling: Profiling job is running to discover optimal configurations
|
||
3. Ready: Profiling complete, generated DGD spec available in status
|
||
4. Deploying: DGD is being created and rolled out (when autoApply=true)
|
||
5. Deployed: DGD is running and healthy
|
||
6. Failed: An unrecoverable error occurred
|
||
|
||
|
||
|
||
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `apiVersion` _string_ | `nvidia.com/v1beta1` | | |
|
||
| `kind` _string_ | `DynamoGraphDeploymentRequest` | | |
|
||
| `metadata` _[ObjectMeta](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#objectmeta-v1-meta)_ | Refer to Kubernetes API documentation for fields of `metadata`. | | |
|
||
| `spec` _[DynamoGraphDeploymentRequestSpec](#v1beta1-dynamographdeploymentrequestspec)_ | Spec defines the desired state for this deployment request. | | |
|
||
| `status` _[DynamoGraphDeploymentRequestStatus](#v1beta1-dynamographdeploymentrequeststatus)_ | Status reflects the current observed state of this deployment request. | | |
|
||
|
||
|
||
#### v1beta1 DynamoGraphDeploymentRequestSpec
|
||
|
||
|
||
|
||
DynamoGraphDeploymentRequestSpec defines the desired state of a DynamoGraphDeploymentRequest.
|
||
Only the Model field is required; all other fields are optional and have sensible defaults.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequest](#v1beta1-dynamographdeploymentrequest)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `model` _string_ | Model specifies the model to deploy (e.g., "Qwen/Qwen3-0.6B", "meta-llama/Llama-3-70b").<br />Can be a HuggingFace ID or a private model name. | | MinLength: 1 <br />Required: \{\} <br /> |
|
||
| `backend` _[BackendType](#backendtype)_ | Backend specifies the inference backend to use for profiling and deployment. | auto | Enum: [auto sglang trtllm vllm] <br />Optional: \{\} <br /> |
|
||
| `image` _string_ | Image is the container image reference for the profiling job (frontend image).<br />Example: "nvcr.io/nvidia/ai-dynamo/dynamo-frontend:1.0.0". | | Optional: \{\} <br /> |
|
||
| `modelCache` _[ModelCacheSpec](#modelcachespec)_ | ModelCache provides optional PVC configuration for pre-downloaded model weights.<br />When provided, weights are loaded from the PVC instead of downloading from HuggingFace. | | Optional: \{\} <br /> |
|
||
| `hardware` _[HardwareSpec](#hardwarespec)_ | Hardware describes the hardware resources available for profiling and deployment.<br />Typically auto-filled by the operator from cluster discovery. | | Optional: \{\} <br /> |
|
||
| `workload` _[WorkloadSpec](#workloadspec)_ | Workload defines the expected workload characteristics for SLA-based profiling. | | Optional: \{\} <br /> |
|
||
| `sla` _[SLASpec](#slaspec)_ | SLA defines service-level agreement targets that drive profiling optimization. | | Optional: \{\} <br /> |
|
||
| `overrides` _[OverridesSpec](#overridesspec)_ | Overrides allows customizing the profiling job and the generated DynamoGraphDeployment. | | Optional: \{\} <br /> |
|
||
| `features` _[FeaturesSpec](#featuresspec)_ | Features controls optional Dynamo platform features in the generated deployment. | | Optional: \{\} <br /> |
|
||
| `searchStrategy` _[SearchStrategy](#searchstrategy)_ | SearchStrategy controls the profiling search depth.<br />"rapid" performs a fast sweep; "thorough" explores more configurations. | rapid | Enum: [rapid thorough] <br />Optional: \{\} <br /> |
|
||
| `autoApply` _boolean_ | AutoApply indicates whether to automatically create a DynamoGraphDeployment<br />after profiling completes. If false, the generated spec is stored in status<br />for manual review and application. | true | Optional: \{\} <br /> |
|
||
|
||
|
||
#### v1beta1 DynamoGraphDeploymentRequestStatus
|
||
|
||
|
||
|
||
DynamoGraphDeploymentRequestStatus represents the observed state of a DynamoGraphDeploymentRequest.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequest](#v1beta1-dynamographdeploymentrequest)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `phase` _[DGDRPhase](#dgdrphase)_ | Phase is the high-level lifecycle phase of the deployment request. | | Enum: [Pending Profiling Ready Deploying Deployed Failed] <br />Optional: \{\} <br /> |
|
||
| `profilingPhase` _[ProfilingPhase](#profilingphase)_ | ProfilingPhase indicates the current sub-phase of the profiling pipeline.<br />Only meaningful when Phase is "Profiling". Cleared when profiling completes or fails. | | Enum: [Initializing SweepingPrefill SweepingDecode SelectingConfig BuildingCurves GeneratingDGD Done] <br />Optional: \{\} <br /> |
|
||
| `dgdName` _string_ | DGDName is the name of the generated or created DynamoGraphDeployment. | | Optional: \{\} <br /> |
|
||
| `profilingJobName` _string_ | ProfilingJobName is the name of the Kubernetes Job running the profiler. | | Optional: \{\} <br /> |
|
||
| `conditions` _[Condition](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#condition-v1-meta) array_ | Conditions contains the latest observed conditions of the deployment request.<br />Standard condition types include: Succeeded, Validation, Profiling, SpecGenerated, DeploymentReady. | | Optional: \{\} <br /> |
|
||
| `profilingResults` _[ProfilingResultsStatus](#profilingresultsstatus)_ | ProfilingResults contains the output of the profiling process including<br />Pareto-optimal configurations and the selected deployment configuration. | | Optional: \{\} <br /> |
|
||
| `deploymentInfo` _[DeploymentInfoStatus](#deploymentinfostatus)_ | DeploymentInfo tracks the state of the deployed DynamoGraphDeployment.<br />Populated when a DGD has been created (either via autoApply or manually). | | Optional: \{\} <br /> |
|
||
| `observedGeneration` _integer_ | ObservedGeneration is the most recent generation observed by the controller. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### FeaturesSpec
|
||
|
||
|
||
|
||
FeaturesSpec controls optional Dynamo platform features in the generated deployment.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestSpec](#v1beta1-dynamographdeploymentrequestspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `planner` _[RawExtension](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#rawextension-runtime-pkg)_ | Planner is the raw SLA planner configuration passed to the planner service.<br />Its schema is defined by dynamo.planner.utils.planner_config.PlannerConfig.<br />Go treats this as opaque bytes; the Planner service validates it at startup.<br />The presence of this field (non-null) enables the planner in the generated DGD. | | Type: object <br />Optional: \{\} <br /> |
|
||
| `mocker` _[MockerSpec](#mockerspec)_ | Mocker configures the simulated (mocker) backend for testing without GPUs. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### GPUSKUType
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
GPUSKUType is the AIC hardware system identifier for a supported GPU.
|
||
|
||
_Validation:_
|
||
- Enum: [gb200_sxm h200_sxm h100_sxm b200_sxm a100_sxm l40s]
|
||
|
||
_Appears in:_
|
||
- [HardwareSpec](#hardwarespec)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `gb200_sxm` | |
|
||
| `h200_sxm` | |
|
||
| `h100_sxm` | |
|
||
| `b200_sxm` | |
|
||
| `a100_sxm` | |
|
||
| `l40s` | |
|
||
|
||
|
||
#### HardwareSpec
|
||
|
||
|
||
|
||
HardwareSpec describes the hardware resources available for profiling and deployment.
|
||
These fields are typically auto-filled by the operator from cluster discovery.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestSpec](#v1beta1-dynamographdeploymentrequestspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `gpuSku` _[GPUSKUType](#gpuskutype)_ | GPUSKU is the AIC hardware system identifier for the GPU.<br />When omitted, the operator auto-detects this via InferHardwareSystem from cluster GPU node labels. | | Enum: [gb200_sxm h200_sxm h100_sxm b200_sxm a100_sxm l40s] <br />Optional: \{\} <br /> |
|
||
| `vramMb` _float_ | VRAMMB is the VRAM per GPU in MiB. | | Optional: \{\} <br /> |
|
||
| `totalGpus` _integer_ | TotalGPUs is the total number of GPUs available in the cluster. | | Optional: \{\} <br /> |
|
||
| `numGpusPerNode` _integer_ | NumGPUsPerNode is the number of GPUs per node. | | Optional: \{\} <br /> |
|
||
|
||
|
||
|
||
|
||
#### MockerSpec
|
||
|
||
|
||
|
||
MockerSpec configures the simulated (mocker) backend.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [FeaturesSpec](#featuresspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `enabled` _boolean_ | Enabled indicates whether to deploy mocker workers instead of real inference workers.<br />Useful for large-scale testing without GPUs. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### ModelCacheSpec
|
||
|
||
|
||
|
||
ModelCacheSpec references a PVC containing pre-downloaded model weights.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestSpec](#v1beta1-dynamographdeploymentrequestspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `pvcName` _string_ | PVCName is the name of the PersistentVolumeClaim containing model weights.<br />The PVC must exist in the same namespace as the DGDR. | | Optional: \{\} <br /> |
|
||
| `pvcModelPath` _string_ | PVCModelPath is the path to the model checkpoint directory within the PVC<br />(e.g. "deepseek-r1" or "models/Llama-3.1-405B-FP8"). | | Optional: \{\} <br /> |
|
||
| `pvcMountPath` _string_ | PVCMountPath is the mount path for the PVC inside the container. | /opt/model-cache | Optional: \{\} <br /> |
|
||
|
||
|
||
#### OverridesSpec
|
||
|
||
|
||
|
||
OverridesSpec allows customizing the profiling job and the generated DynamoGraphDeployment.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestSpec](#v1beta1-dynamographdeploymentrequestspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `profilingJob` _[JobSpec](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#jobspec-v1-batch)_ | ProfilingJob allows overriding the profiling Job specification.<br />Fields set here are merged into the controller-generated Job spec. | | Optional: \{\} <br /> |
|
||
| `dgd` _[RawExtension](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#rawextension-runtime-pkg)_ | DGD allows providing a full or partial nvidia.com/v1alpha1 DynamoGraphDeployment<br />to use as the base for the generated deployment. Fields from profiling results<br />are merged on top. Use this to override backend worker images.<br />The field is stored as a raw embedded resource rather than a typed<br />*v1alpha1.DynamoGraphDeployment to avoid a circular import: v1alpha1 already<br />imports v1beta1 as the conversion hub and Go does not allow import cycles.<br />The EmbeddedResource marker tells the API server to validate that the value is a<br />well-formed Kubernetes object (has apiVersion/kind), but does not enforce that it<br />is specifically a DynamoGraphDeployment. Full type validation (correct apiVersion,<br />kind, and field schema) is performed by the controller during reconciliation. | | EmbeddedResource: \{\} <br />Optional: \{\} <br /> |
|
||
|
||
|
||
#### ParetoConfig
|
||
|
||
|
||
|
||
ParetoConfig represents a single Pareto-optimal deployment configuration
|
||
discovered during profiling.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [ProfilingResultsStatus](#profilingresultsstatus)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `config` _[RawExtension](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#rawextension-runtime-pkg)_ | Config is the full deployment configuration for this Pareto point. | | Type: object <br /> |
|
||
|
||
|
||
#### ProfilingPhase
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
ProfilingPhase represents a sub-phase within the profiling pipeline.
|
||
When the DGDR Phase is "Profiling", this value indicates which step
|
||
of the profiling pipeline is currently executing.
|
||
|
||
_Validation:_
|
||
- Enum: [Initializing SweepingPrefill SweepingDecode SelectingConfig BuildingCurves GeneratingDGD Done]
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestStatus](#v1beta1-dynamographdeploymentrequeststatus)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `Initializing` | Profiler is loading the DGD template, detecting GPU hardware,<br />and resolving the model architecture from HuggingFace.<br /> |
|
||
| `SweepingPrefill` | Sweeping parallelization strategies (TP/TEP/DEP) across GPU counts<br />for prefill, measuring TTFT at each configuration.<br /> |
|
||
| `SweepingDecode` | Sweeping parallelization strategies and concurrency levels<br />for decode, measuring ITL at each configuration.<br /> |
|
||
| `SelectingConfig` | Filtering results against SLA targets and selecting the most<br />cost-efficient configuration that meets TTFT/ITL requirements.<br /> |
|
||
| `BuildingCurves` | Building detailed interpolation curves (ISL→TTFT for prefill,<br />KV-usage×context-length→ITL for decode) using the selected configs.<br /> |
|
||
| `GeneratingDGD` | Packaging profiling data into a ConfigMap and generating<br />the final DGD YAML with planner integration.<br /> |
|
||
| `Done` | Profiling pipeline finished successfully.<br /> |
|
||
|
||
|
||
#### ProfilingResultsStatus
|
||
|
||
|
||
|
||
ProfilingResultsStatus contains the output of the profiling process.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestStatus](#v1beta1-dynamographdeploymentrequeststatus)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `pareto` _[ParetoConfig](#paretoconfig) array_ | Pareto is the list of Pareto-optimal deployment configurations discovered during profiling.<br />Each entry represents a different cost/performance trade-off. | | Optional: \{\} <br /> |
|
||
| `selectedConfig` _[RawExtension](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#rawextension-runtime-pkg)_ | SelectedConfig is the recommended configuration chosen by the profiler<br />based on the SLA targets. This is the configuration used for deployment<br />when autoApply is true. | | Type: object <br />Optional: \{\} <br /> |
|
||
|
||
|
||
#### SLASpec
|
||
|
||
|
||
|
||
SLASpec defines the service-level agreement targets for profiling optimization.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestSpec](#v1beta1-dynamographdeploymentrequestspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `ttft` _float_ | TTFT is the Time To First Token target in milliseconds. | | Optional: \{\} <br /> |
|
||
| `itl` _float_ | ITL is the Inter-Token Latency target in milliseconds. | | Optional: \{\} <br /> |
|
||
| `e2eLatency` _float_ | E2ELatency is the target end-to-end request latency in milliseconds.<br />Alternative to specifying TTFT + ITL. | | Optional: \{\} <br /> |
|
||
|
||
|
||
#### SearchStrategy
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
SearchStrategy controls the profiling search depth.
|
||
|
||
_Validation:_
|
||
- Enum: [rapid thorough]
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestSpec](#v1beta1-dynamographdeploymentrequestspec)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `rapid` | |
|
||
| `thorough` | |
|
||
|
||
|
||
#### WorkloadSpec
|
||
|
||
|
||
|
||
WorkloadSpec defines the workload characteristics for SLA-based profiling.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DynamoGraphDeploymentRequestSpec](#v1beta1-dynamographdeploymentrequestspec)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `isl` _integer_ | ISL is the Input Sequence Length (number of tokens). | 4000 | Optional: \{\} <br /> |
|
||
| `osl` _integer_ | OSL is the Output Sequence Length (number of tokens). | 1000 | Optional: \{\} <br /> |
|
||
| `concurrency` _float_ | Concurrency is the target concurrency level.<br />Required (or RequestRate) when the planner is disabled. | | Optional: \{\} <br /> |
|
||
| `requestRate` _float_ | RequestRate is the target request rate (req/s).<br />Required (or Concurrency) when the planner is disabled. | | Optional: \{\} <br /> |
|
||
|
||
|
||
|
||
## operator.config.dynamo.nvidia.com/v1alpha1
|
||
|
||
|
||
### Resource Types
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
|
||
|
||
#### CertProvisionMode
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
CertProvisionMode controls how webhook TLS certificates are managed.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [WebhookServer](#webhookserver)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `auto` | CertProvisionModeAuto uses the built-in cert-controller to generate and rotate certificates.<br /> |
|
||
| `manual` | CertProvisionModeManual expects certificates to be provided externally (e.g., cert-manager, admin).<br /> |
|
||
|
||
|
||
#### CheckpointConfiguration
|
||
|
||
|
||
|
||
CheckpointConfiguration holds checkpoint/restore settings.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `enabled` _boolean_ | Enabled indicates if checkpoint functionality is enabled | | |
|
||
| `readyForCheckpointFilePath` _string_ | ReadyForCheckpointFilePath signals model readiness for checkpoint jobs | /tmp/ready-for-checkpoint | |
|
||
| `storage` _[CheckpointStorageConfiguration](#checkpointstorageconfiguration)_ | Storage holds storage backend configuration | | |
|
||
|
||
|
||
#### CheckpointOCIConfig
|
||
|
||
|
||
|
||
CheckpointOCIConfig holds OCI registry storage configuration.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [CheckpointStorageConfiguration](#checkpointstorageconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `uri` _string_ | URI is the OCI URI (oci://registry/repository) | | |
|
||
| `credentialsSecretRef` _string_ | CredentialsSecretRef is the name of the docker config secret | | |
|
||
|
||
|
||
#### CheckpointPVCConfig
|
||
|
||
|
||
|
||
CheckpointPVCConfig holds PVC storage configuration.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [CheckpointStorageConfiguration](#checkpointstorageconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `pvcName` _string_ | PVCName is the name of the PVC | snapshot-pvc | |
|
||
| `basePath` _string_ | BasePath is the base directory within the PVC | /checkpoints | |
|
||
|
||
|
||
#### CheckpointS3Config
|
||
|
||
|
||
|
||
CheckpointS3Config holds S3 storage configuration.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [CheckpointStorageConfiguration](#checkpointstorageconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `uri` _string_ | URI is the S3 URI (s3://[endpoint/]bucket/prefix) | | |
|
||
| `credentialsSecretRef` _string_ | CredentialsSecretRef is the name of the credentials secret | | |
|
||
|
||
|
||
#### CheckpointStorageConfiguration
|
||
|
||
|
||
|
||
CheckpointStorageConfiguration holds storage backend configuration for checkpoints.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [CheckpointConfiguration](#checkpointconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `type` _string_ | Type is the storage backend type: pvc, s3, or oci | pvc | |
|
||
| `pvc` _[CheckpointPVCConfig](#checkpointpvcconfig)_ | PVC configuration (used when Type=pvc) | | |
|
||
| `s3` _[CheckpointS3Config](#checkpoints3config)_ | S3 configuration (used when Type=s3) | | |
|
||
| `oci` _[CheckpointOCIConfig](#checkpointociconfig)_ | OCI configuration (used when Type=oci) | | |
|
||
|
||
|
||
#### DiscoveryBackend
|
||
|
||
_Underlying type:_ _string_
|
||
|
||
DiscoveryBackend is the type for the discovery backend.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [DiscoveryConfiguration](#discoveryconfiguration)
|
||
|
||
| Field | Description |
|
||
| --- | --- |
|
||
| `kubernetes` | DiscoveryBackendKubernetes is the Kubernetes discovery backend<br /> |
|
||
| `etcd` | DiscoveryBackendEtcd is the etcd discovery backend<br /> |
|
||
|
||
|
||
#### DiscoveryConfiguration
|
||
|
||
|
||
|
||
DiscoveryConfiguration holds discovery backend settings.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `backend` _[DiscoveryBackend](#discoverybackend)_ | Backend is the discovery backend: "kubernetes" or "etcd" | kubernetes | |
|
||
|
||
|
||
#### GPUConfiguration
|
||
|
||
|
||
|
||
GPUConfiguration holds GPU discovery settings.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `discoveryEnabled` _boolean_ | DiscoveryEnabled indicates whether GPU discovery is enabled | true | |
|
||
|
||
|
||
#### GroveConfiguration
|
||
|
||
|
||
|
||
GroveConfiguration holds Grove orchestrator settings.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OrchestratorConfiguration](#orchestratorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `enabled` _boolean_ | Enabled overrides auto-detection. nil = auto-detect. | | |
|
||
| `terminationDelay` _[Duration](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#duration-v1-meta)_ | TerminationDelay configures the termination delay for Grove PodCliqueSets | 15m | |
|
||
|
||
|
||
#### InfrastructureConfiguration
|
||
|
||
|
||
|
||
InfrastructureConfiguration holds service mesh and backend addresses.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `natsAddress` _string_ | NATSAddress is the address of the NATS server | | |
|
||
| `etcdAddress` _string_ | ETCDAddress is the address of the etcd server | | |
|
||
| `modelExpressURL` _string_ | ModelExpressURL is the URL of the Model Express server to inject into all pods | | |
|
||
| `prometheusEndpoint` _string_ | PrometheusEndpoint is the URL of the Prometheus endpoint to use for metrics | | |
|
||
|
||
|
||
#### IngressConfiguration
|
||
|
||
|
||
|
||
IngressConfiguration holds ingress settings.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `virtualServiceGateway` _string_ | VirtualServiceGateway is the name of the Istio virtual service gateway | | |
|
||
| `controllerClassName` _string_ | ControllerClassName is the ingress controller class name | | |
|
||
| `controllerTLSSecretName` _string_ | ControllerTLSSecretName is the TLS secret for the ingress controller | | |
|
||
| `hostSuffix` _string_ | HostSuffix is the suffix for ingress hostnames | | |
|
||
|
||
|
||
#### KaiSchedulerConfiguration
|
||
|
||
|
||
|
||
KaiSchedulerConfiguration holds Kai-scheduler settings.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OrchestratorConfiguration](#orchestratorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `enabled` _boolean_ | Enabled overrides auto-detection. nil = auto-detect. | | |
|
||
|
||
|
||
#### LWSConfiguration
|
||
|
||
|
||
|
||
LWSConfiguration holds LWS orchestrator settings.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OrchestratorConfiguration](#orchestratorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `enabled` _boolean_ | Enabled overrides auto-detection. nil = auto-detect. | | |
|
||
|
||
|
||
#### LeaderElectionConfiguration
|
||
|
||
|
||
|
||
LeaderElectionConfiguration holds leader election settings.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `enabled` _boolean_ | Enabled enables leader election for controller manager | false | |
|
||
| `id` _string_ | ID is the leader election resource identity | | |
|
||
| `namespace` _string_ | Namespace is the namespace for the leader election resource | | |
|
||
|
||
|
||
#### LoggingConfiguration
|
||
|
||
|
||
|
||
LoggingConfiguration holds logging settings.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `level` _string_ | Level is the log level (e.g., "info", "debug") | info | |
|
||
| `format` _string_ | Format is the log format (e.g., "json", "text") | json | |
|
||
|
||
|
||
#### MPIConfiguration
|
||
|
||
|
||
|
||
MPIConfiguration holds MPI SSH secret settings.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `sshSecretName` _string_ | SSHSecretName is the name of the secret containing the SSH key for MPI | | |
|
||
| `sshSecretNamespace` _string_ | SSHSecretNamespace is the namespace where the MPI SSH secret is located | | |
|
||
|
||
|
||
#### MetricsServer
|
||
|
||
|
||
|
||
MetricsServer extends Server with secure serving option.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [ServerConfiguration](#serverconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `bindAddress` _string_ | BindAddress is the address the server binds to | | |
|
||
| `port` _integer_ | Port is the port the server listens on | | |
|
||
| `secure` _boolean_ | Secure enables secure serving for the metrics endpoint.<br />nil = default to true (secure by default). | | |
|
||
|
||
|
||
#### NamespaceConfiguration
|
||
|
||
|
||
|
||
NamespaceConfiguration determines operator namespace mode.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `restricted` _string_ | Restricted is the namespace to restrict to. Empty = cluster-wide mode. | | |
|
||
| `scope` _[NamespaceScopeConfiguration](#namespacescopeconfiguration)_ | Scope holds namespace scope lease settings (namespace-restricted mode only) | | |
|
||
|
||
|
||
#### NamespaceScopeConfiguration
|
||
|
||
|
||
|
||
NamespaceScopeConfiguration holds lease settings for namespace-restricted mode.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [NamespaceConfiguration](#namespaceconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `leaseDuration` _[Duration](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#duration-v1-meta)_ | LeaseDuration is the duration of namespace scope marker lease before expiration | 30s | |
|
||
| `leaseRenewInterval` _[Duration](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#duration-v1-meta)_ | LeaseRenewInterval is the interval for renewing namespace scope marker lease | 10s | |
|
||
|
||
|
||
#### OperatorConfiguration
|
||
|
||
|
||
|
||
OperatorConfiguration is the Schema for the operator configuration.
|
||
|
||
|
||
|
||
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `apiVersion` _string_ | `operator.config.dynamo.nvidia.com/v1alpha1` | | |
|
||
| `kind` _string_ | `OperatorConfiguration` | | |
|
||
| `server` _[ServerConfiguration](#serverconfiguration)_ | Server configuration (metrics, health probes, webhooks) | | |
|
||
| `leaderElection` _[LeaderElectionConfiguration](#leaderelectionconfiguration)_ | Leader election configuration | | |
|
||
| `namespace` _[NamespaceConfiguration](#namespaceconfiguration)_ | Namespace configuration (restricted vs cluster-wide) | | |
|
||
| `orchestrators` _[OrchestratorConfiguration](#orchestratorconfiguration)_ | Orchestrator configuration with optional overrides | | |
|
||
| `infrastructure` _[InfrastructureConfiguration](#infrastructureconfiguration)_ | Service mesh and infrastructure addresses | | |
|
||
| `ingress` _[IngressConfiguration](#ingressconfiguration)_ | Ingress configuration | | |
|
||
| `rbac` _[RBACConfiguration](#rbacconfiguration)_ | RBAC configuration for cross-namespace resource management (cluster-wide mode) | | |
|
||
| `mpi` _[MPIConfiguration](#mpiconfiguration)_ | MPI SSH secret configuration | | |
|
||
| `checkpoint` _[CheckpointConfiguration](#checkpointconfiguration)_ | Checkpoint/restore configuration | | |
|
||
| `discovery` _[DiscoveryConfiguration](#discoveryconfiguration)_ | Discovery backend configuration | | |
|
||
| `gpu` _[GPUConfiguration](#gpuconfiguration)_ | GPU discovery configuration | | |
|
||
| `logging` _[LoggingConfiguration](#loggingconfiguration)_ | Logging configuration | | |
|
||
| `security` _[SecurityConfiguration](#securityconfiguration)_ | HTTP/2 and TLS settings | | |
|
||
|
||
|
||
#### OrchestratorConfiguration
|
||
|
||
|
||
|
||
OrchestratorConfiguration holds orchestrator override settings.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `grove` _[GroveConfiguration](#groveconfiguration)_ | Grove orchestrator configuration | | |
|
||
| `lws` _[LWSConfiguration](#lwsconfiguration)_ | LWS orchestrator configuration | | |
|
||
| `kaiScheduler` _[KaiSchedulerConfiguration](#kaischedulerconfiguration)_ | KaiScheduler configuration | | |
|
||
|
||
|
||
#### RBACConfiguration
|
||
|
||
|
||
|
||
RBACConfiguration holds RBAC settings for cluster-wide mode.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `plannerClusterRoleName` _string_ | PlannerClusterRoleName is the ClusterRole for planner | | |
|
||
| `dgdrProfilingClusterRoleName` _string_ | DGDRProfilingClusterRoleName is the ClusterRole for DGDR profiling jobs | | |
|
||
| `eppClusterRoleName` _string_ | EPPClusterRoleName is the ClusterRole for EPP | | |
|
||
|
||
|
||
#### SecurityConfiguration
|
||
|
||
|
||
|
||
SecurityConfiguration holds HTTP/2 and TLS settings.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `enableHTTP2` _boolean_ | EnableHTTP2 enables HTTP/2 for metrics and webhook servers | false | |
|
||
|
||
|
||
#### Server
|
||
|
||
|
||
|
||
Server holds a bind address and port.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [MetricsServer](#metricsserver)
|
||
- [ServerConfiguration](#serverconfiguration)
|
||
- [WebhookServer](#webhookserver)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `bindAddress` _string_ | BindAddress is the address the server binds to | | |
|
||
| `port` _integer_ | Port is the port the server listens on | | |
|
||
|
||
|
||
#### ServerConfiguration
|
||
|
||
|
||
|
||
ServerConfiguration holds server bind addresses and ports.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [OperatorConfiguration](#operatorconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `metrics` _[MetricsServer](#metricsserver)_ | Metrics server configuration | \{ bindAddress:0.0.0.0 port:8080 secure:true \} | |
|
||
| `healthProbe` _[Server](#server)_ | Health probe server configuration | \{ bindAddress:0.0.0.0 port:8081 \} | |
|
||
| `webhook` _[WebhookServer](#webhookserver)_ | Webhook server configuration | \{ certDir:/tmp/k8s-webhook-server/serving-certs host:0.0.0.0 port:9443 \} | |
|
||
|
||
|
||
#### WebhookServer
|
||
|
||
|
||
|
||
WebhookServer extends Server with host and certificate directory.
|
||
|
||
|
||
|
||
_Appears in:_
|
||
- [ServerConfiguration](#serverconfiguration)
|
||
|
||
| Field | Description | Default | Validation |
|
||
| --- | --- | --- | --- |
|
||
| `bindAddress` _string_ | BindAddress is the address the server binds to | | |
|
||
| `port` _integer_ | Port is the port the server listens on | | |
|
||
| `host` _string_ | Host is the address the webhook server binds to | | |
|
||
| `certDir` _string_ | CertDir is the directory containing TLS certificates | | |
|
||
| `certProvisionMode` _[CertProvisionMode](#certprovisionmode)_ | CertProvisionMode controls certificate management: "auto" (built-in cert-controller) or "manual" (external) | auto | |
|
||
| `secretName` _string_ | SecretName is the name of the Kubernetes Secret holding webhook TLS certificates | webhook-server-cert | |
|
||
| `serviceName` _string_ | ServiceName is the name of the Kubernetes Service fronting the webhook server.<br />Used to generate certificate SANs. Set by the Helm chart. | | |
|
||
|
||
|
||
# Operator Default Values Injection
|
||
|
||
The Dynamo operator automatically applies default values to various fields when they are not explicitly specified in your deployments. These defaults include:
|
||
|
||
- **Health Probes**: Startup, liveness, and readiness probes are configured differently for frontend, worker, and planner components. For example, worker components receive a startup probe with a 2-hour timeout (720 failures × 10 seconds) to accommodate long model loading times.
|
||
|
||
- **Security Context**: All components receive `fsGroup: 1000` by default to ensure proper file permissions for mounted volumes. This can be overridden via the `extraPodSpec.securityContext` field.
|
||
|
||
- **Shared Memory**: All components receive an 8Gi shared memory volume mounted at `/dev/shm` by default (can be disabled or resized via the `sharedMemory` field).
|
||
|
||
- **Environment Variables**: Components automatically receive environment variables like `DYN_NAMESPACE`, `DYN_PARENT_DGD_K8S_NAME`, `DYNAMO_PORT`, and backend-specific variables.
|
||
|
||
- **Pod Configuration**: Default `terminationGracePeriodSeconds` of 60 seconds and `restartPolicy: Always`.
|
||
|
||
- **Autoscaling**: When enabled without explicit metrics, defaults to CPU-based autoscaling with 80% target utilization.
|
||
|
||
- **Backend-Specific Behavior**: For multinode deployments, probes are automatically modified or removed for worker nodes depending on the backend framework (VLLM, SGLang, or TensorRT-LLM).
|
||
|
||
## Pod Specification Defaults
|
||
|
||
All components receive the following pod-level defaults unless overridden:
|
||
|
||
- **`terminationGracePeriodSeconds`**: `60` seconds
|
||
- **`restartPolicy`**: `Always`
|
||
|
||
## Security Context
|
||
|
||
The operator automatically applies default security context settings to all components to ensure proper file permissions, particularly for mounted volumes:
|
||
|
||
- **`fsGroup`**: `1000` - Sets the group ownership of mounted volumes and any files created in those volumes
|
||
|
||
This default ensures that non-root containers can write to mounted volumes (like model caches or persistent storage) without permission issues. The `fsGroup` setting is particularly important for:
|
||
- Model downloads and caching
|
||
- Compilation cache directories
|
||
- Persistent volume claims (PVCs)
|
||
- SSH key generation in multinode deployments
|
||
|
||
### Overriding Security Context
|
||
|
||
To override the default security context, specify your own `securityContext` in the `extraPodSpec` of your component:
|
||
|
||
```yaml
|
||
services:
|
||
YourWorker:
|
||
extraPodSpec:
|
||
securityContext:
|
||
fsGroup: 2000 # Custom group ID
|
||
runAsUser: 1000
|
||
runAsGroup: 1000
|
||
runAsNonRoot: true
|
||
```
|
||
|
||
**Important**: When you provide *any* `securityContext` object in `extraPodSpec`, the operator will not inject any defaults. This gives you complete control over the security context, including the ability to run as root (by omitting `runAsNonRoot` or setting it to `false`).
|
||
|
||
### OpenShift and Security Context Constraints
|
||
|
||
In OpenShift environments with Security Context Constraints (SCCs), you may need to omit explicit UID/GID values to allow OpenShift's admission controllers to assign them dynamically:
|
||
|
||
```yaml
|
||
services:
|
||
YourWorker:
|
||
extraPodSpec:
|
||
securityContext:
|
||
# Omit fsGroup to let OpenShift assign it based on SCC
|
||
# OpenShift will inject the appropriate UID range
|
||
```
|
||
|
||
Alternatively, if you want to keep the default `fsGroup: 1000` behavior and are certain your cluster allows it, you don't need to specify anything - the operator defaults will work.
|
||
|
||
## Shared Memory Configuration
|
||
|
||
Shared memory is enabled by default for all components:
|
||
|
||
- **Enabled**: `true` (unless explicitly disabled via `sharedMemory.disabled`)
|
||
- **Size**: `8Gi`
|
||
- **Mount Path**: `/dev/shm`
|
||
- **Volume Type**: `emptyDir` with `memory` medium
|
||
|
||
To disable shared memory or customize the size, use the `sharedMemory` field in your component specification.
|
||
|
||
## Health Probes by Component Type
|
||
|
||
The operator applies different default health probes based on the component type.
|
||
|
||
### Frontend Components
|
||
|
||
Frontend components receive the following probe configurations:
|
||
|
||
**Liveness Probe:**
|
||
- **Type**: HTTP GET
|
||
- **Path**: `/health`
|
||
- **Port**: `http` (8000)
|
||
- **Initial Delay**: 60 seconds
|
||
- **Period**: 60 seconds
|
||
- **Timeout**: 30 seconds
|
||
- **Failure Threshold**: 10
|
||
|
||
**Readiness Probe:**
|
||
- **Type**: Exec command
|
||
- **Command**: `curl -s http://localhost:${DYNAMO_PORT}/health | jq -e ".status == \"healthy\""`
|
||
- **Initial Delay**: 60 seconds
|
||
- **Period**: 60 seconds
|
||
- **Timeout**: 30 seconds
|
||
- **Failure Threshold**: 10
|
||
|
||
### Worker Components
|
||
|
||
Worker components receive the following probe configurations:
|
||
|
||
**Liveness Probe:**
|
||
- **Type**: HTTP GET
|
||
- **Path**: `/live`
|
||
- **Port**: `system` (9090)
|
||
- **Period**: 5 seconds
|
||
- **Timeout**: 30 seconds
|
||
- **Failure Threshold**: 1
|
||
|
||
**Readiness Probe:**
|
||
- **Type**: HTTP GET
|
||
- **Path**: `/health`
|
||
- **Port**: `system` (9090)
|
||
- **Period**: 10 seconds
|
||
- **Timeout**: 30 seconds
|
||
- **Failure Threshold**: 60
|
||
|
||
**Startup Probe:**
|
||
- **Type**: HTTP GET
|
||
- **Path**: `/live`
|
||
- **Port**: `system` (9090)
|
||
- **Period**: 10 seconds
|
||
- **Timeout**: 5 seconds
|
||
- **Failure Threshold**: 720 (allows up to 2 hours for startup: 10s × 720 = 7200s)
|
||
|
||
:::{note}
|
||
For larger models (typically >70B parameters) or slower storage systems, you may need to increase the `failureThreshold` to allow more time for model loading. Calculate the required threshold based on your expected startup time: `failureThreshold = (expected_startup_seconds / period)`. Override the startup probe in your component specification if the default 2-hour window is insufficient.
|
||
:::
|
||
|
||
### Multinode Deployment Probe Modifications
|
||
|
||
For multinode deployments, the operator modifies probes based on the backend framework and node role:
|
||
|
||
#### VLLM Backend
|
||
|
||
The operator automatically selects between two deployment modes based on parallelism configuration:
|
||
|
||
**Tensor/Pipeline Parallel Mode** (when `world_size > GPUs_per_node`):
|
||
- Uses Ray for distributed execution (`--distributed-executor-backend ray`)
|
||
- **Leader nodes**: Starts Ray head and runs vLLM; all probes remain active
|
||
- **Worker nodes**: Run Ray agents only; all probes (liveness, readiness, startup) are removed
|
||
|
||
**Data Parallel Mode** (when `world_size × data_parallel_size > GPUs_per_node`):
|
||
- **Worker nodes**: All probes (liveness, readiness, startup) are removed
|
||
- **Leader nodes**: All probes remain active
|
||
|
||
#### SGLang Backend
|
||
- **Worker nodes**: All probes (liveness, readiness, startup) are removed
|
||
|
||
#### TensorRT-LLM Backend
|
||
- **Leader nodes**: All probes remain unchanged
|
||
- **Worker nodes**:
|
||
- Liveness and startup probes are removed
|
||
- Readiness probe is replaced with a TCP socket check on SSH port (2222):
|
||
- **Initial Delay**: 20 seconds
|
||
- **Period**: 20 seconds
|
||
- **Timeout**: 5 seconds
|
||
- **Failure Threshold**: 10
|
||
|
||
## Environment Variables
|
||
|
||
The operator automatically injects environment variables into component containers based on component type, backend framework, and operator configuration. User-provided `envs` values always take precedence over operator defaults.
|
||
|
||
### All Components
|
||
|
||
These environment variables are injected into every component container regardless of type.
|
||
|
||
| Variable | Purpose | Default | Type | Source |
|
||
| --- | --- | --- | --- | --- |
|
||
| `DYN_NAMESPACE` | Dynamo service namespace used for service discovery and routing | Derived from DGD spec | `string` | Downward API annotation on checkpoint-restored pods |
|
||
| `DYN_COMPONENT` | Identifies the component type for runtime behavior | One of: `frontend`, `worker`, `prefill`, `decode`, `planner`, `epp` | `string` | Set from component spec |
|
||
| `DYN_PARENT_DGD_K8S_NAME` | Kubernetes name of the parent DynamoGraphDeployment resource | — | `string` | Set from DGD metadata |
|
||
| `DYN_PARENT_DGD_K8S_NAMESPACE` | Kubernetes namespace of the parent DynamoGraphDeployment resource | — | `string` | Set from DGD metadata |
|
||
| `POD_NAME` | Current pod name | — | `string` | Downward API (`metadata.name`) |
|
||
| `POD_NAMESPACE` | Current pod namespace | — | `string` | Downward API (`metadata.namespace`) |
|
||
| `POD_UID` | Current pod UID | — | `string` | Downward API (`metadata.uid`) |
|
||
| `DYN_DISCOVERY_BACKEND` | Service discovery backend for inter-component communication | `kubernetes` | `string` | Options: `kubernetes`, `etcd` |
|
||
|
||
### Infrastructure (Conditional)
|
||
|
||
These are injected into all components when the corresponding infrastructure service is configured in the operator's `OperatorConfiguration`.
|
||
|
||
| Variable | Purpose | Default | Type | Condition |
|
||
| --- | --- | --- | --- | --- |
|
||
| `NATS_SERVER` | NATS messaging server address | — | `string` | Set when `infrastructure.natsAddress` is configured |
|
||
| `ETCD_ENDPOINTS` | etcd endpoint addresses for distributed state | — | `string` | Set when `infrastructure.etcdAddress` is configured |
|
||
| `MODEL_EXPRESS_URL` | Model Express service URL for model management | — | `string` | Set when `infrastructure.modelExpressURL` is configured |
|
||
| `PROMETHEUS_ENDPOINT` | Prometheus endpoint for metrics collection | — | `string` | Set when `infrastructure.prometheusEndpoint` is configured |
|
||
|
||
### Frontend Components
|
||
|
||
| Variable | Purpose | Default | Type |
|
||
| --- | --- | --- | --- |
|
||
| `DYNAMO_PORT` | HTTP port the frontend listens on | `8000` | `int` |
|
||
| `DYN_HTTP_PORT` | HTTP port for the frontend service (alias) | `8000` | `int` |
|
||
| `DYN_NAMESPACE_PREFIX` | Namespace prefix used for frontend request routing | Same as `DYN_NAMESPACE` | `string` |
|
||
|
||
### Worker Components
|
||
|
||
| Variable | Purpose | Default | Type |
|
||
| --- | --- | --- | --- |
|
||
| `DYN_SYSTEM_ENABLED` | Enables the system HTTP server for health checks and metrics | `true` | `string` (boolean) |
|
||
| `DYN_SYSTEM_USE_ENDPOINT_HEALTH_STATUS` | Endpoints whose health status is used for readiness | `["generate"]` | `string` (JSON array) |
|
||
| `DYN_SYSTEM_PORT` | Port for the system HTTP server (health, metrics) | `9090` | `int` |
|
||
| `DYN_HEALTH_CHECK_ENABLED` | Disables the legacy health check mechanism in favor of the system server | `false` | `string` (boolean) |
|
||
| `NIXL_TELEMETRY_ENABLE` | Enables or disables NIXL telemetry collection | `n` | `string` | Options: `y`, `n` |
|
||
| `NIXL_TELEMETRY_EXPORTER` | Telemetry exporter format for NIXL metrics | `prometheus` | `string` |
|
||
| `NIXL_TELEMETRY_PROMETHEUS_PORT` | Port for NIXL Prometheus metrics endpoint | `19090` | `int` |
|
||
| `DYN_NAMESPACE_WORKER_SUFFIX` | Hash suffix appended to worker namespace for rolling updates | — | `string` | Only set during rolling update transitions |
|
||
|
||
### Planner Components
|
||
|
||
| Variable | Purpose | Default | Type |
|
||
| --- | --- | --- | --- |
|
||
| `PLANNER_PROMETHEUS_PORT` | Port for the planner's Prometheus metrics endpoint | `9085` | `int` |
|
||
|
||
### EPP (Endpoint Picker Plugin) Components
|
||
|
||
| Variable | Purpose | Default | Type |
|
||
| --- | --- | --- | --- |
|
||
| `USE_STREAMING` | Enables streaming mode for inference request proxying | `true` | `string` (boolean) |
|
||
| `RUST_LOG` | Rust log level and filter configuration | `debug,dynamo_llm::kv_router=trace` | `string` |
|
||
|
||
### VLLM Backend
|
||
|
||
| Variable | Purpose | Default | Type | Condition |
|
||
| --- | --- | --- | --- | --- |
|
||
| `VLLM_CACHE_ROOT` | Directory for vLLM compilation cache artifacts | — | `string` | Set when a volume mount has `useAsCompilationCache: true` |
|
||
| `VLLM_NIXL_SIDE_CHANNEL_HOST` | Host IP for the NIXL side channel in multiprocessing mode | Pod IP | `string` | Multinode mp backend only (Downward API: `status.podIP`) |
|
||
|
||
### TensorRT-LLM Backend
|
||
|
||
| Variable | Purpose | Default | Type | Condition |
|
||
| --- | --- | --- | --- | --- |
|
||
| `OMPI_MCA_orte_keep_fqdn_hostnames` | Instructs OpenMPI to preserve FQDN hostnames for inter-node communication | `1` | `string` | Multinode deployments only |
|
||
|
||
### Checkpoint / Restore
|
||
|
||
These environment variables are injected when checkpoint/restore is enabled for a component.
|
||
|
||
| Variable | Purpose | Default | Type | Condition |
|
||
| --- | --- | --- | --- | --- |
|
||
| `DYN_CHECKPOINT_PATH` | Base directory where checkpoint data is stored | From operator checkpoint config `storage.pvc.basePath` | `string` | PVC storage type |
|
||
| `DYN_CHECKPOINT_LOCATION` | Full checkpoint URI (for non-PVC backends) | — | `string` | S3 or OCI storage type |
|
||
| `DYN_CHECKPOINT_HASH` | Identity hash that uniquely identifies the checkpoint | — | `string` | Always set when checkpoint is enabled |
|
||
| `SKIP_WAIT_FOR_CHECKPOINT` | Skips the checkpoint readiness polling loop; checks once and proceeds | — | `string` | Set on restored and DGD pods |
|
||
|
||
## Service Accounts
|
||
|
||
The following component types automatically receive dedicated service accounts:
|
||
|
||
- **Planner**: `planner-serviceaccount`
|
||
- **EPP**: `epp-serviceaccount`
|
||
|
||
## Image Pull Secrets
|
||
|
||
The operator automatically discovers and injects image pull secrets for container images. When a component specifies a container image, the operator:
|
||
|
||
1. Scans all Kubernetes secrets of type `kubernetes.io/dockerconfigjson` in the component's namespace
|
||
2. Extracts the docker registry server URLs from each secret's authentication configuration
|
||
3. Matches the container image's registry host against the discovered registry URLs
|
||
4. Automatically injects matching secrets as `imagePullSecrets` in the pod specification
|
||
|
||
This eliminates the need to manually specify image pull secrets for each component. The operator maintains an internal index of docker secrets and their associated registries, refreshing this index periodically.
|
||
|
||
**To disable automatic image pull secret discovery** for a specific component, add the following annotation:
|
||
|
||
```yaml
|
||
annotations:
|
||
nvidia.com/disable-image-pull-secret-discovery: "true"
|
||
```
|
||
|
||
## Autoscaling Defaults
|
||
|
||
When autoscaling is enabled but no metrics are specified, the operator applies:
|
||
|
||
- **Default Metric**: CPU utilization
|
||
- **Target Average Utilization**: `80%`
|
||
|
||
## Port Configurations
|
||
|
||
Default container ports are configured based on component type:
|
||
|
||
### Frontend Components
|
||
- **Port**: 8000
|
||
- **Protocol**: TCP
|
||
- **Name**: `http`
|
||
|
||
### Worker Components
|
||
- **Port**: 9090 (system)
|
||
- **Protocol**: TCP
|
||
- **Name**: `system`
|
||
- **Port**: 19090 (NIXL)
|
||
- **Protocol**: TCP
|
||
- **Name**: `nixl`
|
||
|
||
### Planner Components
|
||
- **Port**: 9085
|
||
- **Protocol**: TCP
|
||
- **Name**: `metrics`
|
||
|
||
### EPP Components
|
||
- **Port**: 9002 (gRPC)
|
||
- **Protocol**: TCP
|
||
- **Name**: `grpc`
|
||
- **Port**: 9003 (gRPC health)
|
||
- **Protocol**: TCP
|
||
- **Name**: `grpc-health`
|
||
- **Port**: 9090 (metrics)
|
||
- **Protocol**: TCP
|
||
- **Name**: `metrics`
|
||
|
||
## Backend-Specific Configurations
|
||
|
||
### VLLM
|
||
- **Ray Head Port**: 6379 (for Ray cluster coordination in multinode TP/PP deployments)
|
||
- **Data Parallel RPC Port**: 13445 (for data parallel multinode deployments)
|
||
|
||
### SGLang
|
||
- **Distribution Init Port**: 29500 (for multinode deployments)
|
||
|
||
### TensorRT-LLM
|
||
- **SSH Port**: 2222 (for multinode MPI communication)
|
||
- **OpenMPI Environment**: `OMPI_MCA_orte_keep_fqdn_hostnames=1`
|
||
|
||
## Implementation Reference
|
||
|
||
For users who want to understand the implementation details or contribute to the operator, the default values described in this document are set in the following source files:
|
||
|
||
- **Health Probes, Security Context & Pod Specifications**: [`internal/dynamo/graph.go`](https://github.com/ai-dynamo/dynamo/blob/main/deploy/operator/internal/dynamo/graph.go) - Contains the main logic for applying default probes, security context, environment variables, shared memory, and pod configurations
|
||
- **Component-Specific Defaults**:
|
||
- [`internal/dynamo/component_common.go`](https://github.com/ai-dynamo/dynamo/blob/main/deploy/operator/internal/dynamo/component_common.go) - Base container and pod spec shared by all component types
|
||
- [`internal/dynamo/component_frontend.go`](https://github.com/ai-dynamo/dynamo/blob/main/deploy/operator/internal/dynamo/component_frontend.go)
|
||
- [`internal/dynamo/component_worker.go`](https://github.com/ai-dynamo/dynamo/blob/main/deploy/operator/internal/dynamo/component_worker.go)
|
||
- [`internal/dynamo/component_planner.go`](https://github.com/ai-dynamo/dynamo/blob/main/deploy/operator/internal/dynamo/component_planner.go)
|
||
- [`internal/dynamo/component_epp.go`](https://github.com/ai-dynamo/dynamo/blob/main/deploy/operator/internal/dynamo/component_epp.go)
|
||
- **Image Pull Secrets**: [`internal/secrets/docker.go`](https://github.com/ai-dynamo/dynamo/blob/main/deploy/operator/internal/secrets/docker.go) - Implements the docker secret indexer and automatic discovery
|
||
- **Backend-Specific Behavior**:
|
||
- [`internal/dynamo/backend_vllm.go`](https://github.com/ai-dynamo/dynamo/blob/main/deploy/operator/internal/dynamo/backend_vllm.go)
|
||
- [`internal/dynamo/backend_sglang.go`](https://github.com/ai-dynamo/dynamo/blob/main/deploy/operator/internal/dynamo/backend_sglang.go)
|
||
- [`internal/dynamo/backend_trtllm.go`](https://github.com/ai-dynamo/dynamo/blob/main/deploy/operator/internal/dynamo/backend_trtllm.go)
|
||
- **Checkpoint / Restore**: [`internal/checkpoint/dgd_integration.go`](https://github.com/ai-dynamo/dynamo/blob/main/deploy/operator/internal/checkpoint/dgd_integration.go) - Checkpoint env var injection and volume setup
|
||
- **Constants & Annotations**: [`internal/consts/consts.go`](https://github.com/ai-dynamo/dynamo/blob/main/deploy/operator/internal/consts/consts.go) - Defines annotation keys and other constants
|
||
|
||
## Notes
|
||
|
||
- All these defaults can be overridden by explicitly specifying values in your DynamoComponentDeployment or DynamoGraphDeployment resources
|
||
- User-specified probes (via `livenessProbe`, `readinessProbe`, or `startupProbe` fields) take precedence over operator defaults
|
||
- For security context, if you provide *any* `securityContext` in `extraPodSpec`, no defaults will be injected, giving you full control
|
||
- For multinode deployments, some defaults are modified or removed as described above to accommodate distributed execution patterns
|
||
- The `extraPodSpec.mainContainer` field can be used to override probe configurations set by the operator
|