156 lines
4.5 KiB
Markdown
156 lines
4.5 KiB
Markdown
---
|
|
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
|
# SPDX-License-Identifier: Apache-2.0
|
|
title: Quickstart
|
|
---
|
|
|
|
This guide covers running Dynamo **using the CLI on your local machine or VM**.
|
|
|
|
<Info>
|
|
**Looking to deploy on Kubernetes instead?**
|
|
See the [Kubernetes Installation Guide](../kubernetes/installation-guide.md)
|
|
and [Kubernetes Quickstart](../kubernetes/README.md) for cluster deployments.
|
|
</Info>
|
|
|
|
## Install Dynamo
|
|
|
|
**Option A: Containers (Recommended)**
|
|
|
|
Containers have all dependencies pre-installed. No setup required.
|
|
|
|
```bash
|
|
# SGLang
|
|
docker run --gpus all --network host --rm -it nvcr.io/nvidia/ai-dynamo/sglang-runtime:0.8.1
|
|
|
|
# TensorRT-LLM
|
|
docker run --gpus all --network host --rm -it nvcr.io/nvidia/ai-dynamo/tensorrtllm-runtime:0.8.1
|
|
|
|
# vLLM
|
|
docker run --gpus all --network host --rm -it nvcr.io/nvidia/ai-dynamo/vllm-runtime:0.8.1
|
|
```
|
|
|
|
<Tip>
|
|
To run frontend and worker in the same container, either:
|
|
|
|
- Run processes in background with `&` (see Run Dynamo section below), or
|
|
- Open a second terminal and use `docker exec -it <container_id> bash`
|
|
</Tip>
|
|
|
|
See [Release Artifacts](../reference/release-artifacts.md#container-images) for available
|
|
versions and backend guides for run instructions: [SGLang](../backends/sglang/README.md) |
|
|
[TensorRT-LLM](../backends/trtllm/README.md) | [vLLM](../backends/vllm/README.md)
|
|
|
|
**Option B: Install from PyPI**
|
|
|
|
```bash
|
|
# Install uv (recommended Python package manager)
|
|
curl -LsSf https://astral.sh/uv/install.sh | sh
|
|
|
|
# Create virtual environment
|
|
uv venv venv
|
|
source venv/bin/activate
|
|
uv pip install pip
|
|
```
|
|
|
|
Install system dependencies and the Dynamo wheel for your chosen backend:
|
|
|
|
**SGLang**
|
|
|
|
```bash
|
|
sudo apt install python3-dev
|
|
uv pip install --prerelease=allow "ai-dynamo[sglang]"
|
|
```
|
|
|
|
<Note>
|
|
For CUDA 13 (B300/GB300), the container is recommended. See
|
|
[SGLang install docs](https://docs.sglang.io/get_started/install.html) for details.
|
|
</Note>
|
|
|
|
**TensorRT-LLM**
|
|
|
|
```bash
|
|
sudo apt install python3-dev
|
|
pip install torch==2.9.0 torchvision --index-url https://download.pytorch.org/whl/cu130
|
|
pip install --pre --extra-index-url https://pypi.nvidia.com "ai-dynamo[trtllm]"
|
|
```
|
|
|
|
<Note>
|
|
TensorRT-LLM requires `pip` due to a transitive Git URL dependency that
|
|
`uv` doesn't resolve. We recommend using the TensorRT-LLM container for
|
|
broader compatibility. See the [TRT-LLM backend guide](../backends/trtllm/README.md)
|
|
for details.
|
|
</Note>
|
|
|
|
**vLLM**
|
|
|
|
```bash
|
|
sudo apt install python3-dev libxcb1
|
|
uv pip install --prerelease=allow "ai-dynamo[vllm]"
|
|
```
|
|
|
|
## Run Dynamo
|
|
|
|
<Tip>
|
|
**(Optional)** Before running Dynamo, verify your system configuration:
|
|
`python3 deploy/sanity_check.py`
|
|
</Tip>
|
|
|
|
Start the frontend, then start a worker for your chosen backend.
|
|
|
|
<Tip>
|
|
To run in a single terminal (useful in containers), append `> logfile.log 2>&1 &`
|
|
to run processes in background. Example: `python3 -m dynamo.frontend --discovery-backend file > dynamo.frontend.log 2>&1 &`
|
|
</Tip>
|
|
|
|
```bash
|
|
# Start the OpenAI compatible frontend (default port is 8000)
|
|
# --discovery-backend file avoids needing etcd (frontend and workers must share a disk)
|
|
python3 -m dynamo.frontend --discovery-backend file
|
|
```
|
|
|
|
In another terminal (or same terminal if using background mode), start a worker:
|
|
|
|
**SGLang**
|
|
|
|
```bash
|
|
python3 -m dynamo.sglang --model-path Qwen/Qwen3-0.6B --discovery-backend file
|
|
```
|
|
|
|
**TensorRT-LLM**
|
|
|
|
```bash
|
|
python3 -m dynamo.trtllm --model-path Qwen/Qwen3-0.6B --discovery-backend file
|
|
```
|
|
|
|
**vLLM**
|
|
|
|
```bash
|
|
python3 -m dynamo.vllm --model Qwen/Qwen3-0.6B --discovery-backend file \
|
|
--kv-events-config '{"enable_kv_cache_events": false}'
|
|
```
|
|
|
|
<Note>
|
|
For dependency-free local development, disable KV event publishing (avoids NATS):
|
|
|
|
- **vLLM:** Add `--kv-events-config '{"enable_kv_cache_events": false}'`
|
|
- **SGLang:** No flag needed (KV events disabled by default)
|
|
- **TensorRT-LLM:** No flag needed (KV events disabled by default)
|
|
|
|
**TensorRT-LLM only:** The warning `Cannot connect to ModelExpress server/transport error. Using direct download.`
|
|
is expected and can be safely ignored.
|
|
</Note>
|
|
|
|
<Note>
|
|
**Deprecation notice:** vLLM automatically enables KV event publishing when prefix caching is active. In a future release, this will change — KV events will be disabled by default for all backends. Start using `--kv-events-config` explicitly to prepare.
|
|
</Note>
|
|
|
|
## Test Your Deployment
|
|
|
|
```bash
|
|
curl localhost:8000/v1/chat/completions \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"model": "Qwen/Qwen3-0.6B",
|
|
"messages": [{"role": "user", "content": "Hello!"}],
|
|
"max_tokens": 50}'
|
|
```
|