Example:
```
dynamo-run out=<engine> <model> --kv-cache-block-size 64
```
In a distributed system this goes on the worker node and is propagated to ingress via the model deployment card.
Previously hard coded to 16, which is now the default.
- Load context_length from model. Closes#1172
- Store context length and KV cache block size in Model Deployment Card #1170
. New mistralrs and llamacpp version
. mistralrs: Handle Gemma 3 and Llama 4 as vision models
. Update the dynamo-run docs to use Qwen 3
. Our pre-processor now supports Llama 4's newer multi-modal `config.json`
. Upgrade minijinja to handle Qwen 3's prompt template
For Llama 4 we'll need to limit the max seq len. vllm says:
> To serve at least one request with the models's max seq len (10485760), (240.00 GiB KV cache is needed,...
I was able to run Llama 4 with llamacpp and a quantized GGUF, with Dynamo doing the pre-processing.