AI Deployment & Optimization
Choose the guide that matches your model and inference runtime. These are two deployment options, so you do not need to complete the TensorRT guide before using llama.cpp.
Choose a deployment path
| Your goal | Guide | Focus |
|---|---|---|
| Convert an ONNX model for NVIDIA inference | TensorRT Model Conversion | Driver and runtime preparation, engine conversion, and inference with trtexec. |
| Run an LLM locally on an edge device | LLM Edge Deployment | Building llama.cpp, preparing a GGUF model, establishing a baseline, and tuning inference performance. |
Before you start
Check the target platform and software versions in the selected guide. TensorRT conversion applies to supported models beyond VLMs and LLMs; the LLM guide uses its own llama.cpp workflow.
For the ASR-D501 camera-inference demonstration using Qualcomm AI Hub, see Autonomous Systems.