Skip to main content

AI Deployment & Optimization

Choose the guide that matches your model and inference runtime. These are two deployment options, so you do not need to complete the TensorRT guide before using llama.cpp.

Choose a deployment path​

Your goalGuideFocus
Convert an ONNX model for NVIDIA inferenceTensorRT Model ConversionDriver and runtime preparation, engine conversion, and inference with trtexec.
Run an LLM locally on an edge deviceLLM Edge DeploymentBuilding llama.cpp, preparing a GGUF model, establishing a baseline, and tuning inference performance.

Before you start​

Check the target platform and software versions in the selected guide. TensorRT conversion applies to supported models beyond VLMs and LLMs; the LLM guide uses its own llama.cpp workflow.

For the ASR-D501 camera-inference demonstration using Qualcomm AI Hub, see Autonomous Systems.


← Solution Overview