Configure GPU Acceleration
The MoveIt Pro CLI detects supported NVIDIA and AMD GPUs and Qualcomm NPUs automatically after the required host driver and device access are configured. NVIDIA also requires its Container Toolkit. This guide covers those host dependencies and verification steps.
MoveIt Pro has two different integrations of GPU support:
- GPU Acceleration, where MoveIt Pro will utilize GPU resources when rendering simulators and cameras.
- GPU Inference, where MoveIt Pro will utilize GPU resources for Machine Learning models.
GPU Acceleration can be used without GPU Inference, but GPU Inference requires GPU Acceleration to be enabled.
The NVIDIA workflow below enables GPU acceleration for simulation and camera rendering and GPU inference for ML models. Supported AMD GPUs accelerate simulation, camera rendering, and VLA inference through ROCm. The SAM3 examples currently run inference on the CPU on AMD.
Please read and follow ALL the steps below carefully.
These generic NVIDIA instructions do not apply to Jetson. Starting with MoveIt Pro 10.0, the qualified ROS Jazzy Jetson configuration uses MOVEIT_TARGET=-jetson-cuda13.2-cudnn9 with JetPack 7.2.x / L4T r39.2.x. Newer valid L4T releases may run with a compatibility warning. The image is available after the release and Jetson hardware qualification complete. See Configure NVIDIA Jetson.
NVIDIA GPUs
NVIDIA Drivers
Please ensure that you have NVIDIA drivers installed for your system:
nvidia-smi
You should get something similar to the following output:
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 570.195.03 Driver Version: 570.195.03 CUDA Version: 12.8 |
|-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA GeForce RTX 4060 ... Off | 00000000:01:00.0 On | N/A |
| N/A 40C P8 3W / 115W | 806MiB / 8188MiB | 30% Default |
| | | N/A |
+-----------------------------------------+------------------------+----------------------+
+-----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| No running processes found |
+-----------------------------------------------------------------------------------------+
If the command isn't found then you likely do not have the NVIDIA driver installed. To install the NVIDIA driver on Ubuntu, see the instructions below:
sudo add-apt-repository multiverse
sudo apt update
sudo apt install ubuntu-drivers-common ubuntu-restricted-extras
sudo apt update
sudo apt install nvidia-driver-570
sudo reboot
For Debian or other Debian or Ubuntu derivatives, please follow this guide and install nvidia-driver-570.
NVIDIA GPU inference requires NVIDIA driver versions >= 560. See the CUDA Toolkit Compatibility Matrix for more details. "Proprietary" versions of the driver are preferred (as opposed to "Open" or "Open Kernel" versions) but some hardware may require the "Open" version.
For real-time NVIDIA driver support, please follow this guide
NVIDIA Container Toolkit
Next, install the nvidia-container-toolkit following this guide.
Please ensure that you follow the steps under Configuring Docker to configure and restart the Docker daemon.
To verify that the toolkit is installed properly, please run the following sample container:
docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
And observe that the output matches the above nvidia-smi command.
GPU Acceleration and GPU Inference
If both the NVIDIA driver and nvidia-container-toolkit are installed, moveit_pro commands automatically detect the GPU, select the appropriate image, and make the device available for acceleration and inference.
When using the MoveIt Pro CLI, this is done by automatically including an additional Docker Compose file installed on your system at /opt/moveit_pro/nvidia-compose.yaml.
If you are running raw docker compose commands instead of the moveit_pro CLI, you must explicitly include /opt/moveit_pro/nvidia-compose.yaml yourself, as described in not using the MoveIt Pro CLI.
Finding cuda and cudnn in the base image verifies only that the libraries are installed. Running nvidia-smi inside moveit_pro shell, or an equivalent device check, verifies that the Runtime can see the NVIDIA driver and GPU. Neither result proves that simulation or an ML model executed on the GPU.
In MoveIt Pro 10.2 or later, open Debugging → GPU Status to verify observed acceleration. Check the simulation renderer and ML provider details there; a configured GPU session can still fall back to CPU execution.
To test GPU inference functionality, you can utilize the core Behavior "GetMasks2DFromPointQuery" in an Objective.
The segmentation with inference should be significantly faster than without and you should notice higher resource usage in nvidia-smi.
The GPU Status pane reports whether inference ran through an accelerator provider or fell back to CPU. ONNX GPU acceleration remains unknown until a successful inference produces accelerator-provider evidence. CPU fallback detected during model loading can appear earlier. Run an ML Objective, then refresh GPU Status.
AMD GPUs
AMD Drivers and Device Access
Confirm that the host meets AMD's ROCm system requirements, including a supported operating system and GPU driver. ROCm userspace runs in the container; it does not need to be installed on the host.
The driver must expose /dev/kfd and /dev/dri/renderD*. The MoveIt Pro CLI detects AMD GPUs even when inxi is unavailable, selects the ROCm image, and passes the device nodes and their group IDs to the inference server automatically. When the CLI selects the image automatically, missing compute devices stop startup with driver guidance, before any vla_serving.yaml setting applies; an explicit MOVEIT_TARGET skips this check. If the ROCm 7.2.2 images do not support the GPU's architecture, as with some integrated Radeon GPUs, the CLI prints a warning that names the architecture and uses the CPU images; set MOVEIT_TARGET=-rocm7.2.2 to try ROCm anyway. The AMD Container Toolkit is optional.
To run on CPU on an AMD host, set an empty MOVEIT_TARGET for one command (MOVEIT_TARGET= moveit_pro run), or add MOVEIT_TARGET: "" to ~/.config/moveit_pro/moveit_pro_config.9.yaml to keep that choice. Either one selects the CPU images for the Runtime and the inference server. Remove the line to return to automatic selection.
Check the devices before launching:
ls -l /dev/kfd /dev/dri/renderD*
On a supported amd64 host, use a current example workspace and launch VLA simulation with its inference server:
moveit_pro run -c vla_sim --with-inference-server
The server uses ROCm for the policy and refuses silent CPU fallback if the GPU is unavailable. PyTorch calls its device cuda on both AMD and NVIDIA; the server startup log also reports the Torch, HIP, and CUDA versions. Once the ROCm image is running, an explicit device: cpu in vla_serving.yaml selects CPU inference. See Connect a VLA Policy for checkpoint configuration and credentials.
The ARM64 AMD Runtime image provides Mesa display acceleration and CPU inference; the ROCm VLA serving image described here targets amd64.
Qualcomm NPUs
On ARM64 hosts with a Qualcomm Snapdragon SoC, the MoveIt Pro CLI selects the -qnn image. It runs ONNX models on the Hexagon NPU through ONNX Runtime's QNN execution provider. The suffix carries no version because the image ships the QNN libraries itself. The host provides the FastRPC driver and DSP firmware files that reach the DSP. Simulation, camera rendering, and LibTorch models such as VLA policies run on the CPU.
Qualcomm FastRPC Access
Install Qualcomm's qcom-fastrpc1 package from the ppa:ubuntu-qcom-iot/qcom-ppa PPA, along with the DSP firmware package for your board. Together they provide the FastRPC driver libcdsprpc.so.1 and its libdmabufheap.so.0 dependency, the /dev/fastrpc-cdsp and /dev/dma_heap/system device nodes, and the compute DSP's process shell, which the cdsprpcd service links into /usr/lib/dsp/cdsp. Check them before launching:
ls -l /dev/fastrpc-cdsp* /dev/dma_heap/system \
/usr/lib/aarch64-linux-gnu/libcdsprpc.so.1 /usr/lib/aarch64-linux-gnu/libdmabufheap.so.0 \
/usr/lib/dsp/cdsp/fastrpc_shell_unsigned_3
The CLI detects the SoC from /sys/devices/soc0/family. When it finds a Snapdragon SoC with any of these missing, startup stops and names what is missing. An explicit MOVEIT_TARGET skips the check. The CLI includes /opt/moveit_pro/qnn-compose.yaml, which mounts both libraries, /usr/lib/dsp, and /usr/share/qcom (the target of the /usr/lib/dsp links) into the Runtime read-only. These files must match the host kernel and DSP firmware, so the image does not ship them. If libcdsprpc.so.1 lives elsewhere on your host, export MOVEIT_QNN_CDSPRPC_LIBRARY with its path in the shell that runs moveit_pro, so the CLI checks the same file it mounts. The Runtime loads this library, so point it only at the root-owned file Qualcomm's package installs.
A container does not inherit the host's device ACLs, so the container user joins fastrpc and dmaheap groups aligned to the groups that own /dev/fastrpc-cdsp and /dev/dma_heap/system. If /dev/dma_heap/system belongs to the root group and grants access only through an ACL, the CLI reports it. Give the node a dedicated group with a udev rule.
To run on CPU, set an empty MOVEIT_TARGET as described for AMD above.
Models for the NPU
The Hexagon NPU runs models compiled for it ahead of time. An ordinary ONNX model leaves most of its operators on the CPU even when the NPU is available, and some models that run on the CPU fail on the NPU. MoveIt Pro therefore sends a model to the NPU automatically only when it is precompiled for QNN, meaning its graph holds EPContext nodes. Every other model runs on the CPU, exactly as on the ARM64 CPU image. This includes the SAM2 and SAM3 models the segmentation Behaviors ship with.
Qualcomm AI Hub exports precompiled models for a specific chipset. Install the qai-hub-models package with Python 3.10 to 3.12, configure your AI Hub API token with qai-hub configure --api_token <token>, and export with the precompiled_qnn_onnx target runtime:
qai-hub-models export <model> --target-runtime precompiled_qnn_onnx --chipset <chipset>
qai-hub list-devices lists the chipset names; for the Arduino VENTUNO Q it is qualcomm-qcs8275. The export writes each graph as a small .onnx file next to a _qairt_context.bin file; keep them together. The context binary is compiled code that runs on the NPU, so load precompiled models only from sources you trust. An AI Hub export splits a model into its own set of graphs, so check that the graphs match what the loading code expects before pointing a Behavior at them.
Open Debugging → GPU Status after running an ML Objective. An NPU session reports QNNExecutionProvider as its active provider. Operators the QNN HTP backend does not support run on CPUExecutionProvider in the same session.