# It can do it! Making VLA Models Fit an 8 GB Jetson Orin Nano - Savonia AMK

> Fitting VLA models onto 8 GB Jetson Orin Nano.

## The Challenge of Edge Robotics
Vision Language Action (VLA) models are highly effective for robotics, but they are typically demonstrated on high-end desktop GPUs. To be useful in a real-world application, these models must run on the computer carried by the robot—in this case, an 8 GB Jetson Orin Nano Super—alongside cameras, sensors, and control software. With a shared memory pool for CPU and GPU, there is very little room for waste.

## The Build Obstacle
Initial attempts to convert policies into a single TensorRT engine failed because the compiler required significant temporary working memory during the build process, exceeding the 8 GB limit. While PyTorch could run the models, the inference speed was insufficient for real-time robotics.

## The Solution: Splitting the Policy
The solution was to divide the policy at architectural boundaries and execute the pieces through ONNX Runtime (ORT) using its TensorRT execution provider.

*   **SmolVLA:** Divided into nine ONNX graphs.
*   **X-VLA:** Divided into twelve graphs.
*   **EVO-1:** Divided into eleven graphs (vision, language, and action stages, with the token-embedding table remaining on the CPU).

Engines are built one at a time in separate processes, allowing temporary memory to be released after each build. Completed engines are then cached on the Jetson for reuse.

## Optimization and Performance
Splitting the graphs introduced a new challenge: moving intermediate data between the CPU and GPU. This was mitigated by:
*   **I/O Binding:** Keeping cached attention data on the GPU to avoid repeated copying.
*   **GPU Offloading:** Moving small, frequently called action and time projections onto the GPU to reduce overhead.

### Benchmarks
The deployment path (splitting, FP16 precision, TensorRT optimization, and reduced data movement) resulted in significant improvements:

*   **Memory Usage:** SmolVLA peak process memory dropped from 3,826 MB to 2,089 MB.
*   **Energy Efficiency:** Energy per inference dropped from 18.14 to 3.45 joules for SmolVLA, and from 50.69 to 8.22 joules for X-VLA.
*   **Accuracy:** Numerical parity was maintained, with SmolVLA showing a 2.05% difference and X-VLA showing a 0.071% difference compared to FP32 baselines.

## Real-World Application
This approach was used to control a miniature hydraulic excavator. The `kaivuriprokkis` software connects camera observations and joint measurements to the policy, passing 30 Hz action chunks to the excavator controller. The split X-VLA model successfully commands the excavator at a steady 3 Hz.

## Resources
*   [Code, portable ONNX bundles, and measurement records](https://github.com/savonia/jetson-orin-nano-vla)
*   [kaivuriprokkis software](https://github.com/savonia/kaivuriprokkis)

## References
*   Lin, T. et al. 2025. *Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment*. arXiv:2511.04555.
*   Miettinen, E. 2026a. *jetson-orin-nano-vla*. Source code and retained benchmark results. Revision 341eddb.
*   Miettinen, E. 2026b. *kaivuriprokkis: LeRobot integration for the MASI excavator*. Source code and documentation.
*   NVIDIA. n.d. *Jetson Orin Nano Super Developer Kit*.
*   ONNX Runtime. n.d.-a. *TensorRT Execution Provider*.
*   ONNX Runtime. n.d.-b. *I/O Binding*.
*   Shukor, M. et al. 2025. *SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics*. arXiv:2506.01844.
*   Zheng, J. et al. 2025. *X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model*. arXiv:2510.10274.

***

*This work was carried out as part of the AI-MaSi project, supported by Pohjois-Savon liitto and the European Regional Development Fund (EAKR).*