Lähikuva valkoisesta kaivinkoneen varresta, jossa on hydrauliikkaputkia ja SAVONIA-teksti. Tausta on sumea, mikä nostaa koneen keskiöön.

Savonia Article Pro: It can do it! Making VLA Models Fit an 8 GB Jetson Orin Nano

Savonia Article Pro is a collection of multidisciplinary Savonia expertise on various topics.

This work is licensed under CC BY-SA 4.0Creative Commons logoCreative Commons Attribution logoCreative Commons Share Alike logo

I have become really interested in Vision Language Action (VLA) models. They are one of the most exciting ways to get robots doing useful things, but most of them are shown off on big desktop GPUs. I wanted to find out whether they could actually run on a small, affordable board that fits on a real machine.

A robot needs more than a model that runs on a workstation. It needs one that runs on the computer it carries, alongside cameras, sensors and control software. In this article that computer is an 8 GB Jetson Orin Nano Super.

Jetson Orin Nano Super Developer Kit. Image: NVIDIA.

Eight gigabytes fills up quickly

The Jetson’s CPU and GPU share the same 8 GB memory pool. Model weights are only part of that budget: intermediate calculations, runtime libraries, the operating system and the robot software also need space. The board is small enough for our miniature excavator, but there is little room for waste.

The first obstacle appeared before inference. I tried converting a whole policy into one TensorRT engine, an optimized fast executable for NVIDIA GPUs. Building it exceeded the available memory. The compiler needed temporary working memory in addition to the model weights, so a model that looked manageable on disk could still fail during conversion.

All the models could run through PyTorch. The problem was getting the faster TensorRT route to build on this board, as with PyTorch the inference is just too slow for anything.

Build it in pieces

The solution was to divide the policy at its architectural boundaries and execute the pieces through ONNX Runtime (ORT), using its TensorRT execution provider for GPU acceleration. This is the `ort-split` path in the project repository. ORT manages execution, while TensorRT optimizes the supported computations.

SmolVLA becomes nine ONNX graphs covering vision, language and action generation. X-VLA needs twelve, with its larger components divided further. EVO-1 lands in between with eleven: vision, language and action stages, plus its large token-embedding table, which simply stays on the CPU. Engines are built one at a time in separate processes. When a build process exits, its temporary memory is released before the next begins. The completed engines are cached on the Jetson for reuse. Takes time, but works. And you only need to do it once!

Vertailu monoliittisten ja jaettujen ONNX-moottorirakenteiden välillä: monoliittiset rakenteet kuluttavat muistin loppuun, kun taas jakaminen pienempiin graafeihin mahtuu 8 GB:n muistirajaan, jolloin kukin moottori voidaan rakentaa erikseen.
Monolithic engine build needs the compiler’s workspace on top of everything else. Splitting the policy lets each piece build within the 8 GB budget. Image: just released Claude Opus 5.5.

Splitting solved the build problem, though only barely: X-VLA still leaves under 500 MB free, even with 4 GB of swap and a headless OS. Graph boundaries also introduced another cost: moving intermediate data between CPU and GPU. In SmolVLA, the same cached attention data is reused throughout ten action-refinement steps. Keeping that cache on the GPU with ORT’s I/O Binding avoids repeatedly copying it back and forth. Moving the small, repeatedly called action and time projections onto the GPU also reduced overhead.

What did we gain?

The retained benchmarks compare each public base checkpoint in PyTorch FP32 with its split FP16 deployment. They use the same Jetson in MAXN_SUPER mode with fixed clocks and repeatable synthetic observations. SmolVLA uses two camera views and a 50-action chunk; X-VLA uses three views and a 30-action chunk. Both retain ten refinement steps. These are the default settings the models ship with.

Taulukossa verrataan SmoVLA-, X-VLA- ja EVO-1-mallien päättelyaikoja Jetson Orin Nano Super -laitteella; SmoVLA:n nopeuskerroin on suurin, 6,15-kertainen, ja taulukossa on esitetty päättelyajat PyTorch FP32:lle ja ORT split FP16:lle.
Model benchmark table. Image: Opus 5.5.

*As of writing this, no public base checkpoint or PyTorch baseline is available for EVO-1, so it was measured with the LIBERO fine-tuned checkpoint: zuoxingdong/evo1_libero.*

Full measurements can be found here: measurements

The speedup comes from the complete deployment: splitting, FP16 reduced precision, TensorRT optimization and less data movement. Splitting alone is not a sixfold acceleration switch, but it is what makes the rest possible on this board.

SmolVLA’s peak process memory fell from 3,826 MB to 2,089 MB, a reduction of about 45%. X-VLA’s peak remained close to 5,100 MB in both runtimes. For X-VLA, the main gains were making the accelerated build possible and completing inference much sooner. EVO-1 peaked at about 5,300 MB.

Energy per inference also fell: from 18.14 to 3.45 joules for SmolVLA and from 50.69 to 8.22 joules for X-VLA, while EVO-1 used 8.40 joules. These are whole-board measurements excluding any robots, screens etc. connected. SmolVLA actually drew more power while running, but finished quickly enough to use about 81% less energy per inference.

These are comparisons between deployment paths, not rankings of model intelligence. They exclude live camera capture and physical task performance. They also compare FP32 with FP16, so numerical differences must be checked (i.e. parity). SmolVLA’s largest full-chunk difference was 2.05% of the reference action range, while X-VLA’s was just 0.071%! EVO-1’s actions matched the reference outputs shipped with its bundle to a cosine similarity of 0.99999. Speed measurements alone do not establish equivalent robot behaviour, but in my tests the faster inference easily outweighed the tiny loss in accuracy.

Using the models

I used this split-runtime approach to drive a miniature hydraulic excavator. The models were fine-tuned on demonstrations recorded from the physical machine. The kaivuriprokkis software connects camera observations and joint measurements to the policy and passes its 30 Hz action chunks to the excavator controller, which runs separately at 100 Hz.

With a single camera, the split X-VLA manages to command the excavator with a steady 3 Hz stream of action chunks!

split X-VLA controlling the miniature excavator on a Jetson Orin Nano Super. Video: Eetu Miettinen.

A great tool for edge robotics

This work made the accelerated runtime buildable on the available hardware and brought inference latency down substantially. For other robotics developers, the practical lesson is to budget for the engine build as well as execution, and to examine data movement wherever a model is split. The code, portable ONNX bundles and measurement records provide a great starting point for applying this approach to other small or budget-minded robotic platforms.

Orin Nanos can probably be found at almost any university that dabbles in AI or computer vision, and an older Orin Nano can be flashed to Orin Nano Super mode without any hardware mods 🙂

Have fun!

This work was carried out as part of the AI-MaSi project, supported by Pohjois-Savon liitto and the European Regional Development Fund (EAKR).


Authors

Eetu Miettinen, Research Engineer at AI-MaSi, Savonia University of Applied Sciences.

This article was prepared with the help of AI.


References

Miettinen, E. 2026a. jetson-orin-nano-vla. Source code and retained benchmark results. Revision 341eddb. Accessed 11 September 2026.

Miettinen, E. 2026b. kaivuriprokkis: LeRobot integration for the MASI excavator. Source code and documentation. Accessed 11 September 2026.

Lin, T. et al. 2025. Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment. arXiv:2511.04555.

NVIDIA. n.d. Jetson Orin Nano Super Developer Kit. Accessed 11 September 2026.

ONNX Runtime. n.d.-a. TensorRT Execution Provider. Accessed 11 September 2026.

ONNX Runtime. n.d.-b. I/O Binding. Accessed 11 September 2026.

Shukor, M. et al. 2025. SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics. arXiv:2506.01844.

Zheng, J. et al. 2025. X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model. arXiv:2510.10274.