Open-Source APXInf Hits 38Hz on Jetson Thor, Closing the Robot Deployment Latency Gap
The Rust-based engine cuts VLA inference from 278ms to 26ms on NVIDIA Jetson Thor using FP8.

On September 14, 2026, Infinigence-AI pushed a GitHub repository that answers one of the quieter engineering problems in embodied robotics: how to get a billion-parameter Vision-Language-Action model to run on the compute module inside a robot fast enough to actually control it. The project, APXInf, reduces inference latency for Physical Intelligence's π0.5 VLA model from roughly 278 milliseconds to 26 milliseconds on NVIDIA's Jetson Thor module using FP8 precision — a compression factor of more than ten — without retraining the model or measurably degrading task accuracy on the LIBERO-10 manipulation benchmark. At 38 cycles per second, the system crosses the threshold that robotics engineers generally consider necessary for fluid, continuous manipulation.
The release comes from a collaboration between Infinigence-AI, Tsinghua University, and Shanghai Jiao Tong University. Infinigence-AI, a Beijing-based AI inference startup, had previously developed Mizar, an edge inference engine with reported deployment in AI PCs, before applying its hardware expertise toward the robotics domain. APXInf is effectively the next step applied to a more demanding problem: a physical system that must perceive, decide, and act in a continuous loop, with no tolerance for the pauses that cloud latency introduces.
The Deployment Gap That Keeps Robots Tethered to the Cloud
For nearly two years, the gap between what VLA models can do in a research lab and what they can do inside a commercial robot has been one of the sector's defining constraints. A VLA model takes visual observations — typically camera images from one or more robot-mounted cameras — and a natural-language task instruction, and outputs motor commands directly. The architecture eliminates the separate perception, planning, and control modules required by traditional robotic systems, which makes VLA models remarkably flexible: the same model can, in principle, generalize to new objects, environments, and task descriptions without reprogramming.
The engineering problem is that running these models at inference time is expensive. π0.5, Physical Intelligence's current VLA model, operates at the multi-billion-parameter scale and uses a diffusion-style policy head that executes multiple forward passes through the backbone for each predicted action sequence. On desktop-class hardware with discrete GPUs and high-bandwidth GPU memory, inference is fast. On a Jetson module — the class of edge compute that fits inside a commercial robot's chassis — the constraints compound. Jetson Thor and Jetson Orin use LPDDR5 memory, which delivers roughly four to eight times less bandwidth than the HBM memory inside a discrete GPU, and both modules operate within power envelopes that are five to ten times tighter than datacenter cards. Under these constraints, an unoptimized π0.5 inference step takes approximately 278 milliseconds, which is equivalent to roughly 3.6 cycles per second — too slow for even coarse manipulation tasks, let alone the kind of dexterous, reactive motion that makes VLA demos compelling.
The stopgap has been cloud inference: the robot captures sensor data, transmits it to a cloud server running the model, receives the predicted action, and executes. The round-trip latency from network transmission alone typically adds 50 to 200 milliseconds, depending on the network and the data center's proximity. More fundamentally, a robot relying on cloud connectivity for its core decision loop becomes unavailable whenever the network does. The economics also scale badly: each robot in a fleet requires a dedicated inference server allocation to maintain the latency budget.
CUDA Graphs, FP8, and Rust: How APXInf Compresses 278 Milliseconds into 26
APXInf's speedup comes from a layered set of optimizations, each targeting a different source of inefficiency in standard inference pipelines.
The foundational technique is CUDA Graph capture. NVIDIA's CUDA API allows developers to record a sequence of GPU kernel launches as a graph, then replay that graph with a single CPU-side call at steady state, eliminating the per-kernel launch overhead that accumulates when a CPU dispatches GPU operations one by one. For standard language model inference, input shapes change constantly — different prompt lengths, different batch sizes — which makes CUDA Graph capture difficult to apply. VLA inference at the edge has a different profile: the robot operates with a fixed number of camera inputs at fixed resolution, a fixed batch size of one, and a fixed number of diffusion steps. These constraints make the computation graph static, which is precisely the condition CUDA Graph requires. APXInf captures the full inference graph on the first call and replays it on every subsequent control cycle, reducing CPU scheduling overhead to near zero.
FP8 quantization provides a second tier of speedup on Jetson Thor specifically. NVIDIA's Blackwell architecture, which powers the Jetson Thor module, includes native FP8 Tensor Cores that process eight-bit floating-point operations at roughly twice the throughput of sixteen-bit BF16 operations on the same hardware. APXInf applies FP8 precision to the π0.5 model's weights and activations, halving the model's effective memory footprint and doubling compute throughput without requiring model retraining. Jetson Orin, which uses NVIDIA's older Ampere architecture, does not include FP8 Tensor Cores; on Orin, APXInf uses INT8 quantization instead, achieving a different but still substantial speedup at lower precision.
A third layer — one-step action generation — addresses a structural inefficiency in diffusion-based policy heads. π0.5 normally generates an action sequence over multiple diffusion steps (the full-graph configuration uses ten), each requiring a complete forward pass through the transformer backbone. APXInf's one-step variant prunes the action horizon to a single step, reducing the number of backbone forward passes per control cycle from ten to one. The accuracy cost on LIBERO-10 is small enough to be practically negligible, as detailed below.
The runtime itself is written in Rust rather than Python or C++. Rust enforces memory safety through an ownership type system at compile time, eliminating the class of memory corruption bugs that are common in C and C++ systems, while avoiding the garbage-collection pauses that affect Python runtimes. For a real-time control loop where latency spikes can translate directly into physical errors, the predictability of Rust's memory management is an operational property rather than a stylistic preference. The codebase limits unsafe Rust to foreign function interface boundaries with CUDA libraries, keeping the main inference framework body under the ownership model.
APXInf builds CUDA kernels only for the target GPU's specific compute capability — SM87 for Jetson Orin and SM110 for Jetson Thor — avoiding the JIT recompilation overhead present in frameworks that must support multiple architectures at runtime. The repository vendors CUTLASS and FlashAttention directly, so there are no external framework dependencies beyond a standard Linux installation with the NVIDIA driver and CUDA toolkit. The build system does not require Docker.
Benchmark Results: Faster Inference, Same Task Accuracy
APXInf reports its performance on LIBERO-10, a simulated robotic manipulation benchmark widely used in the VLA research community. The benchmark covers ten task families, each evaluated across fifty rollouts with a fixed random seed, giving 500 total episodes. The tasks involve tabletop manipulation — arranging objects, opening containers, and stacking items — in a physics simulator.
The reference π0.5 model achieves 92.4% task success on LIBERO-10 under standard conditions. APXInf's Jetson Thor FP8 one-step configuration achieves 92.2%, a difference of 0.2 percentage points. Thor BF16 (full ten-step, CUDA Graph) achieves 92.8%. Orin BF16 achieves 92.0%. The quantization and action-pruning techniques that produce the latency reduction cost, at most, 0.4 percentage points of task success — within the range of measurement noise for a 500-episode benchmark.
The latency numbers from the GitHub README, measured as CUDA Graph steady-state P50 latency at batch size one with two 224×224 input images and ten flow steps, show: Jetson Thor BF16 full runs at 72.45ms (13.8Hz); Thor FP8 full at 41.16ms (24.3Hz); Orin BF16 full at 165.67ms (6.0Hz). Adding one-step pruning reduces Thor BF16 to 44.05ms (22.7Hz) and Thor FP8 to 26.32ms (38.0Hz) — the latter being the headline figure. For comparison, a desktop RTX 4090 in BF16 full-graph mode reaches 31.38ms (31.9Hz), and the INT8 one-step variant reaches 25.99ms (38.5Hz).
These results are company-reported and have not been independently reproduced as of this writing. LIBERO-10 is a simulated benchmark with fixed seeds; its results do not directly predict performance in uncontrolled real-world environments, where object diversity, lighting variation, and contact dynamics differ substantially from the simulator.
Why Cloud Inference Frameworks Cannot Fill This Gap
The two most widely deployed open-source LLM inference frameworks — vLLM and SGLang — are optimized for a fundamentally different operating condition. Both are designed to maximize throughput across hundreds or thousands of concurrent requests on datacenter hardware, using techniques like continuous batching, PagedAttention, and speculative decoding to amortize computational cost across many users simultaneously. The result is high tokens-per-second at scale. The cost is that single-sample batch-size-one latency, which is irrelevant in the datacenter context, is not a design target. Deploying vLLM or SGLang on a Jetson module to serve a single robot does not produce the latency profile APXInf achieves, because the architectural choices that make those frameworks efficient at scale work against the requirements of a single-stream real-time control loop.
NVIDIA's TensorRT-LLM is more relevant — it applies hardware-specific optimization to transformer inference on NVIDIA GPUs — but it is designed for language models and vision-language models rather than for VLA architectures with diffusion policy heads and action chunking outputs. Adapting TensorRT-LLM to π0.5's specific inference graph requires substantial engineering effort from developers who understand both the model architecture and the TensorRT optimization pipeline.
The closest public alternative is embodied.cpp, published in July 2026 by researchers from Southeast University, Nanjing University, Microsoft Research, and the Institute for AI Industry Research at Tsinghua University. Like APXInf, embodied.cpp is a C++ runtime targeting the specific latency requirements of embodied AI inference in closed-loop control. On π0.5, embodied.cpp achieves 56.85 milliseconds of amortized step latency with a 91% LIBERO success rate — a credible result that independently demonstrates the feasibility of sub-60ms edge VLA inference. APXInf, on the same model on Jetson Thor FP8, reaches 26.32ms with a 92.2% success rate, approximately 2.2 times lower latency at marginally higher accuracy.
The architectural trade-off between the two projects is explicit: embodied.cpp prioritizes portability across heterogeneous hardware — Jetson, Rockchip-based boards, and x86 edge boxes — while APXInf concentrates optimization on the Jetson platform specifically. For a robotics team building on Jetson Thor, APXInf currently offers better performance. For a team building on non-Jetson hardware, embodied.cpp is the more immediately applicable option.
The Agent-Driven Workflow Solving the Model Churn Problem
The inference speedup is the most immediately legible achievement in the APXInf release, but there is a second component that may prove more consequential over time: an agentic model-porting pipeline embedded in the repository as skills/model-port-workflow.
The VLA field in 2026 is moving faster than most inference teams can track. In the eight days between September 14 and September 22 alone, five new robotic AI systems were announced by three different companies, and Black Forest Labs released FLUX 3 Action, a new seven-billion-parameter VLA architecture with its own inference characteristics. Each new model architecture requires a new round of kernel engineering to achieve the kind of performance APXInf demonstrates for π0.5: identifying the model's computational graph, writing or adapting CUDA kernels for its specific operations, running Autotune to select optimal kernel configurations for the deployment GPU, and validating the result against a reference implementation. On a well-staffed team, this process takes weeks. On a robotics company without a dedicated inference engineering team, it may simply not happen.
Read more: Black Forest Labs FLUX 3 Action: Open 7B Robot Model Tops RoboLab-120 Benchmark
APXInf's agent-driven pipeline automates most of this process. The workflow uses AI agents to read a PyTorch reference model implementation, generate a ledger of the model's operators and their shapes, implement a static execution path in Rust, run per-operator verification against the PyTorch reference, execute the Autotune kernel selection procedure, and produce documentation. The human role in this process is to define the acceptance criteria — what accuracy and latency targets the ported model must hit — and to review the output. The agentic pipeline handles the kernel engineering steps.
This design reflects a realistic assessment of where the VLA field is heading. The inference speedup for π0.5 will eventually be replicated by other teams, built into commercial toolchains, or superseded by new hardware. The durable competitive advantage for a robotics team is the ability to adopt new model architectures as they emerge without rebuilding the inference stack from scratch. Infinigence-AI's prior work on Mizar for AI PCs gave the team direct experience with the problem of tracking rapidly iterating model families on fixed edge hardware; APXInf extends that experience into the robotics domain.
Where APXInf Still Falls Short
The tool is fifteen days old and its limitations are as significant as its achievements.
The LIBERO-10 benchmark, while standard in the VLA community, operates entirely in simulation with fixed random seeds. It does not test the distribution shifts that appear when a trained model encounters real objects with different surface textures, real lighting conditions with shadows and reflections, or real contact dynamics where fingers slip, objects tip, and surfaces are not perfectly flat. No published data exists yet showing APXInf-accelerated π0.5 inference performing manipulation tasks on a physical robot in an uncontrolled environment. The gap between simulated benchmark success and real-world task reliability is one of the persistent open problems in embodied AI; APXInf does not close it.
Hardware support is currently limited to NVIDIA Jetson Thor, Jetson Orin, and NVIDIA RTX-series desktop GPUs. Support for AMD GPUs and domestic Chinese AI accelerators (such as those from Cambricon and Biren) is listed on the project roadmap but has not shipped. Teams deploying on non-Jetson hardware will need to continue using alternative frameworks.
Model support beyond π0.5 is planned — WallOSS, Llama, and Qwen3-VL appear in the repository as roadmap items — but the depth of optimization for these models has not been established in the same way as for π0.5. The LIBERO-10 results pertain specifically to π0.5; accuracy and latency for other VLA architectures may differ substantially.
The agent-driven porting workflow is a novel capability that has not been independently validated in a production robotics environment. Its effectiveness at adapting new, architecturally diverse VLA models remains to be demonstrated at scale.
An Open-Source Proof of Concept for Real-Time Robot Control
The significance of APXInf as an open-source release is worth separating from its performance results. The performance numbers demonstrate that 10x+ inference speedup on Jetson hardware is achievable for π0.5-class VLA models using techniques that are, individually, not new — CUDA Graph has been available since CUDA 10, FP8 Tensor Cores are part of Blackwell's public specification, and Rust has been used for systems programming for years. What APXInf contributes is their integration into a coherent, open-source pipeline specifically engineered for the VLA edge inference problem, with the build system, API layers, and porting tooling included.
For a robotics engineer evaluating whether to attempt Jetson-based VLA deployment today, the alternative to APXInf is either reproducing the optimization work from scratch — a multi-week engineering effort — or accepting cloud inference with its latency and connectivity dependencies. The open-source release removes that choice: the optimized inference path is available to clone, build, and benchmark without a dedicated inference engineering team.
The question that the embodied AI sector will be watching over the next six to twelve months is whether APXInf's simulated gains transfer to physical robots performing complex tasks in production environments, and whether the agent-driven porting pipeline proves fast enough to track the VLA field as it continues to accelerate. If both hold, the infrastructure gap between impressive VLA demonstrations and deployed robotic products will have narrowed measurably.