pi0
Candidate3B params
Overview
Needs `huggingface-cli login` + accepting the PaliGemma license before first use. Cheapest per-step compute of all 7 archs, but least accurate — long chunk replay lets the scene drift before replanning.
Architecture
- Vision backbone
- SigLIP-So400m
- Language backbone
- Gemma-2B
- Action head
- flow-matching joint-attention expert
- Action steps per chunk
- 32
- Solver steps
- 10
- Camera views
- 2 — observation.images.image, observation.images.image2
Prerequisites
Gated tokenizer access
Requires accepting the license for google/paligemma-3b-pt-224 on Hugging Face before this model's tokenizer can be downloaded.
Hardware fit
Fits the 8GB Orin Nano target
split only — server+client+sim together exceed 8GB; benchmarked with sim offloaded to a second machine, so latency includes network hop
Benchmarks
| Hardware | Success rate | Step time | Inference time | Memory |
|---|---|---|---|---|
| RTX 3060 | 87.5% | 9.74 ms | 207.2 ms | 5548 MiB VRAM |
| Orin Nano 8GB (split) | — | 39.1 ms | — | 6068 MiB RSS |