#sudo R1: Teaching Robots to Act, Starting from Simulation Alone
Introduce #sudo R1
Embodied AI has learned to think, and is beginning to act — but not yet reliably. Large language models plan multi-step tasks, parse complex instructions, and reason about the physical world. But manipulation remains fragile: no existing system can reliably grasp unfamiliar objects across the full diversity of real-world conditions. Until that changes, the vast economic promise of physical AI — in manufacturing, logistics, agriculture, and eldercare — stays locked behind a capability that no amount of high-level intelligence can substitute for.
We introduce #sudo R1, a fully integrated robot system with self-developed hardware and software, powered by a manipulation-centric foundation model focused on object picking — the gateway primitive of physical manipulation.
Our key results:
- 60-Minute Uncut Evaluation
Continuous execution across diverse objects. No cherry-picked clips. - Zero Real-World Data
Trained on simulation data alone. No real-world demonstrations required. - Strong Zero-Shot Generalizability
Unseen objects covering a wide spectrum of real-world variation: rigid and deformable, opaque and transparent, matte and reflective. - High Robustness
~98% first-attempt success and nearly 100% within two attempts. - True Closed-Loop Agility
Every action is conditioned on the robot's latest observation at 15-25 Hz. - Adaptive Spatial Intelligence
3D obstacle awareness and viable-space reasoning.
True production-grade performance remains ahead. Achieving any one of generalizability, agility, robustness, or spatial intelligence in isolation is already hard; achieving them simultaneously in a single policy is a fundamentally different challenge — and that is what #sudo R1 is built to pursue.
#sudo R1 evaluation — 60 minutes continuous uncut execution.
What #sudo R1 Delivers — and Why It’s Harder Than It Looks
Zero-Shot Generalization Across Diverse Objects with Near-Perfect Robustness
Real environments — warehouses, kitchens, factory floors — present an effectively open-ended distribution of objects that no model has seen before. #sudo R1 successfully picks diverse objects never encountered during training, spanning rigid and deformable, opaque and transparent, matte and reflective — including transparent glass, soft fabric, reflective metal, and irregularly shaped items.
Reliable deployment requires consistent performance regardless of visual conditions. We evaluated #sudo R1 under controlled lighting variations and used a TV screen behind the workspace to simulate a wide range of dynamic backgrounds. Across these conditions, pick success rates remained near-identical.
True Closed-Loop Agility
#sudo R1 has a fully closed-loop policy in which every control step is conditioned on the robot’s latest observation. Compared with many existing VLA models, this design enables more frequent feedback during execution. This enables tracking a target object as it moves, recovering from a perturbation mid-grasp, and adapting the movement trajectory when the scene changes during execution.
Spatial Intelligence
#sudo R1 adapts its trajectory when surrounding objects or structures constrain the feasible approach — navigating around obstacles, avoiding collisions, and exploiting available free space.
Why Simulation Is the Answer That Existing Systems Miss
The field has made significant progress on generalization, dexterity, robustness, and high-frequency control individually. Simulation removes the data constraint by scaling along all dimensions at once.
#sudo R1 is trained entirely on simulation data — no real-world demonstrations.
Training Across the Full Distribution
#sudo R1 demonstrates that simulation alone — with no real-world demonstrations — can produce manipulation policies that approach production-grade reliability across generalizability, agility, robustness, and spatial intelligence simultaneously.