Skill scope
Selected contact and task decisions are summarized here. During deployment, current observations ground the program while learned memories remain fixed.
Agentic Skill Learning for Zero-Shot
Sim-to-Real Robot Manipulation
Learn reusable robot skills in simulation.
Transfer them to the real world.
Keep the skill memories frozen.
01 / The idea
Simulation is where the skills are learned. The real robot receives the frozen hierarchy, using the same API semantics across both worlds.
Skill2Real learns reusable robot skills through a shared robot API. A Proposer–Verifier–Governor loop turns simulation experience into validated skill knowledge. The Cerebellum first learns local manipulation; the Brain then learns task-level composition with the Cerebellum frozen. At deployment, the robot uses both frozen memories, public observations, and API returns—without real-world task-policy fine-tuning or skill-memory updates.
02 / See it in action
A closer look at the learning loop and recorded robot executions, across objects, backgrounds, and manipulation tasks.
03 / Task recordings
Explore manipulation tasks in simulation and the real world, with videos and code.
58 recordings across simulation and the real world. Choose a suite to explore.
Place the tomato sauce in the tray.
The Proposer combines these operations with frozen Cerebellum and Brain memories to perform each task.
get_scene_imageget_topdown_imageget_wrist_viewget_topdown_stereo_viewget_foundation_stereo_depthget_sam_mask_from_boxmeasure_visible_geometryget_topdown_object_candidatespreprocessget_observation_geometryget_previous_proposer_stategrasp_center_to_tool0_posecompile_poseget_end_effector_posegohomegoto_posegripper_gotogoto_pose_both04 / Learned skill library
The Proposer reads Brain and Cerebellum memories, binds them to the current scene, and writes the next program. Explore selected decisions from frozen training memories.
When a rule applies and which decision it changes.
Procedural context for the Proposer’s program generation.
What to bind, carry forward and check in fresh observations.
Local knowledge for perception, grasping, placement and contact interaction.
Selected contact and task decisions are summarized here. During deployment, current observations ground the program while learned memories remain fixed.
126 selected decision summaries from the LIBERO-90 Cerebellum C3, six task-family Brain snapshots, and the archived Robosuite R3 collection. These summaries highlight learned choices rather than reproduce complete execution policies.
The collections remain separate from the newer B4 results. Surface placement retains a partial third round; the Robosuite archive is a ten-task campaign, distinct from the paper's seven-task result series.
05 / How it works
A shared robot API connects simulation and reality. The learning loop validates what enters memory; the hierarchy separates local manipulation from task-level composition.

Writes and executes programs from public observations, then proposes reusable skills from rollout evidence.

Uses privileged simulation evidence to diagnose outcomes, translating it into feedback grounded in public observations.

Validates candidate updates and admits only the skill knowledge supported by validation rollouts.
The Cerebellum acquires reusable manipulation knowledge.
The Brain learns task programs with the Cerebellum fixed.
The Proposer uses both frozen memories through the real API.
The real robot executes with frozen skill memories, without real-world task demonstrations, task-policy fine-tuning, reinforcement learning, or online skill-memory updates. Robot-specific API implementation and calibration are still necessary. The simulator-only evaluator, Verifier, and Governor are removed at deployment.
06 / The evidence
Source-suite skill learning, held-out target tasks, and real-world transfer are evaluated separately. Skill memories stay frozen throughout evaluation.
Four manipulation tasks · 20 trials per task
Astra executes frozen Sol-trained skills.
Sol trains on LIBERO-90; Astra evaluates frozen checkpoints on Pro Long.
Seven tasks with an independently
Opus-trained Robosuite skill library.
Real-world transfer
The full hierarchy improves mean task completion over either skill memory alone, with the same Astra Proposer and API.
Equal mean across pick-and-place, sorting, equation assembly, and drawer manipulation. Frozen Sol-trained C3/B3 libraries.
| Method | Pick & place | Sorting | Equation | Drawer |
|---|---|---|---|---|
| CaP-Agent0 (CaP-X) | 15 | 20 | 10 | 0 |
| Astra, no learned skills | 40 | 35 | 35 | 0 |
| Astra, Brain only | 50 | 60 | 40 | 0 |
| Astra, Cerebellum only | 80 | 80 | 65 | 0 |
| Astra, full Skill2Real | 95 | 85 | 85 | 50 |
In the Pro Long checkpoint series, Overall is the equal mean of position and instruction perturbations. Astra reaches 56.25% at B4 (shown as 56.3%); Opus 5 reaches 49.0%. The shared evaluation protocol does not establish that the two models use the same libraries. Real-world results use the separately recorded fixed C3/B3 libraries, rather than the B4 checkpoint series.