Learning in simulation. Acting in reality.

Skill2Real

Agentic Skill Learning for Zero-Shot
Sim-to-Real Robot Manipulation

Learn reusable robot skills in simulation.
Transfer them to the real world.
Keep the skill memories frozen.

Simulation-learned skills.
Real-world execution.
Public observations + a shared robot APINo real-world task-policy fine-tuning

01 / The idea

Many worlds.
Skills that transfer.

Simulation is where the skills are learned. The real robot receives the frozen hierarchy, using the same API semantics across both worlds.

Zero-shot real-world transfer of simulation-learned skills. Skill2Real learns transferable, skill-based knowledge across diverse simulation environments. The resulting frozen skills are evaluated on real-world tasks without additional learning, demonstrating robustness to variations in lighting, materials, and dynamics.
Research overview

Skill2Real learns reusable robot skills through a shared robot API. A Proposer–Verifier–Governor loop turns simulation experience into validated skill knowledge. The Cerebellum first learns local manipulation; the Brain then learns task-level composition with the Cerebellum frozen. At deployment, the robot uses both frozen memories, public observations, and API returns—without real-world task-policy fine-tuning or skill-memory updates.

02 / See it in action

From learned skills
to physical action.

A closer look at the learning loop and recorded robot executions, across objects, backgrounds, and manipulation tasks.

01:34 · English narration · Captions availableDownload video

03 / Task recordings

One interface.
Many tasks.

Explore manipulation tasks in simulation and the real world, with videos and code.

Robot interface
LIBERO-90

Can into tray

Place the tomato sauce in the tray.

Ready to play
Simulation recordingVideo ↓
ObserveRead skillsWrite codeExecute
can_in_tray.py↓

Robot interfaceThe complete interface, as defined in the paper8 perception · 6 state & pose · 4 control

The Proposer combines these operations with frozen Cerebellum and Brain memories to perform each task.

Public perception 8

get_scene_image
Current front scene image.
get_topdown_image
Current top-down image for visual target selection.
get_wrist_view
Wrist image and available camera calibration.
get_topdown_stereo_view
Calibrated stereo RGB pair.
get_foundation_stereo_depth
Metric stereo depth bound to its source observation and calibration.
get_sam_mask_from_box
Segment a supplied box on a public image.
measure_visible_geometry
Measure visible 3D geometry from a mask, depth, and calibration.
get_topdown_object_candidates
Public source candidates with geometry and validity checks.

State & pose construction 6

preprocess
Refresh public perception outputs.
get_observation_geometry
Geometry grounded in a supplied public description.
get_previous_proposer_state
Prior public decision state and runtime carry geometry.
grasp_center_to_tool0_pose
Convert a chosen grasp center into a tool-frame pose.
compile_pose
Construct grasp or placement poses without moving the robot.
get_end_effector_pose
Current end-effector pose for an arm.

Robot control 4

gohome
Command an arm to its home configuration.
goto_pose
Command a position and XYZW orientation.
gripper_goto
Command jaw opening in meters.
goto_pose_both
Command both arms to their respective poses.
Shared Robot API and Execution ContractDownload reference ↓

04 / Learned skill library

Learned knowledge.
Two levels of memory.

The Proposer reads Brain and Cerebellum memories, binds them to the current scene, and writes the next program. Explore selected decisions from frozen training memories.

01Natural-language guidance

When a rule applies and which decision it changes.

02API-level template

Procedural context for the Proposer’s program generation.

03Compact state description

What to bind, carry forward and check in fresh observations.

Local knowledge for perception, grasping, placement and contact interaction.

65 decisionsDecision summaries · frozen memory

Skill scope

Selected contact and task decisions are summarized here. During deployment, current observations ground the program while learned memories remain fixed.

Collections & paper alignment

126 selected decision summaries from the LIBERO-90 Cerebellum C3, six task-family Brain snapshots, and the archived Robosuite R3 collection. These summaries highlight learned choices rather than reproduce complete execution policies.

The collections remain separate from the newer B4 results. Surface placement retains a partial third round; the Robosuite archive is a ten-task campaign, distinct from the paper's seven-task result series.

05 / How it works

Turn experience
into reusable skills.

A shared robot API connects simulation and reality. The learning loop validates what enters memory; the hierarchy separates local manipulation from task-level composition.

Proposer

Writes and executes programs from public observations, then proposes reusable skills from rollout evidence.

Verifier

Uses privileged simulation evidence to diagnose outcomes, translating it into feedback grounded in public observations.

Governor

Validates candidate updates and admits only the skill knowledge supported by validation rollouts.

Skill2Real pipeline. All in-context skill learning occurs in simulated scenes. Stage I learns reusable Cerebellum skills for local manipulation; Stage II freezes the Cerebellum and learns Brain skills that compositionally organize the local library into task programs. In both stages, the Proposer generates programs from public observations, the Verifier converts privileged execution evidence into publicly grounded feedback, and the Governor admits only updates supported by validation rollouts.
  1. Stage I

    Learn local skills

    The Cerebellum acquires reusable manipulation knowledge.

  2. Stage II

    Learn compositions

    The Brain learns task programs with the Cerebellum fixed.

  3. Deployment

    Freeze & transfer

    The Proposer uses both frozen memories through the real API.

What does “zero-shot” mean here?

The real robot executes with frozen skill memories, without real-world task demonstrations, task-policy fine-tuning, reinforcement learning, or online skill-memory updates. Robot-specific API implementation and calibration are still necessary. The simulator-only evaluator, Verifier, and Governor are removed at deployment.

06 / The evidence

Learn in one setting.
Evaluate beyond it.

Source-suite skill learning, held-out target tasks, and real-world transfer are evaluated separately. Skill memories stay frozen throughout evaluation.

78.75%

Real-world task completion

Four manipulation tasks · 20 trials per task
Astra executes frozen Sol-trained skills.

2.0to56.3%

LIBERO-Pro Long success

Sol trains on LIBERO-90; Astra evaluates frozen checkpoints on Pro Long.

89.4%

Robosuite mean success

Seven tasks with an independently
Opus-trained Robosuite skill library.

Frozen-skill evaluation across training iterations. Left: LIBERO-Pro Long evaluation of frozen LIBERO-90 checkpoints (20 tasks × 10 seeds). Right: seven-task Robosuite evaluation with independently trained libraries. C1–C3 denote Cerebellum learning; B1–B4 denote Brain learning with the Cerebellum fixed. Pro Long is never used for training or skill updates.

Real-world transfer

Both levels
contribute.

The full hierarchy improves mean task completion over either skill memory alone, with the same Astra Proposer and API.

Equal mean across pick-and-place, sorting, equation assembly, and drawer manipulation. Frozen Sol-trained C3/B3 libraries.

Evaluation details and per-task results
Real-world completion (%) · 20 trials per method and task
MethodPick & placeSortingEquationDrawer
CaP-Agent0 (CaP-X)1520100
Astra, no learned skills4035350
Astra, Brain only5060400
Astra, Cerebellum only8080650
Astra, full Skill2Real95858550

In the Pro Long checkpoint series, Overall is the equal mean of position and instruction perturbations. Astra reaches 56.25% at B4 (shown as 56.3%); Opus 5 reaches 49.0%. The shared evaluation protocol does not establish that the two models use the same libraries. Real-world results use the separately recorded fixed C3/B3 libraries, rather than the B4 checkpoint series.

Experience becomes knowledge.
Knowledge becomes action.

Watch Skill2Real

Figure