>90 %
real-world success, across bar geometries
0
real-world training samples
1.4 mm
slot clearance
A single frozen policy, deployed zero-shot on hardware.
The stacked views on the right are the policy's inputs; the main view is for visualization only.
Four insertions on different bars and start poses, then recovery from perturbations and a cluttered scene.
One uncut take. 49 attempts in a row, 45 of them succeed — the counter runs in the corner. Six bars of very different geometry are swapped in, held in different in-hand poses (highlighted in EP 15–23), and placed by hand at a random spot near the slot before every attempt. The 3D models of these bars are in Example bars from real-world tests.
Policies recover from perturbations and retry after collision.
Policies succeed under dramatic background changes.
Every modern concrete structure depends on it. Construction workers lift heavy steel bars, align them with the rack slots, and force them down — thousands of times a day, at a steady cost in chronic injury. And no two bars are alike: appearance varies with rust and mill finish, nominal size with the design, and even one design drifts within its production tolerances — bend angle, cut length, out-of-plane twist, hook shape.
Train an end-to-end visuomotor policy on raw RGB, in simulation at scale — across diverse variations, with the right training recipe — and it inserts real bars robustly. No real-world data, no perception stack.
Three phases, one controller. A privileged teacher learns the task with RL on procedurally generated geometry, a multi-view visual student is distilled from it under heavy randomization, and the student deploys zero-shot with frozen weights.
Phase 1
State-based policy over privileged proprioception, object and contact state, trained across procedurally generated rebar and rack geometries.
Phase 2
Shared ResNet-18 encoder over calibrated views plus proprioception. DAgger + BC (β = 0.5) under appearance and system-identified dynamics randomization.
Phase 3
Frozen weights, same controller, same cameras. No segmentation model, no pose estimator, no fine-tuning.
The student policy inside the simulator it was trained in.
Procedurally generated bars and racks, randomized appearance and initial states.
Drag to rotate.
Real-robot rollouts of the same frozen policy, played at 1×.
The failures above are the main failure modes, and they occur both in simulation and in the real world.