Aloha Hand Over¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the MuJoCo Playground (manipulation) runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 2m | 10-01 06:30 | 2m | — | 394k / 16k | $0.016 | deterministic 1, grader_error 0, n_actions 1036, video_rendered 1 |
| L✗ 1h 00m | 10-01 08:05 | 1h 00m | — | 26.0M / 193k | $0.387 | deterministic 1, grader_error 0, live_success 0, video_rendered 1 |
1 other run(s) of MuJoCo Playground (manipulation)
Codex + GPT-6 Luna, reasoning xhigh (Azure OpenAI API, 2026-09-30 stress test)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 4m | 09-30 02:41 | 4m | — | 911k / 20k | $0.024 | deterministic 1, grader_error 0, n_actions 890, video_rendered 1 |
| L✗ 1h 00m | 09-30 10:55 | 1h 00m | — | 3.8M / 48k | $0.071 | deterministic 1, grader_error 0, live_success 0, video_rendered 1 |
Task instruction (upstream)
Pick up the box with the left arm, hand it over to the right arm, and hold it at the target.
Recorded by us: our scripted IK oracle (it reads object and target poses from the simulator) replayed from the task's frozen start in plain MuJoCo. Playground ships no demonstrations; its reference solutions are trained RL policies.
What MuJoCo Playground (manipulation) states about this task
| Defined in | mujoco_playground/_src/manipulation/aloha/handover.py |
| Env | AlohaHandOver |
| Robot | ALOHA 2 (two ViperX 300s arms) |
| Control Hz | 50 |
| Episode S | 5.0 |
| Metric | Playground: dense RL reward, no success test |
From https://github.com/google-deepmind/mujoco_playground @ 4057c14.
Measured on the stack we run — MuJoCo 3.3.7 (CPU, bit-exact replay), Playground's scene XML @ 4057c14 + MuJoCo Menagerie 1b86ece, 50 Hz control
| Harness Task | mujoco_playground_aloha_hand_over |
| Success Test | ours: box centre within 3 cm of the target, touching the right gripper and nothing of the left arm, off the table (z >= 0.05 m), at the end of the trajectory |
| Oracle | scripted IK oracle (privileged state), replayed in the verifier: success 1, deterministic 1 |
Read from run in plain MuJoCo (robot_coding_bench tasks/mujoco_playground_aloha_hand_over, image rcb-mujoco 0.1.1), 2026-09-28.
Tags¶
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.