Panda Open Cabinet¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the MuJoCo Playground (manipulation) runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 7m | 10-01 08:43 | 7m | — | 1.8M / 49k | $0.051 | deterministic 1, grader_error 0, n_actions 1059, video_rendered 1 |
| L✓ 12m | 10-01 08:19 | 12m | — | 4.0M / 61k | $0.079 | deterministic 1, grader_error 0, live_success 1, video_rendered 1 |
1 other run(s) of MuJoCo Playground (manipulation)
Codex + GPT-6 Luna, reasoning xhigh (Azure OpenAI API, 2026-09-30 stress test)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 14m | 09-30 12:20 | 14m | — | 2.3M / 41k | $0.051 | deterministic 1, grader_error 0, n_actions 2420, video_rendered 1 |
| L✗ 59m | 09-30 11:25 | 59m | — | 13.6M / 107k | $0.207 | deterministic 1, grader_error 0, live_success 0, video_rendered 1 |
Task instruction (upstream)
Grasp the handle and pull it along its slide to the target.
Recorded by us: our scripted IK oracle (it reads object and target poses from the simulator) replayed from the task's frozen start in plain MuJoCo. Playground ships no demonstrations; its reference solutions are trained RL policies.
What MuJoCo Playground (manipulation) states about this task
| Defined in | mujoco_playground/_src/manipulation/franka_emika_panda/open_cabinet.py |
| Env | PandaOpenCabinet |
| Robot | Franka Panda |
| Control Hz | 50 |
| Episode S | 3.0 |
| Metric | Playground: dense RL reward, no success test |
From https://github.com/google-deepmind/mujoco_playground @ 4057c14.
Measured on the stack we run — MuJoCo 3.3.7 (CPU, bit-exact replay), Playground's scene XML @ 4057c14 + MuJoCo Menagerie 1b86ece, 50 Hz control
| Harness Task | mujoco_playground_panda_open_cabinet |
| Success Test | ours: handle within 1.5 cm of the target along its slide, at the end of the trajectory |
| Oracle | scripted IK oracle (privileged state), replayed in the verifier: success 1, deterministic 1 |
Read from run in plain MuJoCo (robot_coding_bench tasks/mujoco_playground_panda_open_cabinet, image rcb-mujoco 0.1.1), 2026-09-28.
Tags¶
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.