Single-stage · PnPCabToCounter¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the RoboCasa runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default runclosed
Not in this run: not in the sample (robot_coding_bench PR
Task instruction (upstream)
Pick the hot dog from the cabinet and place it on the counter.
Playback speed
Recorded by us: our reference solution, replayed by the verifier (the scene camera).
What RoboCasa states about this task
| Success criteria | 1. The object touches the counter below the cabinet. 2. The gripper's grip site (between the fingertips) is more than 25 cm from the object's centre: let go and move the hand away. A "centre" is a body's origin; "touches" means a MuJoCo contact. 3. unlimited (privileged): both fresh-process replays of the handed-in trajectory end in the same state, and the check holds on it 4. limited (standard): the one episode (no reset) is recorded by the service and replays to the same state; the check holds on it live and in the replay |
| Family | robocasa/PnPCabToCounter |
| Robot | Franka Panda on an Omron mobile base with a torso lift (PandaOmron) |
| Category | Single-stage · kitchen pnp |
| Instance | the initial state of official demonstration episode 0 (demo_1, v0.1/single_stage/kitchen_pnp/PnPCabToCounter/2024-04-24); MuJoCo state, model arrays, task state and RNG frozen in instance.npz (SHA-256 c9ebef28995d…, checked on load) |
| Deliverable | (T, 12) native controller commands, 1 ≤ T ≤ 3000, 20 Hz |
| Reference Solution | The demonstration's own actions (original: 276 actions), executed from the frozen scene by the robot's own controller: 276 actions in solution/oracle.npz. The demonstration is a reference solution: it is not in the image and the agent cannot reach it. |
| Limited Mode | Standard mode (robot as a service, eai-standard/2.1) of robocasa-pnp-cab-to-counter-i00-privileged: the same frozen instance and success check, served by the sim sidecar (images/robocasa/standard/server_casa.py, tool module casa_tool), which records the episode, replays it in a fresh simulator and judges it; the verifier grades the sidecar's record in a container of its own. The agent sees the robot's cameras (RGB-D, calibrated), its own joints and end effector, and the instruction. |
| Oracle | full |
| Base Image | ghcr.io/mll-lab-nu/eai-robocasa:0.1.0 |
| Agent Budget | 3600 s of wall clock per mode |
| Environment Source | https://github.com/robocasa/robocasa/blob/756598a5be52e052339bb2d957426e39015c2afb/robocasa/environments/kitchen/single_stage/kitchen_pnp.py#L142 |
| Task Dirs | robocasa-pnp-cab-to-counter-i00-privileged, robocasa-pnp-cab-to-counter-i00-standard |
From https://github.com/robocasa/robocasa @ 756598a (v0.2), as defined in our task definitions @ 8c5594a43.
Tags¶
CapabilityControl & Coordination
Task DomainManipulation
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.