Single-stage · PnPCounterToCab¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the RoboCasa runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default runclosed
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 15m | 10-01 19:18 | 15m | — | 1.6M / 23k | $0.034 | deterministic 1, grader_error 0, missing_trajectory 0, n_actions 321, replay_success 1, video_rendered 1 |
| L✗ 57m | 10-01 19:18 | 57m | — | 35.4M / 135k | $0.468 | deterministic 1, grader_error 0, live_success 0, missing_log 0, replay_success 0, video_rendered 1 |
Task instruction (upstream)
Pick the cake from the counter and place it in the cabinet.
Playback speed
Recorded by us: our reference solution, replayed by the verifier (the scene camera).
What RoboCasa states about this task
| Success criteria | 1. The object is entirely inside the cabinet (all eight corners of its bounding box within the interior, with a 5 cm tolerance). 2. The gripper's grip site (between the fingertips) is more than 25 cm from the object's centre: let go and move the hand away. A "centre" is a body's origin; "touches" means a MuJoCo contact. 3. unlimited (privileged): both fresh-process replays of the handed-in trajectory end in the same state, and the check holds on it 4. limited (standard): the one episode (no reset) is recorded by the service and replays to the same state; the check holds on it live and in the replay |
| Family | robocasa/PnPCounterToCab |
| Robot | Franka Panda on an Omron mobile base with a torso lift (PandaOmron) |
| Category | Single-stage · kitchen pnp |
| Instance | the initial state of official demonstration episode 0 (demo_1, v0.1/single_stage/kitchen_pnp/PnPCounterToCab/2024-04-24); MuJoCo state, model arrays, task state and RNG frozen in instance.npz (SHA-256 e95a84172b1b…, checked on load) |
| Deliverable | (T, 12) native controller commands, 1 ≤ T ≤ 3000, 20 Hz |
| Reference Solution | The demonstration's own actions (original: 217 actions), executed from the frozen scene by the robot's own controller: 217 actions in solution/oracle.npz. The demonstration is a reference solution: it is not in the image and the agent cannot reach it. |
| Limited Mode | Standard mode (robot as a service, eai-standard/2.1) of robocasa-pnp-counter-to-cab-i00-privileged: the same frozen instance and success check, served by the sim sidecar (images/robocasa/standard/server_casa.py, tool module casa_tool), which records the episode, replays it in a fresh simulator and judges it; the verifier grades the sidecar's record in a container of its own. The agent sees the robot's cameras (RGB-D, calibrated), its own joints and end effector, and the instruction. |
| Oracle | full |
| Base Image | ghcr.io/mll-lab-nu/eai-robocasa:0.1.0 |
| Agent Budget | 3600 s of wall clock per mode |
| Environment Source | https://github.com/robocasa/robocasa/blob/756598a5be52e052339bb2d957426e39015c2afb/robocasa/environments/kitchen/single_stage/kitchen_pnp.py#L24 |
| Task Dirs | robocasa-pnp-counter-to-cab-i00-privileged, robocasa-pnp-counter-to-cab-i00-standard |
From https://github.com/robocasa/robocasa @ 756598a (v0.2), as defined in our task definitions @ 8c5594a43.
Tags¶
CapabilityControl & Coordination
Task DomainManipulation
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.