Skip to content

Multi-stage · ArrangeVegetables

keeprobocasa_kitchen—0m 27s@changhechen

Agent runs

Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the RoboCasa runs for every task of a run.

Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default runclosed

TrialStartedAgent timeModel requestsTokens in / outEst. costGrade
U✓ 40m10-01 19:1840m—5.9M / 66k$0.120deterministic 1, grader_error 0, missing_trajectory 0, n_actions 1266, replay_success 1, video_rendered 1
L✗ 59m10-01 19:1859m—20.9M / 102k$0.289deterministic 1, grader_error 0, live_success 0, missing_log 0, replay_success 0, video_rendered 1

Task instruction (upstream)

Pick the vegetables from the sink and place them on the cutting board.

Playback speed

Recorded by us: our reference solution, replayed by the verifier (the scene camera).

What RoboCasa states about this task
Success criteria 1. Each of the two vegetables touches the cutting board and its centre is within 0.7 × the board's horizontal radius of the board's centre (horizontal distance).
2. The gripper's grip site (between the fingertips) is more than 25 cm from the cutting board's centre: let go and move the hand away. A "centre" is a body's origin; "touches" means a MuJoCo contact.
3. unlimited (privileged): both fresh-process replays of the handed-in trajectory end in the same state, and the check holds on it
4. limited (standard): the one episode (no reset) is recorded by the service and replays to the same state; the check holds on it live and in the replay
Family robocasa/ArrangeVegetables
Robot Franka Panda on an Omron mobile base with a torso lift (PandaOmron)
Category Multi-stage · chopping food
Instance the initial state of official demonstration episode 0 (demo_1, v0.1/multi_stage/chopping_food/ArrangeVegetables/2024-05-11); MuJoCo state, model arrays, task state and RNG frozen in instance.npz (SHA-256 512f1d546771…, checked on load)
Deliverable (T, 12) native controller commands, 1 ≤ T ≤ 3000, 20 Hz
Reference Solution The demonstration's own actions (original: 544 actions), executed from the frozen scene by the robot's own controller: 544 actions in solution/oracle.npz. The demonstration is a reference solution: it is not in the image and the agent cannot reach it.
Limited Mode Standard mode (robot as a service, eai-standard/2.1) of robocasa-arrange-vegetables-i00-privileged: the same frozen instance and success check, served by the sim sidecar (images/robocasa/standard/server_casa.py, tool module casa_tool), which records the episode, replays it in a fresh simulator and judges it; the verifier grades the sidecar's record in a container of its own. The agent sees the robot's cameras (RGB-D, calibrated), its own joints and end effector, and the instruction.
Oracle full
Base Image ghcr.io/mll-lab-nu/eai-robocasa:0.1.0
Agent Budget 3600 s of wall clock per mode
Environment Source https://github.com/robocasa/robocasa/blob/756598a5be52e052339bb2d957426e39015c2afb/robocasa/environments/kitchen/multi_stage/chopping_food/arrange_vegetables.py#L4
Task Dirs robocasa-arrange-vegetables-i00-privileged, robocasa-arrange-vegetables-i00-standard

From https://github.com/robocasa/robocasa @ 756598a (v0.2), as defined in our task definitions @ 8c5594a43.

Tags

Task DomainManipulation

Why this task is interesting

Not yet written.

Capability notes

Not yet written.

Oracle demo review

Not yet reviewed.

Discussion