Manipulation · heat_food¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the VLABench runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default runclosed
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 35m | 10-01 18:26 | 35m | — | 11.5M / 54k | $0.164 | deterministic 1, grader_error 0, missing_trajectory 0, n_actions 3420, video_rendered 1 |
| L✗ 1h 00m | 10-01 18:26 | 1h 00m | — | 15.4M / 127k | $0.240 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 6113, replay_success 0, video_rendered 1 |
Task instruction (upstream)
Please heat cooked_food from tray_seen: put it into microwave_seen, close the microwave's door and press its start button (heating itself is not simulated).
Playback speed
Recorded by us: our reference solution, replayed by the verifier (the front camera).
What VLABench states about this task
| Success criteria | 1. VLABench's own success check for heat_food (the task's termination condition). The episode ends the moment it holds. 2. This states exactly what VLABench's success check requires. 3. unlimited (privileged): both fresh-process replays of the handed-in trajectory end in the same state, and the check holds on it 4. limited (standard): the one episode (no reset) is recorded by the service and replays to the same state; the check holds on it live and in the replay |
| Family | vlabench/heat_food |
| Robot | Franka Panda (7-DoF arm, two-finger gripper) |
| Category | Manipulation |
| Instance | seed 0 |
| Reference Solution | 793 control steps on this instance (solution/oracle.npy), replayed by solution/solve.sh. Source: PR #4's first-author reference: VLABench's general skills driven by VLABench's expert sequence or by a skill sequence the PR #4 author wrote, recorded as control steps during PR #4's authoring; only the recorded steps were committed (PR #4 @ ad3c510, refs/pull/4/head). VLABench's experts are removed from the image. Re-verified by scripts/vlabench/freeze.sh: the replay ends with the task's success at its last step. |
| Limited Mode | VLABench's heat_food as a robot service: a Franka Panda arm at a table, instructed "Please heat cooked_food from tray_seen: put it into microwave_seen, close the microwave's door and press its start button (heating itself is not simulated)." (a direct command), on the same frozen instance as vlabench-heat-food-i00-privileged. The simulator runs in the sim sidecar (environment/sim/server.py on main's rcb_service.py); the agent only has the client: four cameras (RGB 320×320, depth on request, calibrated), the robot's own state, and joint-position chunks. One episode, no reset, no object poses; nothing to hand in. |
| Oracle | full |
| Base Image | ghcr.io/mll-lab-nu/eai-vlabench:0.1.0 |
| Agent Budget | 3600 s of wall clock per mode |
| Task Dirs | vlabench-heat-food-i00-privileged, vlabench-heat-food-i00-standard |
From https://github.com/OpenMOSS/VLABench @ cf588fe, as defined in our task definitions @ 8c5594a43.
Tags¶
Task DomainManipulation
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.