Preparing Lunch Box¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the BEHAVIOR-1K runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 4h 00m 0% | 09-27 07:29 | 4h 00m | 416 | 47.2M / 283k | $0.694 | grader_error 0, missing_trajectory 1 |
| L✗ 4h 00m 0% | 09-27 11:32 | 4h 00m | 528 | 64.7M / 284k | $0.915 | deterministic 1, grader_error 0, live_success 0, n_actions 9143, replay_success 0, video_rendered 1 |
Task instruction (upstream)
Put both apple halves, the club sandwich, and the chocolate chip cookie from the chopping board on the kitchen countertop into the packing box on the countertop. Then take the bottle of tea out of the refrigerator, put it into the same box, and close the refrigerator when you're done.
Measured on the stack we run — BEHAVIOR-1K v3.9.2 / OmniGibson 3.9.2 / Isaac Sim 5.1
| Goal predicates | inside open |
| Quantifiers | forall |
| Goal clauses | 5 |
| Objects named in the goal | bottle__of__tea.n.01 chocolate_chip_cookie.n.01 club_sandwich.n.01 electric_refrigerator.n.01 half__apple.n.01 packing_box.n.02 |
| Objects in the problem | 11 in 10 categories |
| Rooms loaded | corridor_0 dining_room_0 entryway_0 garden_0 kitchen_0 living_room_0 living_room_1 |
| Demonstrations | 200 teleoperated episodes |
| Mean episode | 4m 35s (8,245 control steps at 30 Hz) |
| Base travel, mean | 21.4 m |
| Gripper travel, mean | left 14.09 m · right 17.24 m |
| Evaluation instances | 20 public test instances (ids 301–320) |
Read from BDDL activity definitions + 2026-challenge-task-instances (licensed download), 2026-09-21.
Tags¶
Task DomainMobile / Whole-body Manipulation
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.
