Setting The Table¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the BEHAVIOR-1K runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 4h 00m 0% | 09-27 06:10 | 4h 00m | 391 | 58.7M / 250k | $0.820 | deterministic 1, grader_error 0, instance_exact 0, instance_ok 1, n_actions 3093, video_rendered 1 |
| L✗ 4h 00m 0% | 09-27 06:10 | 4h 00m | 884 | 123.3M / 470k | $1.66 | deterministic 0, grader_error 0, live_success 0, n_actions 14364, replay_success 0, video_rendered 0 |
No instruction published upstream
The 2025 carryover tasks ship without instruction text. Watch the demo and write the goal in your own words below, clearly marked as a reconstruction rather than the official goal.
Measured on the stack we run — BEHAVIOR-1K v3.9.2 / OmniGibson 3.9.2 / Isaac Sim 5.1
| Goal predicates | nextto ontop |
| Quantifiers | forall forpairs |
| Goal clauses | 4 |
| Objects named in the goal | breakfast_table.n.01 cupcake.n.01 plate.n.04 table_knife.n.01 tablefork.n.01 |
| Objects in the problem | 14 in 10 categories |
| Rooms loaded | corridor_0 dining_room_0 entryway_0 garden_0 kitchen_0 living_room_0 living_room_1 |
| Demonstrations | 200 teleoperated episodes |
| Mean episode | 9m 54s (17,808 control steps at 30 Hz) |
| Base travel, mean | 55.46 m |
| Gripper travel, mean | left 38.13 m · right 46.91 m |
| Evaluation instances | 20 public test instances (ids 301–320) |
Read from BDDL activity definitions + 2026-challenge-task-instances (licensed download), 2026-09-21.
Tags¶
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.
