Outfit A Basic Toolbox¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the BEHAVIOR-1K runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 4h 00m 0% | 09-27 19:18 | 4h 00m | 507 | 73.5M / 277k | $1.01 | deterministic 1, grader_error 0, instance_exact 1, instance_ok 1, n_actions 4186, video_rendered 1 |
| L✗ 4h 00m 0% | 09-27 15:07 | 4h 00m | 716 | 98.4M / 476k | $1.37 | deterministic 1, grader_error 0, live_success 0, n_actions 10252, replay_success 0, video_rendered 1 |
Task instruction (upstream)
In the utility room, put the drill, pliers, flashlight, Allen wrench, and screwdriver from the tabletop into the toolbox, keep the toolbox on the tabletop, and close the toolbox.
Measured on the stack we run — BEHAVIOR-1K v3.9.2 / OmniGibson 3.9.2 / Isaac Sim 5.1
| Goal predicates | inside ontop open |
| Goal clauses | 7 |
| Objects named in the goal | allen_wrench.n.01 drill.n.01 flashlight.n.01 pliers.n.01 screwdriver.n.01 tabletop.n.01 toolbox.n.01 |
| Objects in the problem | 9 in 9 categories |
| Rooms loaded | corridor_0 dining_room_0 entryway_0 garden_0 utility_room_0 |
| Demonstrations | 200 teleoperated episodes |
| Mean episode | 5m 55s (10,638 control steps at 30 Hz) |
| Base travel, mean | 31.68 m |
| Gripper travel, mean | left 18.18 m · right 20.07 m |
| Evaluation instances | 20 public test instances (ids 301–320) |
Read from BDDL activity definitions + 2026-challenge-task-instances (licensed download), 2026-09-21.
Tags¶
Task DomainMobile / Whole-body Manipulation
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.
