Picking Up Toys¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the BEHAVIOR-1K runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run
Not in this run: not among the 30 chosen for the first round (2026-09-26): mostly particles, cutting and liquids, or long tidy-ups like chosen ones
Task instruction (upstream)
Put all the toys in the child's room - the three board games (two on the bed and one on the table), the two jigsaw puzzles on the table, and the tennis ball on the table - inside the toy box on the table in the child's room.
Measured on the stack we run — BEHAVIOR-1K v3.9.2 / OmniGibson 3.9.2 / Isaac Sim 5.1
| Goal predicates | inside |
| Quantifiers | exists forall |
| Goal clauses | 3 |
| Objects named in the goal | board_game.n.01 jigsaw_puzzle.n.01 tennis_ball.n.01 toy_box.n.01 |
| Objects in the problem | 11 in 8 categories |
| Rooms loaded | childs_room_0 childs_room_1 childs_room_2 corridor_0 dining_room_0 entryway_0 garden_0 |
| Demonstrations | 200 teleoperated episodes |
| Mean episode | 10m 30s (18,890 control steps at 30 Hz) |
| Base travel, mean | 47.05 m |
| Gripper travel, mean | left 20.41 m · right 21.55 m |
| Evaluation instances | 20 public test instances (ids 301–320) |
Read from BDDL activity definitions + 2026-challenge-task-instances (licensed download), 2026-09-21.
Tags¶
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.
