Setting Mousetraps¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the BEHAVIOR-1K runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 4h 00m 0% | 09-27 06:06 | 4h 00m | 457 | 59.6M / 199k | $0.781 | grader_error 0, missing_trajectory 1 |
| L✗ 4h 00m 33% | 09-27 06:05 | 4h 00m | 941 | 132.2M / 469k | $1.71 | deterministic 1, grader_error 0, live_success 0, n_actions 14119, replay_success 0, video_rendered 1 |
Task instruction (upstream)
Take the four mousetraps from the cabinet in the bathroom and place them on the bathroom floor. Make sure all four end up on the same floor surface, and ensure that at least two of them are either under or directly next to the same bathroom sink.
Measured on the stack we run — BEHAVIOR-1K v3.9.2 / OmniGibson 3.9.2 / Isaac Sim 5.1
| Goal predicates | nextto ontop under |
| Quantifiers | exists forall forn |
| Counted quantities | 2 |
| Goal clauses | 2 |
| Objects named in the goal | floor.n.01 mousetrap.n.01 sink.n.01 |
| Objects in the problem | 10 in 5 categories |
| Rooms loaded | bathroom_0 |
| Demonstrations | 200 teleoperated episodes |
| Mean episode | 5m 40s (10,196 control steps at 30 Hz) |
| Base travel, mean | 31.6 m |
| Gripper travel, mean | left 13.15 m · right 12.38 m |
| Evaluation instances | 20 public test instances (ids 301–320) |
Read from BDDL activity definitions + 2026-challenge-task-instances (licensed download), 2026-09-21.
Tags¶
Task DomainMobile / Whole-body Manipulation
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.
