Turning On Radio¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the BEHAVIOR-1K runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 1h 16m | 09-27 06:16 | 1h 16m | 190 | 21.8M / 94k | $0.292 | deterministic 1, grader_error 0, instance_exact 0, instance_ok 1, n_actions 1544, video_rendered 1 |
| L✗ 4h 00m 0% | 09-27 06:15 | 4h 00m | 572 | 75.5M / 347k | $1.07 | deterministic 1, grader_error 0, live_success 0, n_actions 4735, replay_success 0, video_rendered 1 |
Task instruction (upstream)
Turn on the radio receiver that's on the table in the living room.
Measured on the stack we run — BEHAVIOR-1K v3.9.2 / OmniGibson 3.9.2 / Isaac Sim 5.1
| Goal predicates | toggled_on |
| Goal clauses | 1 |
| Objects named in the goal | radio_receiver.n.01 |
| Objects in the problem | 4 in 4 categories |
| Rooms loaded | corridor_0 garden_0 kitchen_0 living_room_0 |
| Demonstrations | 200 teleoperated episodes |
| Mean episode | 1m 12s (2,150 control steps at 30 Hz) |
| Base travel, mean | 5.7 m |
| Gripper travel, mean | left 3.26 m · right 4.07 m |
| Evaluation instances | 20 public test instances (ids 301–320) |
Read from BDDL activity definitions + 2026-challenge-task-instances (licensed download), 2026-09-21.
Tags¶
Skill primitives in the demo: move toturn on switch
Why this task is interesting¶
The simplest task in the suite — 72 seconds, one room, one object, one state change — which is exactly what makes it valuable. It is the natural smoke test for the evaluation harness: if this does not run end to end, nothing will, and the failure will be in the plumbing rather than in the policy. Keep it for that reason even though it discriminates poorly between strong agents.
Note that the goal is a state change (toggled_on), not a pose. There is no rearrangement to score, so partial credit is close to binary here.
Capability notes¶
mob.room-scale-nav— the radio is on a table in the living room; the base must be positioned before the arm can reach. Single room only, so notmob.multi-room-nav.mp.toggle-press— the actual goal: actuate the receiver's control.per.small-object— a power control on a radio is a few centimetres across, and must be located in the head camera at working distance. This is the part most likely to fail in practice.
Deliberately not tagged:
mp.pick-place— nothing is transported. The radio stays where it is.rea.long-horizon— two steps.per.occluded-search— the instruction states the location outright.
Oracle demo review¶
Not yet watched by anyone on the team. The tagging above was derived from the published instruction text alone. Watching the demo is the remaining work on this page — in particular, check whether the control is reachable without torso actuation, which would add mob.height-variation.
Discussion¶
- Suggested as the harness smoke test because it is the shortest demo in the suite and needs no manipulation of transported objects.
