Skip to content

Boxing Food After Dinner

pendingmediumhouse_single_floorDining Room, Kitchen3m 59sunowned

Agent runs

Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the BEHAVIOR-1K runs for every task of a run.

Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run

TrialStartedAgent timeModel requestsTokens in / outEst. costGrade
U✗ 4h 00m 0%09-27 06:114h 00m38742.5M / 214k$0.609grader_error 0, missing_trajectory 1
L✗ 4h 00m 0%09-27 06:114h 00m48562.7M / 279k$0.852deterministic 1, grader_error 0, live_success 0, n_actions 9184, replay_success 0, video_rendered 1

No instruction published upstream

The 2025 carryover tasks ship without instruction text. Watch the demo and write the goal in your own words below, clearly marked as a reconstruction rather than the official goal.

Measured on the stack we run — BEHAVIOR-1K v3.9.2 / OmniGibson 3.9.2 / Isaac Sim 5.1
Goal predicates inside ontop open
Quantifiers forall
Goal clauses 5
Objects named in the goal electric_refrigerator.n.01 plate.n.04 sink.n.01 taco.n.02 tupperware.n.01
Objects in the problem 11 in 9 categories
Rooms loaded corridor_0 dining_room_0 entryway_0 garden_0 kitchen_0 living_room_0 living_room_1
Demonstrations 200 teleoperated episodes
Mean episode 3m 43s (6,698 control steps at 30 Hz)
Base travel, mean 16.11 m
Gripper travel, mean left 11.16 m · right 17.7 m
Evaluation instances 20 public test instances (ids 301–320)

Read from BDDL activity definitions + 2026-challenge-task-instances (licensed download), 2026-09-21.

Tags

Why this task is interesting

Not yet written.

Capability notes

Not yet written.

Oracle demo review

Not yet reviewed.

Discussion