Skip to content

Manipulation · Room

keephumanoidbench_manip——@williamzhangNU

Agent runs

Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the HumanoidBench runs for every task of a run.

Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run

TrialStartedAgent timeModel requestsTokens in / outEst. costGrade
U✗ 1h 00m10-04 06:241h 00m—25.1M / 247k$0.430deterministic 1, grader_error 0, n_actions 1000, video_rendered 1
L✗ 2m10-04 06:262m—244k / 13k$0.012deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 123, replay_success 0, video_rendered 1
3 other run(s) of HumanoidBench

Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)

TrialStartedAgent timeModel requestsTokens in / outEst. costBilledGrade
U✗ 1h 00m10-01 04:501h 00m1069.6M / 121k$2.54$2.64deterministic 1, grader_error 0, n_actions 1000, video_rendered 1
L✗ 54m10-01 00:1954m12712.3M / 87k$2.44$2.53deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 1000, replay_success 0, video_rendered 1

Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)

Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)

Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed

Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)

Task instruction (upstream)

Six loose objects — a chair, a trophy, a pair of headphones, two packages and a snow globe — are scattered around the room (the table and the bookshelf are fixed). Tidy up: gather them together, so they end up close to one another rather than spread out.

humanoidbench_manip scene
No oracle video published upstream — this is the scene it starts from.
What HumanoidBench states about this task
Success criteria 1. the summed per-step reward over one episode reaches 400, HumanoidBench's own success bar (a total of rewards, not a number of steps; an episode is at most 1000 control steps)
2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state
3. limited: the run passes the moment its one episode reaches it (no reset in our runs since 2026-10-04; the recorded episode replays to the same state)
Env Id h1hand-room-v0
Robot Unitree H1 with two Shadow hands (61 actuators)
Category Manipulation
Capability Class C8 · mobile manipulation
Role scored
Scoring Unusually, this task gives you no target positions at all. The dominant term measures how spread out the six objects are — the variance of their horizontal coordinates — and rewards making that small; a smaller term pays for standing steadily. Any arrangement that brings the objects close together scores, wherever in the room you do it.
Ends Early The episode ends if the pelvis drops near the ground.
Zero Action Return 8.84
Action Dim 61
Control Rate Hz 50
Limited Mode head cameras (RGB 256×256), joint angles and velocities; a pelvis IMU and a camera fixed in the room when the run turns them on. Not where the robot is in the room, not where the objects are, not the reward
Agent Budget 3600 s of wall clock per mode

From https://github.com/carlosferrazza/humanoid-bench @ cb11890, as defined in our task definitions @ 841996303.

Tags

Why this task is interesting

Six loose objects lie scattered around the room, and the score rewards gathering them close together, so the robot has to walk over and carry them. Standing through the episode earns less than half the bar.

Capability notes

Not yet written.

Oracle demo review

No demo.

Discussion

Blocked on walking: the robot stands the whole episode in unlimited mode but never reaches an object; without resets GPT-6 Luna's limited episode (2026-10-04) ended in a fall at step 123. (@williamzhangNU)