Manipulation · Room¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the HumanoidBench runs for every task of a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 1h 00m | 10-04 06:24 | 1h 00m | — | 25.1M / 247k | $0.430 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✗ 2m | 10-04 06:26 | 2m | — | 244k / 13k | $0.012 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 123, replay_success 0, video_rendered 1 |
3 other run(s) of HumanoidBench
Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Billed | Grade |
|---|---|---|---|---|---|---|---|
| U✗ 1h 00m | 10-01 04:50 | 1h 00m | 106 | 9.6M / 121k | $2.54 | $2.64 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✗ 54m | 10-01 00:19 | 54m | 127 | 12.3M / 87k | $2.44 | $2.53 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 1000, replay_success 0, video_rendered 1 |
Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)
Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed
Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Task instruction (upstream)
Six loose objects — a chair, a trophy, a pair of headphones, two packages and a snow globe — are scattered around the room (the table and the bookshelf are fixed). Tidy up: gather them together, so they end up close to one another rather than spread out.

What HumanoidBench states about this task
| Success criteria | 1. the summed per-step reward over one episode reaches 400, HumanoidBench's own success bar (a total of rewards, not a number of steps; an episode is at most 1000 control steps) 2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state 3. limited: the run passes the moment its one episode reaches it (no reset in our runs since 2026-10-04; the recorded episode replays to the same state) |
| Env Id | h1hand-room-v0 |
| Robot | Unitree H1 with two Shadow hands (61 actuators) |
| Category | Manipulation |
| Capability Class | C8 · mobile manipulation |
| Role | scored |
| Scoring | Unusually, this task gives you no target positions at all. The dominant term measures how spread out the six objects are — the variance of their horizontal coordinates — and rewards making that small; a smaller term pays for standing steadily. Any arrangement that brings the objects close together scores, wherever in the room you do it. |
| Ends Early | The episode ends if the pelvis drops near the ground. |
| Zero Action Return | 8.84 |
| Action Dim | 61 |
| Control Rate Hz | 50 |
| Limited Mode | head cameras (RGB 256×256), joint angles and velocities; a pelvis IMU and a camera fixed in the room when the run turns them on. Not where the robot is in the room, not where the objects are, not the reward |
| Agent Budget | 3600 s of wall clock per mode |
From https://github.com/carlosferrazza/humanoid-bench @ cb11890, as defined in our task definitions @ 841996303.
Tags¶
Why this task is interesting¶
Six loose objects lie scattered around the room, and the score rewards gathering them close together, so the robot has to walk over and carry them. Standing through the episode earns less than half the bar.
Capability notes¶
Not yet written.
Oracle demo review¶
No demo.
Discussion¶
Blocked on walking: the robot stands the whole episode in unlimited mode but never reaches an object; without resets GPT-6 Luna's limited episode (2026-10-04) ended in a fall at step 123. (@williamzhangNU)