Manipulation · Bookshelf simple¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the HumanoidBench runs for every task of a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 1h 00m | 10-04 06:31 | 1h 00m | — | 29.5M / 264k | $0.487 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✗ 7m | 10-04 06:34 | 7m | — | 2.2M / 36k | $0.045 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 301, replay_success 0, video_rendered 1 |
3 other run(s) of HumanoidBench
Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Billed | Grade |
|---|---|---|---|---|---|---|---|
| U✗ 1h 00m | 10-01 03:43 | 1h 00m | 137 | 14.8M / 120k | $3.08 | $3.18 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✗ 1h 00m | 09-30 23:22 | 1h 00m | 168 | 18.3M / 92k | $3.16 | $3.26 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 100, replay_success 0, video_rendered 1 |
Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)
Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed
Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Task instruction (upstream)
Put each of the five objects onto its assigned place on the bookshelf, one at a time, in order.

What HumanoidBench states about this task
| Success criteria | 1. the summed per-step reward over one episode reaches 2000, HumanoidBench's own success bar (a total of rewards, not a number of steps; an episode is at most 1000 control steps) 2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state 3. limited: the run passes the moment its one episode reaches it (no reset in our runs since 2026-10-04; the recorded episode replays to the same state) |
| Env Id | h1hand-bookshelf_simple-v0 |
| Robot | Unitree H1 with two Shadow hands (61 actuators) |
| Category | Manipulation |
| Capability Class | C9 · long-horizon multi-subtask |
| Role | scored |
| Scoring | Only the object you are currently on is scored: each step pays for standing steadily, for a hand being near that object and for its progress towards its assigned slot. Placing it pays a one-off bonus that grows with how many you have placed — together the bonuses are most of the bar — and moves the scoring on to the next one. So the return is cut into five plateaus and stalling on the first caps the episode low. |
| Ends Early | The episode ends early if the pelvis drops too low or if the object currently being scored goes below about half a metre, and it ends once all five are placed. |
| Zero Action Return | 30.45 |
| Action Dim | 61 |
| Control Rate Hz | 50 |
| Limited Mode | head cameras (RGB 256×256), joint angles and velocities; a pelvis IMU and a camera fixed in the room when the run turns them on. Not where the robot is in the room, not where the objects or their places are, not the reward |
| Agent Budget | 3600 s of wall clock per mode |
From https://github.com/carlosferrazza/humanoid-bench @ cb11890, as defined in our task definitions @ 841996303.
Tags¶
Why this task is interesting¶
Five objects go onto assigned places on a bookshelf, one at a time, in order. The robot has to stand close to the shelf before it can lift the first object from the second shelf to the top one.
Capability notes¶
Not yet written.
Oracle demo review¶
No demo.
Discussion¶
Only the first object gets placed, by GPT-6.1 Sol in unlimited mode; GPT-6 Luna (2026-10-04) brings a palm within 7 cm of it without moving it, and its limited episode ended in a fall at step 301. (@williamzhangNU)