Skip to content

Manipulation · Bookshelf simple

keephumanoidbench_manip——@williamzhangNU

Agent runs

Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the HumanoidBench runs for every task of a run.

Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run

TrialStartedAgent timeModel requestsTokens in / outEst. costGrade
U✗ 1h 00m10-04 06:311h 00m—29.5M / 264k$0.487deterministic 1, grader_error 0, n_actions 1000, video_rendered 1
L✗ 7m10-04 06:347m—2.2M / 36k$0.045deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 301, replay_success 0, video_rendered 1
3 other run(s) of HumanoidBench

Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)

TrialStartedAgent timeModel requestsTokens in / outEst. costBilledGrade
U✗ 1h 00m10-01 03:431h 00m13714.8M / 120k$3.08$3.18deterministic 1, grader_error 0, n_actions 1000, video_rendered 1
L✗ 1h 00m09-30 23:221h 00m16818.3M / 92k$3.16$3.26deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 100, replay_success 0, video_rendered 1

Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)

Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)

Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed

Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)

Task instruction (upstream)

Put each of the five objects onto its assigned place on the bookshelf, one at a time, in order.

humanoidbench_manip scene
No oracle video published upstream — this is the scene it starts from.
What HumanoidBench states about this task
Success criteria 1. the summed per-step reward over one episode reaches 2000, HumanoidBench's own success bar (a total of rewards, not a number of steps; an episode is at most 1000 control steps)
2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state
3. limited: the run passes the moment its one episode reaches it (no reset in our runs since 2026-10-04; the recorded episode replays to the same state)
Env Id h1hand-bookshelf_simple-v0
Robot Unitree H1 with two Shadow hands (61 actuators)
Category Manipulation
Capability Class C9 · long-horizon multi-subtask
Role scored
Scoring Only the object you are currently on is scored: each step pays for standing steadily, for a hand being near that object and for its progress towards its assigned slot. Placing it pays a one-off bonus that grows with how many you have placed — together the bonuses are most of the bar — and moves the scoring on to the next one. So the return is cut into five plateaus and stalling on the first caps the episode low.
Ends Early The episode ends early if the pelvis drops too low or if the object currently being scored goes below about half a metre, and it ends once all five are placed.
Zero Action Return 30.45
Action Dim 61
Control Rate Hz 50
Limited Mode head cameras (RGB 256×256), joint angles and velocities; a pelvis IMU and a camera fixed in the room when the run turns them on. Not where the robot is in the room, not where the objects or their places are, not the reward
Agent Budget 3600 s of wall clock per mode

From https://github.com/carlosferrazza/humanoid-bench @ cb11890, as defined in our task definitions @ 841996303.

Tags

Why this task is interesting

Five objects go onto assigned places on a bookshelf, one at a time, in order. The robot has to stand close to the shelf before it can lift the first object from the second shelf to the top one.

Capability notes

Not yet written.

Oracle demo review

No demo.

Discussion

Only the first object gets placed, by GPT-6.1 Sol in unlimited mode; GPT-6 Luna (2026-10-04) brings a palm within 7 cm of it without moving it, and its limited episode ended in a fall at step 301. (@williamzhangNU)