Skip to content

Locomotion · Balance simple

excludedkeephumanoidbench_loco——@williamzhangNU

Excluded from the benchmark

not among the hardest (robot_coding_bench 2026-10-04): passed in unlimited mode

Greyed and last in the task list, left out of the benchmark's numbers; its demo and agent runs stay (excluded: in state/tasks/humanoidbench.yml).

Agent runs

Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the HumanoidBench runs for every task of a run.

Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run

TrialStartedAgent timeModel requestsTokens in / outEst. costGrade
U✗ 59m09-28 16:1859m—22.9M / 157k$0.328deterministic 1, grader_error 0, n_actions 248, video_rendered 1
L✗ 16m10-01 14:1716m—5.0M / 62k$0.092deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 95, replay_success 0, video_rendered 1
3 other run(s) of HumanoidBench

Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)

TrialStartedAgent timeModel requestsTokens in / outEst. costBilledGrade
U✓ 1h 00m10-01 02:391h 00m13710.5M / 68k$1.99$2.06deterministic 1, grader_error 0, n_actions 1000, video_rendered 1
L✗ 52m09-30 22:2552m1179.1M / 106k$2.28$2.36deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 58, replay_success 0, video_rendered 1

Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)

Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)

Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed

Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)

Task instruction (upstream)

Stand on the balance board and stay on it, upright and still, for the whole episode.

humanoidbench_loco scene
No oracle video published upstream — this is the scene it starts from.
What HumanoidBench states about this task
Success criteria 1. the summed per-step reward over one episode reaches 800, HumanoidBench's own success bar (a total of rewards, not a number of steps; an episode is at most 1000 control steps)
2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state
3. limited: the run passes the moment a live episode reaches it (the recorded episode replays to the same state); otherwise the last episode is graded
Env Id h1-balance_simple-v0
Robot Unitree H1 (19 actuators)
Category Locomotion
Capability Class C3 · quasi-static stabilisation
Role scored
Scoring A good step asks for only three things, multiplied together — low actuator force, staying upright (head at full standing height on the board, torso vertical), and near-zero horizontal velocity — so scoring well for one step is not the difficulty. The difficulty is that the episode ends the moment the pelvis drops too low, anything but the fulcrum sphere touches the ground — the board included — or the sphere touches anything other than the ground and the board, and every step after that is worth nothing.
Zero Action Return 37.05
Action Dim 19
Control Rate Hz 50
Limited Mode head cameras (RGB 256×256), joint angles and velocities; a pelvis IMU and a camera fixed in the room when the run turns them on. Not where the robot is in the room, not the board's pose, not the reward
Agent Budget 3600 s of wall clock per mode

From https://github.com/carlosferrazza/humanoid-bench @ cb11890, as defined in our task definitions @ 7d6a7a44.

Tags

Why this task is interesting

Standing on a seesaw board is easy for a step or two; the task is to keep the board level and the robot upright and still for nearly the whole 1000-step episode.

Capability notes

Not yet written.

Oracle demo review

No demo.

Discussion

Passes in unlimited mode with centre-of-mass and board-state feedback tuned by CMA-ES, standing all 1000 steps by 11 minutes; limited mode never stays up past 99 steps. (@williamzhangNU)