Locomotion · Balance hard¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the HumanoidBench runs for every task of a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 1h 00m | 10-04 05:26 | 1h 00m | — | 31.3M / 216k | $0.478 | deterministic 1, grader_error 0, n_actions 78, video_rendered 1 |
| L✗ 2m | 10-04 05:27 | 2m | — | 398k / 15k | $0.015 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 43, replay_success 0, video_rendered 1 |
3 other run(s) of HumanoidBench
Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Billed | Grade |
|---|---|---|---|---|---|---|---|
| U✗ 1h 00m | 10-01 03:44 | 1h 00m | 123 | 12.5M / 125k | $2.90 | $3.01 | deterministic 1, grader_error 0, n_actions 277, video_rendered 1 |
| L✗ 57m | 09-30 22:25 | 57m | 127 | 10.2M / 112k | $2.49 | $2.58 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 58, replay_success 0, video_rendered 1 |
Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)
Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed
Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Task instruction (upstream)
Stand on the balance board and stay on it, upright and still, for the whole episode.

What HumanoidBench states about this task
| Success criteria | 1. the summed per-step reward over one episode reaches 800, HumanoidBench's own success bar (a total of rewards, not a number of steps; an episode is at most 1000 control steps) 2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state 3. limited: the run passes the moment its one episode reaches it (no reset in our runs since 2026-10-04; the recorded episode replays to the same state) |
| Env Id | h1-balance_hard-v0 |
| Robot | Unitree H1 (19 actuators) |
| Category | Locomotion |
| Capability Class | C3 · quasi-static stabilisation |
| Role | scored |
| Scoring | A good step asks for only three things, multiplied together — low actuator force, staying upright (head at full standing height on the board, torso vertical), and near-zero horizontal velocity — so scoring well for one step is not the difficulty. The difficulty is that the episode ends the moment the pelvis drops too low, anything but the rolling sphere touches the ground — the board included — or the sphere touches anything other than the ground and the board, and every step after that is worth nothing. |
| Zero Action Return | 29.03 |
| Action Dim | 19 |
| Control Rate Hz | 50 |
| Limited Mode | head cameras (RGB 256×256), joint angles and velocities; a pelvis IMU and a camera fixed in the room when the run turns them on. Not where the robot is in the room, not the board's or the roller's pose, not the reward |
| Agent Budget | 3600 s of wall clock per mode |
From https://github.com/carlosferrazza/humanoid-bench @ cb11890, as defined in our task definitions @ 841996303.
Tags¶
Why this task is interesting¶
The board rests on a rolling ball, so the robot has to keep itself and the board balanced at once, for the whole episode.
Capability notes¶
Not yet written.
Oracle demo review¶
No demo.
Discussion¶
Still out of reach: the board stays up at most 277 steps in unlimited mode (GPT-6.1 Sol's CMA-ES-tuned board and centre-of-mass feedback; GPT-6 Luna's, 2026-10-04, 78) and 69 in limited mode; without resets Luna's episode ended in a fall at step 43. (@williamzhangNU)