Locomotion · Walk¶
Excluded from the benchmark
not among the hardest (robot_coding_bench 2026-10-04): passed in unlimited mode
Greyed and last in the task list, left out of the benchmark's numbers; its demo and agent runs stay (excluded: in state/tasks/humanoidbench.yml).
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the HumanoidBench runs for every task of a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 1h 00m | 09-28 16:17 | 1h 00m | — | 25.2M / 193k | $0.407 | deterministic 1, grader_error 0, n_actions 491, video_rendered 1 |
| L✗ 18m | 10-01 14:33 | 18m | — | 3.9M / 77k | $0.088 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 47, replay_success 0, video_rendered 1 |
3 other run(s) of HumanoidBench
Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Billed | Grade |
|---|---|---|---|---|---|---|---|
| U✓ 1h 00m | 10-01 02:39 | 1h 00m | 133 | 11.9M / 87k | $2.35 | $2.42 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✗ 1h 00m | 09-30 22:31 | 1h 00m | 106 | 8.6M / 116k | $2.32 | $2.40 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 441, replay_success 0, video_rendered 1 |
Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 11m | 09-28 19:23 | 11m | — | 574k / 11k | $0.726 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✗ 14m | 09-28 17:21 | 14m | — | 4.6M / 50k | $4.14 | deterministic 1, grader_error 0, live_success 0, n_actions 1000, replay_success 0, video_rendered 1 |
Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 58m | 09-28 20:10 | 58m | — | 17.7M / 83k | $4.74 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✗ 12m | 09-28 17:39 | 12m | — | 4.2M / 21k | $1.26 | deterministic 1, grader_error 0, live_success 0, n_actions 223, replay_success 0, video_rendered 1 |
Task instruction (upstream)
Walk the robot forward and keep walking, without falling.

What HumanoidBench states about this task
| Success criteria | 1. the summed per-step reward over one episode reaches 700, HumanoidBench's own success bar (a total of rewards, not a number of steps; an episode is at most 1000 control steps) 2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state 3. limited: the run passes the moment a live episode reaches it (the recorded episode replays to the same state); otherwise the last episode is graded |
| Env Id | h1-walk-v0 |
| Robot | Unitree H1 (19 actuators) |
| Category | Locomotion |
| Capability Class | C1 · flat-ground periodic gait |
| Role | scored |
| Scoring | Each step is judged on three things at once — how fast the centre of mass is moving forward, how upright the robot is (head at standing height, torso vertical), and how little actuator force it is using — and they multiply. Not being upright is what makes a step worth nothing; heavy actuator force costs part of it. The forward-speed target is on the order of 1 m/s. |
| Ends Early | The episode ends early if the pelvis drops near the ground. |
| Zero Action Return | 6.56 |
| Action Dim | 19 |
| Control Rate Hz | 50 |
| Limited Mode | head cameras (RGB 256×256), joint angles and velocities; a pelvis IMU and a camera fixed in the room when the run turns them on. Not where the robot is in the room, not how fast it is moving, not the reward |
| Agent Budget | 3600 s of wall clock per mode |
From https://github.com/carlosferrazza/humanoid-bench @ cb11890, as defined in our task definitions @ 7d6a7a44.
Tags¶
Why this task is interesting¶
Walking has to be a limit cycle: the first step out of a standing pose is locally worse than not taking it, because the centre of mass must leave the support polygon before the robot gains anything. The reward multiplies forward speed, uprightness and low actuator force, so standing still earns little.
Capability notes¶
Not yet written.
Oracle demo review¶
No demo.
Discussion¶
Passes in unlimited mode with a whole-body inverse-dynamics walker at about 1 m/s; in limited mode the best episode walked 738 steps before falling and the graded one fell at step 441. (@williamzhangNU)