Locomotion · Stair¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the HumanoidBench runs for every task of a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 48m ⓘ | 10-04 05:24 | 48m | — | 27.6M / 226k | $0.448 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✗ 2m | 10-04 05:24 | 2m | — | 525k / 14k | $0.016 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 47, replay_success 0, video_rendered 1 |
3 other run(s) of HumanoidBench
Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Billed | Grade |
|---|---|---|---|---|---|---|---|
| U✗ 1h 00m | 10-01 03:44 | 1h 00m | 132 | 12.4M / 125k | $2.85 | $2.95 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✗ 1h 00m | 09-30 22:25 | 1h 00m | 94 | 7.5M / 123k | $2.33 | $2.43 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 110, replay_success 0, video_rendered 1 |
Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)
Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed
Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Task instruction (upstream)
Walk the robot forward over the stairs, up each flight and down the other side, without falling.

What HumanoidBench states about this task
| Success criteria | 1. the summed per-step reward over one episode reaches 700, HumanoidBench's own success bar (a total of rewards, not a number of steps; an episode is at most 1000 control steps) 2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state 3. limited: the run passes the moment its one episode reaches it (no reset in our runs since 2026-10-04; the recorded episode replays to the same state) |
| Env Id | h1-stair-v0 |
| Robot | Unitree H1 (19 actuators) |
| Category | Locomotion |
| Capability Class | C2 · constrained gait |
| Role | scored |
| Scoring | The same three multiplied requirements as walking on the flat — forward speed (the target is on the order of 1 m/s), staying upright, low actuator force — with one change that matters: the height part of "upright" is measured as the head's height above the feet, not above the ground, because the ground rises and falls under you and so cannot be the yardstick, and the torso tilt it accepts is looser than on the flat. |
| Ends Early | Termination is also different: the episode ends only when the torso tips close to horizontal, rather than when it drops to a fixed height. |
| Zero Action Return | 8.63 |
| Action Dim | 19 |
| Control Rate Hz | 50 |
| Limited Mode | head cameras (RGB 256×256), joint angles and velocities; a pelvis IMU and a camera fixed in the room when the run turns them on. Not where the robot is in the room, not the stairs' geometry, not the reward |
| Agent Budget | 3600 s of wall clock per mode |
From https://github.com/carlosferrazza/humanoid-bench @ cb11890, as defined in our task definitions @ 841996303.
Tags¶
Why this task is interesting¶
The robot walks up a flight of stairs and down again at about 1 m/s. Each step up asks for a foot lifted onto the next tread without losing balance, so a flat-ground gait is not enough.
Capability notes¶
Not yet written.
Oracle demo review¶
No demo.
Discussion¶
Out of reach in both modes: GPT-6.1 Sol's gaits never got past the third stair transfer, GPT-6 Luna's best unlimited trajectory (2026-10-04) only reaches the first riser, and without resets Luna's limited episode ended in a fall at step 47. (@williamzhangNU)