Skip to content

Locomotion · Maze

excludedkeephumanoidbench_loco——@williamzhangNU

Excluded from the benchmark

not among the hardest (robot_coding_bench 2026-10-04): passed in unlimited mode

Greyed and last in the task list, left out of the benchmark's numbers; its demo and agent runs stay (excluded: in state/tasks/humanoidbench.yml).

Agent runs

Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the HumanoidBench runs for every task of a run.

Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run

TrialStartedAgent timeModel requestsTokens in / outEst. costGrade
U✗ 1h 00m09-28 17:441h 00m—27.1M / 203k$0.430deterministic 1, grader_error 0, n_actions 1000, video_rendered 1
L✗ 45m10-01 14:1745m—20.8M / 134k$0.308deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 1000, replay_success 0, video_rendered 1
3 other run(s) of HumanoidBench

Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)

TrialStartedAgent timeModel requestsTokens in / outEst. costBilledGrade
U✓ 1h 00m10-01 03:451h 00m11712.3M / 84k$2.39$2.48deterministic 1, grader_error 0, n_actions 1000, video_rendered 1
L✗ 54m09-30 22:2554m1089.6M / 108k$2.39$2.48deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 181, replay_success 0, video_rendered 1

Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)

Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)

Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed

Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)

Task instruction (upstream)

Walk the robot along the maze's corridors from checkpoint to checkpoint, in order, turning at each corner, without hitting the walls.

humanoidbench_loco scene
No oracle video published upstream — this is the scene it starts from.
What HumanoidBench states about this task
Success criteria 1. the summed per-step reward over one episode reaches 1200, HumanoidBench's own success bar (a total of rewards, not a number of steps; an episode is at most 1000 control steps)
2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state
3. limited: the run passes the moment a live episode reaches it (the recorded episode replays to the same state); otherwise the last episode is graded
Env Id h1-maze-v0
Robot Unitree H1 (19 actuators)
Category Locomotion
Capability Class C4 · navigation and passage
Role scored
Scoring Most of the bar is not accumulated, it is jumped to: reaching a checkpoint pays a one-off bonus, and the bonuses grow with how many you have already reached, so the later corners are worth far more than the first. Between checkpoints each step pays a small amount for posture, for moving in the current corridor's direction (the speed target is on the order of 2 m/s) and for nearing the next checkpoint; touching a wall multiplies that per-step amount down to a fraction (the bonuses are not reduced). Walking correctly but never reaching a checkpoint leaves the episode below the bar.
Ends Early The episode ends early if the pelvis drops near the ground.
Zero Action Return 120.59
Action Dim 19
Control Rate Hz 50
Limited Mode head cameras (RGB 256×256), joint angles and velocities; a pelvis IMU and a camera fixed in the room when the run turns them on. Not where the robot is in the maze, not where the checkpoints are, not the reward
Agent Budget 3600 s of wall clock per mode

From https://github.com/carlosferrazza/humanoid-bench @ cb11890, as defined in our task definitions @ 7d6a7a44.

Tags

Why this task is interesting

The robot walks the corridors of a maze through checkpoints at three corners. Most of the reward is the checkpoint bonuses, which grow along the way, so upright walking with turns comes before anything else.

Capability notes

Not yet written.

Oracle demo review

No demo.

Discussion

Passes in unlimited mode with whole-body control walking a planned route through every checkpoint without touching a wall; limited mode never gets past the first stage. (@williamzhangNU)