HumanoidBench runs¶
Agent runs on the tasks of HumanoidBench, one run at a time.
How to read this page
A run is one agent configuration, harness + model + settings; each task in it runs once per mode. Pick a run above; the table lists every task of it, one row per mode, in the default order (finished first, then running, easier tasks first). A pill links to that trial's log page. Filter the table with the chips above it.
successfailederror / stalledrunninggradingqueuednot run
Bars count trials, one per task and mode, by state. Progress is the partial credit beside success, in [0, 1]: BEHAVIOR's 2026 challenge q_score, elsewhere the grader's final_reward. A success always counts as full progress, 1.0, whatever that key says (RoboLab's final_reward is 0 for a task without scored subtasks); mean progress averages it over the graded trials. Agent time adds up finished trials; running ones are shown apart. A benchmark can also declare grader metrics, continuous scores its grader writes (IoU, F1, a distance): they get a column each, a sort, and a run figure, the mean or the median of the graded trials.
Tokens are the model gateway's count of every request (input includes the cached part, output the reasoning part). ≈ $ is those tokens at the list prices in data/prices.yml, an estimate and never a bill; billed is what a provider charged, where it says. The two are never added together.
A benchmark's owner can exclude tasks from the final benchmark (excluded: in state/tasks/<benchmark>.yml). Their trials stay, last in a run's table under Excluded from the benchmark; the numbers count the benchmark's own tasks, and the Count switch adds the excluded ones.
Runs are registered per benchmark: how to register a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default
0 requests282.8M in (98% cached)2.2M out≈ $4.56 at list pricebilled: —
0 requests953.8M in (98% cached)6.6M out≈ $14.40 at list pricebilled: —
Run details
- Harness
- Two harnesses, by protocol. Before robot_coding_bench's protocol v1.0: Codex CLI 0.157.0 (HumanoidBench) or 0.158.0–0.159.2 (KinDER) as harbor_agents.codex_openrouter:CodexOpenRouter (dev/pingyue, commit c9449ce), Codex over OpenRouter with web search off, without the model gateway. On protocol v1.0 (2026-10-01): Harbor's built-in codex agent, Codex CLI 0.159.2, with OpenRouter registered as its own provider and web search disabled (scripts/<benchmark>/configs/codex-openrouter.yaml); from 2026-10-04 (HumanoidBench and KinDER's hardest tasks) started by scripts/run_agent_batch.sh, with the same Codex config passed as AGENT_CONFIG and no usage gateway. state/runs/<benchmark>.yml says which batch is which.
- Model
openrouter/openai/gpt-6-luna- Route
- OpenRouter, with the owner's API key (billed per token)
- Settings
- reasoning_effort medium, version 0.157.0–0.159.2
- Modes
- unlimited, limited
- Batches
hb-luna-0928hb-luna-v1-1001hb-luna-1004
A later batch replaces an earlier one's run of the tasks it reruns; the earlier one stays as history.- Scope
- 9 tasks in the run + 15 excluded from the benchmark (last in the table) · 0 not in it · 0 removed (lists at the end)
- Progress
- the grader's none, in [0, 1]; a success counts as 1.0
- Billing
- openrouter: billed per request by OpenRouter (the charge it reports in every response)
- List price
- openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
- Run id
codex-gpt6_luna-medium-openrouter·data/agents/codex-gpt6_luna-medium-openrouter.yml,state/runs/humanoidbench.yml- Results
- as collected 10-04 08:02, published with
make publish-runs
- Locomotion · Stair unlimited the agent ended itself at 48 min, its trajectory saved, judging that stair traversal stayed unsolved
| Task | Trial (its log page) | Progress | Resets used | Agent time | Requests | Tokens in / out | Est. cost |
|---|---|---|---|---|---|---|---|
| Locomotion · Balance hard | U✗ 1h 00m prev | — | — | 1h 00m | — | 31.3M / 216k | $0.478 |
| L✗ 2m prev | — | 0 | 2m | — | 398k / 15k | $0.015 | |
| Locomotion · Hurdle | U✗ 59m prev | — | — | 59m | — | 51.5M / 268k | $0.738 |
| L✗ 1m prev | — | 0 | 1m | — | 246k / 6k | $0.0080 | |
| Locomotion · Stair | U✗ 48m ⓘ prev | — | — | 48m | — | 27.6M / 226k | $0.448 |
| L✗ 2m prev | — | 0 | 2m | — | 525k / 14k | $0.016 | |
| Manipulation · Bookshelf simple | U✗ 1h 00m prev | — | — | 1h 00m | — | 29.5M / 264k | $0.487 |
| L✗ 7m prev | — | 0 | 7m | — | 2.2M / 36k | $0.045 | |
| Manipulation · Highbar simple | U✗ 59m prev | — | — | 59m | — | 24.5M / 162k | $0.377 |
| L✗ 7m prev | — | 0 | 7m | — | 3.0M / 40k | $0.059 | |
| Manipulation · Package | U✗ 1h 00m prev | — | — | 1h 00m | — | 25.8M / 242k | $0.469 |
| L✗ 8m prev | — | 0 | 8m | — | 3.7M / 42k | $0.066 | |
| Manipulation · Powerlift | U✗ 1h 00m prev | — | — | 1h 00m | — | 34.5M / 281k | $0.580 |
| L✗ 4m prev | — | 0 | 4m | — | 953k / 23k | $0.026 | |
| Manipulation · Room | U✗ 1h 00m prev | — | — | 1h 00m | — | 25.1M / 247k | $0.430 |
| L✗ 2m prev | — | 0 | 2m | — | 244k / 13k | $0.012 | |
| Manipulation · Truck | U✗ 59m prev | — | — | 59m | — | 19.6M / 110k | $0.266 |
| L✗ 7m prev | — | 0 | 7m | — | 2.2M / 38k | $0.047 | |
| excluded Excluded from the benchmark · 15 tasks, kept for the record; the benchmark's numbers leave them out | |||||||
| Locomotion · Balance simple | U✗ 59m | — | — | 59m | — | 22.9M / 157k | $0.328 |
| L✗ 16m prev | — | 50 | 16m | — | 5.0M / 62k | $0.092 | |
| Locomotion · Maze | U✗ 1h 00m | — | — | 1h 00m | — | 27.1M / 203k | $0.430 |
| L✗ 45m prev | — | 50 | 45m | — | 20.8M / 134k | $0.308 | |
| Locomotion · Reach | U✓ 59m | 1.00 | — | 59m | — | 26.5M / 151k | $0.362 |
| L✗ 20m prev | — | 0 | 20m | — | 5.5M / 71k | $0.107 | |
| Locomotion · Sit hard | U✓ 58m | 1.00 | — | 58m | — | 16.9M / 78k | $0.222 |
| L✓ 6m prev | 1.00 | 4 | 6m | — | 1.1M / 21k | $0.027 | |
| Locomotion · Sit simple | U✓ 59m | 1.00 | — | 59m | — | 32.7M / 111k | $0.403 |
| L✓ 3m | 1.00 | 0 | 3m | — | 583k / 15k | $0.017 | |
| Locomotion · Walk | U✗ 1h 00m | — | — | 1h 00m | — | 25.2M / 193k | $0.407 |
| L✗ 18m prev | — | 50 | 18m | — | 3.9M / 77k | $0.088 | |
| Manipulation · Basketball | U✗ 1h 00m | — | — | 1h 00m | — | 24.4M / 219k | $0.407 |
| L✗ 1h 00m prev | — | 43 | 1h 00m | — | 29.7M / 175k | $0.406 | |
| Manipulation · Cabinet | U✗ 58m | — | — | 58m | — | 18.0M / 177k | $0.317 |
| L✗ 1h 00m prev | — | 33 | 1h 00m | — | 36.1M / 141k | $0.451 | |
| Manipulation · Cube | U✗ 58m | — | — | 58m | — | 33.2M / 137k | $0.446 |
| L✗ 36m prev | — | 50 | 36m | — | 15.8M / 126k | $0.238 | |
| Manipulation · Door | U✗ 1h 00m | — | — | 1h 00m | — | 24.2M / 238k | $0.419 |
| L✗ 1h 00m prev | — | 36 | 1h 00m | — | 30.0M / 171k | $0.433 | |
| Manipulation · Insert normal | U✗ 1h 00m | — | — | 1h 00m | — | 28.4M / 232k | $0.458 |
| L✗ 1h 00m prev | — | 24 | 1h 00m | — | 32.3M / 139k | $0.411 | |
| Manipulation · Kitchen | U✗ 57m | — | — | 57m | — | 22.2M / 212k | $0.385 |
| L✗ 1h 00m prev | — | 35 | 1h 00m | — | 30.7M / 169k | $0.414 | |
| Manipulation · Push | U✓ 36m prev | 1.00 | — | 36m | — | 30.0M / 131k | $0.387 |
| L✓ 24m prev | 1.00 | 10 | 24m | — | 8.1M / 55k | $0.120 | |
| Manipulation · Spoon | U✗ 1h 00m | — | — | 1h 00m | — | 21.5M / 239k | $0.391 |
| L✗ 1h 00m prev | — | 24 | 1h 00m | — | 32.7M / 162k | $0.455 | |
| Manipulation · Window | U✗ 58m | — | — | 58m | — | 32.6M / 228k | $0.494 |
| L✗ 1h 00m prev | — | 37 | 1h 00m | — | 32.9M / 133k | $0.414 | |
Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)
2,199 requests217.7M in (98% cached)1.9M out (1.3M reasoning)≈ $49.40 at list pricebilled: $51.64
5,725 requests555.6M in (98% cached)4.5M out (3.1M reasoning)≈ $119.63 at list pricebilled: $124.58
Run details
- Harness
- Codex CLI 0.159.2 as harbor_agents.codex_openrouter:CodexOpenRouter (robot_coding_bench dev/pingyue, commit e1f8660f): Codex over OpenRouter with web search off; the API key stays in the model gateway, the agent's container only sees a placeholder
- Model
openrouter/openai/gpt-6.1-sol- Route
- OpenRouter, with the owner's API key (billed per token), through our model gateway
- Settings
- reasoning_effort medium, version 0.159.2
- Modes
- unlimited, limited
- Batches
hb-sol61-0930- Scope
- 9 tasks in the run + 15 excluded from the benchmark (last in the table) · 0 not in it · 0 removed (lists at the end)
- Progress
- the grader's none, in [0, 1]; a success counts as 1.0
- Billing
- openrouter: billed per request by OpenRouter (the charge it reports in every response)
- List price
- openai/gpt-6.1-sol: $2.0 in / $0.1 cached / $10.0 out per 1M tokens, as of 2026-09-30
- Run id
codex-gpt6_1_sol-medium·data/agents/codex-gpt6_1_sol-medium.yml,state/runs/humanoidbench.yml- Results
- as collected 10-04 08:02, published with
make publish-runs
| Task | Trial (its log page) | Progress | Resets used | Agent time | Requests | Tokens in / out | Est. cost | Billed |
|---|---|---|---|---|---|---|---|---|
| Locomotion · Balance hard | U✗ 1h 00m | — | — | 1h 00m | 123 | 12.5M / 125k | $2.90 | $3.01 |
| L✗ 57m | — | 50 | 57m | 127 | 10.2M / 112k | $2.49 | $2.58 | |
| Locomotion · Hurdle | U✗ 59m | — | — | 59m | 107 | 8.7M / 96k | $2.15 | $2.23 |
| L✗ 30m | — | 50 | 30m | 73 | 5.5M / 62k | $1.45 | $1.52 | |
| Locomotion · Stair | U✗ 1h 00m | — | — | 1h 00m | 132 | 12.4M / 125k | $2.85 | $2.95 |
| L✗ 1h 00m | — | 49 | 1h 00m | 94 | 7.5M / 123k | $2.33 | $2.43 | |
| Manipulation · Bookshelf simple | U✗ 1h 00m | — | — | 1h 00m | 137 | 14.8M / 120k | $3.08 | $3.18 |
| L✗ 1h 00m | — | 50 | 1h 00m | 168 | 18.3M / 92k | $3.16 | $3.26 | |
| Manipulation · Highbar simple | U✗ 1h 00m | — | — | 1h 00m | 110 | 9.7M / 109k | $2.41 | $2.50 |
| L✗ 1h 00m | — | 50 | 1h 00m | 137 | 14.4M / 112k | $3.66 | $3.95 | |
| Manipulation · Package | U✗ 1h 00m | — | — | 1h 00m | 130 | 12.4M / 123k | $2.87 | $2.98 |
| L✗ 1h 00m | — | 44 | 1h 00m | 124 | 15.0M / 104k | $2.98 | $3.10 | |
| Manipulation · Powerlift | U✗ 1h 00m | — | — | 1h 00m | 137 | 14.0M / 111k | $2.89 | $2.99 |
| L✗ 1h 00m | — | 48 | 1h 00m | 145 | 16.5M / 95k | $3.53 | $3.78 | |
| Manipulation · Room | U✗ 1h 00m | — | — | 1h 00m | 106 | 9.6M / 121k | $2.54 | $2.64 |
| L✗ 54m | — | 50 | 54m | 127 | 12.3M / 87k | $2.44 | $2.53 | |
| Manipulation · Truck | U✗ 1h 00m | — | — | 1h 00m | 120 | 13.3M / 108k | $2.80 | $2.91 |
| L✗ 50m | — | 50 | 50m | 102 | 10.4M / 87k | $2.87 | $3.12 | |
| excluded Excluded from the benchmark · 15 tasks, kept for the record; the benchmark's numbers leave them out | ||||||||
| Locomotion · Balance simple | U✓ 1h 00m | 1.00 | — | 1h 00m | 137 | 10.5M / 68k | $1.99 | $2.06 |
| L✗ 52m | — | 50 | 52m | 117 | 9.1M / 106k | $2.28 | $2.36 | |
| Locomotion · Maze | U✓ 1h 00m | 1.00 | — | 1h 00m | 117 | 12.3M / 84k | $2.39 | $2.48 |
| L✗ 54m | — | 50 | 54m | 108 | 9.6M / 108k | $2.39 | $2.48 | |
| Locomotion · Reach | U✓ 59m | 1.00 | — | 59m | 124 | 10.5M / 107k | $2.46 | $2.55 |
| L✓ 16m | 1.00 | 29 | 16m | 56 | 4.6M / 32k | $1.06 | $1.13 | |
| Locomotion · Sit hard | U✓ 1h 00m | 1.00 | — | 1h 00m | 106 | 7.8M / 88k | $1.94 | $2.02 |
| L✓ 7m | 1.00 | 11 | 7m | 43 | 1.4M / 15k | $0.395 | $0.421 | |
| Locomotion · Sit simple | U✓ 1h 00m | 1.00 | — | 1h 00m | 148 | 13.2M / 73k | $2.32 | $2.39 |
| L✓ 1m | 1.00 | 0 | 1m | 8 | 165k / 2k | $0.087 | $0.101 | |
| Locomotion · Walk | U✓ 1h 00m | 1.00 | — | 1h 00m | 133 | 11.9M / 87k | $2.35 | $2.42 |
| L✗ 1h 00m | — | 49 | 1h 00m | 106 | 8.6M / 116k | $2.32 | $2.40 | |
| Manipulation · Basketball | U✓ 1h 00m | 1.00 | — | 1h 00m | 157 | 14.1M / 105k | $2.79 | $2.87 |
| L✗ 58m | — | 50 | 58m | 160 | 14.6M / 81k | $2.58 | $2.67 | |
| Manipulation · Cabinet | U✗ 1h 00m | — | — | 1h 00m | 130 | 15.1M / 107k | $3.02 | $3.13 |
| L✗ 1h 00m | — | 27 | 1h 00m | 131 | 12.6M / 101k | $2.69 | $2.79 | |
| Manipulation · Cube | U✓ 1h 00m | 1.00 | — | 1h 00m | 121 | 10.6M / 99k | $2.37 | $2.45 |
| L✗ 1h 00m | — | 48 | 1h 00m | 171 | 19.5M / 100k | $3.37 | $3.48 | |
| Manipulation · Door | U✗ 1h 00m | — | — | 1h 00m | 124 | 15.3M / 109k | $3.09 | $3.21 |
| L✗ 1h 00m | — | 50 | 1h 00m | 121 | 12.7M / 117k | $3.38 | $3.63 | |
| Manipulation · Insert normal | U✗ 1h 00m | — | — | 1h 00m | 125 | 14.2M / 117k | $3.00 | $3.11 |
| L✗ 1h 00m | — | 50 | 1h 00m | 133 | 13.6M / 96k | $2.70 | $2.80 | |
| Manipulation · Kitchen | U✓ 1h 00m | 1.00 | — | 1h 00m | 135 | 13.8M / 102k | $2.76 | $2.85 |
| L✗ 1h 00m | — | 16 | 1h 00m | 131 | 12.9M / 94k | $2.62 | $2.73 | |
| Manipulation · Push | U✓ 1h 00m | 1.00 | — | 1h 00m | 119 | 9.9M / 88k | $2.16 | $2.24 |
| L✓ 14m | 1.00 | 15 | 14m | 43 | 2.2M / 24k | $0.625 | $0.670 | |
| Manipulation · Spoon | U✗ 1h 00m | — | — | 1h 00m | 117 | 13.7M / 110k | $2.87 | $2.98 |
| L✗ 1h 00m | — | 40 | 1h 00m | 161 | 17.2M / 81k | $2.94 | $3.05 | |
| Manipulation · Window | U✗ 59m | — | — | 59m | 121 | 14.1M / 112k | $2.95 | $3.06 |
| L✗ 51m | — | 50 | 51m | 123 | 12.0M / 78k | $2.32 | $2.41 | |
Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)
0 requests26.6M in (98% cached)307k out≈ $24.32 at list pricebilled: —
Run details
- Harness
- Claude Code 2.1.283 as harbor_agents.claude_code_openrouter:ClaudeCodeOpenRouter (robot_coding_bench dev/pingyue, commit c9449ce): Claude Code over OpenRouter, without the model gateway; in limited mode the adapter ends the trial when the session ends (harbor_agents/limited.py)
- Model
openrouter/anthropic/claude-opus-5.5- Route
- OpenRouter, with the owner's API key (billed per token)
- Settings
- reasoning_effort medium, version 2.1.283
- Modes
- unlimited, limited
- Batches
hb-opus-0928- Scope
- 0 tasks in the run + 5 excluded from the benchmark (last in the table) · 19 not in it · 0 removed (lists at the end)
- Progress
- the grader's none, in [0, 1]; a success counts as 1.0
- Billing
- openrouter: billed per request by OpenRouter (the charge it reports in every response)
- List price
- anthropic/claude-opus-5.5: $6.25 in / $0.5 cached / $25.0 out per 1M tokens, as of 2026-09-28
- Run id
claude_code-claude_opus5_5-medium·data/agents/claude_code-claude_opus5_5-medium.yml,state/runs/humanoidbench.yml- Results
- as collected 10-04 08:02, published with
make publish-runs
- Manipulation · Cube unlimited the agent ended itself at 17 min, believing its time was nearly up; it never checked the clock
- Manipulation · Cube limited the agent ended itself at 6 min with 13 of its 50 resets left
- Manipulation · Door limited `spec()` placed the hand 0.31 m beyond the real one (a massless link's centre of mass; fixed since); taking it for the fingertip, the agent concluded its arms pass through the door and ended itself at 14 min with 16 resets left
| Task | Trial (its log page) | Progress | Resets used | Agent time | Requests | Tokens in / out | Est. cost |
|---|---|---|---|---|---|---|---|
| excluded Excluded from the benchmark · 5 tasks, kept for the record; the benchmark's numbers leave them out | |||||||
| Locomotion · Sit hard | U✓ 5m | 1.00 | — | 5m | — | 829k / 14k | $0.939 |
| L✓ 6m | 1.00 | 21 | 6m | — | 1.5M / 22k | $1.55 | |
| Locomotion · Walk | U✓ 11m | 1.00 | — | 11m | — | 574k / 11k | $0.726 |
| L✗ 14m | — | 50 | 14m | — | 4.6M / 50k | $4.14 | |
| Manipulation · Cube | U✗ 17m ⓘ | — | — | 17m | — | 1.5M / 23k | $1.57 |
| L✗ 6m ⓘ | — | 37 | 6m | — | 1.4M / 20k | $1.44 | |
| Manipulation · Door | U✓ 49m | 1.00 | — | 49m | — | 5.8M / 55k | $4.81 |
| L✗ 14m ⓘ | — | 34 | 14m | — | 4.4M / 55k | $4.07 | |
| Manipulation · Push | U✓ 4m | 1.00 | — | 4m | — | 863k / 14k | $0.992 |
| L✓ 13m | 1.00 | 9 | 13m | — | 5.1M / 43k | $4.09 | |
19 task(s) not in this run
| Task | Why |
|---|---|
| Locomotion · Balance hard | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Locomotion · Hurdle | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Locomotion · Stair | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Manipulation · Bookshelf simple | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Manipulation · Highbar simple | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Manipulation · Package | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Manipulation · Powerlift | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Manipulation · Room | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Manipulation · Truck | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Locomotion · Balance simple | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Locomotion · Maze | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Locomotion · Reach | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Locomotion · Sit simple | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Manipulation · Basketball | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Manipulation · Cabinet | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Manipulation · Insert normal | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Manipulation · Kitchen | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Manipulation · Spoon | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Manipulation · Window | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed
0 requests56.9M in (99% cached)302k out≈ $16.15 at list pricebilled: —
Run details
- Harness
- Codex CLI 0.157.0 as harbor_agents.codex_openrouter:CodexOpenRouter (robot_coding_bench dev/pingyue, commit c9449ce): Codex over OpenRouter with web search off, without the model gateway; in limited mode the adapter ends the trial when the session ends (harbor_agents/limited.py)
- Model
openrouter/openai/gpt-6-sol- Route
- OpenRouter, with the owner's API key (billed per token)
- Settings
- reasoning_effort medium, version 0.157.0
- Modes
- unlimited, limited
- Batches
hb-sol-0928- Scope
- 0 tasks in the run + 5 excluded from the benchmark (last in the table) · 19 not in it · 0 removed (lists at the end)
- Progress
- the grader's none, in [0, 1]; a success counts as 1.0
- Billing
- openrouter: billed per request by OpenRouter (the charge it reports in every response)
- List price
- openai/gpt-6-sol: $2.5 in / $0.2 cached / $10.0 out per 1M tokens, as of 2026-09-28
- Closed
- the OpenRouter credit ran out on 2026-09-28 before the unlimited runs on door, push and cube
- Run id
codex-gpt6_sol-medium·data/agents/codex-gpt6_sol-medium.yml,state/runs/humanoidbench.yml- Results
- as collected 10-04 08:02, published with
make publish-runs
- Manipulation · Door limited `spec()` placed the hand 0.31 m beyond the real one (a massless link's centre of mass; fixed since), and the agent aimed its hand at the handle with it
| Task | Trial (its log page) | Progress | Resets used | Agent time | Requests | Tokens in / out | Est. cost |
|---|---|---|---|---|---|---|---|
| excluded Excluded from the benchmark · 5 tasks, kept for the record; the benchmark's numbers leave them out | |||||||
| Locomotion · Sit hard | U✓ 1h 00m | 1.00 | — | 1h 00m | — | 18.6M / 65k | $4.86 |
| L✗ 10m | — | 50 | 10m | — | 1.6M / 18k | $0.616 | |
| Locomotion · Walk | U✗ 58m | — | — | 58m | — | 17.7M / 83k | $4.74 |
| L✗ 12m | — | 50 | 12m | — | 4.2M / 21k | $1.26 | |
| Manipulation · Cube | Unot run | — | — | — | — | — | — |
| L✗ 11m | — | 50 | 11m | — | 1.7M / 22k | $0.668 | |
| Manipulation · Door | Unot run | — | — | — | — | — | — |
| L✗ 36m ⓘ | — | 50 | 36m | — | 8.3M / 66k | $2.60 | |
| Manipulation · Push | Unot run | — | — | — | — | — | — |
| L✓ 18m | 1.00 | 7 | 18m | — | 4.8M / 26k | $1.41 | |
19 task(s) not in this run
| Task | Why |
|---|---|
| Locomotion · Balance hard | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Locomotion · Hurdle | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Locomotion · Stair | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Manipulation · Bookshelf simple | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Manipulation · Highbar simple | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Manipulation · Package | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Manipulation · Powerlift | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Manipulation · Room | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Manipulation · Truck | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) |
| Locomotion · Balance simple | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Locomotion · Maze | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Locomotion · Reach | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Locomotion · Sit simple | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Manipulation · Basketball | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Manipulation · Cabinet | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Manipulation · Insert normal | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Manipulation · Kitchen | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Manipulation · Spoon | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |
| Manipulation · Window | not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark |