Runs¶
Every agent run, added up over the benchmarks that take part in it.
How to read this page
A run is one agent configuration, harness + model + settings; each task in it runs once per mode. Each card is one run, added up over the benchmarks that take part in it; a benchmark's name opens its page for that run.
successfailederror / stalledrunninggradingqueuednot run
Bars count trials, one per task and mode, by state. Progress is the partial credit beside success, in [0, 1]: BEHAVIOR's 2026 challenge q_score, elsewhere the grader's final_reward. A success always counts as full progress, 1.0, whatever that key says (RoboLab's final_reward is 0 for a task without scored subtasks); mean progress averages it over the graded trials. Agent time adds up finished trials; running ones are shown apart. A benchmark can also declare grader metrics, continuous scores its grader writes (IoU, F1, a distance): they get a column each, a sort, and a run figure, the mean or the median of the graded trials.
Tokens are the model gateway's count of every request (input includes the cached part, output the reasoning part). ≈ $ is those tokens at the list prices in data/prices.yml, an estimate and never a bill; billed is what a provider charged, where it says. The two are never added together.
A benchmark's owner can exclude tasks from the final benchmark (excluded: in state/tasks/<benchmark>.yml). Their trials stay, last in a run's table under Excluded from the benchmark; the numbers count the benchmark's own tasks, and the Count switch adds the excluded ones.
Runs are registered per benchmark: how to register a run.
Results as of 10-04 13:42
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)
102,689 requests13.85B in (98% cached)80.9M out (52.1M reasoning)≈ $203.86 at list pricebilled: none (subscription)
103,210 requests13.87B in (98% cached)81.4M out (52.4M reasoning)≈ $204.46 at list pricebilled: none (subscription)
Codex + GPT-6 Luna, reasoning xhigh (Azure OpenAI API, 2026-09-30 stress test)
0 requests1.29B in (98% cached)12.5M out≈ $21.37 at list pricebilled: —
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)
0 requests502.1M in (98% cached)4.3M out≈ $8.25 at list pricebilled: —
0 requests1.36B in (98% cached)10.1M out≈ $20.94 at list pricebilled: —
Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)
2,800 requests257.8M in (98% cached)2.2M out (1.5M reasoning)≈ $57.73 at list pricebilled: $60.39
6,741 requests610.2M in (98% cached)4.9M out (3.3M reasoning)≈ $131.93 at list pricebilled: $137.64
Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)
0 requests26.6M in (98% cached)307k out≈ $24.32 at list pricebilled: —
Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed
0 requests56.9M in (99% cached)302k out≈ $16.15 at list pricebilled: —
Codex + GPT-6 Luna, reasoning medium (ChatGPT login)closed
0 requests120.0M in (97% cached)784k out≈ $1.90 at list pricebilled: none (subscription)
Runs by benchmark¶
Every number: each benchmark in each run
| Benchmark | Run | In run | Not in / removed | Excluded (not counted) | Unlimited done / success | Limited done / success | Running | Queued | Failed grading / stalled | Mean progress | Grader metrics | Agent-hours | Requests | Tokens in / cached | Tokens out / reasoning | Est. cost | Billed | Results as of |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BEHAVIOR-1K | Codex · GPT-6 Luna · xhigh · ChatGPT login | 49 | 51 / 0 | — | 49/49 · 9 (18%) | 49/49 · 0 (0%) | 0 | 0 | 0 | 0.13 | — | 360.4 h | 52,654 | 6.81B / 6.69B | 28.2M / 21.5M | $92.76 | none (subscription) | 09-28 21:10 |
| DexToolBench | Codex · GPT-6 Luna · xhigh · ChatGPT login | 21 | 3 / 0 | — | 21/21 · 1 (5%) | 21/21 · 0 (0%) | 0 | 0 | 0 | 0.06 | Goals reached 3 | 40.1 h | 0 | 699.1M / 684.5M | 6.0M / 0 | $11.30 | none (subscription) | 10-01 17:29 |
| DexToolBench | Codex · GPT-6 Luna · xhigh · Azure API stress test | 21 | 3 / 0 | — | 21/21 · 1 (5%) | 21/21 · 0 (0%) | 0 | 0 | 0 | 0.06 | Goals reached 2 | 40.1 h | 0 | 761.8M / 747.6M | 7.5M / 0 | $12.66 | — | 10-01 17:29 |
| HumanoidBench | Codex CLI 0.157–0.159 · GPT-6 Luna · medium · OpenRouter | 9 | 0 / 0 | 15 tasks · 30/30 done · 7 ✓ · $9.84 est. | 9/9 · 0 (0%) | 9/9 · 0 (0%) | 0 | 0 | 0 | — | Resets used 0 | 9.4 h | 0 | 282.8M / 276.0M | 2.2M / 0 | $4.56 | — | 10-04 08:02 |
| HumanoidBench | Codex CLI 0.159 · GPT-6.1 Sol · medium · OpenRouter | 9 | 0 / 0 | 15 tasks · 30/30 done · 14 ✓ · $70.23 est. | 9/9 · 0 (0%) | 9/9 · 0 (0%) | 0 | 0 | 0 | — | Resets used 49 | 17.2 h | 2,199 | 217.7M / 213.2M | 1.9M / 1.3M | $49.40 | $51.64 | 10-04 08:02 |
| HumanoidBench | Claude Code 2.1 · Claude Opus 5.5 · medium · OpenRouter | 0 | 19 / 0 | 5 tasks · 10/10 done · 6 ✓ · $24.32 est. | — | — | 0 | 0 | 0 | — | — | — | 0 | 0 / 0 | 0 / 0 | — | — | 10-04 08:02 |
| HumanoidBench | Codex CLI 0.157 · GPT-6 Sol · medium · OpenRouter closed | 0 | 19 / 0 | 5 tasks · 7/10 done · 2 ✓ · $16.15 est. | — | — | 0 | 0 | 0 | — | — | — | 0 | 0 / 0 | 0 / 0 | — | — | 10-04 08:02 |
| KinDER | Codex CLI 0.157–0.159 · GPT-6 Luna · medium · OpenRouter | 4 | 0 / 0 | 7 tasks · 14/14 done · 8 ✓ · $2.85 est. | 4/4 · 2 (50%) | 4/4 · 0 (0%) | 0 | 0 | 0 | 1.00 | — | 6.0 h | 0 | 154.8M / 151.3M | 1.4M / 0 | $2.57 | — | 10-04 08:03 |
| KinDER | Codex CLI 0.159 · GPT-6.1 Sol · medium · OpenRouter | 4 | 0 / 0 | 7 tasks · 14/14 done · 14 ✓ · $3.97 est. | 4/4 · 4 (100%) | 4/4 · 4 (100%) | 0 | 0 | 0 | 1.00 | — | 3.3 h | 601 | 40.1M / 39.3M | 273k / 166k | $8.33 | $8.75 | 10-04 08:03 |
| MetaWorld+ | Codex · GPT-6 Luna · xhigh · ChatGPT login closed | 4 | 46 / 0 | — | 4/4 · 4 (100%) | 4/4 · 1 (25%) | 0 | 0 | 0 | 1.00 | — | 2.7 h | 0 | 30.3M / 29.0M | 387k / 0 | $0.607 | none (subscription) | 10-02 15:07 |
| MolmoSpaces | Codex CLI 0.157–0.159 · GPT-6 Luna · medium · OpenRouter closed | 8 | 0 / 0 | — | 8/8 · 7 (88%) | 8/8 · 0 (0%) | 0 | 0 | 0 | 1.00 | — | 4.1 h | 0 | 64.5M / 63.0M | 666k / 0 | $1.11 | — | 10-04 13:42 |
| MuJoCo Playground (manipulation) | Codex · GPT-6 Luna · xhigh · ChatGPT login | 4 | 6 / 0 | — | 4/4 · 3 (75%) | 4/4 · 1 (25%) | 0 | 0 | 0 | 1.00 | — | 4.7 h | 0 | 82.0M / 80.2M | 839k / 0 | $1.40 | none (subscription) | 10-01 17:29 |
| MuJoCo Playground (manipulation) | Codex · GPT-6 Luna · xhigh · Azure API stress test | 4 | 6 / 0 | — | 4/4 · 4 (100%) | 4/4 · 0 (0%) | 0 | 0 | 0 | 1.00 | — | 4.6 h | 0 | 39.1M / 38.2M | 467k / 0 | $0.704 | — | 10-01 17:29 |
| RoboCasa-GR1 | Codex · GPT-6 Luna · xhigh · ChatGPT login closed | 4 | 20 / 0 | — | 4/4 · 2 (50%) | 4/4 · 0 (0%) | 0 | 0 | 0 | 1.00 | — | 6.6 h | 0 | 113.8M / 111.4M | 840k / 0 | $1.78 | none (subscription) | 10-02 15:07 |
| RoboCasa | Codex · GPT-6 Luna · xhigh · ChatGPT login closed | 4 | 25 / 0 | — | 4/4 · 4 (100%) | 4/4 · 0 (0%) | 0 | 0 | 0 | 1.00 | — | 5.7 h | 0 | 103.6M / 101.5M | 649k / 0 | $1.55 | none (subscription) | 10-02 15:07 |
| RoboCasa365 | Codex · GPT-6 Luna · xhigh · ChatGPT login closed | 4 | 46 / 0 | — | 4/4 · 3 (75%) | 4/4 · 0 (0%) | 0 | 0 | 0 | 1.00 | — | 5.4 h | 0 | 71.9M / 69.9M | 660k / 0 | $1.22 | none (subscription) | 10-02 15:07 |
| RoboLab | Codex · GPT-6 Luna · xhigh · ChatGPT login | 88 | 0 / 32 | — | 88/88 · 49 (56%) | 88/88 · 25 (28%) | 0 | 0 | 0 | 0.84 | — | 183.3 h | 25,767 | 2.89B / 2.82B | 16.8M / 12.9M | $42.95 | none (subscription) | 09-30 05:08 |
| RoboPaint | Codex · GPT-6 Luna · xhigh · ChatGPT login | 82 | 0 / 0 | 9 tasks · 18/18 done · 18 ✓ · $0.564 est. | 82/82 · 49 (60%) | 82/82 · 1 (1%) | 0 | 0 | 0 | 0.71 | Final reward 0.654 · IoU 0.706 | 166.6 h | 16,037 | 1.75B / 1.70B | 17.0M / 13.0M | $29.81 | none (subscription) | 09-30 05:08 |
| RoboTwin 2.0 | Codex · GPT-6 Luna · xhigh · ChatGPT login | 25 | 25 / 0 | — | 25/25 · 25 (100%) | 25/25 · 8 (32%) | 0 | 0 | 0 | 1.00 | — | 22.5 h | 0 | 371.3M / 363.8M | 2.8M / 0 | $5.79 | none (subscription) | 10-01 17:29 |
| RoboTwin 2.0 | Codex · GPT-6 Luna · medium · ChatGPT login closed | 22 | 28 / 0 | — | 21/22 · 11 (52%) | 22/22 · 3 (14%) | 0 | 1 | 0 | 1.00 | — | 8.9 h | 0 | 120.0M / 116.5M | 784k / 0 | $1.90 | none (subscription) | 10-01 17:29 |
| RoboTwin 2.0 | Codex · GPT-6 Luna · xhigh · Azure API stress test | 47 | 3 / 0 | — | 47/47 · 42 (89%) | 47/47 · 20 (43%) | 0 | 0 | 0 | 1.00 | — | 36.2 h | 0 | 489.0M / 479.2M | 4.5M / 0 | $8.00 | — | 10-01 17:29 |
| RoboWits | Codex · GPT-6 Luna · xhigh · ChatGPT login | 29 | 0 / 0 | 1 tasks · 2/2 done · 0 ✓ · $0.033 est. | 27/27 · 16 (59%) | 29/29 · 8 (28%) | 0 | 0 | 0 | 0.87 | — | 63.6 h | 8,231 | 893.8M / 873.8M | 6.3M / 4.8M | $13.90 | none (subscription) | 09-30 05:08 |
| VLABench | Codex · GPT-6 Luna · xhigh · ChatGPT login closed | 4 | 32 / 0 | — | 4/4 · 4 (100%) | 4/4 · 2 (50%) | 0 | 0 | 0 | 1.00 | — | 3.6 h | 0 | 47.0M / 45.8M | 404k / 0 | $0.784 | none (subscription) | 10-02 15:07 |