Skip to content

Runs

Every agent run, added up over the benchmarks that take part in it.

How to read this page

A run is one agent configuration, harness + model + settings; each task in it runs once per mode. Each card is one run, added up over the benchmarks that take part in it; a benchmark's name opens its page for that run.

successfailederror / stalledrunninggradingqueuednot run

Bars count trials, one per task and mode, by state. Progress is the partial credit beside success, in [0, 1]: BEHAVIOR's 2026 challenge q_score, elsewhere the grader's final_reward. A success always counts as full progress, 1.0, whatever that key says (RoboLab's final_reward is 0 for a task without scored subtasks); mean progress averages it over the graded trials. Agent time adds up finished trials; running ones are shown apart. A benchmark can also declare grader metrics, continuous scores its grader writes (IoU, F1, a distance): they get a column each, a sort, and a run figure, the mean or the median of the graded trials.

Tokens are the model gateway's count of every request (input includes the cached part, output the reasoning part). ≈ $ is those tokens at the list prices in data/prices.yml, an estimate and never a bill; billed is what a provider charged, where it says. The two are never added together.

A benchmark's owner can exclude tasks from the final benchmark (excluded: in state/tasks/<benchmark>.yml). Their trials stay, last in a run's table under Excluded from the benchmark; the numbers count the benchmark's own tasks, and the Count switch adds the excluded ones.

Runs are registered per benchmark: how to register a run.

Results as of 10-04 13:42

Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)

12 benchmarks · 318 tasks + 10 excluded · codex-gpt6_luna-xhigh

CountRoboPaint: 9 excluded from the benchmark · RoboWits: 1 excluded from the benchmark
Unlimited169 / 316 success
169 success147 failed316 done · 53%
Limited46 / 318 success
46 success272 failed318 done · 14%
Mean progress0.61
q_score / final_reward · 483 graded
Agent time865.1 h
finished trials

102,689 requests13.85B in (98% cached)80.9M out (52.1M reasoning)≈ $203.86 at list pricebilled: none (subscription)

↻ 1 follow-up arranged⊘ 2 not counted

Codex + GPT-6 Luna, reasoning xhigh (Azure OpenAI API, 2026-09-30 stress test)

3 benchmarks · 72 tasks · codex-gpt6_luna-xhigh-azure

Unlimited47 / 72 success
47 success25 failed72 done · 65%
Limited20 / 72 success
20 success52 failed72 done · 28%
Mean progress0.63
progress · 108 graded
Agent time81.0 h
finished trials

0 requests1.29B in (98% cached)12.5M out≈ $21.37 at list pricebilled: —

Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)

3 benchmarks · 21 tasks + 22 excluded · codex-gpt6_luna-medium-openrouter

CountHumanoidBench: 15 excluded from the benchmark · KinDER: 7 excluded from the benchmark
Unlimited9 / 21 success
9 success12 failed21 done · 43%
Limited0 / 21 success
21 failed21 done · 0%
Mean progress1.00
none · 9 graded
Agent time19.5 h
finished trials

0 requests502.1M in (98% cached)4.3M out≈ $8.25 at list pricebilled: —

Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)

2 benchmarks · 13 tasks + 22 excluded · codex-gpt6_1_sol-medium

CountHumanoidBench: 15 excluded from the benchmark · KinDER: 7 excluded from the benchmark
Unlimited4 / 13 success
4 success9 failed13 done · 31%
Limited4 / 13 success
4 success9 failed13 done · 31%
Mean progress1.00
none · 8 graded
Agent time20.5 h
finished trials

2,800 requests257.8M in (98% cached)2.2M out (1.5M reasoning)≈ $57.73 at list pricebilled: $60.39

Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)

1 benchmark · 0 tasks + 5 excluded · claude_code-claude_opus5_5-medium

CountHumanoidBench: 5 excluded from the benchmark

Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed

1 benchmark · 0 tasks + 5 excluded · codex-gpt6_sol-medium

CountHumanoidBench: 5 excluded from the benchmark

Codex + GPT-6 Luna, reasoning medium (ChatGPT login)closed

1 benchmark · 22 tasks · codex-gpt6_luna-medium

Unlimited11 / 22 success
11 success10 failed1 not run21 done · 52%
Limited3 / 22 success
3 success19 failed22 done · 14%
Mean progress1.00
final_reward · 14 graded
Agent time8.9 h
finished trials

0 requests120.0M in (97% cached)784k out≈ $1.90 at list pricebilled: none (subscription)

Runs by benchmark

BenchmarkCodex · GPT-6 Luna · xhigh · ChatGPT loginCodex · GPT-6 Luna · xhigh · Azure API stress testCodex CLI 0.157–0.159 · GPT-6 Luna · medium · OpenRouterCodex CLI 0.159 · GPT-6.1 Sol · medium · OpenRouterClaude Code 2.1 · Claude Opus 5.5 · medium · OpenRouterCodex CLI 0.157 · GPT-6 Sol · medium · OpenRouterCodex · GPT-6 Luna · medium · ChatGPT login
BEHAVIOR-1K9 ✓ · 98/98default——————
DexToolBench1 ✓ · 42/42default1 ✓ · 42/42—————
HumanoidBench——0 ✓ · 18/18+ 15 excludeddefault0 ✓ · 18/18+ 15 excluded0 ✓ · 0/0+ 5 excluded0 ✓ · 0/0+ 5 excludedclosed—
KinDER——2 ✓ · 8/8+ 7 excludeddefault8 ✓ · 8/8+ 7 excluded———
MetaWorld+5 ✓ · 8/8defaultclosed——————
MolmoSpaces——7 ✓ · 16/16defaultclosed————
MuJoCo Playground (manipulation)4 ✓ · 8/8default4 ✓ · 8/8—————
RoboCasa-GR12 ✓ · 8/8defaultclosed——————
RoboCasa4 ✓ · 8/8defaultclosed——————
RoboCasa3653 ✓ · 8/8defaultclosed——————
RoboLab74 ✓ · 176/176default——————
RoboPaint50 ✓ · 164/164+ 9 excludeddefault——————
RoboTwin 2.033 ✓ · 50/50default62 ✓ · 94/94————14 ✓ · 43/44closed
RoboWits24 ✓ · 56/56+ 1 excludeddefault——————
VLABench6 ✓ · 8/8defaultclosed——————
Every number: each benchmark in each run
BenchmarkRunIn runNot in / removedExcluded (not counted)Unlimited done / successLimited done / successRunningQueuedFailed grading / stalledMean progressGrader metricsAgent-hoursRequestsTokens in / cachedTokens out / reasoningEst. costBilledResults as of
BEHAVIOR-1KCodex · GPT-6 Luna · xhigh · ChatGPT login4951 / 0—49/49 · 9 (18%)49/49 · 0 (0%)0000.13—360.4 h52,6546.81B / 6.69B28.2M / 21.5M$92.76none (subscription)09-28 21:10
DexToolBenchCodex · GPT-6 Luna · xhigh · ChatGPT login213 / 0—21/21 · 1 (5%)21/21 · 0 (0%)0000.06Goals reached 340.1 h0699.1M / 684.5M6.0M / 0$11.30none (subscription)10-01 17:29
DexToolBenchCodex · GPT-6 Luna · xhigh · Azure API stress test213 / 0—21/21 · 1 (5%)21/21 · 0 (0%)0000.06Goals reached 240.1 h0761.8M / 747.6M7.5M / 0$12.66—10-01 17:29
HumanoidBenchCodex CLI 0.157–0.159 · GPT-6 Luna · medium · OpenRouter90 / 015 tasks · 30/30 done · 7 ✓ · $9.84 est.9/9 · 0 (0%)9/9 · 0 (0%)000—Resets used 09.4 h0282.8M / 276.0M2.2M / 0$4.56—10-04 08:02
HumanoidBenchCodex CLI 0.159 · GPT-6.1 Sol · medium · OpenRouter90 / 015 tasks · 30/30 done · 14 ✓ · $70.23 est.9/9 · 0 (0%)9/9 · 0 (0%)000—Resets used 4917.2 h2,199217.7M / 213.2M1.9M / 1.3M$49.40$51.6410-04 08:02
HumanoidBenchClaude Code 2.1 · Claude Opus 5.5 · medium · OpenRouter019 / 05 tasks · 10/10 done · 6 ✓ · $24.32 est.——000———00 / 00 / 0——10-04 08:02
HumanoidBenchCodex CLI 0.157 · GPT-6 Sol · medium · OpenRouter closed019 / 05 tasks · 7/10 done · 2 ✓ · $16.15 est.——000———00 / 00 / 0——10-04 08:02
KinDERCodex CLI 0.157–0.159 · GPT-6 Luna · medium · OpenRouter40 / 07 tasks · 14/14 done · 8 ✓ · $2.85 est.4/4 · 2 (50%)4/4 · 0 (0%)0001.00—6.0 h0154.8M / 151.3M1.4M / 0$2.57—10-04 08:03
KinDERCodex CLI 0.159 · GPT-6.1 Sol · medium · OpenRouter40 / 07 tasks · 14/14 done · 14 ✓ · $3.97 est.4/4 · 4 (100%)4/4 · 4 (100%)0001.00—3.3 h60140.1M / 39.3M273k / 166k$8.33$8.7510-04 08:03
MetaWorld+Codex · GPT-6 Luna · xhigh · ChatGPT login closed446 / 0—4/4 · 4 (100%)4/4 · 1 (25%)0001.00—2.7 h030.3M / 29.0M387k / 0$0.607none (subscription)10-02 15:07
MolmoSpacesCodex CLI 0.157–0.159 · GPT-6 Luna · medium · OpenRouter closed80 / 0—8/8 · 7 (88%)8/8 · 0 (0%)0001.00—4.1 h064.5M / 63.0M666k / 0$1.11—10-04 13:42
MuJoCo Playground (manipulation)Codex · GPT-6 Luna · xhigh · ChatGPT login46 / 0—4/4 · 3 (75%)4/4 · 1 (25%)0001.00—4.7 h082.0M / 80.2M839k / 0$1.40none (subscription)10-01 17:29
MuJoCo Playground (manipulation)Codex · GPT-6 Luna · xhigh · Azure API stress test46 / 0—4/4 · 4 (100%)4/4 · 0 (0%)0001.00—4.6 h039.1M / 38.2M467k / 0$0.704—10-01 17:29
RoboCasa-GR1Codex · GPT-6 Luna · xhigh · ChatGPT login closed420 / 0—4/4 · 2 (50%)4/4 · 0 (0%)0001.00—6.6 h0113.8M / 111.4M840k / 0$1.78none (subscription)10-02 15:07
RoboCasaCodex · GPT-6 Luna · xhigh · ChatGPT login closed425 / 0—4/4 · 4 (100%)4/4 · 0 (0%)0001.00—5.7 h0103.6M / 101.5M649k / 0$1.55none (subscription)10-02 15:07
RoboCasa365Codex · GPT-6 Luna · xhigh · ChatGPT login closed446 / 0—4/4 · 3 (75%)4/4 · 0 (0%)0001.00—5.4 h071.9M / 69.9M660k / 0$1.22none (subscription)10-02 15:07
RoboLabCodex · GPT-6 Luna · xhigh · ChatGPT login880 / 32—88/88 · 49 (56%)88/88 · 25 (28%)0000.84—183.3 h25,7672.89B / 2.82B16.8M / 12.9M$42.95none (subscription)09-30 05:08
RoboPaintCodex · GPT-6 Luna · xhigh · ChatGPT login820 / 09 tasks · 18/18 done · 18 ✓ · $0.564 est.82/82 · 49 (60%)82/82 · 1 (1%)0000.71Final reward 0.654 · IoU 0.706166.6 h16,0371.75B / 1.70B17.0M / 13.0M$29.81none (subscription)09-30 05:08
RoboTwin 2.0Codex · GPT-6 Luna · xhigh · ChatGPT login2525 / 0—25/25 · 25 (100%)25/25 · 8 (32%)0001.00—22.5 h0371.3M / 363.8M2.8M / 0$5.79none (subscription)10-01 17:29
RoboTwin 2.0Codex · GPT-6 Luna · medium · ChatGPT login closed2228 / 0—21/22 · 11 (52%)22/22 · 3 (14%)0101.00—8.9 h0120.0M / 116.5M784k / 0$1.90none (subscription)10-01 17:29
RoboTwin 2.0Codex · GPT-6 Luna · xhigh · Azure API stress test473 / 0—47/47 · 42 (89%)47/47 · 20 (43%)0001.00—36.2 h0489.0M / 479.2M4.5M / 0$8.00—10-01 17:29
RoboWitsCodex · GPT-6 Luna · xhigh · ChatGPT login290 / 01 tasks · 2/2 done · 0 ✓ · $0.033 est.27/27 · 16 (59%)29/29 · 8 (28%)0000.87—63.6 h8,231893.8M / 873.8M6.3M / 4.8M$13.90none (subscription)09-30 05:08
VLABenchCodex · GPT-6 Luna · xhigh · ChatGPT login closed432 / 0—4/4 · 4 (100%)4/4 · 2 (50%)0001.00—3.6 h047.0M / 45.8M404k / 0$0.784none (subscription)10-02 15:07