Skip to content

KinDER runs

Agent runs on the tasks of KinDER, one run at a time.

How to read this page

A run is one agent configuration, harness + model + settings; each task in it runs once per mode. Pick a run above; the table lists every task of it, one row per mode, in the default order (finished first, then running, easier tasks first). A pill links to that trial's log page. Filter the table with the chips above it.

successfailederror / stalledrunninggradingqueuednot run

Bars count trials, one per task and mode, by state. Progress is the partial credit beside success, in [0, 1]: BEHAVIOR's 2026 challenge q_score, elsewhere the grader's final_reward. A success always counts as full progress, 1.0, whatever that key says (RoboLab's final_reward is 0 for a task without scored subtasks); mean progress averages it over the graded trials. Agent time adds up finished trials; running ones are shown apart. A benchmark can also declare grader metrics, continuous scores its grader writes (IoU, F1, a distance): they get a column each, a sort, and a run figure, the mean or the median of the graded trials.

Tokens are the model gateway's count of every request (input includes the cached part, output the reasoning part). ≈ $ is those tokens at the list prices in data/prices.yml, an estimate and never a bill; billed is what a provider charged, where it says. The two are never added together.

A benchmark's owner can exclude tasks from the final benchmark (excluded: in state/tasks/<benchmark>.yml). Their trials stay, last in a run's table under Excluded from the benchmark; the numbers count the benchmark's own tasks, and the Count switch adds the excluded ones.

Runs are registered per benchmark: how to register a run.

Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default

11 tasks × 2 modes · 7 of them excluded from the benchmark · results as of 10-04 08:03

Count7 of the run's tasks are excluded from the benchmark: last in the table
Unlimited2 / 4 success
2 success2 failed4 done · 50%
Limited0 / 4 success
4 failed4 done · 0%
Mean progress1.00
none · 2 graded
Agent time6.0 h
finished trials

0 requests154.8M in (98% cached)1.4M out≈ $2.57 at list pricebilled: —

Run details
Harness
Before v1.0: Codex CLI 0.158.0 (the first six unlimited runs), 0.159.0 (the other five) and 0.159.2 (limited) as harbor_agents.codex_openrouter:CodexOpenRouter, Codex over OpenRouter with web search off, without the model gateway; in limited mode the adapter ends the trial when the session ends (harbor_agents/limited.py). On v1.0: Harbor's codex agent, Codex CLI 0.159.2, web search disabled (scripts/kinder/configs/codex-openrouter.yaml)
Model
openrouter/openai/gpt-6-luna
Route
OpenRouter, with the owner's API key (billed per token)
Settings
reasoning_effort medium, version 0.157.0–0.159.2
Modes
unlimited, limited
Batches
kinder-luna-0930 kinder-luna-v1-1001 kinder-luna-1004
A later batch replaces an earlier one's run of the tasks it reruns; the earlier one stays as history.
Scope
4 tasks in the run + 7 excluded from the benchmark (last in the table) · 0 not in it · 0 removed (lists at the end)
Progress
the grader's none, in [0, 1]; a success counts as 1.0
Billing
openrouter: billed per request by OpenRouter (the charge it reports in every response)
List price
openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
Run id
codex-gpt6_luna-medium-openrouter · data/agents/codex-gpt6_luna-medium-openrouter.yml, state/runs/kinder.yml
Results
as collected 10-04 08:03, published with make publish-runs
TaskTrial (its log page)ProgressAgent timeRequestsTokens in / outEst. cost
Dynamic3D · Constrained cupboardU✗ 1h 00m prev—1h 00m—13.4M / 284k$0.333
L✗ 23m prev—23m—14.4M / 125k$0.225
Dynamic3D · Scoop pourU✗ 1h 00m prev—1h 00m—15.5M / 147k$0.249
L✗ 43m prev—43m—33.8M / 179k$0.480
Dynamic3D · Sweep into drawerU✓ 50m prev1.0050m—10.7M / 176k$0.243
L✗ 43m prev—43m—36.4M / 194k$0.513
Dynamic3D · Sweep simpleU✓ 56m prev1.0056m—13.9M / 185k$0.284
L✗ 25m prev—25m—16.7M / 117k$0.244
excluded Excluded from the benchmark · 7 tasks, kept for the record; the benchmark's numbers leave them out
Dynamic3D · Balance beamU✓ 5m1.005m—1.4M / 23k$0.031
L✗ 32m prev—32m—8.5M / 114k$0.157
Dynamic3D · DynamoU✓ 4m prev1.004m—328k / 8k$0.0095
L✓ 7m1.007m—2.3M / 29k$0.044
Dynamic3D · ShelfU✓ 21m1.0021m—10.3M / 87k$0.161
L✗ 1h 00m prev—1h 00m—32.3M / 166k$0.428
Dynamic3D · Sort blocksU✓ 43m1.0043m—10.3M / 137k$0.192
L✗ 1h 00m prev—1h 00m—20.1M / 183k$0.340
Dynamic3D · TossingU✓ 37m1.0037m—17.0M / 150k$0.290
L✗ 32m prev—32m—20.3M / 83k$0.267
Kinematic3D · Obstruction 3DU✓ 11m1.0011m—3.3M / 41k$0.061
L✗ 1h 00m prev—1h 00m—34.0M / 153k$0.462
Kinematic3D · Packing 3DU✓ 9m1.009m—2.6M / 33k$0.050
L✗ 1h 00m prev—1h 00m—21.5M / 195k$0.362

Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)

11 tasks × 2 modes · 7 of them excluded from the benchmark · results as of 10-04 08:03

Count7 of the run's tasks are excluded from the benchmark: last in the table
Unlimited4 / 4 success
4 success4 done · 100%
Limited4 / 4 success
4 success4 done · 100%
Mean progress1.00
none · 8 graded
Agent time3.3 h
finished trials

601 requests40.1M in (98% cached)273k out (166k reasoning)≈ $8.33 at list pricebilled: $8.75

Run details
Harness
Codex CLI 0.159.2 as harbor_agents.codex_openrouter:CodexOpenRouter (robot_coding_bench dev/pingyue, commit e1f8660f): Codex over OpenRouter with web search off; the API key stays in the model gateway, the agent's container only sees a placeholder
Model
openrouter/openai/gpt-6.1-sol
Route
OpenRouter, with the owner's API key (billed per token), through our model gateway
Settings
reasoning_effort medium, version 0.159.2
Modes
unlimited, limited
Batches
kinder-sol61-0930
Scope
4 tasks in the run + 7 excluded from the benchmark (last in the table) · 0 not in it · 0 removed (lists at the end)
Progress
the grader's none, in [0, 1]; a success counts as 1.0
Billing
openrouter: billed per request by OpenRouter (the charge it reports in every response)
List price
openai/gpt-6.1-sol: $2.0 in / $0.1 cached / $10.0 out per 1M tokens, as of 2026-09-30
Run id
codex-gpt6_1_sol-medium · data/agents/codex-gpt6_1_sol-medium.yml, state/runs/kinder.yml
Results
as collected 10-04 08:03, published with make publish-runs
TaskTrial (its log page)ProgressAgent timeRequestsTokens in / outEst. costBilled
Dynamic3D · Constrained cupboardU✓ 9m1.009m351.8M / 12k$0.430$0.466
L✓ 45m1.0045m1057.8M / 72k$1.78$1.86
Dynamic3D · Scoop pourU✓ 47m1.0047m12011.4M / 73k$2.18$2.27
L✓ 19m1.0019m713.3M / 25k$0.733$0.774
Dynamic3D · Sweep into drawerU✓ 16m1.0016m522.9M / 20k$0.654$0.697
L✓ 13m1.0013m682.3M / 16k$0.504$0.533
Dynamic3D · Sweep simpleU✓ 27m1.0027m907.3M / 32k$1.29$1.35
L✓ 22m1.0022m603.4M / 24k$0.756$0.803
excluded Excluded from the benchmark · 7 tasks, kept for the record; the benchmark's numbers leave them out
Dynamic3D · Balance beamU✓ 4m1.004m16539k / 5k$0.193$0.217
L✓ 5m1.005m22533k / 8k$0.198$0.215
Dynamic3D · DynamoU✓ 1m1.001m9168k / 2k$0.086$0.099
L✓ 4m1.004m31749k / 5k$0.190$0.207
Dynamic3D · ShelfU✓ 6m1.006m24991k / 8k$0.294$0.324
L✓ 15m1.0015m703.0M / 14k$0.581$0.619
Dynamic3D · Sort blocksU✓ 7m1.007m361.3M / 11k$0.348$0.375
L✓ 10m1.0010m441.7M / 16k$0.450$0.481
Dynamic3D · TossingU✓ 6m1.006m27922k / 9k$0.273$0.296
L✓ 12m1.0012m481.8M / 14k$0.444$0.475
Kinematic3D · Obstruction 3DU✓ 3m1.003m16546k / 4k$0.192$0.216
L✓ 5m1.005m25721k / 9k$0.248$0.272
Kinematic3D · Packing 3DU✓ 4m1.004m18718k / 7k$0.243$0.270
L✓ 5m1.005m29753k / 8k$0.232$0.252