Skip to content

MolmoSpaces runs

Agent runs on the tasks of MolmoSpaces, one run at a time.

How to read this page

A run is one agent configuration, harness + model + settings; each task in it runs once per mode. Pick a run above; the table lists every task of it, one row per mode, in the default order (finished first, then running, easier tasks first). A pill links to that trial's log page. Filter the table with the chips above it.

successfailederror / stalledrunninggradingqueuednot run

Bars count trials, one per task and mode, by state. Progress is the partial credit beside success, in [0, 1]: BEHAVIOR's 2026 challenge q_score, elsewhere the grader's final_reward. A success always counts as full progress, 1.0, whatever that key says (RoboLab's final_reward is 0 for a task without scored subtasks); mean progress averages it over the graded trials. Agent time adds up finished trials; running ones are shown apart. A benchmark can also declare grader metrics, continuous scores its grader writes (IoU, F1, a distance): they get a column each, a sort, and a run figure, the mean or the median of the graded trials.

Tokens are the model gateway's count of every request (input includes the cached part, output the reasoning part). ≈ $ is those tokens at the list prices in data/prices.yml, an estimate and never a bill; billed is what a provider charged, where it says. The two are never added together.

A benchmark's owner can exclude tasks from the final benchmark (excluded: in state/tasks/<benchmark>.yml). Their trials stay, last in a run's table under Excluded from the benchmark; the numbers count the benchmark's own tasks, and the Count switch adds the excluded ones.

Runs are registered per benchmark: how to register a run.

Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)defaultclosed

8 tasks × 2 modes · results as of 10-04 13:42

Unlimited7 / 8 success
7 success1 failed8 done · 88%
Limited0 / 8 success
8 failed8 done · 0%
Mean progress1.00
none · 7 graded
Agent time4.1 h
finished trials

0 requests64.5M in (98% cached)666k out≈ $1.11 at list pricebilled: —

Run details
Harness
Codex CLI 0.159.2 as Harbor's codex agent (OpenRouter as its own provider, no model gateway)
Model
openrouter/openai/gpt-6-luna
Route
OpenRouter, with the owner's API key (billed per token)
Settings
reasoning_effort medium, version 0.157.0–0.159.2
Modes
unlimited, limited
Batches
molmospaces-luna-1003
Scope
8 tasks in the run · 0 not in it · 0 removed (lists at the end)
Progress
the grader's none, in [0, 1]; a success counts as 1.0
Billing
openrouter: billed per request by OpenRouter (the charge it reports in every response)
List price
openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
Closed
the selection run of 2026-10-03, run once
Run id
codex-gpt6_luna-medium-openrouter · data/agents/codex-gpt6_luna-medium-openrouter.yml, state/runs/molmospaces.yml
Results
as collected 10-04 13:42, published with make publish-runs
TaskTrial (its log page)ProgressAgent timeRequestsTokens in / outEst. cost
Franka · CloseU✓ 7m1.007m—1.2M / 18k$0.027
L✗ 23m—23m—10.2M / 86k$0.159
Franka · OpenU✓ 9m1.009m—1.0M / 26k$0.029
L✗ 15m—15m—4.9M / 60k$0.089
Franka · PickU✓ 17m1.0017m—3.9M / 35k$0.066
L✗ 10m—10m—3.1M / 47k$0.062
Franka · Pick and placeU✓ 9m1.009m—1.6M / 17k$0.031
L✗ 14m—14m—3.4M / 41k$0.062
Franka · Pick and place by colourU✓ 7m1.007m—1.5M / 14k$0.029
L✗ 11m—11m—2.5M / 36k$0.049
Franka · Pick and place next toU✓ 16m1.0016m—2.9M / 22k$0.047
L✗ 11m—11m—3.0M / 36k$0.055
RB-Y1 · Navigate toU✓ 8m1.008m—1.8M / 13k$0.032
L✗ 8m—8m—1.9M / 33k$0.041
RB-Y1 · Open doorU✗ 1h 00m—1h 00m—14.3M / 122k$0.222
L✗ 20m—20m—7.2M / 60k$0.113