MolmoSpaces runs¶
Agent runs on the tasks of MolmoSpaces, one run at a time.
How to read this page
A run is one agent configuration, harness + model + settings; each task in it runs once per mode. Pick a run above; the table lists every task of it, one row per mode, in the default order (finished first, then running, easier tasks first). A pill links to that trial's log page. Filter the table with the chips above it.
successfailederror / stalledrunninggradingqueuednot run
Bars count trials, one per task and mode, by state. Progress is the partial credit beside success, in [0, 1]: BEHAVIOR's 2026 challenge q_score, elsewhere the grader's final_reward. A success always counts as full progress, 1.0, whatever that key says (RoboLab's final_reward is 0 for a task without scored subtasks); mean progress averages it over the graded trials. Agent time adds up finished trials; running ones are shown apart. A benchmark can also declare grader metrics, continuous scores its grader writes (IoU, F1, a distance): they get a column each, a sort, and a run figure, the mean or the median of the graded trials.
Tokens are the model gateway's count of every request (input includes the cached part, output the reasoning part). ≈ $ is those tokens at the list prices in data/prices.yml, an estimate and never a bill; billed is what a provider charged, where it says. The two are never added together.
A benchmark's owner can exclude tasks from the final benchmark (excluded: in state/tasks/<benchmark>.yml). Their trials stay, last in a run's table under Excluded from the benchmark; the numbers count the benchmark's own tasks, and the Count switch adds the excluded ones.
Runs are registered per benchmark: how to register a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)defaultclosed
0 requests64.5M in (98% cached)666k out≈ $1.11 at list pricebilled: —
Run details
- Harness
- Codex CLI 0.159.2 as Harbor's codex agent (OpenRouter as its own provider, no model gateway)
- Model
openrouter/openai/gpt-6-luna- Route
- OpenRouter, with the owner's API key (billed per token)
- Settings
- reasoning_effort medium, version 0.157.0–0.159.2
- Modes
- unlimited, limited
- Batches
molmospaces-luna-1003- Scope
- 8 tasks in the run · 0 not in it · 0 removed (lists at the end)
- Progress
- the grader's none, in [0, 1]; a success counts as 1.0
- Billing
- openrouter: billed per request by OpenRouter (the charge it reports in every response)
- List price
- openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
- Closed
- the selection run of 2026-10-03, run once
- Run id
codex-gpt6_luna-medium-openrouter·data/agents/codex-gpt6_luna-medium-openrouter.yml,state/runs/molmospaces.yml- Results
- as collected 10-04 13:42, published with
make publish-runs
| Task | Trial (its log page) | Progress | Agent time | Requests | Tokens in / out | Est. cost |
|---|---|---|---|---|---|---|
| Franka · Close | U✓ 7m | 1.00 | 7m | — | 1.2M / 18k | $0.027 |
| L✗ 23m | — | 23m | — | 10.2M / 86k | $0.159 | |
| Franka · Open | U✓ 9m | 1.00 | 9m | — | 1.0M / 26k | $0.029 |
| L✗ 15m | — | 15m | — | 4.9M / 60k | $0.089 | |
| Franka · Pick | U✓ 17m | 1.00 | 17m | — | 3.9M / 35k | $0.066 |
| L✗ 10m | — | 10m | — | 3.1M / 47k | $0.062 | |
| Franka · Pick and place | U✓ 9m | 1.00 | 9m | — | 1.6M / 17k | $0.031 |
| L✗ 14m | — | 14m | — | 3.4M / 41k | $0.062 | |
| Franka · Pick and place by colour | U✓ 7m | 1.00 | 7m | — | 1.5M / 14k | $0.029 |
| L✗ 11m | — | 11m | — | 2.5M / 36k | $0.049 | |
| Franka · Pick and place next to | U✓ 16m | 1.00 | 16m | — | 2.9M / 22k | $0.047 |
| L✗ 11m | — | 11m | — | 3.0M / 36k | $0.055 | |
| RB-Y1 · Navigate to | U✓ 8m | 1.00 | 8m | — | 1.8M / 13k | $0.032 |
| L✗ 8m | — | 8m | — | 1.9M / 33k | $0.041 | |
| RB-Y1 · Open door | U✗ 1h 00m | — | 1h 00m | — | 14.3M / 122k | $0.222 |
| L✗ 20m | — | 20m | — | 7.2M / 60k | $0.113 |