Benchmark task reference¶
Documentation for the benchmarks behind our robotic coding-agent work. For every task: what it asks for, its oracle demonstration, the capabilities it requires, and whether we are keeping it; and how our coding agents do on it.
Benchmarks¶
Agent runs¶
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login) · Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter), on every benchmark: success among graded trials, the mean score (partial credit, a success counts as 1.0), and what the model use would cost at list price. All runs, tasks and trial logs →
| Benchmark | Tasks run | Success, unlimited | Success, limited | Mean score | Est. cost (list price) | Tokens in / out |
|---|---|---|---|---|---|---|
| BEHAVIOR-1K | 49 | 18% 9/49 | 0% 0/49 | 0.13 q_score | $92.76 | 6.81B / 28.2M |
| DexToolBench | 21 | 5% 1/21 | 0% 0/21 | 0.06 progress | $11.30 | 699.1M / 6.0M |
| HumanoidBench | 9 + 15 excluded | 0% 0/9 | 0% 0/9 | — | $4.56 | 282.8M / 2.2M |
| KinDER | 4 + 7 excluded | 50% 2/4 | 0% 0/4 | 1.00 none | $2.57 | 154.8M / 1.4M |
| MetaWorld+ | 4 | 100% 4/4 | 25% 1/4 | 1.00 none | $0.607 | 30.3M / 387k |
| MolmoSpaces | 8 | 88% 7/8 | 0% 0/8 | 1.00 none | $1.11 | 64.5M / 666k |
| MuJoCo Playground (manipulation) | 4 | 75% 3/4 | 25% 1/4 | 1.00 final_reward | $1.40 | 82.0M / 839k |
| RoboCasa-GR1 | 4 | 50% 2/4 | 0% 0/4 | 1.00 none | $1.78 | 113.8M / 840k |
| RoboCasa | 4 | 100% 4/4 | 0% 0/4 | 1.00 none | $1.55 | 103.6M / 649k |
| RoboCasa365 | 4 | 75% 3/4 | 0% 0/4 | 1.00 none | $1.22 | 71.9M / 660k |
| RoboLab | 88 | 56% 49/88 | 28% 25/88 | 0.84 final_reward | $42.95 | 2.89B / 16.8M |
| RoboPaint | 82 + 9 excluded | 60% 49/82 | 1% 1/82 | 0.65 mean final reward* | $29.81 | 1.75B / 17.0M |
| RoboTwin 2.0 | 25 | 100% 25/25 | 32% 8/25 | 1.00 final_reward | $5.79 | 371.3M / 2.8M |
| RoboWits | 29 + 1 excluded | 59% 16/27 | 28% 8/29 | 0.87 final_reward | $13.90 | 893.8M / 6.3M |
| VLABench | 4 | 100% 4/4 | 50% 2/4 | 1.00 none | $0.784 | 47.0M / 404k |
* RoboPaint: the plain mean of the grader's final reward over the graded trials, a success at its own value. Its metric depends on the task family: IoU for lettering, kaishu and xingshu; 1 - mean ΔE / 20 for acrylic and oil.
How to contribute¶
Everything you would change lives in one folder, state/. Task pages under docs/ hold only upstream metadata and prose, so reviewing a day of someone's triage is git diff state/ — not a hundred file diffs.
| I want to… | Edit this | Or use |
|---|---|---|
| Triage a task, set its tags | state/tasks/<benchmark>.yml | the Edit button on any task row |
| Add or reword a display tag | state/display_tags.yml | the Label taxonomy page |
| Add or reword a detailed label | state/taxonomy.yml | — |
| Write up why a task is interesting | docs/benchmarks/<benchmark>/tasks/<task_id>.md | — |
| Record benchmark-level facts | data/benchmarks/<benchmark>.yml | — |
| Change how the task list looks | scripts/gallery.py, docs/javascripts/tasklist.js | — |
| Add a whole new benchmark | walkthrough | — |
Start here¶
git clone <this-repo> && cd <this-repo>
make install # one-time: creates .venv, installs the pinned toolchain
make demos # one-time: downloads 134 oracle demos + posters (~1.2 GB)
make edit # start the site WITH editing enabled
make edit prints the URL it picked (the port is chosen automatically, so a colleague's copy never collides with yours). Open it, and you are on the task list.
Then:
- Set the speed to 2× or 3× once — it is remembered across pages.
- Filter Status to
pendingand pick a short task. - Click Edit on that row, set
ownerto your handle, Save. Now nobody duplicates your work. - Watch the demo in the row, then set status, difficulty and labels.
- Optionally open the task page and write down why — that prose is the part a reviewer actually reads.
make check, then commit. Your whole session is a diff ofstate/.
Editing only works under make edit — not make serve
| Command | Site | Edit buttons | Label taxonomy page |
|---|---|---|---|
make serve | yes | no | read-only |
make edit | yes | yes | editable |
The two commands serve the identical site. The difference is that make edit also starts a small local daemon (scripts/editd.py, port 8009) that is the only thing able to write to state/. A browser page cannot write files on its own.
Under make serve the Edit buttons are still visible — clicking one tells you to run make edit rather than silently doing nothing. If your edits are not saving, this is why: stop the server and run make edit.
A deployed copy of this site is always read-only, by design.
-
Triage a task
The main job. Watch the demo, decide keep or drop, apply labels. How to do it →
-
Argue with the labels
The vocabulary is a first draft, not a standard: tasks show display tags in two groups, Capability and Task Domain, and keep their detailed labels for analysis. Label taxonomy →
-
Reconstruct a missing goal
50 of the 100 tasks ship no instruction text upstream. Watching the demo and writing the goal down is high-value. How →
-
Add a benchmark
A YAML file plus task pages; every view picks it up automatically. Walkthrough →
Scope
Working document, compiled in part from public BEHAVIOR Challenge material. Not affiliated with or endorsed by the BEHAVIOR team.