Skip to content

Benchmark task reference

Documentation for the benchmarks behind our robotic coding-agent work. For every task: what it asks for, its oracle demonstration, the capabilities it requires, and whether we are keeping it; and how our coding agents do on it.

Benchmarks

BEHAVIOR-1KOmniGibson (NVIDIA Isaac Sim)100 tasks · 100 untriaged · 0 untaggedOpen the task list →DexToolBenchMuJoCo 3.3.7 (our port; upstream runs Isaac Sim / Isaac Gym)24 tasks · 3 untriaged · 0 untaggedOpen the task list →HumanoidBenchMuJoCo 3 (HumanoidBench's gymnasium environments)9 tasks + 15 excluded · 0 untriaged · 0 untaggedOpen the task list →KinDERMuJoCo 3.3 (KinDER's Dynamic3D environments); PyBullet for two kinematic tasks now excluded4 tasks + 7 excluded · 0 untriaged · 0 untaggedOpen the task list →MetaWorld+MuJoCo 3.3.0 (Meta-World v3)50 tasks · 0 untriaged · 0 untaggedOpen the task list →MolmoSpacesMuJoCo 3.5.0 (molmo-spaces 0.2.9)8 tasks · 0 untriaged · 0 untaggedOpen the task list →MuJoCo Playground (manipulation)MuJoCo 3.3.7 (plain MuJoCo on the CPU; upstream trains in MJX / MuJoCo Warp)10 tasks · 6 untriaged · 0 untaggedOpen the task list →RoboCasa-GR1robosuite @ 51cc017 (1.5.1, mink whole-body IK) / MuJoCo 3.2.6 (RoboCasa GR1 tabletop)24 tasks · 0 untriaged · 0 untaggedOpen the task list →RoboCasarobosuite 1.5.1 / MuJoCo 3.2.6 (RoboCasa v0.2)29 tasks · 0 untriaged · 0 untaggedOpen the task list →RoboCasa365robosuite 1.5.2 / MuJoCo 3.3.1 (RoboCasa365)50 tasks · 0 untriaged · 0 untaggedOpen the task list →RoboLabIsaac Lab 2.3.2 / Isaac Sim 5.1 (NVIDIA)120 tasks · 120 untriaged · 0 untaggedOpen the task list →RoboPaintManiSkill 3.0.1 / SAPIEN 3.0.3 + Spline-FRIDA stroke model82 tasks + 9 excluded · 82 untriaged · 0 untaggedOpen the task list →RoboTwin 2.0SAPIEN 3 (PhysX) + CuRobo50 tasks · 0 untriaged · 0 untaggedOpen the task list →RoboWitsGenesis 0.4.7 (gs_gym)29 tasks + 1 excluded · 29 untriaged · 0 untaggedOpen the task list →VLABenchMuJoCo 3.2.2 + dm_control 1.0.22 (VLABench)36 tasks · 0 untriaged · 0 untaggedOpen the task list →

Agent runs

Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login) · Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter), on every benchmark: success among graded trials, the mean score (partial credit, a success counts as 1.0), and what the model use would cost at list price. All runs, tasks and trial logs →

BenchmarkTasks runSuccess, unlimitedSuccess, limitedMean scoreEst. cost (list price)Tokens in / out
BEHAVIOR-1K4918% 9/490% 0/490.13 q_score$92.766.81B / 28.2M
DexToolBench215% 1/210% 0/210.06 progress$11.30699.1M / 6.0M
HumanoidBench9 + 15 excluded0% 0/90% 0/9—$4.56282.8M / 2.2M
KinDER4 + 7 excluded50% 2/40% 0/41.00 none$2.57154.8M / 1.4M
MetaWorld+4100% 4/425% 1/41.00 none$0.60730.3M / 387k
MolmoSpaces888% 7/80% 0/81.00 none$1.1164.5M / 666k
MuJoCo Playground (manipulation)475% 3/425% 1/41.00 final_reward$1.4082.0M / 839k
RoboCasa-GR1450% 2/40% 0/41.00 none$1.78113.8M / 840k
RoboCasa4100% 4/40% 0/41.00 none$1.55103.6M / 649k
RoboCasa365475% 3/40% 0/41.00 none$1.2271.9M / 660k
RoboLab8856% 49/8828% 25/880.84 final_reward$42.952.89B / 16.8M
RoboPaint82 + 9 excluded60% 49/821% 1/820.65 mean final reward*$29.811.75B / 17.0M
RoboTwin 2.025100% 25/2532% 8/251.00 final_reward$5.79371.3M / 2.8M
RoboWits29 + 1 excluded59% 16/2728% 8/290.87 final_reward$13.90893.8M / 6.3M
VLABench4100% 4/450% 2/41.00 none$0.78447.0M / 404k

* RoboPaint: the plain mean of the grader's final reward over the graded trials, a success at its own value. Its metric depends on the task family: IoU for lettering, kaishu and xingshu; 1 - mean ΔE / 20 for acrylic and oil.

How to contribute

Everything you would change lives in one folder, state/. Task pages under docs/ hold only upstream metadata and prose, so reviewing a day of someone's triage is git diff state/ — not a hundred file diffs.

I want to… Edit this Or use
Triage a task, set its tags state/tasks/<benchmark>.yml the Edit button on any task row
Add or reword a display tag state/display_tags.yml the Label taxonomy page
Add or reword a detailed label state/taxonomy.yml —
Write up why a task is interesting docs/benchmarks/<benchmark>/tasks/<task_id>.md —
Record benchmark-level facts data/benchmarks/<benchmark>.yml —
Change how the task list looks scripts/gallery.py, docs/javascripts/tasklist.js —
Add a whole new benchmark walkthrough —

Start here

git clone <this-repo> && cd <this-repo>

make install     # one-time: creates .venv, installs the pinned toolchain
make demos       # one-time: downloads 134 oracle demos + posters (~1.2 GB)
make edit        # start the site WITH editing enabled

make edit prints the URL it picked (the port is chosen automatically, so a colleague's copy never collides with yours). Open it, and you are on the task list.

Then:

  1. Set the speed to 2× or 3× once — it is remembered across pages.
  2. Filter Status to pending and pick a short task.
  3. Click Edit on that row, set owner to your handle, Save. Now nobody duplicates your work.
  4. Watch the demo in the row, then set status, difficulty and labels.
  5. Optionally open the task page and write down why — that prose is the part a reviewer actually reads.
  6. make check, then commit. Your whole session is a diff of state/.

Editing only works under make edit — not make serve

Command Site Edit buttons Label taxonomy page
make serve yes no read-only
make edit yes yes editable

The two commands serve the identical site. The difference is that make edit also starts a small local daemon (scripts/editd.py, port 8009) that is the only thing able to write to state/. A browser page cannot write files on its own.

Under make serve the Edit buttons are still visible — clicking one tells you to run make edit rather than silently doing nothing. If your edits are not saving, this is why: stop the server and run make edit.

A deployed copy of this site is always read-only, by design.

  • Triage a task

    The main job. Watch the demo, decide keep or drop, apply labels. How to do it →

  • Argue with the labels

    The vocabulary is a first draft, not a standard: tasks show display tags in two groups, Capability and Task Domain, and keep their detailed labels for analysis. Label taxonomy →

  • Reconstruct a missing goal

    50 of the 100 tasks ship no instruction text upstream. Watching the demo and writing the goal down is high-value. How →

  • Add a benchmark

    A YAML file plus task pages; every view picks it up automatically. Walkthrough →

Scope

Working document, compiled in part from public BEHAVIOR Challenge material. Not affiliated with or endorsed by the BEHAVIOR team.