MuJoCo Playground (manipulation) runs¶
Agent runs on the tasks of MuJoCo Playground (manipulation), one run at a time.
How to read this page
A run is one agent configuration, harness + model + settings; each task in it runs once per mode. Pick a run above; the table lists every task of it, one row per mode, in the default order (finished first, then running, easier tasks first). A pill links to that trial's log page. Filter the table with the chips above it.
successfailederror / stalledrunninggradingqueuednot run
Bars count trials, one per task and mode, by state. Progress is the partial credit beside success, in [0, 1]: BEHAVIOR's 2026 challenge q_score, elsewhere the grader's final_reward. A success always counts as full progress, 1.0, whatever that key says (RoboLab's final_reward is 0 for a task without scored subtasks); mean progress averages it over the graded trials. Agent time adds up finished trials; running ones are shown apart. A benchmark can also declare grader metrics, continuous scores its grader writes (IoU, F1, a distance): they get a column each, a sort, and a run figure, the mean or the median of the graded trials.
Tokens are the model gateway's count of every request (input includes the cached part, output the reasoning part). ≈ $ is those tokens at the list prices in data/prices.yml, an estimate and never a bill; billed is what a provider charged, where it says. The two are never added together.
A benchmark's owner can exclude tasks from the final benchmark (excluded: in state/tasks/<benchmark>.yml). Their trials stay, last in a run's table under Excluded from the benchmark; the numbers count the benchmark's own tasks, and the Count switch adds the excluded ones.
Runs are registered per benchmark: how to register a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default
0 requests82.0M in (98% cached)839k out≈ $1.40 at list pricebilled: none (subscription)
Run details
- Harness
- Since 2026-10-01 (batch mj-luna-xh-f1-1001, the official run): Harbor's built-in codex agent (-a codex, Codex CLI 0.159.0) on Azure OpenAI's API (gpt-6-luna), reasoning effort xhigh with detailed reasoning summaries, no model gateway (token counts are Harbor's per-trial totals); the tasks it has not reached yet still show the earlier result: Harbor's built-in codex agent (-a codex, Codex CLI 0.158.0) on the owner's ChatGPT login, reasoning effort xhigh with detailed reasoning summaries; not the CodexChatGPT wrapper the other benchmarks used, so Codex applies ChatGPT's metadata for gpt-6-luna (its base instructions, verbosity low); token counts are Harbor's per-trial totals
- Model
chatgpt/openai/gpt-6-luna- Route
- the owner's ChatGPT login (a Pro subscription), through our model gateway
- Settings
- reasoning_effort xhigh, version 0.157.0
- Modes
- unlimited, limited
- Batches
mj-luna-xh-0928mj-luna-xh-f1-1001
A later batch replaces an earlier one's run of the tasks it reruns; the earlier one stays as history.- Scope
- 4 tasks in the run · 6 not in it · 0 removed (lists at the end)
- Progress
- the grader's final_reward, in [0, 1]; a success counts as 1.0
- Billing
- subscription: a subscription login: no per-token bill
- List price
- openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
- Notes
- Limited mode: the RoboLab, RoboPaint and RoboWits results are from the robot service's rcb-limited/2.2 protocol; RoboPaint's and RoboWits' earlier limited round is kept as history, RoboLab's (2.0) was withdrawn. BEHAVIOR-1K's limited results are from its earlier round.
- Run id
codex-gpt6_luna-xhigh·data/agents/codex-gpt6_luna-xhigh.yml,state/runs/mujoco-playground.yml· formerlycodex-0.157-gpt6luna-xhigh-cgpt- Results
- as collected 10-01 17:29, published with
make publish-runs
- Panda Pick Cube Orientation unlimited Harbor reports a verifier timeout, but the verdict (success) was written first: the grader then spent 20 min rendering videos of the agent's rollouts on software rendering (an image bug, fixed in rcb-mujoco 0.1.2, plus a render budget in the grader)
| Task | Trial (its log page) | Progress | Agent time | Requests | Tokens in / out | Est. cost |
|---|---|---|---|---|---|---|
| Panda Open Cabinetmedium | U✓ 7m prev | 1.00 | 7m | — | 1.8M / 49k | $0.051 |
| L✓ 12m prev | 1.00 | 12m | — | 4.0M / 61k | $0.079 | |
| Panda Pick Cube Orientationmedium | U✓ 25m ⓘ | 1.00 | 25m | — | 4.7M / 72k | $0.099 |
| L✗ 1h 00m | — | 1h 00m | — | 19.9M / 136k | $0.292 | |
| Aloha Hand Overhard | U✓ 2m prev | 1.00 | 2m | — | 394k / 16k | $0.016 |
| L✗ 1h 00m prev | — | 1h 00m | — | 26.0M / 193k | $0.387 | |
| Aloha Single Peg Insertionhard | U✗ 1h 00m | — | 1h 00m | — | 7.5M / 183k | $0.193 |
| L✗ 55m | — | 55m | — | 17.7M / 129k | $0.284 |
6 task(s) not in this run
| Task | Why |
|---|---|
| Aero Cube Rotate Z Axis | not built yet |
| Leap Cube Reorient | not built yet |
| Leap Cube Rotate Z Axis | not built yet |
| Panda Pick Cube | not built yet |
| Panda Pick Cube Cartesian | not built yet |
| Panda Robotiq Push Cube | not built yet |
Codex + GPT-6 Luna, reasoning xhigh (Azure OpenAI API, 2026-09-30 stress test)
0 requests39.1M in (98% cached)467k out≈ $0.704 at list pricebilled: —
Run details
- Harness
- Harbor's built-in codex agent (-a codex, Codex CLI 0.159.0) on Azure OpenAI's API (gpt-6-luna), reasoning effort xhigh with detailed reasoning summaries; limited mode on rcb-limited/2.1 with crash recovery; image rcb-mujoco 0.1.2. No model gateway: token counts are Harbor's
- Model
azure/gpt-6-luna- Route
- Azure OpenAI deployments of gpt-6-luna, one endpoint per Harbor job: the lab's main resource (1M tokens/min) and two more (333k tokens/min each); each trial talked to one endpoint
- Settings
- reasoning_effort xhigh, reasoning_summary detailed, version 0.159.0
- Modes
- unlimited, limited
- Batches
mj-luna-xh-az-0930mj-luna-xh-az-rerun-0930
A later batch replaces an earlier one's run of the tasks it reruns; the earlier one stays as history.- Scope
- 4 tasks in the run · 6 not in it · 0 removed (lists at the end)
- Progress
- the grader's final_reward, in [0, 1]; a success counts as 1.0
- Billing
- api: billed by the provider's API; we hold no per-trial bill, so only the estimate is shown
- List price
- openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
- Notes
- Not an official result. A large-batch stress run of every task on robot_coding_bench dev/kangrui (RoboTwin 2.0, DexToolBench, MuJoCo Playground), both modes, limited mode on rcb-limited/2.1 with crash recovery. RoboTwin ran mostly on an AWS g6e (L40S); the four RoboTwin tasks whose frozen instance does not rebuild bit-exactly there, and every MuJoCo task, ran on the lab machine.
- Run id
codex-gpt6_luna-xhigh-azure·data/agents/codex-gpt6_luna-xhigh-azure.yml,state/runs/mujoco-playground.yml- Results
- as collected 10-01 17:29, published with
make publish-runs
| Task | Trial (its log page) | Progress | Agent time | Requests | Tokens in / out | Est. cost |
|---|---|---|---|---|---|---|
| Panda Open Cabinetmedium | U✓ 14m | 1.00 | 14m | — | 2.3M / 41k | $0.051 |
| L✗ 59m | — | 59m | — | 13.6M / 107k | $0.207 | |
| Panda Pick Cube Orientationmedium | U✓ 12m | 1.00 | 12m | — | 2.0M / 43k | $0.049 |
| L✗ 1h 00m | — | 1h 00m | — | 11.1M / 128k | $0.193 | |
| Aloha Hand Overhard | U✓ 4m | 1.00 | 4m | — | 911k / 20k | $0.024 |
| L✗ 1h 00m | — | 1h 00m | — | 3.8M / 48k | $0.071 | |
| Aloha Single Peg Insertionhard | U✓ 8m | 1.00 | 8m | — | 1.4M / 38k | $0.040 |
| L✗ 1h 00m | — | 1h 00m | — | 3.9M / 42k | $0.069 |
6 task(s) not in this run
| Task | Why |
|---|---|
| Aero Cube Rotate Z Axis | not built yet |
| Leap Cube Reorient | not built yet |
| Leap Cube Rotate Z Axis | not built yet |
| Panda Pick Cube | not built yet |
| Panda Pick Cube Cartesian | not built yet |
| Panda Robotiq Push Cube | not built yet |