Skip to content

DexToolBench runs

Agent runs on the tasks of DexToolBench, one run at a time.

How to read this page

A run is one agent configuration, harness + model + settings; each task in it runs once per mode. Pick a run above; the table lists every task of it, one row per mode, in the default order (finished first, then running, easier tasks first). A pill links to that trial's log page. Filter the table with the chips above it.

successfailederror / stalledrunninggradingqueuednot run

Bars count trials, one per task and mode, by state. Progress is the partial credit beside success, in [0, 1]: BEHAVIOR's 2026 challenge q_score, elsewhere the grader's final_reward. A success always counts as full progress, 1.0, whatever that key says (RoboLab's final_reward is 0 for a task without scored subtasks); mean progress averages it over the graded trials. Agent time adds up finished trials; running ones are shown apart. A benchmark can also declare grader metrics, continuous scores its grader writes (IoU, F1, a distance): they get a column each, a sort, and a run figure, the mean or the median of the graded trials.

Tokens are the model gateway's count of every request (input includes the cached part, output the reasoning part). ≈ $ is those tokens at the list prices in data/prices.yml, an estimate and never a bill; billed is what a provider charged, where it says. The two are never added together.

A benchmark's owner can exclude tasks from the final benchmark (excluded: in state/tasks/<benchmark>.yml). Their trials stay, last in a run's table under Excluded from the benchmark; the numbers count the benchmark's own tasks, and the Count switch adds the excluded ones.

Runs are registered per benchmark: how to register a run.

Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default

21 tasks × 2 modes · results as of 10-01 17:29

Unlimited1 / 21 success
1 success20 failed21 done · 5%
Limited0 / 21 success
21 failed21 done · 0%
Mean progress0.06
progress · 42 graded
Agent time40.1 h
finished trials

0 requests699.1M in (98% cached)6.0M out≈ $11.30 at list pricebilled: none (subscription)

Grader metrics
Goals reached3mean · median 0 · 33 graded
Run details
Harness
Since 2026-10-01 (batch mj-luna-xh-f1-1001, the official run): Harbor's built-in codex agent (-a codex, Codex CLI 0.159.0) on Azure OpenAI's API (gpt-6-luna), reasoning effort xhigh with detailed reasoning summaries, no model gateway (token counts are Harbor's per-trial totals); the tasks it has not reached yet still show the earlier result: Harbor's built-in codex agent (-a codex, Codex CLI 0.158.0) on the owner's ChatGPT login, reasoning effort xhigh with detailed reasoning summaries; not the CodexChatGPT wrapper the other benchmarks used, so Codex applies ChatGPT's metadata for gpt-6-luna (its base instructions, verbosity low); token counts are Harbor's per-trial totals
Model
chatgpt/openai/gpt-6-luna
Route
the owner's ChatGPT login (a Pro subscription), through our model gateway
Settings
reasoning_effort xhigh, version 0.157.0
Modes
unlimited, limited
Batches
mj-luna-xh-0928 mj-luna-xh-nt-0929 mj-luna-xh-nt-rerun-0929 mj-luna-xh-f1-1001
A later batch replaces an earlier one's run of the tasks it reruns; the earlier one stays as history.
Scope
21 tasks in the run · 3 not in it · 0 removed (lists at the end)
Progress
the grader's progress, in [0, 1]; a success counts as 1.0
Billing
subscription: a subscription login: no per-token bill
List price
openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
Notes
Limited mode: the RoboLab, RoboPaint and RoboWits results are from the robot service's rcb-limited/2.2 protocol; RoboPaint's and RoboWits' earlier limited round is kept as history, RoboLab's (2.0) was withdrawn. BEHAVIOR-1K's limited results are from its earlier round.
Run id
codex-gpt6_luna-xhigh · data/agents/codex-gpt6_luna-xhigh.yml, state/runs/dextoolbench.yml · formerly codex-0.157-gpt6luna-xhigh-cgpt
Results
as collected 10-01 17:29, published with make publish-runs
TaskTrial (its log page)ProgressGoals reachedAgent timeRequestsTokens in / outEst. cost
Claw hammer · swing downhardU✗ 54m 32%0.321254m—15.6M / 137k$0.254
L✗ 1h 00m 0% prev0.0001h 00m—26.1M / 124k$0.364
Handle eraser · wipe smilehardU✗ 1h 00m 0%0.00—1h 00m—7.9M / 103k$0.152
L✗ 56m 0% prev0.00056m—23.8M / 149k$0.348
Long screwdriver · spin horizontalhardU✗ 1h 00m 0% prev0.0001h 00m—15.9M / 241k$0.313
L✗ 58m 0% prev0.00058m—18.6M / 200k$0.324
Mallet hammer · swing downhardU✓ 28m1.003628m—3.0M / 60k$0.071
L✗ 57m 0% prev0.00057m—7.4M / 82k$0.131
Sharpie marker · draw smilehardU✗ 1h 00m 0%0.00—1h 00m—10.6M / 169k$0.230
L✗ 1h 00m 0% prev0.0001h 00m—18.4M / 131k$0.281
Sharpie marker · write chardU✗ 1h 00m 0%0.00—1h 00m—12.0M / 165k$0.247
L✗ 1h 00m 0% prev0.0001h 00m—22.6M / 173k$0.343
Staples marker · draw smilehardU✗ 1h 00m 0%0.00—1h 00m—10.9M / 161k$0.227
L✗ 1h 00m 0% prev0.0001h 00m—20.4M / 124k$0.293
Blue brush · sweep forwardextremeU✗ 1h 00m 0%0.0001h 00m—5.6M / 106k$0.126
L✗ 53m 0% prev0.00053m—25.3M / 136k$0.377
Claw hammer · swing sideextremeU✗ 1h 00m 0%0.00—1h 00m—5.3M / 96k$0.119
L✗ 58m 0% prev0.00058m—25.5M / 157k$0.370
Flat spatula · flip overextremeU✗ 1h 00m 0%0.00—1h 00m—9.4M / 130k$0.186
L✗ 53m 0% prev0.00053m—25.4M / 130k$0.349
Flat spatula · serve plateextremeU✗ 1h 00m 0%0.00—1h 00m—9.1M / 113k$0.169
L✗ 59m 0% prev0.00059m—33.8M / 151k$0.447
Handle eraser · wipe cextremeU✗ 1h 00m 19%0.1961h 00m—11.2M / 118k$0.195
L✗ 1h 00m 0% prev0.0001h 00m—34.1M / 148k$0.461
Long screwdriver · spin verticalextremeU✗ 1h 00m 0%0.00—1h 00m—7.5M / 122k$0.159
L✗ 1h 00m 0% prev0.0001h 00m—18.2M / 120k$0.267
Mallet hammer · swing sideextremeU✗ 1h 00m 0%0.0001h 00m—13.0M / 166k$0.252
L✗ 58m 0% prev0.00058m—24.9M / 114k$0.336
Red brush · sweep forwardextremeU✗ 37m 42% prev0.421637m—12.8M / 178k$0.252
L✗ 59m 0% prev0.00059m—22.2M / 203k$0.353
Red brush · sweep rightextremeU✗ 1h 00m 5%0.0521h 00m—17.7M / 139k$0.274
L✗ 57m 0% prev0.00057m—15.5M / 107k$0.232
Short screwdriver · spin horizontalextremeU✗ 1h 00m 0%0.00—1h 00m—7.6M / 104k$0.154
L✗ 56m 0% prev0.00056m—24.1M / 139k$0.340
Short screwdriver · spin verticalextremeU✗ 1h 00m 14%0.14111h 00m—16.1M / 143k$0.267
L✗ 58m 0% prev0.00058m—18.8M / 138k$0.293
Spoon spatula · flip overextremeU✗ 54m 15%0.15654m—9.3M / 118k$0.173
L✗ 59m 0% prev0.00059m—19.9M / 166k$0.330
Spoon spatula · serve plateextremeU✗ 56m 0% prev0.00056m—16.6M / 284k$0.356
L✗ 58m 0% prev0.00058m—21.5M / 167k$0.323
Staples marker · write cextremeU✗ 1h 00m 3%0.0311h 00m—13.7M / 159k$0.249
L✗ 57m 0% prev0.00057m—21.7M / 118k$0.312
3 task(s) not in this run
TaskWhy
Blue brush · sweep rightnot built: SimToolReal's pretrained policy reaches no goal on these in our MuJoCo port
Flat eraser · wipe cnot built: SimToolReal's pretrained policy reaches no goal on these in our MuJoCo port
Flat eraser · wipe smilenot built: SimToolReal's pretrained policy reaches no goal on these in our MuJoCo port

Codex + GPT-6 Luna, reasoning xhigh (Azure OpenAI API, 2026-09-30 stress test)

21 tasks × 2 modes · results as of 10-01 17:29

Unlimited1 / 21 success
1 success20 failed21 done · 5%
Limited0 / 21 success
21 failed21 done · 0%
Mean progress0.06
progress · 42 graded
Agent time40.1 h
finished trials

0 requests761.8M in (98% cached)7.5M out≈ $12.66 at list pricebilled: —

Grader metrics
Goals reached2mean · median 0 · 37 graded
Run details
Harness
Harbor's built-in codex agent (-a codex, Codex CLI 0.159.0) on Azure OpenAI's API (gpt-6-luna), reasoning effort xhigh with detailed reasoning summaries; limited mode on rcb-limited/2.1 with crash recovery; image rcb-mujoco 0.1.2. No model gateway: token counts are Harbor's
Model
azure/gpt-6-luna
Route
Azure OpenAI deployments of gpt-6-luna, one endpoint per Harbor job: the lab's main resource (1M tokens/min) and two more (333k tokens/min each); each trial talked to one endpoint
Settings
reasoning_effort xhigh, reasoning_summary detailed, version 0.159.0
Modes
unlimited, limited
Batches
mj-luna-xh-az-0930 mj-luna-xh-az-rerun-0930
A later batch replaces an earlier one's run of the tasks it reruns; the earlier one stays as history.
Scope
21 tasks in the run · 3 not in it · 0 removed (lists at the end)
Progress
the grader's progress, in [0, 1]; a success counts as 1.0
Billing
api: billed by the provider's API; we hold no per-trial bill, so only the estimate is shown
List price
openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
Notes
Not an official result. A large-batch stress run of every task on robot_coding_bench dev/kangrui (RoboTwin 2.0, DexToolBench, MuJoCo Playground), both modes, limited mode on rcb-limited/2.1 with crash recovery. RoboTwin ran mostly on an AWS g6e (L40S); the four RoboTwin tasks whose frozen instance does not rebuild bit-exactly there, and every MuJoCo task, ran on the lab machine.
Run id
codex-gpt6_luna-xhigh-azure · data/agents/codex-gpt6_luna-xhigh-azure.yml, state/runs/dextoolbench.yml
Results
as collected 10-01 17:29, published with make publish-runs
ⓘ Notes
  • Red brush · sweep right unlimited rerun: the first trial got no model response: its endpoint (a second Azure resource, *.services.ai.azure.com) is not on the task's egress allowlist
  • Handle eraser · wipe c limited rerun: the first trial ended on Azure's token rate limit (Codex retries its remote context compaction only twice)
  • Claw hammer · swing down unlimited rerun: the first trial ended on Azure's token rate limit (Codex retries its remote context compaction only twice)
  • Claw hammer · swing down limited rerun: the first trial was lost when its task directory was regenerated under the running robot service
  • Claw hammer · swing side unlimited rerun: the first trial got no model response: its endpoint (a second Azure resource, *.services.ai.azure.com) is not on the task's egress allowlist
  • Sharpie marker · draw smile unlimited rerun: the first trial was stopped by us to move the work to another endpoint or machine
  • Sharpie marker · write c limited rerun: the first trial ended on Azure's token rate limit (Codex retries its remote context compaction only twice) (3 earlier attempts in all)
  • Long screwdriver · spin horizontal unlimited rerun: the first trial got no model response: its endpoint (a second Azure resource, *.services.ai.azure.com) is not on the task's egress allowlist (2 earlier attempts in all)
  • Long screwdriver · spin vertical unlimited rerun: the first trial got no model response: its endpoint (a second Azure resource, *.services.ai.azure.com) is not on the task's egress allowlist
  • Flat spatula · flip over limited rerun: the first trial was stopped by us to move the work to another endpoint or machine
  • Flat spatula · serve plate unlimited rerun: the first trial was stopped by us to move the work to another endpoint or machine
  • Flat spatula · serve plate limited rerun: the first trial was stopped by us to move the work to another endpoint or machine (2 earlier attempts in all)
  • Spoon spatula · flip over limited rerun: the first trial was stopped by us to move the work to another endpoint or machine
  • Spoon spatula · serve plate limited rerun: the first trial was stopped by us to move the work to another endpoint or machine
TaskTrial (its log page)ProgressGoals reachedAgent timeRequestsTokens in / outEst. cost
Claw hammer · swing downhardU✗ 1h 00m 3% ⓘ prev0.0311h 00m—15.1M / 159k$0.263
L✗ 59m 0% ⓘ prev0.00059m—19.8M / 126k$0.280
Handle eraser · wipe smilehardU✓ 20m1.002920m—4.0M / 81k$0.094
L✗ 59m 0%0.00059m—28.9M / 154k$0.390
Long screwdriver · spin horizontalhardU✗ 58m 6% ⓘ prev0.06258m—16.2M / 263k$0.350
L✗ 1h 00m 0%0.0001h 00m—22.6M / 112k$0.299
Mallet hammer · swing downhardU✗ 59m 3%0.03159m—15.0M / 196k$0.279
L✗ 1h 00m 0%0.0001h 00m—16.9M / 112k$0.245
Sharpie marker · draw smilehardU✗ 59m 0% ⓘ prev0.00059m—17.8M / 235k$0.355
L✗ 57m 0%0.00057m—18.6M / 235k$0.339
Sharpie marker · write chardU✗ 59m 8%0.08259m—17.7M / 257k$0.348
L✗ 59m 0% ⓘ prev0.00059m—14.2M / 129k$0.232
Staples marker · draw smilehardU✗ 57m 0%0.00057m—17.4M / 215k$0.315
L✗ 59m 0%0.00059m—19.4M / 185k$0.321
Blue brush · sweep forwardextremeU✗ 1h 00m 6%0.0621h 00m—21.3M / 227k$0.360
L✗ 1h 00m 0%0.0001h 00m—11.9M / 84k$0.175
Claw hammer · swing sideextremeU✗ 57m 12% ⓘ prev0.12557m—15.0M / 244k$0.315
L✗ 58m 0%0.00058m—28.7M / 274k$0.470
Flat spatula · flip overextremeU✗ 1h 00m 0%0.00—1h 00m—14.8M / 177k$0.277
L✗ 53m 0% ⓘ prev0.00053m—19.9M / 124k$0.280
Flat spatula · serve plateextremeU✗ 1h 00m 0% ⓘ prev0.0001h 00m—17.5M / 232k$0.330
L✗ 1h 00m 0% ⓘ prev0.0001h 00m—16.3M / 162k$0.267
Handle eraser · wipe cextremeU✗ 43m 13%0.13443m—13.3M / 189k$0.257
L✗ 1h 00m 0% ⓘ prev0.0001h 00m—26.8M / 208k$0.403
Long screwdriver · spin verticalextremeU✗ 1h 00m 0% ⓘ prev0.00—1h 00m—21.9M / 273k$0.403
L✗ 1h 00m 0%0.0001h 00m—23.0M / 129k$0.312
Mallet hammer · swing sideextremeU✗ 1h 00m 25%0.2581h 00m—20.9M / 218k$0.354
L✗ 59m 0%0.00059m—16.4M / 126k$0.247
Red brush · sweep forwardextremeU✗ 1h 00m 0%0.00—1h 00m—11.3M / 170k$0.229
L✗ 58m 0%0.00058m—13.7M / 88k$0.194
Red brush · sweep rightextremeU✗ 48m 13% ⓘ prev0.13548m—10.1M / 163k$0.207
L✗ 57m 0%0.00057m—26.5M / 131k$0.352
Short screwdriver · spin horizontalextremeU✗ 55m 8%0.08355m—15.3M / 203k$0.284
L✗ 1h 00m 0%0.0001h 00m—26.6M / 162k$0.371
Short screwdriver · spin verticalextremeU✗ 1h 00m 0%0.00—1h 00m—18.1M / 299k$0.392
L✗ 59m 0%0.00059m—16.7M / 122k$0.249
Spoon spatula · flip overextremeU✗ 1h 00m 0%0.0001h 00m—17.7M / 188k$0.304
L✗ 58m 0% ⓘ prev0.00058m—24.5M / 210k$0.382
Spoon spatula · serve plateextremeU✗ 1h 00m 0%0.00—1h 00m—11.4M / 203k$0.252
L✗ 57m 0% ⓘ prev0.00057m—22.0M / 206k$0.355
Staples marker · write cextremeU✗ 1h 00m 38%0.38111h 00m—7.8M / 116k$0.156
L✗ 1h 00m 0%0.0001h 00m—28.6M / 139k$0.377
3 task(s) not in this run
TaskWhy
Blue brush · sweep rightnot built: SimToolReal's pretrained policy reaches no goal on these in our MuJoCo port
Flat eraser · wipe cnot built: SimToolReal's pretrained policy reaches no goal on these in our MuJoCo port
Flat eraser · wipe smilenot built: SimToolReal's pretrained policy reaches no goal on these in our MuJoCo port