Skip to content

RoboWits runs

Agent runs on the tasks of RoboWits, one run at a time.

How to read this page

A run is one agent configuration, harness + model + settings; each task in it runs once per mode. Pick a run above; the table lists every task of it, one row per mode, in the default order (finished first, then running, easier tasks first). A pill links to that trial's log page. Filter the table with the chips above it.

successfailederror / stalledrunninggradingqueuednot run

Bars count trials, one per task and mode, by state. Progress is the partial credit beside success, in [0, 1]: BEHAVIOR's 2026 challenge q_score, elsewhere the grader's final_reward. A success always counts as full progress, 1.0, whatever that key says (RoboLab's final_reward is 0 for a task without scored subtasks); mean progress averages it over the graded trials. Agent time adds up finished trials; running ones are shown apart. A benchmark can also declare grader metrics, continuous scores its grader writes (IoU, F1, a distance): they get a column each, a sort, and a run figure, the mean or the median of the graded trials.

Tokens are the model gateway's count of every request (input includes the cached part, output the reasoning part). ≈ $ is those tokens at the list prices in data/prices.yml, an estimate and never a bill; billed is what a provider charged, where it says. The two are never added together.

A benchmark's owner can exclude tasks from the final benchmark (excluded: in state/tasks/<benchmark>.yml). Their trials stay, last in a run's table under Excluded from the benchmark; the numbers count the benchmark's own tasks, and the Count switch adds the excluded ones.

Runs are registered per benchmark: how to register a run.

Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default

30 tasks × 2 modes · 1 of them excluded from the benchmark · results as of 09-30 05:08

Count1 of the run's tasks is excluded from the benchmark: last in the table
Unlimited16 / 27 success
16 success11 failed27 done · 59%
Limited8 / 29 success
8 success21 failed29 done · 28%
Mean progress0.87
final_reward · 29 graded
Agent time63.6 h
finished trials

8,231 requests893.8M in (98% cached)6.3M out (4.8M reasoning)≈ $13.90 at list pricebilled: none (subscription)

Run details
Harness
Codex CLI 0.157.0 as harbor_agents.codex_chatgpt:CodexChatGPT (robot_coding_bench commit 8cfd5cb): the OpenRouter client with our model gateway pointed at ChatGPT's backend
Model
chatgpt/openai/gpt-6-luna
Route
the owner's ChatGPT login (a Pro subscription), through our model gateway
Settings
reasoning_effort xhigh, version 0.157.0
Modes
unlimited, limited
Batches
robowits-lunaxh-0925-2042 robowits-lunaxh-gpu-0926 robowits-lunaxh-rerun-0927 robowits-lunaxh-limited-0929 robowits-lunaxh-limited-rerun-0929
A later batch replaces an earlier one's run of the tasks it reruns; the earlier one stays as history.
Scope
29 tasks in the run + 1 excluded from the benchmark (last in the table) · 0 not in it · 0 removed (lists at the end)
Progress
the grader's final_reward, in [0, 1]; a success counts as 1.0
Billing
subscription: a subscription login: no per-token bill
List price
openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
Notes
Limited mode: the RoboLab, RoboPaint and RoboWits results are from the robot service's rcb-limited/2.2 protocol; RoboPaint's and RoboWits' earlier limited round is kept as history, RoboLab's (2.0) was withdrawn. BEHAVIOR-1K's limited results are from its earlier round.
Run id
codex-gpt6_luna-xhigh · data/agents/codex-gpt6_luna-xhigh.yml, state/runs/robowits.yml · formerly codex-0.157-gpt6luna-xhigh-cgpt
Results
as collected 09-30 05:08, published with make publish-runs
TaskTrial (its log page)ProgressAgent timeRequestsTokens in / outEst. cost
Stack BowlseasyU✓ 1h 24m1.001h 24m14617.7M / 134k$0.272
L✗ 1h 30m prev—1h 30m14317.0M / 119k$0.256
Cover With LidmediumU✓ 1h 05m1.001h 05m15519.0M / 115k$0.276
L✗ 1h 30m prev—1h 30m21824.0M / 201k$0.403
Cylinder Through HolemediumU✗ 21m 24%0.2421m472.6M / 46k$0.066
L✗ 1h 30m prev—1h 30m38342.8M / 214k$0.612
Gap RetrievemediumU✓ 25m1.0025m522.2M / 30k$0.045
L✓ 50m prev1.0050m15821.9M / 141k$0.342
Stack CubesmediumU✓ 47m1.0047m634.3M / 72k$0.093
L✓ 1h 19m prev1.001h 19m28343.3M / 165k$0.574
Align BlockshardU✓ 15m1.0015m432.1M / 24k$0.041
L✓ 18m prev1.0018m734.4M / 52k$0.083
Balance BoardhardU✗ 1h 30m 20%0.201h 30m10413.7M / 153k$0.240
L✗ 1h 30m prev—1h 30m11813.6M / 105k$0.212
Ball Into BottlehardU✓ 1h 06m1.001h 06m1087.8M / 71k$0.129
L✗ 1h 30m prev—1h 30m25532.1M / 168k$0.469
Collect ScrewshardU✓ 10m1.0010m301.1M / 11k$0.022
L✓ 30m prev1.0030m987.8M / 86k$0.139
DominoshardU✓ 11m1.0011m331.1M / 17k$0.026
L✓ 1h 08m prev1.001h 08m15015.8M / 114k$0.240
Hold CuphardU✓ 1h 15m1.001h 15m29336.2M / 145k$0.495
L✓ 1h 30m prev1.001h 30m22929.1M / 176k$0.438
Move CubehardU✓ 25m1.0025m492.5M / 45k$0.057
L✓ 11m prev1.0011m321.5M / 23k$0.033
Pinch CardhardU✗ 1h 30m—1h 30m16517.1M / 181k$0.320
L✗ 1h 30m prev—1h 30m26631.8M / 260k$0.517
Place Tall BoxhardU✓ 31m1.0031m462.1M / 34k$0.047
L✗ 1h 30m prev—1h 30m25428.3M / 234k$0.483
Raise PlatformhardU✓ 1h 01m1.001h 01m10712.0M / 111k$0.199
L✗ 1h 30m prev—1h 30m10514.3M / 162k$0.275
Retrieve CubehardU✓ 57m1.0057m11311.0M / 101k$0.181
L✗ 1h 30m prev—1h 30m15419.4M / 140k$0.294
Retrieve RollhardU✗ 1h 30m—1h 30m13110.7M / 107k$0.193
L✗ 1h 30m prev—1h 30m12712.3M / 114k$0.202
Roll Up BallhardU✗ 11m 6%0.0611m291.3M / 15k$0.026
L✗ 1h 26m prev—1h 26m16622.3M / 155k$0.330
Stand BulbhardU✓ 1h 00m1.001h 00m14116.1M / 121k$0.264
L✗ 1h 30m prev—1h 30m32836.3M / 218k$0.541
Ball Onto TowerextremeU✓ 26m1.0026m422.2M / 41k$0.052
L✓ 31m prev1.0031m321.6M / 41k$0.046
Place BookextremeU✗ 1h 30m—1h 30m14317.5M / 173k$0.316
L✗ 1h 30m prev—1h 30m15821.0M / 153k$0.314
Round Dough SheetextremeU✗ 1h 30m 63%0.631h 30m17715.8M / 78k$0.225
L✗ 1h 30m prev—1h 30m22527.2M / 114k$0.358
Seal ColanderextremeU✗ 1h 30m—1h 30m12011.7M / 90k$0.182
L✗ 1h 30m prev—1h 30m20023.7M / 169k$0.379
Separate Marbles And SandextremeU✗ 1h 30m prev—1h 30m1015.0M / 27k$0.074
L✗ 1h 30m prev—1h 30m18718.1M / 96k$0.255
Stabilize BottleextremeU✓ 58m1.0058m1148.6M / 48k$0.125
L✗ 1h 30m prev—1h 30m23529.6M / 131k$0.398
Stand PagesextremeU✗ 1h 30m 7%0.071h 30m16817.3M / 174k$0.321
L✗ 1h 30m prev—1h 30m15214.1M / 91k$0.208
Water Into MugextremeU✗ 1h 30m—1h 30m16211.1M / 83k$0.174
L✗ 1h 30m prev—1h 30m20621.3M / 184k$0.366
Ball Into JarhardUnot counted0.201h 30m17317.1M / 93k$0.242
L✗ 1h 30m prev—1h 30m21528.3M / 144k$0.385
Differentiate CubesextremeUnot counted0.261h 30m13914.1M / 103k$0.216
L✗ 1h 30m prev—1h 30m19921.2M / 107k$0.292
excluded Excluded from the benchmark · 1 task, kept for the record; the benchmark's numbers leave it out
Align ChopsticksextremeU✗ 15m 67%0.6715m281.0M / 23k$0.029
L✗ 1m—1m7115k / 3k$0.0040