Skip to content

RoboTwin 2.0 runs

Agent runs on the tasks of RoboTwin 2.0, one run at a time.

How to read this page

A run is one agent configuration, harness + model + settings; each task in it runs once per mode. Pick a run above; the table lists every task of it, one row per mode, in the default order (finished first, then running, easier tasks first). A pill links to that trial's log page. Filter the table with the chips above it.

successfailederror / stalledrunninggradingqueuednot run

Bars count trials, one per task and mode, by state. Progress is the partial credit beside success, in [0, 1]: BEHAVIOR's 2026 challenge q_score, elsewhere the grader's final_reward. A success always counts as full progress, 1.0, whatever that key says (RoboLab's final_reward is 0 for a task without scored subtasks); mean progress averages it over the graded trials. Agent time adds up finished trials; running ones are shown apart. A benchmark can also declare grader metrics, continuous scores its grader writes (IoU, F1, a distance): they get a column each, a sort, and a run figure, the mean or the median of the graded trials.

Tokens are the model gateway's count of every request (input includes the cached part, output the reasoning part). ≈ $ is those tokens at the list prices in data/prices.yml, an estimate and never a bill; billed is what a provider charged, where it says. The two are never added together.

A benchmark's owner can exclude tasks from the final benchmark (excluded: in state/tasks/<benchmark>.yml). Their trials stay, last in a run's table under Excluded from the benchmark; the numbers count the benchmark's own tasks, and the Count switch adds the excluded ones.

Runs are registered per benchmark: how to register a run.

Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default

25 tasks × 2 modes · results as of 10-01 17:29

Unlimited25 / 25 success
25 success25 done · 100%
Limited8 / 25 success
8 success17 failed25 done · 32%
Mean progress1.00
final_reward · 33 graded
Agent time22.5 h
finished trials

0 requests371.3M in (98% cached)2.8M out≈ $5.79 at list pricebilled: none (subscription)

Run details
Harness
Since 2026-10-01 (batch rt-luna-xh-f1-1001, the official run): Harbor's built-in codex agent (-a codex, Codex CLI 0.159.0) on Azure OpenAI's API (gpt-6-luna), reasoning effort xhigh with detailed reasoning summaries, no model gateway (token counts are Harbor's per-trial totals); the tasks it has not reached yet still show the earlier result: Harbor's built-in codex agent (-a codex; Codex CLI 0.157.1 limited, 0.158.0 unlimited) on the owner's ChatGPT login, not the CodexChatGPT wrapper the other benchmarks used: with the plain model id Codex applies ChatGPT's metadata for gpt-6-luna (its base instructions, verbosity low), and there is no model gateway (token counts are Harbor's per-trial totals)
Model
chatgpt/openai/gpt-6-luna
Route
the owner's ChatGPT login (a Pro subscription), through our model gateway
Settings
reasoning_effort xhigh, version 0.157.0
Modes
unlimited, limited
Batches
rt-luna-xh-0927 rt-luna-xh-rerun-0927 rt-luna-xh-f1-1001
A later batch replaces an earlier one's run of the tasks it reruns; the earlier one stays as history.
Scope
25 tasks in the run · 25 not in it · 0 removed (lists at the end)
Progress
the grader's final_reward, in [0, 1]; a success counts as 1.0
Billing
subscription: a subscription login: no per-token bill
List price
openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
Notes
Limited mode: the RoboLab, RoboPaint and RoboWits results are from the robot service's rcb-limited/2.2 protocol; RoboPaint's and RoboWits' earlier limited round is kept as history, RoboLab's (2.0) was withdrawn. BEHAVIOR-1K's limited results are from its earlier round.
Run id
codex-gpt6_luna-xhigh · data/agents/codex-gpt6_luna-xhigh.yml, state/runs/robotwin-2.yml · formerly codex-0.157-gpt6luna-xhigh-cgpt
Results
as collected 10-01 17:29, published with make publish-runs
ⓘ Notes
  • Pick dual bottles limited rerun: in the first trial the robot service crashed while starting (the batch's first three trials started together), so the agent could send no action
  • Place object stand limited rerun: in the first trial the robot service crashed while starting (the batch's first three trials started together), so the agent could send no action
  • Press stapler unlimited regraded offline: the verifier ran out of GPU memory (OIDN) while other jobs shared the GPU and hung until Harbor's 1800 s verifier timeout, with no verdict; the same grader rerun on the collected trajectory, alone on the GPU, reads success (deterministic)
  • Put bottles dustbin unlimited Harbor reports a verifier timeout, but the verdict (success, deterministic) was written first; the grader ran out of time rendering the replay video of this 15,347-step trajectory, which the same grader rendered afterwards (scripts/regrade_trials.sh)
  • Shake bottle limited rerun: in the first trial the robot service crashed while starting (the batch's first three trials started together), so the agent could send no action
TaskTrial (its log page)ProgressAgent timeRequestsTokens in / outEst. cost
Adjust bottleeasyU✓ 6m1.006m—1.4M / 12k$0.026
L✗ 52m—52m—21.0M / 104k$0.287
Move can poteasyU✓ 3m1.003m—333k / 7k$0.010
L✗ 12m—12m—1.1M / 20k$0.027
Pick dual bottleseasyU✓ 9m1.009m—1.6M / 13k$0.029
L✗ 11m ⓘ prev—11m—807k / 19k$0.022
Place bread skilleteasyU✓ 21m1.0021m—5.3M / 66k$0.097
L✓ 10m1.0010m—2.5M / 53k$0.060
Place mouse padeasyU✓ 6m1.006m—1.1M / 23k$0.028
L✓ 6m1.006m—1.4M / 21k$0.030
Place object standeasyU✓ 5m1.005m—569k / 7k$0.013
L✓ 51m ⓘ prev1.0051m—13.4M / 102k$0.208
Press staplereasyU✓ 5m ⓘ1.005m—601k / 6k$0.013
L✓ 11m1.0011m—1.2M / 23k$0.029
Stack blocks twoeasyU✓ 16m1.0016m—3.1M / 35k$0.058
L✓ 53m1.0053m—33.6M / 123k$0.431
Handover microphonemediumU✓ 21m1.0021m—2.3M / 33k$0.049
L✗ 36m—36m—9.6M / 96k$0.163
Lift potmediumU✓ 2m1.002m—603k / 4k$0.012
L✗ 1h 00m—1h 00m—14.3M / 89k$0.212
Open laptopmediumU✓ 3m1.003m—876k / 8k$0.017
L✓ 41m1.0041m—9.5M / 77k$0.150
Place a to b rightmediumU✓ 9m1.009m—869k / 14k$0.020
L✗ 57m—57m—17.2M / 121k$0.258
Place burger friesmediumU✓ 3m1.003m—588k / 9k$0.014
L✗ 1h 00m—1h 00m—14.9M / 112k$0.222
Place dual shoesmediumU✓ 25m1.0025m—6.4M / 48k$0.103
L✗ 58m—58m—23.6M / 126k$0.339
Place fanmediumU✓ 4m1.004m—785k / 8k$0.017
L✗ 12m—12m—1.8M / 33k$0.042
Place object basketmediumU✓ 23m1.0023m—7.9M / 52k$0.120
L✗ 1h 00m—1h 00m—15.3M / 83k$0.224
Place phone standmediumU✓ 2m prev1.002m—747k / 7k$0.015
L✗ 21m prev—21m—7.8M / 52k$0.114
Rotate QR codemediumU✓ 9m1.009m—1.7M / 16k$0.032
L✗ 1h 00m—1h 00m—25.9M / 133k$0.371
Shake bottlemediumU✓ 4m1.004m—950k / 6k$0.017
L✓ 17m ⓘ prev1.0017m—3.6M / 45k$0.068
Dump bin big binhardU✓ 4m1.004m—1.0M / 9k$0.019
L✓ 40m1.0040m—6.8M / 81k$0.124
Hanging mughardU✓ 34m1.0034m—5.5M / 50k$0.094
L✗ 1h 00m—1h 00m—9.2M / 82k$0.151
Open microwavehardU✓ 28m1.0028m—4.9M / 51k$0.087
L✗ 1h 00m—1h 00m—10.2M / 74k$0.161
Place can baskethardU✓ 31m prev1.0031m—9.0M / 103k$0.158
L✗ 59m prev—59m—21.4M / 159k$0.316
Put bottles dustbinhardU✓ 38m ⓘ prev1.0038m—10.8M / 129k$0.194
L✗ 1h 00m prev—1h 00m—25.6M / 138k$0.346
Scan objecthardU✓ 18m1.0018m—4.5M / 39k$0.075
L✗ 53m—53m—6.5M / 78k$0.119
25 task(s) not in this run
TaskWhy
Beat block hammernot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Blocks ranking rgbnot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Blocks ranking sizenot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Click alarm clocknot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Click bellnot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Grab rollernot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Handover blocknot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Move pill bottle padnot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Move playing card awaynot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Move stapler padnot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Pick diverse bottlesnot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Place a to b leftnot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Place bread basketnot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Place cans plastic boxnot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Place container platenot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Place empty cupnot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Place object scalenot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Place shoenot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Put object cabinetnot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Shake bottle horizontallynot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Stack blocks threenot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Stack bowls threenot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Stack bowls twonot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Stamp sealnot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks
Turn switchnot in this run yet: the ChatGPT-login batches covered the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two; the official batches add sampled tasks

Codex + GPT-6 Luna, reasoning medium (ChatGPT login)closed

22 tasks × 2 modes · results as of 10-01 17:29

Unlimited11 / 22 success
11 success10 failed1 not run21 done · 52%
Limited3 / 22 success
3 success19 failed22 done · 14%
Mean progress1.00
final_reward · 14 graded
Agent time8.9 h
finished trials

0 requests120.0M in (97% cached)784k out≈ $1.90 at list pricebilled: none (subscription)

Run details
Harness
Harbor's built-in codex agent (-a codex) on the owner's ChatGPT login (CODEX_FORCE_AUTH_JSON=1). Harbor installs the latest Codex CLI per trial, so the version varies by batch (0.156.1 to 0.158.0; each batch's version is in state/runs/<benchmark>.yml). Unlike harbor_agents.codex_chatgpt:CodexChatGPT (run codex-gpt6_luna-xhigh), the model id is plain gpt-6-luna, so Codex applies ChatGPT's own metadata for it (its base instructions, verbosity low), and there is no model gateway: token counts are Harbor's per-trial totals.
Model
openai/gpt-6-luna
Route
the owner's ChatGPT login (a Pro subscription), Codex straight to ChatGPT's backend (no model gateway)
Settings
reasoning_effort medium
Modes
unlimited, limited
Batches
rt-luna-med-0923 rt-luna-med-v20-0925 rt-luna-med-v21-0928
A later batch replaces an earlier one's run of the tasks it reruns; the earlier one stays as history.
Scope
22 tasks in the run · 28 not in it · 0 removed (lists at the end)
Progress
the grader's final_reward, in [0, 1]; a success counts as 1.0
Billing
subscription: a subscription login: no per-token bill
List price
openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
Closed
stack_blocks_two was not in the 2026-09-23 unlimited batch
Run id
codex-gpt6_luna-medium · data/agents/codex-gpt6_luna-medium.yml, state/runs/robotwin-2.yml
Results
as collected 10-01 17:29, published with make publish-runs
ⓘ Notes
  • Open microwave limited regraded offline: the live episode succeeded and the service's verification replay succeeded and was deterministic, but rendering the videos of this 40,384-step episode outlasted the collect hook, so Harbor's grade read 0 (a service bug, since fixed); two fresh replays of the recorded trajectory confirm success
TaskTrial (its log page)ProgressAgent timeRequestsTokens in / outEst. cost
Adjust bottleeasyU✓ 3m1.003m—373k / 2k$0.0068
L✗ 8m prev—8m—1.1M / 15k$0.025
Move can poteasyU✓ 9m1.009m—1.4M / 8k$0.023
L✗ 21m prev—21m—6.1M / 31k$0.089
Pick dual bottleseasyU✓ 3m1.003m—913k / 3k$0.015
L✗ 3m prev—3m—483k / 7k$0.011
Place object standeasyU✗ 6m—6m—644k / 3k$0.011
L✓ 5m prev1.005m—770k / 8k$0.017
Press staplereasyU✓ 4m1.004m—400k / 3k$0.0082
L✗ 12m prev—12m—688k / 11k$0.016
Handover microphonemediumU✓ 21m1.0021m—2.1M / 12k$0.034
L✗ 27m prev—27m—11.6M / 54k$0.166
Lift potmediumU✓ 2m1.002m—331k / 2k$0.0069
L✗ 5m prev—5m—493k / 7k$0.011
Open laptopmediumU✓ 8m1.008m—812k / 4k$0.015
L✓ 4m prev1.004m—580k / 6k$0.012
Place a to b rightmediumU✗ 5m—5m—743k / 4k$0.013
L✗ 4m prev—4m—538k / 5k$0.011
Place dual shoesmediumU✗ 6m—6m—380k / 4k$0.0080
L✗ 7m prev—7m—1.2M / 13k$0.024
Place fanmediumU✗ 7m—7m—726k / 6k$0.014
L✗ 19m prev—19m—3.4M / 37k$0.063
Place object basketmediumU✗ 1h 00m—1h 00m—10.4M / 76k$0.166
L✗ 6m prev—6m—1.4M / 12k$0.025
Place phone standmediumU✗ 4m—4m—427k / 3k$0.0085
L✗ 24m prev—24m—6.3M / 34k$0.095
Rotate QR codemediumU✓ 5m1.005m—558k / 4k$0.012
L✗ 8m prev—8m—1.5M / 13k$0.028
Shake bottlemediumU✓ 2m1.002m—457k / 2k$0.0081
L✗ 8m prev—8m—1.3M / 17k$0.027
Dump bin big binhardU✓ 2m1.002m—201k / 2k$0.0049
L✗ 10m prev—10m—1.9M / 21k$0.036
Hanging mughardU✗ 12m—12m—979k / 7k$0.017
L✗ 18m prev—18m—3.7M / 38k$0.066
Open microwavehardU✗ 4m—4m—487k / 3k$0.010
L✓ 39m ⓘ prev1.0039m—15.5M / 68k$0.210
Place can baskethardU✗ 7m—7m—906k / 7k$0.019
L✗ 12m prev—12m—3.7M / 29k$0.061
Put bottles dustbinhardU✗ 26m—26m—2.2M / 20k$0.040
L✗ 50m prev—50m—23.4M / 110k$0.319
Scan objecthardU✓ 6m1.006m—753k / 5k$0.016
L✗ 29m prev—29m—4.7M / 36k$0.077
Stack blocks twoeasyUnot run—————
L✗ 18m prev—18m—3.2M / 31k$0.057
28 task(s) not in this run
TaskWhy
Beat block hammernot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Blocks ranking rgbnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Blocks ranking sizenot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Click alarm clocknot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Click bellnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Grab rollernot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Handover blocknot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Move pill bottle padnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Move playing card awaynot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Move stapler padnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Pick diverse bottlesnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Place a to b leftnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Place bread basketnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Place bread skilletnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Place burger friesnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Place cans plastic boxnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Place container platenot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Place empty cupnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Place mouse padnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Place object scalenot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Place shoenot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Put object cabinetnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Shake bottle horizontallynot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Stack blocks threenot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Stack bowls threenot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Stack bowls twonot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Stamp sealnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two
Turn switchnot built as a harness task: the run covers the 11 kept tasks, 10 sampled dropped ones and stack_blocks_two

Codex + GPT-6 Luna, reasoning xhigh (Azure OpenAI API, 2026-09-30 stress test)

47 tasks × 2 modes · results as of 10-01 17:29

Unlimited42 / 47 success
42 success5 failed47 done · 89%
Limited20 / 47 success
20 success27 failed47 done · 43%
Mean progress1.00
final_reward · 62 graded
Agent time36.2 h
finished trials

0 requests489.0M in (98% cached)4.5M out≈ $8.00 at list pricebilled: —

Run details
Harness
Harbor's built-in codex agent (-a codex, Codex CLI 0.159.0) on Azure OpenAI's API (gpt-6-luna), reasoning effort xhigh with detailed reasoning summaries; limited mode on rcb-limited/2.1 with crash recovery. Most trials ran on an AWS g6e (L40S); place_a2b_left, place_a2b_right, place_object_scale and place_object_stand, whose frozen instances rebuild to another state hash there, ran on the lab machine. No model gateway: token counts are Harbor's
Model
azure/gpt-6-luna
Route
Azure OpenAI deployments of gpt-6-luna, one endpoint per Harbor job: the lab's main resource (1M tokens/min) and two more (333k tokens/min each); each trial talked to one endpoint
Settings
reasoning_effort xhigh, reasoning_summary detailed, version 0.159.0
Modes
unlimited, limited
Batches
rt-luna-xh-az-0930 rt-luna-xh-az-rerun-0930
A later batch replaces an earlier one's run of the tasks it reruns; the earlier one stays as history.
Scope
47 tasks in the run · 3 not in it · 0 removed (lists at the end)
Progress
the grader's final_reward, in [0, 1]; a success counts as 1.0
Billing
api: billed by the provider's API; we hold no per-trial bill, so only the estimate is shown
List price
openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
Notes
Not an official result. A large-batch stress run of every task on robot_coding_bench dev/kangrui (RoboTwin 2.0, DexToolBench, MuJoCo Playground), both modes, limited mode on rcb-limited/2.1 with crash recovery. RoboTwin ran mostly on an AWS g6e (L40S); the four RoboTwin tasks whose frozen instance does not rebuild bit-exactly there, and every MuJoCo task, ran on the lab machine.
Run id
codex-gpt6_luna-xhigh-azure · data/agents/codex-gpt6_luna-xhigh-azure.yml, state/runs/robotwin-2.yml
Results
as collected 10-01 17:29, published with make publish-runs
ⓘ Notes
  • Handover microphone limited rerun: the first trial ended on Azure's token rate limit (Codex retries its remote context compaction only twice)
  • Hanging mug limited rerun: the first trial lost its robot service, which crashed 7 times while starting (another user's job held half the GPU)
  • Lift pot limited rerun: the first trial was stopped by us to move the work to another endpoint or machine
  • Open microwave limited rerun: the first trial lost its robot service, which crashed 4 times while starting (GPU out of memory under load) and ran out of restarts
  • Place a to b left limited rerun: the first trial lost its robot service (infra_error)
  • Place a to b right unlimited rerun: the first trial was stopped by us to move the work to another endpoint or machine
  • Place bread skillet unlimited rerun: the first trial was stopped by us to move the work to another endpoint or machine
  • Place burger fries limited rerun: the first trial was stopped by us to move the work to another endpoint or machine
  • Place dual shoes unlimited regraded offline: the verifier ran out of GPU memory (cuRobo) while other jobs shared one of our machines machine's GPU, so it wrote grader_error; the same grader rerun on the collected trajectory, alone on the GPU, reads success (deterministic)
  • Place object stand unlimited rerun: the first trial was stopped by us to move the work to another endpoint or machine (2 earlier attempts in all)
  • Place phone stand limited rerun: the first trial was stopped by us to move the work to another endpoint or machine
  • Press stapler unlimited regraded offline: the verifier ran out of GPU memory (cuRobo) while other jobs shared one of our machines machine's GPU, so it wrote grader_error; the same grader rerun on the collected trajectory, alone on the GPU, reads success (deterministic)
  • Put bottles dustbin limited rerun: the first trial was stopped by us to move the work to another endpoint or machine
TaskTrial (its log page)ProgressAgent timeRequestsTokens in / outEst. cost
Adjust bottleeasyU✓ 3m1.003m—790k / 8k$0.016
L✗ 42m—42m—2.7M / 44k$0.055
Grab rollereasyU✓ 4m1.004m—831k / 7k$0.015
L✓ 4m1.004m—810k / 21k$0.023
Move can poteasyU✓ 7m1.007m—839k / 12k$0.023
L✓ 16m1.0016m—1.3M / 21k$0.028
Move pill bottle padeasyU✓ 8m1.008m—1.2M / 13k$0.023
L✗ 16m—16m—1.1M / 22k$0.026
Move playing card awayeasyU✓ 6m1.006m—604k / 13k$0.016
L✗ 18m—18m—6.0M / 50k$0.093
Move stapler padeasyU✓ 9m1.009m—1.6M / 13k$0.026
L✗ 10m—10m—1.9M / 33k$0.041
Pick dual bottleseasyU✓ 4m1.004m—467k / 8k$0.012
L✗ 10m—10m—827k / 23k$0.024
Place bread skilleteasyU✓ 18m ⓘ prev1.0018m—4.1M / 50k$0.076
L✗ 13m—13m—3.4M / 65k$0.076
Place container plateeasyU✓ 28m1.0028m—1.7M / 21k$0.033
L✓ 21m1.0021m—2.2M / 34k$0.046
Place empty cupeasyU✗ 1h 00m—1h 00m—16.9M / 91k$0.230
L✓ 11m1.0011m—2.2M / 43k$0.051
Place mouse padeasyU✓ 7m1.007m—753k / 13k$0.019
L✓ 9m1.009m—2.0M / 36k$0.044
Place object scaleeasyU✓ 3m1.003m—464k / 7k$0.011
L✓ 27m1.0027m—6.0M / 50k$0.094
Place object standeasyU✓ 14m ⓘ prev1.0014m—630k / 8k$0.014
L✓ 29m1.0029m—9.8M / 60k$0.147
Place shoeeasyU✓ 10m1.0010m—418k / 8k$0.011
L✗ 21m—21m—6.6M / 67k$0.111
Press staplereasyU✓ 7m ⓘ1.007m—1.9M / 21k$0.036
L✓ 3m1.003m—625k / 14k$0.017
Stack blocks twoeasyU✓ 16m1.0016m—2.6M / 35k$0.050
L✓ 37m1.0037m—9.6M / 100k$0.160
Stack bowls twoeasyU✓ 9m1.009m—1.1M / 14k$0.023
L✗ 18m—18m—3.5M / 48k$0.067
Turn switcheasyU✓ 8m1.008m—752k / 16k$0.020
L✓ 33m1.0033m—8.8M / 94k$0.151
Beat block hammermediumU✓ 4m1.004m—878k / 9k$0.018
L✗ 41m—41m—12.0M / 93k$0.179
Blocks ranking rgbmediumU✓ 8m1.008m—536k / 13k$0.016
L✓ 6m1.006m—1.1M / 18k$0.025
Blocks ranking sizemediumU✓ 12m1.0012m—2.3M / 21k$0.040
L✗ 1h 00m—1h 00m—7.6M / 54k$0.124
Handover blockmediumU✓ 11m1.0011m—2.1M / 24k$0.039
L✗ 1h 00m—1h 00m—12.1M / 111k$0.194
Handover microphonemediumU✓ 21m1.0021m—3.1M / 37k$0.058
L✗ 17m ⓘ prev—17m—5.5M / 62k$0.096
Lift potmediumU✓ 7m1.007m—1.0M / 12k$0.021
L✗ 23m ⓘ prev—23m—4.1M / 57k$0.086
Open laptopmediumU✓ 5m1.005m—1.3M / 10k$0.023
L✓ 51m1.0051m—14.9M / 172k$0.260
Pick diverse bottlesmediumU✓ 3m1.003m—590k / 6k$0.012
L✗ 7m—7m—1.7M / 25k$0.036
Place a to b leftmediumU✓ 4m1.004m—846k / 12k$0.019
L✓ 21m ⓘ prev1.0021m—3.1M / 43k$0.060
Place a to b rightmediumU✓ 7m ⓘ prev1.007m—589k / 7k$0.013
L✗ 23m—23m—2.0M / 53k$0.054
Place bread basketmediumU✓ 13m1.0013m—3.4M / 44k$0.066
L✓ 21m1.0021m—3.5M / 41k$0.064
Place burger friesmediumU✓ 5m1.005m—361k / 9k$0.012
L✗ 1h 00m ⓘ prev—1h 00m—13.4M / 96k$0.196
Place cans plastic boxmediumU✓ 18m1.0018m—3.5M / 38k$0.063
L✗ 59m—59m—14.3M / 83k$0.198
Place dual shoesmediumU✓ 12m ⓘ1.0012m—3.5M / 36k$0.062
L✓ 39m1.0039m—4.0M / 58k$0.078
Place fanmediumU✓ 26m1.0026m—3.4M / 31k$0.056
L✗ 1h 00m—1h 00m—21.0M / 107k$0.283
Place object basketmediumU✓ 23m1.0023m—4.4M / 42k$0.073
L✗ 43m—43m—13.8M / 99k$0.202
Place phone standmediumU✓ 7m1.007m—1.2M / 11k$0.022
L✓ 21m ⓘ prev1.0021m—2.8M / 41k$0.055
Rotate QR codemediumU✓ 4m1.004m—551k / 10k$0.015
L✗ 1h 00m—1h 00m—19.1M / 110k$0.272
Shake bottlemediumU✓ 4m1.004m—525k / 6k$0.011
L✓ 8m1.008m—1.7M / 28k$0.036
Shake bottle horizontallymediumU✓ 6m1.006m—1.1M / 12k$0.022
L✓ 17m1.0017m—5.5M / 49k$0.088
Stack blocks threemediumU✓ 57m1.0057m—11.0M / 120k$0.187
L✗ 1h 00m—1h 00m—29.1M / 140k$0.395
Stack bowls threemediumU✗ 1h 00m—1h 00m—11.3M / 72k$0.163
L✗ 1h 00m—1h 00m—15.9M / 84k$0.213
Stamp sealmediumU✓ 9m1.009m—731k / 9k$0.015
L✗ 38m—38m—5.3M / 51k$0.087
Dump bin big binhardU✓ 7m1.007m—1.4M / 17k$0.027
L✗ 12m—12m—3.2M / 49k$0.065
Hanging mughardU✗ 1h 00m—1h 00m—19.6M / 131k$0.283
L✗ 45m ⓘ prev—45m—16.7M / 130k$0.250
Open microwavehardU✓ 20m1.0020m—2.0M / 39k$0.048
L✓ 27m ⓘ prev1.0027m—7.5M / 133k$0.161
Place can baskethardU✗ 1h 00m—1h 00m—17.3M / 132k$0.259
L✓ 10m1.0010m—3.1M / 47k$0.062
Put bottles dustbinhardU✗ 1h 00m—1h 00m—8.5M / 79k$0.147
L✗ 1h 00m ⓘ prev—1h 00m—10.7M / 72k$0.160
Scan objecthardU✓ 23m1.0023m—3.4M / 47k$0.066
L✗ 49m—49m—17.0M / 137k$0.259
3 task(s) not in this run
TaskWhy
Click alarm clocknot built as a harness task: click_bell (success is one instant of contact, which a replay does not reproduce), click_alarmclock and put_object_cabinet (the expert demonstration fails on all 6 seeds)
Click bellnot built as a harness task: click_bell (success is one instant of contact, which a replay does not reproduce), click_alarmclock and put_object_cabinet (the expert demonstration fails on all 6 seeds)
Put object cabinetnot built as a harness task: click_bell (success is one instant of contact, which a replay does not reproduce), click_alarmclock and put_object_cabinet (the expert demonstration fails on all 6 seeds)