Skip to content

HumanoidBench runs

Agent runs on the tasks of HumanoidBench, one run at a time.

How to read this page

A run is one agent configuration, harness + model + settings; each task in it runs once per mode. Pick a run above; the table lists every task of it, one row per mode, in the default order (finished first, then running, easier tasks first). A pill links to that trial's log page. Filter the table with the chips above it.

successfailederror / stalledrunninggradingqueuednot run

Bars count trials, one per task and mode, by state. Progress is the partial credit beside success, in [0, 1]: BEHAVIOR's 2026 challenge q_score, elsewhere the grader's final_reward. A success always counts as full progress, 1.0, whatever that key says (RoboLab's final_reward is 0 for a task without scored subtasks); mean progress averages it over the graded trials. Agent time adds up finished trials; running ones are shown apart. A benchmark can also declare grader metrics, continuous scores its grader writes (IoU, F1, a distance): they get a column each, a sort, and a run figure, the mean or the median of the graded trials.

Tokens are the model gateway's count of every request (input includes the cached part, output the reasoning part). ≈ $ is those tokens at the list prices in data/prices.yml, an estimate and never a bill; billed is what a provider charged, where it says. The two are never added together.

A benchmark's owner can exclude tasks from the final benchmark (excluded: in state/tasks/<benchmark>.yml). Their trials stay, last in a run's table under Excluded from the benchmark; the numbers count the benchmark's own tasks, and the Count switch adds the excluded ones.

Runs are registered per benchmark: how to register a run.

Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default

24 tasks × 2 modes · 15 of them excluded from the benchmark · results as of 10-04 08:02

Count15 of the run's tasks are excluded from the benchmark: last in the table
Unlimited0 / 9 success
9 failed9 done · 0%
Limited0 / 9 success
9 failed9 done · 0%
Mean progress—
none · 0 graded
Agent time9.4 h
finished trials

0 requests282.8M in (98% cached)2.2M out≈ $4.56 at list pricebilled: —

Grader metrics
Resets used0mean · median 0 · 9 graded
Run details
Harness
Two harnesses, by protocol. Before robot_coding_bench's protocol v1.0: Codex CLI 0.157.0 (HumanoidBench) or 0.158.0–0.159.2 (KinDER) as harbor_agents.codex_openrouter:CodexOpenRouter (dev/pingyue, commit c9449ce), Codex over OpenRouter with web search off, without the model gateway. On protocol v1.0 (2026-10-01): Harbor's built-in codex agent, Codex CLI 0.159.2, with OpenRouter registered as its own provider and web search disabled (scripts/<benchmark>/configs/codex-openrouter.yaml); from 2026-10-04 (HumanoidBench and KinDER's hardest tasks) started by scripts/run_agent_batch.sh, with the same Codex config passed as AGENT_CONFIG and no usage gateway. state/runs/<benchmark>.yml says which batch is which.
Model
openrouter/openai/gpt-6-luna
Route
OpenRouter, with the owner's API key (billed per token)
Settings
reasoning_effort medium, version 0.157.0–0.159.2
Modes
unlimited, limited
Batches
hb-luna-0928 hb-luna-v1-1001 hb-luna-1004
A later batch replaces an earlier one's run of the tasks it reruns; the earlier one stays as history.
Scope
9 tasks in the run + 15 excluded from the benchmark (last in the table) · 0 not in it · 0 removed (lists at the end)
Progress
the grader's none, in [0, 1]; a success counts as 1.0
Billing
openrouter: billed per request by OpenRouter (the charge it reports in every response)
List price
openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
Run id
codex-gpt6_luna-medium-openrouter · data/agents/codex-gpt6_luna-medium-openrouter.yml, state/runs/humanoidbench.yml
Results
as collected 10-04 08:02, published with make publish-runs
ⓘ Notes
  • Locomotion · Stair unlimited the agent ended itself at 48 min, its trajectory saved, judging that stair traversal stayed unsolved
TaskTrial (its log page)ProgressResets usedAgent timeRequestsTokens in / outEst. cost
Locomotion · Balance hardU✗ 1h 00m prev——1h 00m—31.3M / 216k$0.478
L✗ 2m prev—02m—398k / 15k$0.015
Locomotion · HurdleU✗ 59m prev——59m—51.5M / 268k$0.738
L✗ 1m prev—01m—246k / 6k$0.0080
Locomotion · StairU✗ 48m ⓘ prev——48m—27.6M / 226k$0.448
L✗ 2m prev—02m—525k / 14k$0.016
Manipulation · Bookshelf simpleU✗ 1h 00m prev——1h 00m—29.5M / 264k$0.487
L✗ 7m prev—07m—2.2M / 36k$0.045
Manipulation · Highbar simpleU✗ 59m prev——59m—24.5M / 162k$0.377
L✗ 7m prev—07m—3.0M / 40k$0.059
Manipulation · PackageU✗ 1h 00m prev——1h 00m—25.8M / 242k$0.469
L✗ 8m prev—08m—3.7M / 42k$0.066
Manipulation · PowerliftU✗ 1h 00m prev——1h 00m—34.5M / 281k$0.580
L✗ 4m prev—04m—953k / 23k$0.026
Manipulation · RoomU✗ 1h 00m prev——1h 00m—25.1M / 247k$0.430
L✗ 2m prev—02m—244k / 13k$0.012
Manipulation · TruckU✗ 59m prev——59m—19.6M / 110k$0.266
L✗ 7m prev—07m—2.2M / 38k$0.047
excluded Excluded from the benchmark · 15 tasks, kept for the record; the benchmark's numbers leave them out
Locomotion · Balance simpleU✗ 59m——59m—22.9M / 157k$0.328
L✗ 16m prev—5016m—5.0M / 62k$0.092
Locomotion · MazeU✗ 1h 00m——1h 00m—27.1M / 203k$0.430
L✗ 45m prev—5045m—20.8M / 134k$0.308
Locomotion · ReachU✓ 59m1.00—59m—26.5M / 151k$0.362
L✗ 20m prev—020m—5.5M / 71k$0.107
Locomotion · Sit hardU✓ 58m1.00—58m—16.9M / 78k$0.222
L✓ 6m prev1.0046m—1.1M / 21k$0.027
Locomotion · Sit simpleU✓ 59m1.00—59m—32.7M / 111k$0.403
L✓ 3m1.0003m—583k / 15k$0.017
Locomotion · WalkU✗ 1h 00m——1h 00m—25.2M / 193k$0.407
L✗ 18m prev—5018m—3.9M / 77k$0.088
Manipulation · BasketballU✗ 1h 00m——1h 00m—24.4M / 219k$0.407
L✗ 1h 00m prev—431h 00m—29.7M / 175k$0.406
Manipulation · CabinetU✗ 58m——58m—18.0M / 177k$0.317
L✗ 1h 00m prev—331h 00m—36.1M / 141k$0.451
Manipulation · CubeU✗ 58m——58m—33.2M / 137k$0.446
L✗ 36m prev—5036m—15.8M / 126k$0.238
Manipulation · DoorU✗ 1h 00m——1h 00m—24.2M / 238k$0.419
L✗ 1h 00m prev—361h 00m—30.0M / 171k$0.433
Manipulation · Insert normalU✗ 1h 00m——1h 00m—28.4M / 232k$0.458
L✗ 1h 00m prev—241h 00m—32.3M / 139k$0.411
Manipulation · KitchenU✗ 57m——57m—22.2M / 212k$0.385
L✗ 1h 00m prev—351h 00m—30.7M / 169k$0.414
Manipulation · PushU✓ 36m prev1.00—36m—30.0M / 131k$0.387
L✓ 24m prev1.001024m—8.1M / 55k$0.120
Manipulation · SpoonU✗ 1h 00m——1h 00m—21.5M / 239k$0.391
L✗ 1h 00m prev—241h 00m—32.7M / 162k$0.455
Manipulation · WindowU✗ 58m——58m—32.6M / 228k$0.494
L✗ 1h 00m prev—371h 00m—32.9M / 133k$0.414

Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)

24 tasks × 2 modes · 15 of them excluded from the benchmark · results as of 10-04 08:02

Count15 of the run's tasks are excluded from the benchmark: last in the table
Unlimited0 / 9 success
9 failed9 done · 0%
Limited0 / 9 success
9 failed9 done · 0%
Mean progress—
none · 0 graded
Agent time17.2 h
finished trials

2,199 requests217.7M in (98% cached)1.9M out (1.3M reasoning)≈ $49.40 at list pricebilled: $51.64

Grader metrics
Resets used49mean · median 50 · 9 graded
Run details
Harness
Codex CLI 0.159.2 as harbor_agents.codex_openrouter:CodexOpenRouter (robot_coding_bench dev/pingyue, commit e1f8660f): Codex over OpenRouter with web search off; the API key stays in the model gateway, the agent's container only sees a placeholder
Model
openrouter/openai/gpt-6.1-sol
Route
OpenRouter, with the owner's API key (billed per token), through our model gateway
Settings
reasoning_effort medium, version 0.159.2
Modes
unlimited, limited
Batches
hb-sol61-0930
Scope
9 tasks in the run + 15 excluded from the benchmark (last in the table) · 0 not in it · 0 removed (lists at the end)
Progress
the grader's none, in [0, 1]; a success counts as 1.0
Billing
openrouter: billed per request by OpenRouter (the charge it reports in every response)
List price
openai/gpt-6.1-sol: $2.0 in / $0.1 cached / $10.0 out per 1M tokens, as of 2026-09-30
Run id
codex-gpt6_1_sol-medium · data/agents/codex-gpt6_1_sol-medium.yml, state/runs/humanoidbench.yml
Results
as collected 10-04 08:02, published with make publish-runs
TaskTrial (its log page)ProgressResets usedAgent timeRequestsTokens in / outEst. costBilled
Locomotion · Balance hardU✗ 1h 00m——1h 00m12312.5M / 125k$2.90$3.01
L✗ 57m—5057m12710.2M / 112k$2.49$2.58
Locomotion · HurdleU✗ 59m——59m1078.7M / 96k$2.15$2.23
L✗ 30m—5030m735.5M / 62k$1.45$1.52
Locomotion · StairU✗ 1h 00m——1h 00m13212.4M / 125k$2.85$2.95
L✗ 1h 00m—491h 00m947.5M / 123k$2.33$2.43
Manipulation · Bookshelf simpleU✗ 1h 00m——1h 00m13714.8M / 120k$3.08$3.18
L✗ 1h 00m—501h 00m16818.3M / 92k$3.16$3.26
Manipulation · Highbar simpleU✗ 1h 00m——1h 00m1109.7M / 109k$2.41$2.50
L✗ 1h 00m—501h 00m13714.4M / 112k$3.66$3.95
Manipulation · PackageU✗ 1h 00m——1h 00m13012.4M / 123k$2.87$2.98
L✗ 1h 00m—441h 00m12415.0M / 104k$2.98$3.10
Manipulation · PowerliftU✗ 1h 00m——1h 00m13714.0M / 111k$2.89$2.99
L✗ 1h 00m—481h 00m14516.5M / 95k$3.53$3.78
Manipulation · RoomU✗ 1h 00m——1h 00m1069.6M / 121k$2.54$2.64
L✗ 54m—5054m12712.3M / 87k$2.44$2.53
Manipulation · TruckU✗ 1h 00m——1h 00m12013.3M / 108k$2.80$2.91
L✗ 50m—5050m10210.4M / 87k$2.87$3.12
excluded Excluded from the benchmark · 15 tasks, kept for the record; the benchmark's numbers leave them out
Locomotion · Balance simpleU✓ 1h 00m1.00—1h 00m13710.5M / 68k$1.99$2.06
L✗ 52m—5052m1179.1M / 106k$2.28$2.36
Locomotion · MazeU✓ 1h 00m1.00—1h 00m11712.3M / 84k$2.39$2.48
L✗ 54m—5054m1089.6M / 108k$2.39$2.48
Locomotion · ReachU✓ 59m1.00—59m12410.5M / 107k$2.46$2.55
L✓ 16m1.002916m564.6M / 32k$1.06$1.13
Locomotion · Sit hardU✓ 1h 00m1.00—1h 00m1067.8M / 88k$1.94$2.02
L✓ 7m1.00117m431.4M / 15k$0.395$0.421
Locomotion · Sit simpleU✓ 1h 00m1.00—1h 00m14813.2M / 73k$2.32$2.39
L✓ 1m1.0001m8165k / 2k$0.087$0.101
Locomotion · WalkU✓ 1h 00m1.00—1h 00m13311.9M / 87k$2.35$2.42
L✗ 1h 00m—491h 00m1068.6M / 116k$2.32$2.40
Manipulation · BasketballU✓ 1h 00m1.00—1h 00m15714.1M / 105k$2.79$2.87
L✗ 58m—5058m16014.6M / 81k$2.58$2.67
Manipulation · CabinetU✗ 1h 00m——1h 00m13015.1M / 107k$3.02$3.13
L✗ 1h 00m—271h 00m13112.6M / 101k$2.69$2.79
Manipulation · CubeU✓ 1h 00m1.00—1h 00m12110.6M / 99k$2.37$2.45
L✗ 1h 00m—481h 00m17119.5M / 100k$3.37$3.48
Manipulation · DoorU✗ 1h 00m——1h 00m12415.3M / 109k$3.09$3.21
L✗ 1h 00m—501h 00m12112.7M / 117k$3.38$3.63
Manipulation · Insert normalU✗ 1h 00m——1h 00m12514.2M / 117k$3.00$3.11
L✗ 1h 00m—501h 00m13313.6M / 96k$2.70$2.80
Manipulation · KitchenU✓ 1h 00m1.00—1h 00m13513.8M / 102k$2.76$2.85
L✗ 1h 00m—161h 00m13112.9M / 94k$2.62$2.73
Manipulation · PushU✓ 1h 00m1.00—1h 00m1199.9M / 88k$2.16$2.24
L✓ 14m1.001514m432.2M / 24k$0.625$0.670
Manipulation · SpoonU✗ 1h 00m——1h 00m11713.7M / 110k$2.87$2.98
L✗ 1h 00m—401h 00m16117.2M / 81k$2.94$3.05
Manipulation · WindowU✗ 59m——59m12114.1M / 112k$2.95$3.06
L✗ 51m—5051m12312.0M / 78k$2.32$2.41

Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)

5 tasks × 2 modes · 5 of them excluded from the benchmark · results as of 10-04 08:02

Count5 of the run's tasks are excluded from the benchmark: last in the table
No task of the benchmark proper in this run yet: its 5 tasks are excluded from the benchmark.
Run details
Harness
Claude Code 2.1.283 as harbor_agents.claude_code_openrouter:ClaudeCodeOpenRouter (robot_coding_bench dev/pingyue, commit c9449ce): Claude Code over OpenRouter, without the model gateway; in limited mode the adapter ends the trial when the session ends (harbor_agents/limited.py)
Model
openrouter/anthropic/claude-opus-5.5
Route
OpenRouter, with the owner's API key (billed per token)
Settings
reasoning_effort medium, version 2.1.283
Modes
unlimited, limited
Batches
hb-opus-0928
Scope
0 tasks in the run + 5 excluded from the benchmark (last in the table) · 19 not in it · 0 removed (lists at the end)
Progress
the grader's none, in [0, 1]; a success counts as 1.0
Billing
openrouter: billed per request by OpenRouter (the charge it reports in every response)
List price
anthropic/claude-opus-5.5: $6.25 in / $0.5 cached / $25.0 out per 1M tokens, as of 2026-09-28
Run id
claude_code-claude_opus5_5-medium · data/agents/claude_code-claude_opus5_5-medium.yml, state/runs/humanoidbench.yml
Results
as collected 10-04 08:02, published with make publish-runs
ⓘ Notes
  • Manipulation · Cube unlimited the agent ended itself at 17 min, believing its time was nearly up; it never checked the clock
  • Manipulation · Cube limited the agent ended itself at 6 min with 13 of its 50 resets left
  • Manipulation · Door limited `spec()` placed the hand 0.31 m beyond the real one (a massless link's centre of mass; fixed since); taking it for the fingertip, the agent concluded its arms pass through the door and ended itself at 14 min with 16 resets left
TaskTrial (its log page)ProgressResets usedAgent timeRequestsTokens in / outEst. cost
excluded Excluded from the benchmark · 5 tasks, kept for the record; the benchmark's numbers leave them out
Locomotion · Sit hardU✓ 5m1.00—5m—829k / 14k$0.939
L✓ 6m1.00216m—1.5M / 22k$1.55
Locomotion · WalkU✓ 11m1.00—11m—574k / 11k$0.726
L✗ 14m—5014m—4.6M / 50k$4.14
Manipulation · CubeU✗ 17m ⓘ——17m—1.5M / 23k$1.57
L✗ 6m ⓘ—376m—1.4M / 20k$1.44
Manipulation · DoorU✓ 49m1.00—49m—5.8M / 55k$4.81
L✗ 14m ⓘ—3414m—4.4M / 55k$4.07
Manipulation · PushU✓ 4m1.00—4m—863k / 14k$0.992
L✓ 13m1.00913m—5.1M / 43k$4.09
19 task(s) not in this run
TaskWhy
Locomotion · Balance hardnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Locomotion · Hurdlenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Locomotion · Stairnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Manipulation · Bookshelf simplenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Manipulation · Highbar simplenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Manipulation · Packagenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Manipulation · Powerliftnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Manipulation · Roomnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Manipulation · Trucknot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Locomotion · Balance simplenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Locomotion · Mazenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Locomotion · Reachnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Locomotion · Sit simplenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Manipulation · Basketballnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Manipulation · Cabinetnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Manipulation · Insert normalnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Manipulation · Kitchennot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Manipulation · Spoonnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Manipulation · Windownot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark

Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed

5 tasks × 2 modes · 5 of them excluded from the benchmark · results as of 10-04 08:02

Count5 of the run's tasks are excluded from the benchmark: last in the table
No task of the benchmark proper in this run yet: its 5 tasks are excluded from the benchmark.
Run details
Harness
Codex CLI 0.157.0 as harbor_agents.codex_openrouter:CodexOpenRouter (robot_coding_bench dev/pingyue, commit c9449ce): Codex over OpenRouter with web search off, without the model gateway; in limited mode the adapter ends the trial when the session ends (harbor_agents/limited.py)
Model
openrouter/openai/gpt-6-sol
Route
OpenRouter, with the owner's API key (billed per token)
Settings
reasoning_effort medium, version 0.157.0
Modes
unlimited, limited
Batches
hb-sol-0928
Scope
0 tasks in the run + 5 excluded from the benchmark (last in the table) · 19 not in it · 0 removed (lists at the end)
Progress
the grader's none, in [0, 1]; a success counts as 1.0
Billing
openrouter: billed per request by OpenRouter (the charge it reports in every response)
List price
openai/gpt-6-sol: $2.5 in / $0.2 cached / $10.0 out per 1M tokens, as of 2026-09-28
Closed
the OpenRouter credit ran out on 2026-09-28 before the unlimited runs on door, push and cube
Run id
codex-gpt6_sol-medium · data/agents/codex-gpt6_sol-medium.yml, state/runs/humanoidbench.yml
Results
as collected 10-04 08:02, published with make publish-runs
ⓘ Notes
  • Manipulation · Door limited `spec()` placed the hand 0.31 m beyond the real one (a massless link's centre of mass; fixed since), and the agent aimed its hand at the handle with it
TaskTrial (its log page)ProgressResets usedAgent timeRequestsTokens in / outEst. cost
excluded Excluded from the benchmark · 5 tasks, kept for the record; the benchmark's numbers leave them out
Locomotion · Sit hardU✓ 1h 00m1.00—1h 00m—18.6M / 65k$4.86
L✗ 10m—5010m—1.6M / 18k$0.616
Locomotion · WalkU✗ 58m——58m—17.7M / 83k$4.74
L✗ 12m—5012m—4.2M / 21k$1.26
Manipulation · CubeUnot run——————
L✗ 11m—5011m—1.7M / 22k$0.668
Manipulation · DoorUnot run——————
L✗ 36m ⓘ—5036m—8.3M / 66k$2.60
Manipulation · PushUnot run——————
L✓ 18m1.00718m—4.8M / 26k$1.41
19 task(s) not in this run
TaskWhy
Locomotion · Balance hardnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Locomotion · Hurdlenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Locomotion · Stairnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Manipulation · Bookshelf simplenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Manipulation · Highbar simplenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Manipulation · Packagenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Manipulation · Powerliftnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Manipulation · Roomnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Manipulation · Trucknot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Locomotion · Balance simplenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Locomotion · Mazenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Locomotion · Reachnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Locomotion · Sit simplenot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Manipulation · Basketballnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Manipulation · Cabinetnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Manipulation · Insert normalnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Manipulation · Kitchennot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Manipulation · Spoonnot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark
Manipulation · Windownot in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard) · excluded from the benchmark