Skip to content

VLABench runs

Agent runs on the tasks of VLABench, one run at a time.

How to read this page

A run is one agent configuration, harness + model + settings; each task in it runs once per mode. Pick a run above; the table lists every task of it, one row per mode, in the default order (finished first, then running, easier tasks first). A pill links to that trial's log page. Filter the table with the chips above it.

successfailederror / stalledrunninggradingqueuednot run

Bars count trials, one per task and mode, by state. Progress is the partial credit beside success, in [0, 1]: BEHAVIOR's 2026 challenge q_score, elsewhere the grader's final_reward. A success always counts as full progress, 1.0, whatever that key says (RoboLab's final_reward is 0 for a task without scored subtasks); mean progress averages it over the graded trials. Agent time adds up finished trials; running ones are shown apart. A benchmark can also declare grader metrics, continuous scores its grader writes (IoU, F1, a distance): they get a column each, a sort, and a run figure, the mean or the median of the graded trials.

Tokens are the model gateway's count of every request (input includes the cached part, output the reasoning part). ≈ $ is those tokens at the list prices in data/prices.yml, an estimate and never a bill; billed is what a provider charged, where it says. The two are never added together.

A benchmark's owner can exclude tasks from the final benchmark (excluded: in state/tasks/<benchmark>.yml). Their trials stay, last in a run's table under Excluded from the benchmark; the numbers count the benchmark's own tasks, and the Count switch adds the excluded ones.

Runs are registered per benchmark: how to register a run.

Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)defaultclosed

4 tasks × 2 modes · results as of 10-02 15:07

Unlimited4 / 4 success
4 success4 done · 100%
Limited2 / 4 success
2 success2 failed4 done · 50%
Mean progress1.00
none · 6 graded
Agent time3.6 h
finished trials

0 requests47.0M in (97% cached)404k out≈ $0.784 at list pricebilled: none (subscription)

Run details
Harness
Codex CLI 0.159.3 as Harbor's codex agent (ChatGPT login, no model gateway)
Model
chatgpt/openai/gpt-6-luna
Route
the owner's ChatGPT login (a Pro subscription), through our model gateway
Settings
reasoning_effort xhigh, version 0.157.0
Modes
unlimited, limited
Batches
vlabench-lunaxh-1001
Scope
4 tasks in the run · 32 not in it · 0 removed (lists at the end)
Progress
the grader's none, in [0, 1]; a success counts as 1.0
Billing
subscription: a subscription login: no per-token bill
List price
openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
Closed
the sample model trial of PR
Notes
Limited mode: the RoboLab, RoboPaint and RoboWits results are from the robot service's rcb-limited/2.2 protocol; RoboPaint's and RoboWits' earlier limited round is kept as history, RoboLab's (2.0) was withdrawn. BEHAVIOR-1K's limited results are from its earlier round.
Run id
codex-gpt6_luna-xhigh · data/agents/codex-gpt6_luna-xhigh.yml, state/runs/vlabench.yml · formerly codex-0.157-gpt6luna-xhigh-cgpt
Results
as collected 10-02 15:07, published with make publish-runs
TaskTrial (its log page)ProgressAgent timeRequestsTokens in / outEst. cost
Manipulation · heat_foodU✓ 35m1.0035m—11.5M / 54k$0.164
L✗ 1h 00m—1h 00m—15.4M / 127k$0.240
Manipulation · insert_flowerU✓ 11m1.0011m—949k / 13k$0.035
L✗ 55m—55m—10.2M / 102k$0.173
Manipulation · select_fruitU✓ 2m1.002m—214k / 3k$0.0057
L✓ 30m1.0030m—6.3M / 54k$0.104
Physical QA · friction_qaU✓ 6m1.006m—274k / 7k$0.0094
L✓ 15m1.0015m—2.2M / 43k$0.053
32 task(s) not in this run
TaskWhy
Common sense · select_billiards_common_sensenot in the sample (robot_coding_bench PR
Common sense · select_chemistry_tube_common_sensenot in the sample (robot_coding_bench PR
Common sense · select_fruit_common_sensenot in the sample (robot_coding_bench PR
Manipulation · add_condimentnot in the sample (robot_coding_bench PR
Manipulation · book_rearrangenot in the sample (robot_coding_bench PR
Manipulation · find_unseen_objectnot in the sample (robot_coding_bench PR
Manipulation · hammer_nail_and_hang_picturenot in the sample (robot_coding_bench PR
Manipulation · insert_bloom_flowernot in the sample (robot_coding_bench PR
Manipulation · play_math_gamenot in the sample (robot_coding_bench PR
Manipulation · put_box_on_paintingnot in the sample (robot_coding_bench PR
Manipulation · select_billiardsnot in the sample (robot_coding_bench PR
Manipulation · select_chemistry_tubenot in the sample (robot_coding_bench PR
Manipulation · select_drinknot in the sample (robot_coding_bench PR
Manipulation · select_mahjongnot in the sample (robot_coding_bench PR
Manipulation · select_nth_largest_pokernot in the sample (robot_coding_bench PR
Manipulation · select_paintingnot in the sample (robot_coding_bench PR
Manipulation · select_pokernot in the sample (robot_coding_bench PR
Manipulation · select_specific_type_booknot in the sample (robot_coding_bench PR
Manipulation · select_toynot in the sample (robot_coding_bench PR
Manipulation · select_unique_type_mahjongnot in the sample (robot_coding_bench PR
Manipulation · take_out_cool_drinknot in the sample (robot_coding_bench PR
Manipulation · texas_holdemnot in the sample (robot_coding_bench PR
Physical QA · magnetism_qanot in the sample (robot_coding_bench PR
Physical QA · reflection_qanot in the sample (robot_coding_bench PR
Semantic · select_chemistry_tube_semanticnot in the sample (robot_coding_bench PR
Semantic · select_mahjong_semanticnot in the sample (robot_coding_bench PR
Semantic · select_poker_semanticnot in the sample (robot_coding_bench PR
Spatial · select_billiards_spatialnot in the sample (robot_coding_bench PR
Spatial · select_book_spatialnot in the sample (robot_coding_bench PR
Spatial · select_chemistry_tube_spatialnot in the sample (robot_coding_bench PR
Spatial · select_mahjong_spatialnot in the sample (robot_coding_bench PR
Spatial · select_poker_spatialnot in the sample (robot_coding_bench PR