VLABench runs¶
Agent runs on the tasks of VLABench, one run at a time.
How to read this page
A run is one agent configuration, harness + model + settings; each task in it runs once per mode. Pick a run above; the table lists every task of it, one row per mode, in the default order (finished first, then running, easier tasks first). A pill links to that trial's log page. Filter the table with the chips above it.
successfailederror / stalledrunninggradingqueuednot run
Bars count trials, one per task and mode, by state. Progress is the partial credit beside success, in [0, 1]: BEHAVIOR's 2026 challenge q_score, elsewhere the grader's final_reward. A success always counts as full progress, 1.0, whatever that key says (RoboLab's final_reward is 0 for a task without scored subtasks); mean progress averages it over the graded trials. Agent time adds up finished trials; running ones are shown apart. A benchmark can also declare grader metrics, continuous scores its grader writes (IoU, F1, a distance): they get a column each, a sort, and a run figure, the mean or the median of the graded trials.
Tokens are the model gateway's count of every request (input includes the cached part, output the reasoning part). ≈ $ is those tokens at the list prices in data/prices.yml, an estimate and never a bill; billed is what a provider charged, where it says. The two are never added together.
A benchmark's owner can exclude tasks from the final benchmark (excluded: in state/tasks/<benchmark>.yml). Their trials stay, last in a run's table under Excluded from the benchmark; the numbers count the benchmark's own tasks, and the Count switch adds the excluded ones.
Runs are registered per benchmark: how to register a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)defaultclosed
0 requests47.0M in (97% cached)404k out≈ $0.784 at list pricebilled: none (subscription)
Run details
- Harness
- Codex CLI 0.159.3 as Harbor's codex agent (ChatGPT login, no model gateway)
- Model
chatgpt/openai/gpt-6-luna- Route
- the owner's ChatGPT login (a Pro subscription), through our model gateway
- Settings
- reasoning_effort xhigh, version 0.157.0
- Modes
- unlimited, limited
- Batches
vlabench-lunaxh-1001- Scope
- 4 tasks in the run · 32 not in it · 0 removed (lists at the end)
- Progress
- the grader's none, in [0, 1]; a success counts as 1.0
- Billing
- subscription: a subscription login: no per-token bill
- List price
- openai/gpt-6-luna: $0.1 in / $0.01 cached / $0.5 out per 1M tokens, as of 2026-09-25
- Closed
- the sample model trial of PR
- Notes
- Limited mode: the RoboLab, RoboPaint and RoboWits results are from the robot service's rcb-limited/2.2 protocol; RoboPaint's and RoboWits' earlier limited round is kept as history, RoboLab's (2.0) was withdrawn. BEHAVIOR-1K's limited results are from its earlier round.
- Run id
codex-gpt6_luna-xhigh·data/agents/codex-gpt6_luna-xhigh.yml,state/runs/vlabench.yml· formerlycodex-0.157-gpt6luna-xhigh-cgpt- Results
- as collected 10-02 15:07, published with
make publish-runs
| Task | Trial (its log page) | Progress | Agent time | Requests | Tokens in / out | Est. cost |
|---|---|---|---|---|---|---|
| Manipulation · heat_food | U✓ 35m | 1.00 | 35m | — | 11.5M / 54k | $0.164 |
| L✗ 1h 00m | — | 1h 00m | — | 15.4M / 127k | $0.240 | |
| Manipulation · insert_flower | U✓ 11m | 1.00 | 11m | — | 949k / 13k | $0.035 |
| L✗ 55m | — | 55m | — | 10.2M / 102k | $0.173 | |
| Manipulation · select_fruit | U✓ 2m | 1.00 | 2m | — | 214k / 3k | $0.0057 |
| L✓ 30m | 1.00 | 30m | — | 6.3M / 54k | $0.104 | |
| Physical QA · friction_qa | U✓ 6m | 1.00 | 6m | — | 274k / 7k | $0.0094 |
| L✓ 15m | 1.00 | 15m | — | 2.2M / 43k | $0.053 |
32 task(s) not in this run
| Task | Why |
|---|---|
| Common sense · select_billiards_common_sense | not in the sample (robot_coding_bench PR |
| Common sense · select_chemistry_tube_common_sense | not in the sample (robot_coding_bench PR |
| Common sense · select_fruit_common_sense | not in the sample (robot_coding_bench PR |
| Manipulation · add_condiment | not in the sample (robot_coding_bench PR |
| Manipulation · book_rearrange | not in the sample (robot_coding_bench PR |
| Manipulation · find_unseen_object | not in the sample (robot_coding_bench PR |
| Manipulation · hammer_nail_and_hang_picture | not in the sample (robot_coding_bench PR |
| Manipulation · insert_bloom_flower | not in the sample (robot_coding_bench PR |
| Manipulation · play_math_game | not in the sample (robot_coding_bench PR |
| Manipulation · put_box_on_painting | not in the sample (robot_coding_bench PR |
| Manipulation · select_billiards | not in the sample (robot_coding_bench PR |
| Manipulation · select_chemistry_tube | not in the sample (robot_coding_bench PR |
| Manipulation · select_drink | not in the sample (robot_coding_bench PR |
| Manipulation · select_mahjong | not in the sample (robot_coding_bench PR |
| Manipulation · select_nth_largest_poker | not in the sample (robot_coding_bench PR |
| Manipulation · select_painting | not in the sample (robot_coding_bench PR |
| Manipulation · select_poker | not in the sample (robot_coding_bench PR |
| Manipulation · select_specific_type_book | not in the sample (robot_coding_bench PR |
| Manipulation · select_toy | not in the sample (robot_coding_bench PR |
| Manipulation · select_unique_type_mahjong | not in the sample (robot_coding_bench PR |
| Manipulation · take_out_cool_drink | not in the sample (robot_coding_bench PR |
| Manipulation · texas_holdem | not in the sample (robot_coding_bench PR |
| Physical QA · magnetism_qa | not in the sample (robot_coding_bench PR |
| Physical QA · reflection_qa | not in the sample (robot_coding_bench PR |
| Semantic · select_chemistry_tube_semantic | not in the sample (robot_coding_bench PR |
| Semantic · select_mahjong_semantic | not in the sample (robot_coding_bench PR |
| Semantic · select_poker_semantic | not in the sample (robot_coding_bench PR |
| Spatial · select_billiards_spatial | not in the sample (robot_coding_bench PR |
| Spatial · select_book_spatial | not in the sample (robot_coding_bench PR |
| Spatial · select_chemistry_tube_spatial | not in the sample (robot_coding_bench PR |
| Spatial · select_mahjong_spatial | not in the sample (robot_coding_bench PR |
| Spatial · select_poker_spatial | not in the sample (robot_coding_bench PR |