Skip to content

Spatial · select_mahjong_spatial

keepvlabench_table—0m 15s@changhechen

Agent runs

Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the VLABench runs for every task of a run.

Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default runclosed

Not in this run: not in the sample (robot_coding_bench PR

Task instruction (upstream)

Put the 7th mahjong from left to right (as seen from the robot) on the placemat.

Playback speed

Recorded by us: our reference solution, replayed by the verifier (the front camera).

What VLABench states about this task
Success criteria 1. VLABench's own success check for select_mahjong_spatial (the task's termination condition). The episode ends the moment it holds.
2. This states exactly what VLABench's success check requires.
3. unlimited (privileged): both fresh-process replays of the handed-in trajectory end in the same state, and the check holds on it
4. limited (standard): the one episode (no reset) is recorded by the service and replays to the same state; the check holds on it live and in the replay
Family vlabench/select_mahjong_spatial
Robot Franka Panda (7-DoF arm, two-finger gripper)
Category Spatial
Instance seed 0
Reference Solution 149 control steps on this instance (solution/oracle.npy), replayed by solution/solve.sh. Source: VLABench's own expert (get_expert_skill_sequence) driving its general skills, recorded as control steps during PR #4's authoring; only the recorded steps were committed (PR #4 @ ad3c510, refs/pull/4/head). VLABench's experts are removed from the image. Re-verified by scripts/vlabench/freeze.sh: the replay ends with the task's success at its last step.
Limited Mode VLABench's select_mahjong_spatial as a robot service: a Franka Panda arm at a table, instructed "Put the 7th mahjong from left to right (as seen from the robot) on the placemat." (a description of where the target is), on the same frozen instance as vlabench-select-mahjong-spatial-i00-privileged. The simulator runs in the sim sidecar (environment/sim/server.py on main's rcb_service.py); the agent only has the client: four cameras (RGB 320×320, depth on request, calibrated), the robot's own state, and joint-position chunks. One episode, no reset, no object poses; nothing to hand in.
Oracle full
Base Image ghcr.io/mll-lab-nu/eai-vlabench:0.1.0
Agent Budget 3600 s of wall clock per mode
Task Dirs vlabench-select-mahjong-spatial-i00-privileged, vlabench-select-mahjong-spatial-i00-standard

From https://github.com/OpenMOSS/VLABench @ cf588fe, as defined in our task definitions @ 8c5594a43.

Tags

Task DomainManipulation

Why this task is interesting

Not yet written.

Capability notes

Not yet written.

Oracle demo review

Not yet reviewed.

Discussion