Manipulation · play_math_game¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the VLABench runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default runclosed
Not in this run: not in the sample (robot_coding_bench PR
Task instruction (upstream)
Please give the answer of the following question by rearrange the number cube in placemat_seen:A mechanic charges different rates to repair the tires of trucks and cars. For each truck tire that is repaired, the mechanic will charge $60 and for each car tire that is repaired, the mechanic will charge $40. On Thursday, the mechanic repairs 6 truck tires and 4 car tires. On Friday, the mechanic repairs 12 car tries and doesn't repair any truck tires. How much more revenue did the mechanic earn on the day with higher revenue? Put the number cubes that spell the answer on placemat_seen, side by side, read from left to right as seen from the robot.
Recorded by us: our reference solution, replayed by the verifier (the front camera).
What VLABench states about this task
| Success criteria | 1. VLABench's own success check for play_math_game (the task's termination condition). The episode ends the moment it holds. 2. This states exactly what VLABench's success check requires. 3. unlimited (privileged): both fresh-process replays of the handed-in trajectory end in the same state, and the check holds on it 4. limited (standard): the one episode (no reset) is recorded by the service and replays to the same state; the check holds on it live and in the replay |
| Family | vlabench/play_math_game |
| Robot | Franka Panda (7-DoF arm, two-finger gripper) |
| Category | Manipulation |
| Instance | seed 0 |
| Reference Solution | 167 control steps on this instance (solution/oracle.npy), replayed by solution/solve.sh. Source: VLABench's own expert (get_expert_skill_sequence) driving its general skills, recorded as control steps during PR #4's authoring; only the recorded steps were committed (PR #4 @ ad3c510, refs/pull/4/head). VLABench's experts are removed from the image. Re-verified by scripts/vlabench/freeze.sh: the replay ends with the task's success at its last step. |
| Limited Mode | VLABench's play_math_game as a robot service: a Franka Panda arm at a table, instructed "Please give the answer of the following question by rearrange the number cube in placemat_seen:A mechanic charges different rates to repair the tires of trucks and cars. For each truck tire that is repaired, the mechanic will charge $60 and for each car tire that is repaired, the mechanic will charge $40. On Thursday, the mechanic repairs 6 truck tires and 4 car tires. On Friday, the mechanic repairs 12 car tries and doesn't repair any truck tires. How much more revenue did the mechanic earn on the day with higher revenue? Put the number cubes that spell the answer on placemat_seen, side by side, read from left to right as seen from the robot." (a direct command), on the same frozen instance as vlabench-play-math-game-i00-privileged. The simulator runs in the sim sidecar (environment/sim/server.py on main's rcb_service.py); the agent only has the client: four cameras (RGB 320×320, depth on request, calibrated), the robot's own state, and joint-position chunks. One episode, no reset, no object poses; nothing to hand in. |
| Oracle | full |
| Base Image | ghcr.io/mll-lab-nu/eai-vlabench:0.1.0 |
| Agent Budget | 3600 s of wall clock per mode |
| Task Dirs | vlabench-play-math-game-i00-privileged, vlabench-play-math-game-i00-standard |
From https://github.com/OpenMOSS/VLABench @ cf588fe, as defined in our task definitions @ 8c5594a43.
Tags¶
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.