Manipulation · put_box_on_painting¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the VLABench runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default runclosed
Not in this run: not in the sample (robot_coding_bench PR
Task instruction (upstream)
Please put the box on the Dinner and Love on the River.
Playback speed
Recorded by us: our reference solution, replayed by the verifier (the front camera).
What VLABench states about this task
| Success criteria | 1. VLABench's own success check for put_box_on_painting (the task's termination condition). The episode ends the moment it holds. 2. unlimited (privileged): both fresh-process replays of the handed-in trajectory end in the same state, and the check holds on it 3. limited (standard): the one episode (no reset) is recorded by the service and replays to the same state; the check holds on it live and in the replay |
| Family | vlabench/put_box_on_painting |
| Robot | Franka Panda (7-DoF arm, two-finger gripper) |
| Category | Manipulation |
| Instance | seed 0 |
| Reference Solution | 131 control steps on this instance (solution/oracle.npy), replayed by solution/solve.sh. Source: VLABench's own expert (get_expert_skill_sequence) driving its general skills, recorded as control steps during PR #4's authoring; only the recorded steps were committed (PR #4 @ ad3c510, refs/pull/4/head). VLABench's experts are removed from the image. Re-verified by scripts/vlabench/freeze.sh: the replay ends with the task's success at its last step. |
| Limited Mode | VLABench's put_box_on_painting as a robot service: a Franka Panda arm at a table, instructed "Please put the box on the Dinner and Love on the River." (a direct command), on the same frozen instance as vlabench-put-box-on-painting-i00-privileged. The simulator runs in the sim sidecar (environment/sim/server.py on main's rcb_service.py); the agent only has the client: four cameras (RGB 320×320, depth on request, calibrated), the robot's own state, and joint-position chunks. One episode, no reset, no object poses; nothing to hand in. |
| Oracle | full |
| Base Image | ghcr.io/mll-lab-nu/eai-vlabench:0.1.0 |
| Agent Budget | 3600 s of wall clock per mode |
| Task Dirs | vlabench-put-box-on-painting-i00-privileged, vlabench-put-box-on-painting-i00-standard |
From https://github.com/OpenMOSS/VLABench @ cf588fe, as defined in our task definitions @ 8c5594a43.
Tags¶
Task DomainManipulation
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.