Manipulation · find_unseen_object¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the VLABench runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default runclosed
Not in this run: not in the sample (robot_coding_bench PR
Task instruction (upstream)
Find the apple for me and take it out of cabinet_seen.
Playback speed
Recorded by us: our reference solution, replayed by the verifier (the front camera).
What VLABench states about this task
| Success criteria | 1. VLABench's own success check for find_unseen_object (the task's termination condition). The episode ends the moment it holds. 2. This states exactly what VLABench's success check requires. 3. unlimited (privileged): both fresh-process replays of the handed-in trajectory end in the same state, and the check holds on it 4. limited (standard): the one episode (no reset) is recorded by the service and replays to the same state; the check holds on it live and in the replay |
| Family | vlabench/find_unseen_object |
| Robot | Franka Panda (7-DoF arm, two-finger gripper) |
| Category | Manipulation |
| Instance | seed 0 |
| Reference Solution | 407 control steps on this instance (solution/oracle.npy), replayed by solution/solve.sh. Source: PR #4's first-author reference: VLABench's general skills driven by VLABench's expert sequence or by a skill sequence the PR #4 author wrote, recorded as control steps during PR #4's authoring; only the recorded steps were committed (PR #4 @ ad3c510, refs/pull/4/head). VLABench's experts are removed from the image. Re-verified by scripts/vlabench/freeze.sh: the replay ends with the task's success at its last step. |
| Limited Mode | VLABench's find_unseen_object as a robot service: a Franka Panda arm at a table, instructed "Find the apple for me and take it out of cabinet_seen." (a direct command), on the same frozen instance as vlabench-find-unseen-object-i00-privileged. The simulator runs in the sim sidecar (environment/sim/server.py on main's rcb_service.py); the agent only has the client: four cameras (RGB 320×320, depth on request, calibrated), the robot's own state, and joint-position chunks. One episode, no reset, no object poses; nothing to hand in. |
| Oracle | full |
| Base Image | ghcr.io/mll-lab-nu/eai-vlabench:0.1.0 |
| Agent Budget | 3600 s of wall clock per mode |
| Task Dirs | vlabench-find-unseen-object-i00-privileged, vlabench-find-unseen-object-i00-standard |
From https://github.com/OpenMOSS/VLABench @ cf588fe, as defined in our task definitions @ 8c5594a43.
Tags¶
Task DomainManipulation
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.