Franka · Pick and place¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the MolmoSpaces runs for every task of a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default runclosed
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 9m | 10-03 21:11 | 9m | — | 1.6M / 17k | $0.031 | deterministic 1, grader_error 0, n_actions 215, replay_success 1, video_rendered 1 |
| L✗ 14m | 10-03 21:37 | 14m | — | 3.4M / 41k | $0.062 | deterministic 1, grader_error 0, live_success 0, n_actions 606, replay_success 0, video_rendered 1 |
Task instruction (upstream)
Pick up the yellow handheld gps with antenna and place it in or on the shallow round wooden bowl with grain.

What MolmoSpaces states about this task
| Success criteria | 1. the robot does not touch the object; the receptacle supports the object (it carries at least 50% of the object's weight through contact, or the object lies on it by upstream's on-receptacle test, or the object is within 5 mm and 10 degrees, relative to the receptacle, of a pose in which it was supported at an earlier control step); and the receptacle has moved at most 15 cm and tilted at most 60 degrees from its start pose 2. judged at the end of the episode: the trajectory's last row (privileged), done (standard) or 606 control steps (the benchmark's 40 s horizon), whichever comes first3. unlimited (privileged): both fresh-process replays of the handed-in trajectory end in the same state, and the check holds on it 4. limited (standard): the one episode (no reset) is recorded by the service and replays to the same state; the check holds on it live and in the replay |
| Family | pick-and-place |
| Robot | Franka FR3 with a Robotiq 2F-85 gripper on a fixed base (the DROID setup) |
| Category | Franka |
| Instance | molmospaces-bench-v2/20260415, package procthor-objaverse/FrankaPickandPlaceHardBench/FrankaPickandPlaceHardBench_20260206_json_benchmark, episode 0 |
| Deliverable | /app/output/trajectory.npz with actions: float64 (T, 8), one row per control step, the targets of upstream's joint-position controllers |
| Reference Solution | None is shipped. Upstream's scripted experts (molmo_spaces/policy/solvers) are reference solutions and are not in the image; neither are grasp files nor the public MolmoBot trajectories (agent egress: the model APIs only). |
| Limited Mode | Standard-mode twin of molmospaces-pick-and-place-i00-privileged (the same frozen MolmoSpaces episode): Pick up the yellow handheld gps with antenna and place it in or on the shallow round wooden bowl with grain. The agent gets only the eai-standard/2.2 client (docs/STANDARD_MODE_2_2.md); the simulator runs in the sim sidecar (environment/docker-compose.yaml), which owns the episode, serves cameras, proprioception and upstream's kinematic model, and records every executed row. The collect hook (environment/sim/finalize.sh) ends the episode, lets the service exit, replays the trajectory in two fresh processes and writes final.json; the verifier grades those artifacts in a separate sandbox (tests/Dockerfile). |
| Oracle | none — no reference solution (MolmoSpaces' planners and grasp files are not shipped); positive example graded 1 by the separate verifier: molmospaces-pick-and-place-i01-p__ywMYWxy (run codex-gpt6_luna-medium, batch molmospaces-luna-1003; https://embodied-agent-interface-v2-internal.github.io/runs/molmospaces/codex-gpt6_luna-medium-openrouter/pick_and_place/unlimited/); human review in PR #47 |
| Base Image | ghcr.io/mll-lab-nu/eai-molmospaces:0.2.0 |
| Agent Budget | 3600 s of wall clock per mode |
| Task Dirs | molmospaces-pick-and-place-i00-privileged, molmospaces-pick-and-place-i00-standard |
From https://github.com/allenai/molmospaces @ molmo-spaces 0.2.9 (benchmark molmospaces-bench-v2/20260415), as defined in our task definitions @ 030f55607.
Tags¶
Why this task is interesting¶
A small handheld GPS stands upright beside a shallow wooden bowl. It tips over easily, and once it lies flat the grasp has to be planned again at a new height; carrying it into the bowl and letting go without knocking the bowl is the rest.
Capability notes¶
Not yet written.
Oracle demo review¶
No demo.
Discussion¶
Kept because limited mode failed on it in two runs in the same way: the GPS tipped over and the agent did not plan again for its new pose. (@williamzhangNU)