Franka · Close¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the MolmoSpaces runs for every task of a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default runclosed
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 7m | 10-03 19:09 | 7m | — | 1.2M / 18k | $0.027 | deterministic 1, grader_error 0, n_actions 99, replay_success 1, video_rendered 1 |
| L✗ 23m | 10-03 19:57 | 23m | — | 10.2M / 86k | $0.159 | deterministic 1, grader_error 0, live_success 0, n_actions 455, replay_success 0, video_rendered 1 |
Task instruction (upstream)
Close the drawer.

No oracle video published upstream — this is the scene it starts from.
What MolmoSpaces states about this task
| Success criteria | 1. the target's joint is open at most 15% of its range (it starts 50% open) 2. judged at the end of the episode: the trajectory's last row (privileged), done (standard) or 455 control steps (the benchmark's 30 s horizon), whichever comes first3. unlimited (privileged): both fresh-process replays of the handed-in trajectory end in the same state, and the check holds on it 4. limited (standard): the one episode (no reset) is recorded by the service and replays to the same state; the check holds on it live and in the replay |
| Family | close |
| Robot | Franka FR3 with a Robotiq 2F-85 gripper on a fixed base (the DROID setup) |
| Category | Franka |
| Instance | molmospaces-bench-v2/20260415, package ithor/FrankaCloseHardBench/FrankaCloseHardBench_20260206_json_benchmark, episode 327 |
| Deliverable | /app/output/trajectory.npz with actions: float64 (T, 8), one row per control step, the targets of upstream's joint-position controllers |
| Reference Solution | None is shipped. Upstream's scripted experts (molmo_spaces/policy/solvers) are reference solutions and are not in the image; neither are grasp files nor the public MolmoBot trajectories (agent egress: the model APIs only). |
| Limited Mode | Standard-mode twin of molmospaces-close-i00-privileged (the same frozen MolmoSpaces episode): Close the drawer. The agent gets only the eai-standard/2.2 client (docs/STANDARD_MODE_2_2.md); the simulator runs in the sim sidecar (environment/docker-compose.yaml), which owns the episode, serves cameras, proprioception and upstream's kinematic model, and records every executed row. The collect hook (environment/sim/finalize.sh) ends the episode, lets the service exit, replays the trajectory in two fresh processes and writes final.json; the verifier grades those artifacts in a separate sandbox (tests/Dockerfile). |
| Oracle | none — no reference solution (MolmoSpaces' planners and grasp files are not shipped); positive example graded 1 by the separate verifier: molmospaces-close-i00-privileged__U2cQyVE (run codex-gpt6_luna-medium, batch molmospaces-luna-1003; https://embodied-agent-interface-v2-internal.github.io/runs/molmospaces/codex-gpt6_luna-medium-openrouter/close/unlimited/); human review in PR #47 |
| Base Image | ghcr.io/mll-lab-nu/eai-molmospaces:0.2.0 |
| Agent Budget | 3600 s of wall clock per mode |
| Task Dirs | molmospaces-close-i00-privileged, molmospaces-close-i00-standard |
From https://github.com/allenai/molmospaces @ molmo-spaces 0.2.9 (benchmark molmospaces-bench-v2/20260415), as defined in our task definitions @ 030f55607.
Tags¶
CapabilityControl & Coordination
Task DomainManipulation
Why this task is interesting¶
Close a drawer that starts half open, in upstream's Hard configuration. The drawer is easy to see; telling from the cameras which way closes it is not.
Capability notes¶
Not yet written.
Oracle demo review¶
No demo.
Discussion¶
Limited mode pulled the drawer further open, taking the opening direction for the closing one. (@williamzhangNU)