Skip to content

RB-Y1 · Navigate to

keepholodeck-objaverse val 1430——@williamzhangNU

Agent runs

Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the MolmoSpaces runs for every task of a run.

Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default runclosed

TrialStartedAgent timeModel requestsTokens in / outEst. costGrade
U✓ 8m10-03 21:458m—1.8M / 13k$0.032deterministic 1, grader_error 0, n_actions 111, replay_success 1, video_rendered 1
L✗ 8m10-03 21:028m—1.9M / 33k$0.041deterministic 1, grader_error 0, live_success 0, n_actions 164, replay_success 0, video_rendered 1

Task instruction (upstream)

Navigate to any bed (2 available).

holodeck-objaverse val 1430 scene
No oracle video published upstream — this is the scene it starts from.
What MolmoSpaces states about this task
Success criteria 1. for the one of the 2 target objects nearest to the robot base (horizontal distance), the horizontal distance from the base to its position (the origin of its body, not its nearest surface) is less than 1.5 m and at least one of its pixels shows in head_camera
2. judged at the end of the episode: the trajectory's last row (privileged), done (standard) or 500 control steps (the benchmark's 100 s horizon), whichever comes first
3. unlimited (privileged): both fresh-process replays of the handed-in trajectory end in the same state, and the check holds on it
4. limited (standard): the one episode (no reset) is recorded by the service and replays to the same state; the check holds on it live and in the replay
Family navigate-to
Robot Rainbow RB-Y1: holonomic base, 6-joint torso, two 7-joint arms with parallel grippers
Category RB-Y1
Instance molmospaces-bench-v2/20260415, package holodeck-objaverse/NavToObjDataGenConfig/NavToObjHolodeckBench_20260115_json_benchmark, episode 94
Deliverable /app/output/trajectory.npz with actions: float64 (T, 25), one row per control step, the targets of upstream's joint-position controllers
Reference Solution None is shipped. Upstream's scripted experts (molmo_spaces/policy/solvers) are reference solutions and are not in the image; neither are grasp files nor the public MolmoBot trajectories (agent egress: the model APIs only).
Limited Mode Standard-mode twin of molmospaces-navigate-to-i00-privileged (the same frozen MolmoSpaces episode): Navigate to any bed (2 available). The agent gets only the eai-standard/2.2 client (docs/STANDARD_MODE_2_2.md); the simulator runs in the sim sidecar (environment/docker-compose.yaml), which owns the episode, serves cameras, proprioception and upstream's kinematic model, and records every executed row. The collect hook (environment/sim/finalize.sh) ends the episode, lets the service exit, replays the trajectory in two fresh processes and writes final.json; the verifier grades those artifacts in a separate sandbox (tests/Dockerfile).
Oracle none — no reference solution (MolmoSpaces' planners and grasp files are not shipped); positive example graded 1 by the separate verifier: molmospaces-navigate-to-i01-priv__Nxshh8t (run codex-gpt6_luna-medium, batch molmospaces-luna-1003; https://embodied-agent-interface-v2-internal.github.io/runs/molmospaces/codex-gpt6_luna-medium-openrouter/navigate_to/unlimited/); human review in PR #47
Base Image ghcr.io/mll-lab-nu/eai-molmospaces:0.2.0
Agent Budget 3600 s of wall clock per mode
Task Dirs molmospaces-navigate-to-i00-privileged, molmospaces-navigate-to-i00-standard

From https://github.com/allenai/molmospaces @ molmo-spaces 0.2.9 (benchmark molmospaces-bench-v2/20260415), as defined in our task definitions @ 030f55607.

Tags

Why this task is interesting

Drive the RB-Y1 to either of two beds that start more than 7 m away and out of view, through doorways in an unseen Holodeck house, and stop close to one with it in the head camera's view.

Capability notes

Not yet written.

Oracle demo review

No demo.

Discussion

Exercises real multi-room navigation. The success distance is measured to the bed's position, not its edge, which the instruction now states. (@williamzhangNU)