Dynamic3D · Scoop pour¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the KinDER runs for every task of a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 1h 00m | 10-04 05:24 | 1h 00m | — | 15.5M / 147k | $0.249 | deterministic 1, grader_error 0, missing_trajectory 0, n_actions 1000, video_rendered 1 |
| L✗ 43m | 10-04 05:24 | 43m | — | 33.8M / 179k | $0.480 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 1000, replay_success 0, video_rendered 1 |
1 other run(s) of KinDER
Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Billed | Grade |
|---|---|---|---|---|---|---|---|
| U✓ 47m | 10-01 05:55 | 47m | 120 | 11.4M / 73k | $2.18 | $2.27 | deterministic 1, grader_error 0, missing_trajectory 0, n_actions 935, video_rendered 1 |
| L✓ 19m | 10-01 02:43 | 19m | 71 | 3.3M / 25k | $0.733 | $0.774 | deterministic 1, grader_error 0, live_success 1, missing_log 0, n_actions 414, replay_success 1, video_rendered 1 |
Task instruction (upstream)
Move all 30 small dark-red cubes from the yellow bin into the green bin; a scoop is in the scene. The task succeeds the moment every cube's centre is inside the green bin, at least 1 cm from each of its inner walls and at least 0.5 cm below its rim.

What KinDER states about this task
| Success criteria | 1. KinDER's own goal check (the environment's terminated) fires within the episode: at most 1000 control steps (100 s at 10 Hz); no partial credit2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state 3. limited: the one episode reaches it (no reset in our runs); the service records the episode, and it replays to the same state |
| Env Id | kinder/ScoopPour3D-o30-v0 |
| Robot | TidyBot++: holonomic base, Kinova Gen3 7-DoF arm, Robotiq 2F-85 gripper |
| Category | Dynamic3D (MuJoCo) |
| Instance | seed 0: the scene KinDER builds from it, pinned by its digest |
| Holding Still | does not reach the goal (checked by running an episode of zero actions) |
| Action Dim | 11 |
| Actions | base pos and yaw (3), arm joints (7), gripper pos (1); each base and joint delta at most 0.1 a step |
| State Dim | 595 |
| Limited Mode | base RGB and wrist RGB-D cameras (480×640), base odometry, the arm's joint angles and velocities, the fingers' position. Not where any object or the goal region is, not the goal check before the episode ends |
| Agent Budget | 3600 s of wall clock per mode |
From https://github.com/Princeton-Robot-Planning-and-Learning/kindergarden @ 5b2dbac, as defined in our task definitions @ ae0d92a50.
Tags¶
Why this task is interesting¶
Thirty small cubes go from the yellow bin into the green one, each over the bin walls and inside the green bin's margins, on a counter about 0.46 m high. A small scoop is there; loading it is a contact problem, not a grasp.
Capability notes¶
Not yet written.
Oracle demo review¶
No demo.
Discussion¶
GPT-6.1 Sol passed both modes by pouring rather than scooping: limited mode tipped the yellow bin over the green one, unlimited mode moved the green bin to the floor and tipped the yellow bin into it. GPT-6 Luna (2026-10-04) failed both: in unlimited mode its scoop moved 14 of the 30 cubes in a trial run but the saved trajectory lost the load on the lift, and in limited mode it never picked up the scoop and moved single cubes with the gripper. (@williamzhangNU)