Dynamic3D · Tossing¶
Excluded from the benchmark
not among the hardest (robot_coding_bench 2026-10-04): GPT-6 Luna passed it in unlimited mode in 37 min
Greyed and last in the task list, left out of the benchmark's numbers; its demo and agent runs stay (excluded: in state/tasks/kinder.yml).
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the KinDER runs for every task of a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 37m | 09-29 00:02 | 37m | — | 17.0M / 150k | $0.290 | deterministic 1, grader_error 0, missing_trajectory 0, n_actions 644, video_rendered 1 |
| L✗ 32m | 10-01 14:17 | 32m | — | 20.3M / 83k | $0.267 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 1000, replay_success 0, video_rendered 1 |
1 other run(s) of KinDER
Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Billed | Grade |
|---|---|---|---|---|---|---|---|
| U✓ 6m | 10-01 05:54 | 6m | 27 | 922k / 9k | $0.273 | $0.296 | deterministic 1, grader_error 0, missing_trajectory 0, n_actions 584, video_rendered 1 |
| L✓ 12m | 10-01 02:42 | 12m | 48 | 1.8M / 14k | $0.444 | $0.475 | deterministic 1, grader_error 0, live_success 1, missing_log 0, n_actions 327, replay_success 1, video_rendered 1 |
Task instruction (upstream)
Get the green cube lying on the floor into the yellow bin on the far side of the low wall that runs across the room. The task succeeds the moment the cube's centre is inside the bin's goal region, the translucent green box marked in the bin.

What KinDER states about this task
| Success criteria | 1. KinDER's own goal check (the environment's terminated) fires within the episode: at most 1000 control steps (100 s at 10 Hz); no partial credit2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state 3. limited: the one episode reaches it (no reset in our runs); the service records the episode, and it replays to the same state |
| Env Id | kinder/Tossing3D-o1-v0 |
| Robot | TidyBot++: holonomic base, Kinova Gen3 7-DoF arm, Robotiq 2F-85 gripper |
| Category | Dynamic3D (MuJoCo) |
| Instance | seed 0: the scene KinDER builds from it, pinned by its digest |
| Holding Still | does not reach the goal (checked by running an episode of zero actions) |
| Action Dim | 18 |
| Actions | base pos and yaw (3), arm joints (7), gripper pos (1), and arm joint velocity targets (7, rad/s); each base and joint delta at most 0.1 a step |
| State Dim | 61 |
| Limited Mode | base RGB and wrist RGB-D cameras (480×640), base odometry, the arm's joint angles and velocities, the fingers' position. Not where any object or the goal region is, not the goal check before the episode ends |
| Agent Budget | 3600 s of wall clock per mode |
From https://github.com/Princeton-Robot-Planning-and-Learning/kindergarden @ 5b2dbac, as defined in our task definitions @ 7d6a7a44.
Tags¶
Why this task is interesting¶
A low wall runs across the room and stops the base, so the robot has to stay on its side and send the cube over the wall into a bin about 2 m away with its arm. It first has to find where the bin is, then get the release speed and the moment it opens the gripper right: a little off, and the cube hits the bin's rim.
Capability notes¶
Not yet written.
Oracle demo review¶
No demo.
Discussion¶
Both modes pass: limited mode tuned the throw speed over five episodes by watching where the cube landed; unlimited mode searched launch velocities from a simulation snapshot. (@williamzhangNU)