Dynamic3D · Balance beam¶
Excluded from the benchmark
not among the hardest (robot_coding_bench 2026-10-04): GPT-6 Luna passed it in unlimited mode in 5 min
Greyed and last in the task list, left out of the benchmark's numbers; its demo and agent runs stay (excluded: in state/tasks/kinder.yml).
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the KinDER runs for every task of a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 5m | 09-29 00:02 | 5m | — | 1.4M / 23k | $0.031 | deterministic 1, grader_error 0, missing_trajectory 0, n_actions 574, video_rendered 1 |
| L✗ 32m | 10-01 14:17 | 32m | — | 8.5M / 114k | $0.157 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 1000, replay_success 0, video_rendered 1 |
1 other run(s) of KinDER
Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Billed | Grade |
|---|---|---|---|---|---|---|---|
| U✓ 4m | 10-01 05:52 | 4m | 16 | 539k / 5k | $0.193 | $0.217 | deterministic 1, grader_error 0, missing_trajectory 0, n_actions 696, video_rendered 1 |
| L✓ 5m | 10-01 02:41 | 5m | 22 | 533k / 8k | $0.198 | $0.215 | deterministic 1, grader_error 0, live_success 1, missing_log 0, n_actions 558, replay_success 1, video_rendered 1 |
Task instruction (upstream)
Put the two small green cubes and the larger dark-red cube on the seesaw so that it balances. The task succeeds the moment all three cubes' centres are inside the translucent green region marked on the seesaw's beam and the beam is within 5° of level.

What KinDER states about this task
| Success criteria | 1. KinDER's own goal check (the environment's terminated) fires within the episode: at most 1000 control steps (100 s at 10 Hz); no partial credit2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state 3. limited: the one episode reaches it (no reset in our runs); the service records the episode, and it replays to the same state |
| Env Id | kinder/BalanceBeam3D-o3-v0 |
| Robot | TidyBot++: holonomic base, Kinova Gen3 7-DoF arm, Robotiq 2F-85 gripper |
| Category | Dynamic3D (MuJoCo) |
| Instance | seed 0: the scene KinDER builds from it, pinned by its digest |
| Holding Still | does not reach the goal (checked by running an episode of zero actions) |
| Action Dim | 11 |
| Actions | base pos and yaw (3), arm joints (7), gripper pos (1); each base and joint delta at most 0.1 a step |
| State Dim | 86 |
| Limited Mode | base RGB and wrist RGB-D cameras (480×640), base odometry, the arm's joint angles and velocities, the fingers' position. Not where any object or the goal region is, not the goal check before the episode ends |
| Agent Budget | 3600 s of wall clock per mode |
From https://github.com/Princeton-Robot-Planning-and-Learning/kindergarden @ 5b2dbac, as defined in our task definitions @ 7d6a7a44.
Tags¶
Why this task is interesting¶
Two light cubes and one twice as heavy go onto a narrow seesaw beam, all inside its marked region, and the beam must end within 5° of level. Getting them onto the beam is not enough: they have to balance about the pivot.
Capability notes¶
Not yet written.
Oracle demo review¶
No demo.
Discussion¶
Both modes pass: limited mode measured the goal region with wrist depth and placed all three cubes in its third episode; unlimited mode reads the goal region from the state. (@williamzhangNU)