Locomotion · Sit hard¶
Excluded from the benchmark
not among the hardest (robot_coding_bench 2026-10-04): passed in both modes
Greyed and last in the task list, left out of the benchmark's numbers; its demo and agent runs stay (excluded: in state/tasks/humanoidbench.yml).
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the HumanoidBench runs for every task of a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 58m | 09-28 18:21 | 58m | — | 16.9M / 78k | $0.222 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✓ 6m | 10-01 15:22 | 6m | — | 1.1M / 21k | $0.027 | deterministic 1, grader_error 0, live_success 1, missing_log 0, n_actions 940, replay_success 1, video_rendered 1 |
3 other run(s) of HumanoidBench
Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Billed | Grade |
|---|---|---|---|---|---|---|---|
| U✓ 1h 00m | 10-01 04:47 | 1h 00m | 106 | 7.8M / 88k | $1.94 | $2.02 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✓ 7m | 09-30 22:25 | 7m | 43 | 1.4M / 15k | $0.395 | $0.421 | deterministic 1, grader_error 0, live_success 1, missing_log 0, n_actions 924, replay_success 1, video_rendered 1 |
Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 5m | 09-28 19:37 | 5m | — | 829k / 14k | $0.939 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✓ 6m | 09-28 17:25 | 6m | — | 1.5M / 22k | $1.55 | deterministic 1, grader_error 0, live_success 1, n_actions 912, replay_success 1, video_rendered 1 |
Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 1h 00m | 09-28 21:12 | 1h 00m | — | 18.6M / 65k | $4.86 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✗ 10m | 09-28 17:45 | 10m | — | 1.6M / 18k | $0.616 | deterministic 1, grader_error 0, live_success 0, n_actions 788, replay_success 0, video_rendered 1 |
Task instruction (upstream)
The robot starts a short walk in front of the chair, its heading turned by a random angle of up to about 100° from facing straight away from it. Get to the chair, sit down on it, and stay seated upright and still.

What HumanoidBench states about this task
| Success criteria | 1. the summed per-step reward over one episode reaches 750, HumanoidBench's own success bar (a total of rewards, not a number of steps; an episode is at most 1000 control steps) 2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state 3. limited: the run passes the moment a live episode reaches it (the recorded episode replays to the same state); otherwise the last episode is graded |
| Env Id | h1-sit_hard-v0 |
| Robot | Unitree H1 (19 actuators) |
| Category | Locomotion |
| Capability Class | C8 · mobile manipulation |
| Role | scored |
| Scoring | A product: the average of two seat terms — how close the pelvis is to the height it has when seated, and whether the robot is over the seat rather than beside it — multiplied by staying upright, a seated posture (the head's height above the torso), not moving, and low actuator force. Every step spent not yet sitting forfeits the reward that step could have paid, so sitting early matters as much as sitting well. |
| Ends Early | The episode ends early if the pelvis drops below about half its standing height — not far below its height when seated. |
| Zero Action Return | 13.56 |
| Action Dim | 19 |
| Control Rate Hz | 50 |
| Limited Mode | head cameras (RGB 256×256), joint angles and velocities; a pelvis IMU and a camera fixed in the room when the run turns them on. Not where the robot is in the room, not the chair's pose, not the reward |
| Agent Budget | 3600 s of wall clock per mode |
From https://github.com/carlosferrazza/humanoid-bench @ cb11890, as defined in our task definitions @ 7d6a7a44.
Tags¶
Why this task is interesting¶
The chair is about 0.55 m behind the pelvis's start and the robot starts turned (a yaw of 0.37 rad in our instance), so a squat in place lands on the front edge of the seat. The robot has to turn or step to the chair first, and every step before it sits forfeits reward.
Capability notes¶
Not yet written.
Oracle demo review¶
No demo.
Discussion¶
Passes in both modes; limited mode only just (750.77 against 750), by turning the hips toward the chair before sitting. (@williamzhangNU)