Skip to content

Locomotion · Sit simple

excludedkeephumanoidbench_loco——@williamzhangNU

Excluded from the benchmark

a fixture, not scored: it checks the pipeline end to end (subset.toml role = fixture)

Greyed and last in the task list, left out of the benchmark's numbers; its demo and agent runs stay (excluded: in state/tasks/humanoidbench.yml).

Agent runs

Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the HumanoidBench runs for every task of a run.

Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run

TrialStartedAgent timeModel requestsTokens in / outEst. costGrade
U✓ 59m09-28 16:1659m—32.7M / 111k$0.403deterministic 1, grader_error 0, n_actions 1000, video_rendered 1
L✓ 3m09-28 16:163m—583k / 15k$0.017deterministic 1, grader_error 0, live_success 1, n_actions 978, replay_success 1, video_rendered 1
3 other run(s) of HumanoidBench

Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)

TrialStartedAgent timeModel requestsTokens in / outEst. costBilledGrade
U✓ 1h 00m10-01 02:381h 00m14813.2M / 73k$2.32$2.39deterministic 1, grader_error 0, n_actions 1000, video_rendered 1
L✓ 1m09-30 22:251m8165k / 2k$0.087$0.101deterministic 1, grader_error 0, live_success 1, missing_log 0, n_actions 782, replay_success 1, video_rendered 1

Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)

Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)

Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed

Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)

Task instruction (upstream)

Sit the robot down on the chair and keep it seated, upright and still.

humanoidbench_loco scene
No oracle video published upstream — this is the scene it starts from.
What HumanoidBench states about this task
Success criteria 1. the summed per-step reward over one episode reaches 750, HumanoidBench's own success bar (a total of rewards, not a number of steps; an episode is at most 1000 control steps)
2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state
3. limited: the run passes the moment its one episode reaches it (no reset in our runs since 2026-10-04; the recorded episode replays to the same state)
Env Id h1-sit_simple-v0
Robot Unitree H1 (19 actuators)
Category Locomotion
Capability Class fixture (C3 · quasi-static stabilisation, too easy to score)
Role fixture
Scoring A product: the average of two seat terms — how close the pelvis is to the height it has when seated, and whether the robot is over the seat rather than beside it — multiplied by staying upright, a seated posture (the head's height above the torso), not moving, and low actuator force. Every step spent not yet sitting forfeits the reward that step could have paid, so both how fast and how steadily the robot sits matter.
Ends Early The episode ends early if the pelvis drops below about half its standing height — not far below its height when seated.
Zero Action Return 637.64
Action Dim 19
Control Rate Hz 50
Limited Mode head cameras (RGB 256×256), joint angles and velocities; a pelvis IMU and a camera fixed in the room when the run turns them on. Not where the robot is in the room, not the chair's pose, not the reward
Agent Budget 3600 s of wall clock per mode

From https://github.com/carlosferrazza/humanoid-bench @ cb11890, as defined in our task definitions @ 841996303.

Tags

Why this task is interesting

A fixture: the chair is about 0.25 m straight behind the pelvis, so a symmetric squat seats the robot. It checks the whole pipeline end to end and is not scored.

Capability notes

Not yet written.

Oracle demo review

No demo.

Discussion

Works as the pipeline fixture: passes both modes, limited mode in its first episode. (@williamzhangNU)