Manipulation · Powerlift¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the HumanoidBench runs for every task of a run.
Codex CLI 0.157–0.159 + GPT-6 Luna, reasoning medium (OpenRouter)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 1h 00m | 10-04 05:31 | 1h 00m | — | 34.5M / 281k | $0.580 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✗ 4m | 10-04 05:38 | 4m | — | 953k / 23k | $0.026 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 1000, replay_success 0, video_rendered 1 |
3 other run(s) of HumanoidBench
Codex CLI 0.159.2 + GPT-6.1 Sol, reasoning medium (OpenRouter)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Billed | Grade |
|---|---|---|---|---|---|---|---|
| U✗ 1h 00m | 10-01 02:40 | 1h 00m | 137 | 14.0M / 111k | $2.89 | $2.99 | deterministic 1, grader_error 0, n_actions 1000, video_rendered 1 |
| L✗ 1h 00m | 09-30 23:48 | 1h 00m | 145 | 16.5M / 95k | $3.53 | $3.78 | deterministic 1, grader_error 0, live_success 0, missing_log 0, n_actions 182, replay_success 0, video_rendered 1 |
Claude Code 2.1.283 + Claude Opus 5.5, reasoning medium (OpenRouter)
Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Codex CLI 0.157.0 + GPT-6 Sol, reasoning medium (OpenRouter)closed
Not in this run: not in this run (the two dearer models ran five tasks, chosen to span locomotion and manipulation, easy to hard)
Task instruction (upstream)
Lift the dumbbell from the floor to overhead — close to two metres up — and hold it there.

What HumanoidBench states about this task
| Success criteria | 1. the summed per-step reward over one episode reaches 800, HumanoidBench's own success bar (a total of rewards, not a number of steps; an episode is at most 1000 control steps) 2. unlimited: both fresh-process replays of the handed-in trajectory reach it and end in the same state 3. limited: the run passes the moment its one episode reaches it (no reset in our runs since 2026-10-04; the recorded episode replays to the same state) |
| Env Id | h1hand-powerlift-v0 |
| Robot | Unitree H1 with two Shadow hands (61 actuators) |
| Category | Manipulation |
| Capability Class | C5 · whole-body power and momentum |
| Role | scored |
| Scoring | A weighted sum of two things: a small share for standing steadily, and the bulk for how high the dumbbell is, measured against a target band around two metres. Be warned that the height term is broad. Clearing the bar means lifting it high and keeping it there for most of the episode, not improving the height a little. |
| Ends Early | The episode ends early if the pelvis drops near the ground. |
| Zero Action Return | 20.39 |
| Action Dim | 61 |
| Control Rate Hz | 50 |
| Limited Mode | head cameras (RGB 256×256), joint angles and velocities; a pelvis IMU and a camera fixed in the room when the run turns them on. Not where the robot is in the room, not the dumbbell's pose, not the reward |
| Agent Budget | 3600 s of wall clock per mode |
From https://github.com/carlosferrazza/humanoid-bench @ cb11890, as defined in our task definitions @ 841996303.
Tags¶
Why this task is interesting¶
Squat, grip a 104 kg dumbbell and lift it overhead: the robot must not fall while it reaches down to the bar, and must hold on once it gets there.
Capability notes¶
Not yet written.
Oracle demo review¶
No demo.
Discussion¶
Blocked on the grasp: the dumbbell never leaves the floor in either mode; GPT-6 Luna's unlimited run (2026-10-04) never got its fingers onto it. (@williamzhangNU)