Skip to content

Claw hammer · swing down

keephardtabletop—0m 12s@JamesKrW

Agent runs

Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the DexToolBench runs for every task of a run.

Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run

TrialStartedAgent timeModel requestsTokens in / outEst. costGrade
U✗ 54m 32%09-29 02:4654m—15.6M / 137k$0.254deterministic 1, grader_error 0, n_actions 1386, video_rendered 1
L✗ 1h 00m 0%09-29 07:091h 00m—26.1M / 124k$0.364deterministic 1, grader_error 0, video_rendered 1
1 other run(s) of DexToolBench

Codex + GPT-6 Luna, reasoning xhigh (Azure OpenAI API, 2026-09-30 stress test)

TrialStartedAgent timeModel requestsTokens in / outEst. costGrade
U✗ 1h 00m 3% ⓘ09-30 20:461h 00m—15.1M / 159k$0.263deterministic 1, grader_error 0, n_actions 2246, video_rendered 1
L✗ 59m 0% ⓘ09-30 14:3559m—19.8M / 126k$0.280deterministic 1, grader_error 0, video_rendered 1

Task instruction (upstream)

Pick up the claw hammer and move it through the demonstrated motion (a downward hammer swing): bring it to each of the 37 goal poses in order.

Playback speed

Recorded by us: SimToolReal's pretrained RL policy replayed in our MuJoCo port of the task, four views (front, side / top, oblique); it reaches 37 of 37 goals here. The green ghost tool is the current goal.

What DexToolBench states about this task
Defined in dextoolbench/trajectories/hammer/claw_hammer/swing_down.json
Category hammer
Object claw_hammer
Task swing_down
N Goals 37
Tool Model assets/urdf/dextoolbench/hammer/claw_hammer/claw_hammer.urdf
Table table_narrow_nail.urdf (one nail)
Metric task progress: goals reached / goals (8 grasp-box keypoints within 1.5 cm; 10 s per goal)

From https://github.com/tylerlum/simtoolreal @ 313d5ae.

Measured on the stack we run — MuJoCo 3.3.7 (CPU, bit-exact replay), KUKA iiwa 14 (MuJoCo Menagerie 1b86ece) + Sharpa HA4, 600 Hz physics, 60 Hz control
Harness Task dextoolbench_hammer_claw_hammer_swing_down
Oracle SimToolReal's pretrained policy, recorded in this scene and replayed: 37/37 goals

Read from run in our MuJoCo port (robot_coding_bench tasks/dextoolbench_hammer_claw_hammer_swing_down, image rcb-mujoco 0.1.1), 2026-09-28.

Tags

Task DomainManipulation

Why this task is interesting

Not yet written.

Capability notes

Not yet written.

Oracle demo review

Not yet reviewed.

Discussion