Skip to content

Spoon spatula · serve plate

keepextremetabletop—0m 16s@JamesKrW

Agent runs

Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the DexToolBench runs for every task of a run.

Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run

TrialStartedAgent timeModel requestsTokens in / outEst. costGrade
U✗ 56m 0%10-01 07:3556m—16.6M / 284k$0.356deterministic 1, grader_error 0, n_actions 1800, video_rendered 1
L✗ 58m 0%10-01 10:1958m—21.5M / 167k$0.323deterministic 1, grader_error 0, video_rendered 1
1 other run(s) of DexToolBench

Codex + GPT-6 Luna, reasoning xhigh (Azure OpenAI API, 2026-09-30 stress test)

TrialStartedAgent timeModel requestsTokens in / outEst. costGrade
U✗ 1h 00m 0%09-30 08:521h 00m—11.4M / 203k$0.252grader_error 0, missing_trajectory 1
L✗ 57m 0% ⓘ10-01 03:5157m—22.0M / 206k$0.355deterministic 1, grader_error 0, video_rendered 1

Task instruction (upstream)

Pick up the spoon spatula and move it through the demonstrated motion (serving onto a plate): bring it to each of the 48 goal poses in order.

Playback speed

Recorded by us: SimToolReal's pretrained RL policy replayed in our MuJoCo port of the task, four views (front, side / top, oblique); it reaches 13 of 48 goals here. The green ghost tool is the current goal.

What DexToolBench states about this task
Defined in dextoolbench/trajectories/spatula/spoon_spatula/serve_plate.json
Category spatula
Object spoon_spatula
Task serve_plate
N Goals 48
Tool Model assets/urdf/dextoolbench/spatula/spoon_spatula/spoon_spatula.urdf
Table table_narrow_bowl_plate.urdf (a bowl and a plate)
Metric task progress: goals reached / goals (8 grasp-box keypoints within 1.5 cm; 10 s per goal)

From https://github.com/tylerlum/simtoolreal @ 313d5ae.

Measured on the stack we run — MuJoCo 3.3.7 (CPU, bit-exact replay), KUKA iiwa 14 (MuJoCo Menagerie 1b86ece) + Sharpa HA4, 600 Hz physics, 60 Hz control
Harness Task dextoolbench_spatula_spoon_spatula_serve_plate
Oracle SimToolReal's pretrained policy, recorded in this scene and replayed: 13/48 goals

Read from run in our MuJoCo port (robot_coding_bench tasks/dextoolbench_spatula_spoon_spatula_serve_plate, image rcb-mujoco 0.1.1), 2026-09-28.

Tags

Task DomainManipulation

Why this task is interesting

Not yet written.

Capability notes

Not yet written.

Oracle demo review

Not yet reviewed.

Discussion