Staples marker · draw smile¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the DexToolBench runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 1h 00m 0% | 09-28 23:36 | 1h 00m | — | 10.9M / 161k | $0.227 | grader_error 0, missing_trajectory 1 |
| L✗ 1h 00m 0% | 09-29 12:09 | 1h 00m | — | 20.4M / 124k | $0.293 | deterministic 1, grader_error 0, video_rendered 1 |
1 other run(s) of DexToolBench
Codex + GPT-6 Luna, reasoning xhigh (Azure OpenAI API, 2026-09-30 stress test)
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 57m 0% | 09-30 04:00 | 57m | — | 17.4M / 215k | $0.315 | deterministic 1, grader_error 0, n_actions 1800, video_rendered 1 |
| L✗ 59m 0% | 09-30 17:54 | 59m | — | 19.4M / 185k | $0.321 | deterministic 1, grader_error 0, video_rendered 1 |
Task instruction (upstream)
Pick up the staples marker and move it through the demonstrated motion (drawing a smiley face): bring it to each of the 49 goal poses in order.
Recorded by us: SimToolReal's pretrained RL policy replayed in our MuJoCo port of the task, four views (front, side / top, oblique); it reaches 49 of 49 goals here. The green ghost tool is the current goal.
What DexToolBench states about this task
| Defined in | dextoolbench/trajectories/marker/staples_marker/draw_smile.json |
| Category | marker |
| Object | staples_marker |
| Task | draw_smile |
| N Goals | 49 |
| Tool Model | assets/urdf/dextoolbench/marker/staples_marker/staples_marker.urdf |
| Table | table_narrow_whiteboard.urdf (a whiteboard at the table's edge) |
| Metric | task progress: goals reached / goals (8 grasp-box keypoints within 1.5 cm; 10 s per goal) |
From https://github.com/tylerlum/simtoolreal @ 313d5ae.
Measured on the stack we run — MuJoCo 3.3.7 (CPU, bit-exact replay), KUKA iiwa 14 (MuJoCo Menagerie 1b86ece) + Sharpa HA4, 600 Hz physics, 60 Hz control
| Harness Task | dextoolbench_marker_staples_marker_draw_smile |
| Oracle | SimToolReal's pretrained policy, recorded in this scene and replayed: 49/49 goals |
Read from run in our MuJoCo port (robot_coding_bench tasks/dextoolbench_marker_staples_marker_draw_smile, image rcb-mujoco 0.1.1), 2026-09-28.
Tags¶
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.