Line · Cottage¶
Excluded from the benchmark
too easy for the final benchmark (line drawings; kept for the record, results in Runs)
Greyed and last in the task list, left out of the benchmark's numbers; its demo and agent runs stay (excluded: in state/tasks/robopaint.yml).
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the RoboPaint runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 19m | 09-28 07:05 | 19m | 34 | 1.5M / 32k | $0.039 | deterministic 1, grader_error 0, instance_exact 1, instance_ok 1, missing_trajectory 0, n_actions 7376, video_rendered 1 |
| L✓ 21m | 09-28 07:06 | 21m | 43 | 2.4M / 47k | $0.057 | deterministic 1, grader_error 0, instance_exact 1, n_actions 4919, replay_success 1, video_rendered 1 |
Task instruction (upstream)
Paint the coloured line drawing shown in the picture /app/target.png onto the paper sheet in front of the robot, loading the brush with paint from the palette beside the paper.
Recorded by us: our reference solution, replayed by the verifier (canvas camera, scene camera, side camera, exact canvas).
What RoboPaint states about this task
| Success criteria | 1. precision >= 0.90: painted pixels within 1 mm of a line of their colour 2. recall >= 0.95: the drawing's line length with paint of its colour within 1 mm 3. every colour's own recall >= 0.95 4. (reported) every element of the drawing counts as complete at recall >= 0.90 5. both fresh-process replays of the trajectory meet this and end in the same state |
| Family | line |
| Picture | a cottage, a tree, the sun, birds |
| Difficulty | easy |
| Line Complexity | medium |
| Tier | privileged |
| Colours | black red yellow green |
| Line Width | 3 mm |
| Line Length | 1398 mm |
| Line Length By Colour | black — 682 mm red — 191 mm yellow — 144 mm green — 381 mm |
| Continuous Score | F1 |
| Scoring | s.score() is served by a separate scoring container: it replays the actions the agent executed since its last reset on the graders' copy and returns the metrics, never images or target data |
| Agent Budget | 5400 s of wall clock per mode |
From our RoboPaint task definitions.
Tags¶
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.