Stack White Mugs¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the RoboLab runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✗ 1h 30m | 09-27 16:29 | 1h 30m | 108 | 10.5M / 96k | $0.173 | grader_error 0, missing_trajectory 1 |
| L✗ 1h 30m | 09-29 06:52 | 1h 30m | 231 | 28.1M / 171k | $0.429 | deterministic 1, grader_error 0, live_success 0, n_actions 2669, replay_success 0, video_rendered 1 |
Task instruction (upstream)
Stack the white mugs on top of each other.

No oracle video published upstream — this is the scene it starts from.
What RoboLab states about this task
| Success predicate | stacked |
| Predicate arguments | objects: ['sideways_white_mug', 'upright_white_mug'] · order: None |
| Scored subtasks | 1 |
| Subtask predicates | Subtask |
| Objects | red_mug bowl ceramic_mug upright_white_mug sideways_white_mug cordless_drill measuring_cup |
| Upstream attributes | stacking color |
| Upstream difficulty | simple |
| Episode budget | 60 s |
| Other wordings | vague — Stack the white mugs specific — Take a white mug and place it on top of the other white mug |
| Environment class | StackWhiteMugsTask |
| Upstream name | StackWhiteMugsTask |
| Defined in | robolab/tasks/benchmark/stack_white_mugs_task.py |
From https://github.com/NVlabs/RoboLab @ v0.3.1.
Tags¶
Task DomainManipulation
Skill primitives in the demo: colorstacking
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.