Stack Cubes¶
Agent runs
Each mode of this task runs once per run. Est. cost is tokens at list price (data/prices.yml), never a bill; see the RoboWits runs for every task of a run.
Codex + GPT-6 Luna, reasoning xhigh (ChatGPT login)default run
| Trial | Started | Agent time | Model requests | Tokens in / out | Est. cost | Grade |
|---|---|---|---|---|---|---|
| U✓ 47m | 09-26 01:42 | 47m | 63 | 4.3M / 72k | $0.093 | deterministic 1, grader_error 0, instance_exact 1, instance_ok 1, n_actions 2798, video_rendered 1 |
| L✓ 1h 19m | 09-29 20:48 | 1h 19m | 283 | 43.3M / 165k | $0.574 | deterministic 1, grader_error 0, live_success 1, n_actions 7555, replay_success 1, video_rendered 1 |
Task instruction (upstream)
Build a stable two-layer stack so the red cube's center is above the green colored dot.
Playback speed
What RoboWits states about this task
| Success criteria | 1. Base cubes are on the table and touching with >= 50% face overlap 2. Apex cube sits on top of base cubes 3. Apex cube center is within tolerance of target marker 4. Apex-target overlap >= 20% |
| Objects | apex cube apex target base cube 1 base cube 2 |
| Physics | rigid |
| Episode budget | 200 control steps |
| Default control mode | EE_ABS |
| Evaluation scenes | 50 |
| Published mutations | 5 |
| Registry id | robowits/14-stack-cubes-v0 |
| Environment class | StackCubesEnv |
| Defined in | gs_gym/envs/robowits/14_stack_cubes.py |
From https://github.com/UMass-Embodied-AGI/RoboWits @ 9cc30ae.
Tags¶
Task DomainManipulation
Why this task is interesting¶
Not yet written.
Capability notes¶
Not yet written.
Oracle demo review¶
Not yet reviewed.