Skip to content

About RoboLab

120 tabletop tasks on a fixed-base Franka with a Robotiq 2F-85 gripper, simulated in Isaac Lab / Isaac Sim. The manipulation is deliberately ordinary — almost every task is pick-and-place. What varies is how the goal is stated: by colour, by size, by category, by count, by spatial relation, by order, or vaguely enough that the agent has to interpret it.

All 120 tasks Capability coverage

Why we are looking at it

It isolates language grounding from motor difficulty. A suite where the physical act is constant and the instruction changes is the cleanest way to ask whether an agent understood the goal, and it is the axis BEHAVIOR-1K does not test (its instructions are long but each names its objects). Three wordings ship per task — default, vague, specific — which makes "did it understand, or did it pattern-match?" a measurable question rather than a rhetorical one.

The embodiment

Robot Franka Panda + Robotiq 2F-85 (RoboLab's "DROID" configuration)
Action 7 absolute joint-position targets + 1 binary gripper command, at 15 Hz
Cameras Wrist 1280 × 720; an egocentric viewport camera for recording
Privileged state Available (--enable-gt-state); tasks can be run state-based or from pixels
Episode budget 20–300 s, declared per task (episode_length_s)

Tasks are not tied to the embodiment — any Isaac Lab robot can be dropped in, which is unusual and makes the suite reusable if we change robots.

How a task defines success

A task is a dataclass: a scene, three instruction wordings, an episode budget, attribute tags, one success predicate and a list of scored subtasks. Success is a composable predicate over object state — object_in_container (72 tasks), object_on_top (13), object_outside_of_and_on_surface (7), stacked (5), object_upright (4), and spatial relations. All of it is mirrored into each task page's upstream: block.

Partial credit comes from the subtask list: 7 tasks score in fractions, the rest are all-or-nothing.

Upstream's attribute tags, and its difficulty label

RoboLab tags every task from an 11-term vocabulary and derives a difficulty label from it (SKILL_WEIGHTS over the attributes against DIFFICULTY_THRESHOLDS). We mirror both:

Attribute Tasks Attribute Tasks
semantics 60 conjunction 8
spatial 29 counting 7
color 26 vague 7
affordance 12 size 6
sorting 12 stacking 6
reorientation 6

Those tags are what a task page's skills: records, seeded from upstream.

difficulty_label is theirs; difficulty is ours

Upstream's label (108 simple, 8 moderate, 4 complex) is computed from the attribute weights above — it measures how much language the task involves. Our difficulty in state/tasks/robolab.yml asks a different question: how hard is this for a coding agent driving the robot? Do not copy one into the other.

Demo coverage

13 of the 120 tasks have a clip on the project page, downloaded by make demos. RoboLab distributes no demonstration dataset — it evaluates policies rather than shipping teleoperation — so the other 107 rows show the scene the task starts from instead, taken from the scene previews that ship with the repository. A still is not a demonstration, and the row says so; it is the difference between reviewing a task and guessing at it.

One row plays a video that is ours, not upstream's: banana_in_bowl, the task we integrated into robot_coding_bench. The credit line under the player says so.

What we run it on

Pinned to v0.3.1, the tag our image builds from, measured on an RTX 4090 laptop:

Image nvidia/cuda 12.8 + RoboLab v0.3.1 from our uv.lock + its asset library, 14.5 GB
Physics CPU PhysX — bit-exact live, in-process and across processes
GPU PhysX Reproduces a live run in a fresh process, but drifts ~1 mm across in-process replays after reset_to; enable_enhanced_determinism does not change that
Throughput ~15 control steps/s with the wrist and viewport cameras rendering
Startup ~70 s Isaac Sim launch + ~30 s scene build per process, ~5 GB VRAM

Version matters here: v0.3.1 moved the canonical ground plane and made ee_recorder_bodies mandatory, so recordings do not transfer from v0.2.x.