Labelling tasks¶
Labels are what turn a pile of videos into something you can reason about: "we have four tasks that need bimanual manipulation outside the kitchen" is only sayable if someone labelled them.
Two groups on every page¶
What a task page and a task list show, and what the list filters on, are the task's display tags (display_tags: in state/tasks/<benchmark>.yml), in two groups:
| Group | The question | Tags |
|---|---|---|
| Capability | What does the agent have to be able to do? | Perception & Understanding · Planning & Reasoning · Control & Coordination · Feedback & Adaptation |
| Task Domain | What kind of task is it? | Manipulation · Locomotion & Stability · Navigation & Exploration · Mobile / Whole-body Manipulation · Interaction & Collaboration |
Take every tag that applies, and at least one from each group: validation rejects a task tagged in one group but not the other. Pages show the Capability tags first and the Task Domain tags after them, each group in its own colour. The editor (make edit) sets them.
scripts/migrate_tag_groups.py gave every task its first display tags, from its detailed labels where it has them and from a rule per benchmark where it does not; python scripts/migrate_tag_groups.py --table prints which label implies which tag, and why.
The rest of this page is about those detailed labels (labels:): kept as they are, for analysis and scripts/suggest_labels.py, set by hand in the state file, and not shown on task pages.
Eight questions, not forty-five labels¶
The vocabulary is organised into eight facets. Work down them and answer one question each, rather than scanning the whole list:
| Facet | The question | |
|---|---|---|
| Scale | How far does the robot have to go? | pick one |
| Body | What must the body do beyond moving one arm? | |
| Contact | What must the contact do? | |
| Object state | What must change about an object, beyond its pose? | |
| Goal | What makes the goal hard to satisfy, or to read? | |
| Structure | How much has to go right, and can you undo it? | pick one, plus flags |
| Perception | What must be perceived that is not geometry? | |
| Generalisation | What changes between episodes? | pick one, plus a flag |
Three of the eight are ladders: their labels are rungs, and a task takes the highest one that applies, not several. A task is scale-building or scale-room, never both. That is where most of the vocabulary's information density comes from, and it keeps the label count per task down: a typical task ends up with four to six labels even though 45 exist.
A tabletop task answers "nothing" to Body; a locomotion task answers "nothing" to Contact and Object state. See the benchmark landscape for the fourteen suites the facets were written against, and run python scripts/task_skill_digest.py to see what all 250 tasks we hold actually demand — that digest is what the vocabulary was fitted to.
Where the vocabulary comes from¶
Read this before you argue with a label
Each benchmark publishes its own idea of capability, and they do not agree: BEHAVIOR a formal goal language and an object ontology, RoboLab an 11-term attribute list about how the goal is worded, RoboWits a paragraph of prose per task, HumanoidBench nothing but a reward function. None of them annotates long horizon, bimanual coordination or dexterity, which are exactly the axes we want to compare on. The survey of what our three publish is regenerated by scripts/collect_official_labels.py into the untracked tmp/ folder; the wider reading of fourteen suites is the benchmark landscape.
The labels here are ours. Where a label can be derived from something a benchmark publishes, it records that under derived_from, so the derivation is inspectable and scripts/suggest_labels.py can propose labels for a task instead of making you read the list. Where it cannot, the label says so.
This is a starting point, not a standard. Thirteen of the 45 labels are reached by nothing in the three benchmarks we have integrated — they wait for a suite with legs, hands or randomisation. Every one of them is reached by a benchmark in the survey, though: that was the rule for keeping a label at all. Two axes nobody in the field covers (other agents in the scene, a world that changes on its own) are named in the landscape page and deliberately left out of the vocabulary until something can carry them.
The test to apply¶
Could an agent that lacks this still satisfy the goal?
If yes, do not label it. Label what is required, not everything visible.
Two mistakes this rules out:
- Labelling the demo instead of the task. The teleoperator opened a drawer in passing; the goal never mentions it. That belongs in
skills, which is descriptive, not in the labels, which are normative. - Labelling the scenery. A sink in the room is not Fluids and granular media, and a cabinet nobody opens is not Open and close. Label it only when the goal depends on it.
Three to six, not fifteen¶
If you find yourself wanting fifteen, you are describing the video rather than the requirement. One label per facet is the common case; two in a facet is fine when both are genuinely required.
Let the benchmark suggest the labels¶
Two things make proposals so you only have to confirm them. With make edit, the editor marks the display tags a task's annotated BEHAVIOR skills imply — open door suggests Control & Coordination. And python scripts/suggest_labels.py applies every derived_from rule to all 250 tasks at once, proposing detailed labels with the reason for each:
$ python scripts/suggest_labels.py --task water_into_mug
water_into_mug (robowits)
interaction fluid-granular <- physics_materials: sph; paper_insight mentions water, pour
goal physical-reasoning <- every task in this benchmark, by the benchmark's own claim
It writes a report to tmp/, never to state/: a proposal that wrote itself into the review surface would be indistinguishable from a decision someone made.
The derived labels are the easy half. The judgement is in Goal and language and Perception, which live in the wording of the goal rather than anything visible in the video:
| Phrasing in the instruction | Label |
|---|---|
| "exactly two", "all three", "use any shelf" | Counting and quantifiers |
| a long chain of "then … then … and make sure X at the end" | Activity, often with Required order |
| rooms named that are not the same room | Building |
| "cooked", "frozen", "boiled" | Heat, cook, freeze |
| "turned on", "switched off" | Operate a device |
| "sliced", "chopped", "halved" | Cut — and Irreversible with it |
| "the red one", "the smaller bowl", "all the green fruit" | Attribute reference |
| "left of", "behind", "between", "in the centre" | Spatial reference |
| "so that it fits", "without knocking anything over", "through the gap" | Geometric constraint |
| the target is inside or behind something | Partial observability |
Worked example — Can Meat
Open the kitchen cabinet, take out the two hinged jars, open them, place exactly two cooked bratwursts from the chopping board into each jar, then close both jars, put them back inside the cabinet, and close the cabinet.
Ladders: scale-room (a house, but the kitchen only) · horizon-composite (four dependent steps) · the generalisation rung is whatever the instance set gives you.
Flags: articulated (cabinet, two jar lids) · transport · precision-fit (a bratwurst into a jar mouth is tight) · quantified-goal ("exactly two … into each") · ordered (open before filling, close before putting back) · partial-observability (you cannot see inside the cabinet or the jars from the start).
Not labelled: scale-building — one room, so the ladder stops at scale-room. Not thermal either: the bratwursts arrive cooked, nothing is cooked during the episode. That is the kind of call worth writing on the page, since the word "cooked" invites the opposite conclusion.
When nothing fits¶
Do not invent an id by hand — validation rejects unknown ones.
For a display tag, use Add & tag in the editor. It adds your tag to the group you pick in state/display_tags.yml and tags the task in one step. A new tag is a proposal, not a decision: argue for it in a PR, with the tasks that share it as the evidence. A new detailed label goes into state/taxonomy.yml the same way, by hand.