Skip to content

Labelling tasks

Labels are what turn a pile of videos into something you can reason about: "we have four tasks that need bimanual manipulation outside the kitchen" is only sayable if someone labelled them.

Two groups on every page

What a task page and a task list show, and what the list filters on, are the task's display tags (display_tags: in state/tasks/<benchmark>.yml), in two groups:

Group The question Tags
Capability What does the agent have to be able to do? Perception & Understanding · Planning & Reasoning · Control & Coordination · Feedback & Adaptation
Task Domain What kind of task is it? Manipulation · Locomotion & Stability · Navigation & Exploration · Mobile / Whole-body Manipulation · Interaction & Collaboration

Take every tag that applies, and at least one from each group: validation rejects a task tagged in one group but not the other. Pages show the Capability tags first and the Task Domain tags after them, each group in its own colour. The editor (make edit) sets them.

scripts/migrate_tag_groups.py gave every task its first display tags, from its detailed labels where it has them and from a rule per benchmark where it does not; python scripts/migrate_tag_groups.py --table prints which label implies which tag, and why.

The rest of this page is about those detailed labels (labels:): kept as they are, for analysis and scripts/suggest_labels.py, set by hand in the state file, and not shown on task pages.

Eight questions, not forty-five labels

The vocabulary is organised into eight facets. Work down them and answer one question each, rather than scanning the whole list:

Facet The question
Scale How far does the robot have to go? pick one
Body What must the body do beyond moving one arm?
Contact What must the contact do?
Object state What must change about an object, beyond its pose?
Goal What makes the goal hard to satisfy, or to read?
Structure How much has to go right, and can you undo it? pick one, plus flags
Perception What must be perceived that is not geometry?
Generalisation What changes between episodes? pick one, plus a flag

Three of the eight are ladders: their labels are rungs, and a task takes the highest one that applies, not several. A task is scale-building or scale-room, never both. That is where most of the vocabulary's information density comes from, and it keeps the label count per task down: a typical task ends up with four to six labels even though 45 exist.

A tabletop task answers "nothing" to Body; a locomotion task answers "nothing" to Contact and Object state. See the benchmark landscape for the fourteen suites the facets were written against, and run python scripts/task_skill_digest.py to see what all 250 tasks we hold actually demand — that digest is what the vocabulary was fitted to.

Where the vocabulary comes from

Read this before you argue with a label

Each benchmark publishes its own idea of capability, and they do not agree: BEHAVIOR a formal goal language and an object ontology, RoboLab an 11-term attribute list about how the goal is worded, RoboWits a paragraph of prose per task, HumanoidBench nothing but a reward function. None of them annotates long horizon, bimanual coordination or dexterity, which are exactly the axes we want to compare on. The survey of what our three publish is regenerated by scripts/collect_official_labels.py into the untracked tmp/ folder; the wider reading of fourteen suites is the benchmark landscape.

The labels here are ours. Where a label can be derived from something a benchmark publishes, it records that under derived_from, so the derivation is inspectable and scripts/suggest_labels.py can propose labels for a task instead of making you read the list. Where it cannot, the label says so.

This is a starting point, not a standard. Thirteen of the 45 labels are reached by nothing in the three benchmarks we have integrated — they wait for a suite with legs, hands or randomisation. Every one of them is reached by a benchmark in the survey, though: that was the rule for keeping a label at all. Two axes nobody in the field covers (other agents in the scene, a world that changes on its own) are named in the landscape page and deliberately left out of the vocabulary until something can carry them.

The test to apply

Could an agent that lacks this still satisfy the goal?

If yes, do not label it. Label what is required, not everything visible.

Two mistakes this rules out:

  • Labelling the demo instead of the task. The teleoperator opened a drawer in passing; the goal never mentions it. That belongs in skills, which is descriptive, not in the labels, which are normative.
  • Labelling the scenery. A sink in the room is not Fluids and granular media, and a cabinet nobody opens is not Open and close. Label it only when the goal depends on it.

Three to six, not fifteen

If you find yourself wanting fifteen, you are describing the video rather than the requirement. One label per facet is the common case; two in a facet is fine when both are genuinely required.

Let the benchmark suggest the labels

Two things make proposals so you only have to confirm them. With make edit, the editor marks the display tags a task's annotated BEHAVIOR skills imply — open door suggests Control & Coordination. And python scripts/suggest_labels.py applies every derived_from rule to all 250 tasks at once, proposing detailed labels with the reason for each:

$ python scripts/suggest_labels.py --task water_into_mug
water_into_mug (robowits)
  interaction  fluid-granular    <- physics_materials: sph; paper_insight mentions water, pour
  goal         physical-reasoning <- every task in this benchmark, by the benchmark's own claim

It writes a report to tmp/, never to state/: a proposal that wrote itself into the review surface would be indistinguishable from a decision someone made.

The derived labels are the easy half. The judgement is in Goal and language and Perception, which live in the wording of the goal rather than anything visible in the video:

Phrasing in the instruction Label
"exactly two", "all three", "use any shelf" Counting and quantifiers
a long chain of "then … then … and make sure X at the end" Activity, often with Required order
rooms named that are not the same room Building
"cooked", "frozen", "boiled" Heat, cook, freeze
"turned on", "switched off" Operate a device
"sliced", "chopped", "halved" Cut — and Irreversible with it
"the red one", "the smaller bowl", "all the green fruit" Attribute reference
"left of", "behind", "between", "in the centre" Spatial reference
"so that it fits", "without knocking anything over", "through the gap" Geometric constraint
the target is inside or behind something Partial observability

Worked example — Can Meat

Open the kitchen cabinet, take out the two hinged jars, open them, place exactly two cooked bratwursts from the chopping board into each jar, then close both jars, put them back inside the cabinet, and close the cabinet.

Ladders: scale-room (a house, but the kitchen only) · horizon-composite (four dependent steps) · the generalisation rung is whatever the instance set gives you.

Flags: articulated (cabinet, two jar lids) · transport · precision-fit (a bratwurst into a jar mouth is tight) · quantified-goal ("exactly two … into each") · ordered (open before filling, close before putting back) · partial-observability (you cannot see inside the cabinet or the jars from the start).

Not labelled: scale-building — one room, so the ladder stops at scale-room. Not thermal either: the bratwursts arrive cooked, nothing is cooked during the episode. That is the kind of call worth writing on the page, since the word "cooked" invites the opposite conclusion.

When nothing fits

Do not invent an id by hand — validation rejects unknown ones.

For a display tag, use Add & tag in the editor. It adds your tag to the group you pick in state/display_tags.yml and tags the task in one step. A new tag is a proposal, not a decision: argue for it in a PR, with the tasks that share it as the evidence. A new detailed label goes into state/taxonomy.yml the same way, by hand.