Skip to content

Task page guide

A task has two halves, kept in different places on purpose.

Where Who writes it
Upstream metadata frontmatter of docs/benchmarks/<bench>/tasks/<id>.md the importer, never you
Prose body of that same file you
Curated state state/tasks/<bench>.yml you, or the Edit button

The split means a full afternoon of triage is one file's diff, and a re-sync from upstream can never clobber a decision you made.

The task page

---
title: Can Meat
task_id: can_meat
benchmark: behavior-1k

# --- upstream: synced by scripts/import_behavior_tasks.py. Do not hand-edit. ---
upstream:
  source: https://behavior.stanford.edu/challenge/tasks/index.html
  synced: '2026-09-20'
  cohort: 2026-new
  instruction: Open the kitchen cabinet, take out the two hinged jars, ...
  scene_model: house_single_floor
  rooms: [kitchen]
  demo_duration_s: 395
  oracle_video: https://player.vimeo.com/video/1114054618
  oracle_thumbnail: https://vumbnail.com/1114054618.jpg
---

Anything you write in upstream: is destroyed on the next sync. If upstream is wrong, say so in the prose and attribute the correction to us. cohort is derived by the importer from the demo's video host, because the 2025 carryover tasks use YouTube and the 2026 ones use Vimeo.

Below the frontmatter, four prose headings. Delete the HTML comments as you fill them in; leave _Not yet written._ where you have nothing, so the gaps stay visible.

Section What belongs there
Why this task is interesting One paragraph. Why it earns a slot — or why it does not.
Capability notes Justify each label. One bullet each is plenty.
Oracle demo review Does the demo satisfy the goal? Is it clean? Would it mislead an imitation learner?
Discussion Conclusions from the comment thread. Sign your points.

Do not restate the instruction, scene, rooms or duration in prose — all of it renders automatically at the top of the page, and a copy would go stale.

The curated state

state/tasks/<benchmark>.yml, one block per task:

can_meat:
  status: keep
  difficulty: hard
  labels: [articulated, pick-place, insert-attach, long-horizon, counting,
           object-state, search]
  display_tags: [perception-understanding, planning-reasoning, control-coordination,
                 feedback-adaptation, mobile-whole-body-manipulation]
  owner: alice
  note: Dense multi-constraint goal; high discriminative value per episode.
  skills: [open door, open lid, pick up from, place in, close lid, close door]
Field Values
status keep · drop · needs-review · pending
difficulty unrated · easy · medium · hard · extreme — for a coding agent driving this robot, not a human teleoperator
labels tier-2 ids from the taxonomy, the detailed labels, not shown on the page; unknown ids fail validation
display_tags what the page and the task list show: ids from the display tags, at least one Capability and one Task Domain tag; unknown ids, or tags from one group only, fail validation
owner bare GitHub handle, no @
note free text; optional
excluded optional: why the task is not part of the final benchmark — see below
skills BEHAVIOR's own primitives observed in the demo; must match their vocabulary exactly
media extra screenshots or clips — see below

Defaults are omitted from the file, so a block only ever contains decisions someone actually made.

Use the Edit button

With make edit, every row of the task list has an Edit button that writes this file for you, while you are watching the demo. That is when the judgement actually happens.

excluded: out of the final benchmark, still on the site

line_city:
  difficulty: easy
  excluded: "too easy for the final benchmark (line drawings; kept for the record, results in Runs)"

The benchmark owner's call, per benchmark, and separate from status (a benchmark that drops tasks in review keeps doing so, and looks the same). The task keeps its page, demo and agent runs; lists show it greyed, labelled excluded, at the bottom; the benchmark's counts and its Runs numbers leave it out, with the excluded ones shown beside them; the task page opens on a banner with the reason. The value must say why. Set it in the file by hand; the Edit button leaves it alone.

labels versus skills

skills is descriptive — what the demo does, in BEHAVIOR's own vocabulary. labels is normative — what an agent must handle. They come apart often: a teleoperator may open a drawer the goal never required. See Labelling tasks.

media

Never commit media files

Host them elsewhere and reference by URL — drag into a GitHub issue comment for a permanent CDN link, or use a HuggingFace dataset repo for anything systematic. kind is image, video or link; url must be absolute.

The caption is the value: "fails at 2:10, drops the jar while turning" tells a reader whether to click; "rollout 3" does not.

can_meat:
  media:
    - kind: image
      url: https://user-images.githubusercontent.com/12345/grasp-failure.png
      caption: Gripper clips through the jar rim on approach
      credit: "@handle"

Reconstructing a missing instruction

The 50 2025-carryover tasks ship no instruction upstream. We leave it empty — inventing one would launder a guess into a field that reads as official.

Put your reconstruction in Why this task is interesting, clearly marked, and set status: needs-review with a note saying the goal is reconstructed:

!!! note "Reconstructed goal — not official"

    Upstream publishes no instruction for this task. From the demo, the goal
    appears to be: move both frozen fruit packages from the countertop into the
    freezer and close it. Unverified against the BDDL definition. (@handle)

Before you push

make check

Warnings are fine — an untriaged task is all warnings. Errors are not.