Contributing¶
Setup¶
git clone <this-repo> && cd <this-repo>
make install # one-time: .venv + the pinned toolchain
make demos # one-time: 98 oracle demos + posters, ~1.2 GB
make edit # start the site WITH editing enabled
The port is chosen automatically and printed, so two people on one machine do not collide. Pin it with make edit PORT=9000.
make demos is optional but strongly recommended — without local video files the task list shows placeholders. We use a <video> element rather than an embedded YouTube/Vimeo player precisely so playback speed works, and that needs local files.
make serve versus make edit¶
If you intend to change anything, you need make edit
make serve | make edit | |
|---|---|---|
| The site | yes | yes |
| Edit button on a task row | visible, but will not save | saves to state/ |
| Label taxonomy page | read-only | editable |
| Extra process | none | scripts/editd.py on port 8009 |
Both commands serve the same site. The difference is the daemon: a web page cannot write to your filesystem, so every save goes through a small local HTTP API. make edit starts it and stops it with the server.
Under make serve, clicking Edit replaces the button text with run `make edit` rather than doing nothing, so the failure is visible instead of mysterious. If your changes are not sticking, that is why.
The daemon is localhost-only, unauthenticated and single-user — everyone runs their own copy, so there is no locking and no conflict handling. Do not expose it on a network. A deployed build has no daemon and is therefore always read-only.
Everything you change lives in state/¶
All curation — every status, label, owner and note, plus the taxonomy itself — is in one folder:
state/
taxonomy.yml the label vocabulary: 6 facets, 35 labels
display_tags.yml what pages show: Capability and Task Domain tags
tasks/behavior-1k.yml per-task status, difficulty, labels, display tags, owner, notes
Task pages under docs/ hold only upstream metadata and prose. So reviewing a colleague's triage is:
rather than opening a hundred files. Two people on different branches touch one file each instead of a hundred each.
Editing from the page¶
With make edit, every task row has an Edit button. It opens in place, so you decide while watching the demo — which is when the judgement actually happens. You can set:
- status — keep, drop, needs-review, pending
- difficulty — for a coding agent driving this robot, not a human
- display tags — as toggle buttons, under their group: Capability, then Task Domain; a save needs at least one from each. Tags implied by the task's annotated BEHAVIOR skills are marked with a
·, so usually you are confirming a suggestion rather than reading the whole vocabulary.
Saving writes state/tasks/<benchmark>.yml and nothing else — task pages are never touched. Add & tag creates a new tag, in the group you pick, when the existing vocabulary does not fit. The detailed labels (labels:) are not on the page: set them in the state file by hand.
The Label taxonomy page has the same treatment for the display tags themselves: rename groups, add or remove tags, save back to state/display_tags.yml. It refuses to delete a tag that tasks still use, and names them, so the vocabulary cannot silently break your task state.
Your first contribution: triage one task¶
- Open the task list. The default order puts already-labelled work first, so scroll past it — or filter Status to
pending. - Set the playback speed to 2× or 3× once; it is remembered everywhere.
- Click Edit on a task, set
ownerto your GitHub handle, and Save. Now nobody duplicates your work. - Watch the demo in the row. Then set status, difficulty and labels, and Save.
- Optionally open the task page and write down why — see the task page guide. The prose is what a reviewer reads.
make check, then commit. Your whole session is a diff ofstate/.
One task per pull request. A one-file PR merges in minutes; a thirty-file PR sits for a week. Exempt: taxonomy changes, tooling, and upstream re-syncs.
What goes where¶
| You want to… | Edit |
|---|---|
| Triage a task, set tags | state/tasks/<bench>.yml — or the Edit button |
| Leave a task out of the final benchmark, keeping it on the site | excluded: "<why>" in state/tasks/<bench>.yml — see the task page guide |
| Add or reword a display tag | state/display_tags.yml — or the Labels page |
| Add or reword a detailed label | state/taxonomy.yml |
| Write prose about a task | docs/benchmarks/<bench>/tasks/<task_id>.md |
| Record benchmark-level facts | data/benchmarks/<bench>.yml |
| Register an agent run, or your benchmark's part in one | data/agents/<run>.yml, state/runs/<bench>.yml — see Registering runs |
| Change the task list / rows | scripts/gallery.py, docs/javascripts/tasklist.js |
| Change a generated page | scripts/gen_pages.py (the Runs pages: scripts/runpages.py) |
| Change the task-page header | scripts/hooks.py |
Never hand-edit an aggregate page
The task list, coverage matrix and capability reference are generated at build time and not committed. Wanting to edit one means the change belongs in the generator or in task frontmatter. This is the single rule that keeps many contributors from colliding.
Review¶
CI runs make check: scripts/validate.py (schema, taxonomy, registries) plus mkdocs build --strict (no broken links). Errors block a merge; warnings — an incomplete task page — do not. It then runs make public, the read-only build that is published.
Publishing¶
Every push to main publishes the site at https://embodied-agent-interface-v2-internal.github.io/, through the Deploy workflow in GitHub Actions. There is nothing to run by hand.
- Before you push, run
make check, thenmake publicto see the public site as it will be deployed. - The public site is read-only. It shows no Edit buttons. Its agent runs come from the committed snapshot that
make publish-runswrites, without machine names or paths. Runs markedhiddenare not shown. - To take a change back, revert the commit and push.
HANDOFF.md, section "Publishing", has the details: the steps, rollback, and who needs which permission.
A human reviewer checks the things CI cannot:
- Do the capability tags pass the necessity test?
- Is every
dropreason falsifiable? "Too hard" is not; "the goal is satisfiable without touching the target object" is. - Are reconstructed goals labelled as reconstructions?
- Did anyone write into an
upstream:block?
Approving does not mean agreeing with every call — it means the call is stated clearly enough to disagree with later.
When two people disagree¶
Set status: needs-review, put both positions in status_reason, and settle it on the task page's comment thread. Whoever closes it updates the page. Never silently overwrite someone else's triage decision.
Conventions¶
- English only, in pages and commit messages.
- Sign your judgements with your handle.
status: needs-reviewwith a real question beats silence.- Never edit
upstream:— it is overwritten on the next sync.
Re-syncing upstream¶
Only upstream: blocks should change. Open it as its own PR. If prose changed, stop and investigate. Tasks that vanish upstream are reported, never deleted.