Judging protocol
Definitions over vibes.
The lab is an evidence-backed field guide to autonomous work: playable systems are a flagship family alongside builds, precision work, repair, and edge-case probes. Its only special power is restraint: a fixed meaning of one-shot, hashed prompts, explicit evidence requirements, and a refusal to display what cannot be substantiated.
What "one shot" means
A one-shot run is a single execution attempt driven byone initial brief with no human follow-up: no answered questions, no extra hints, no corrections, and no manual fixes between the start of the run and the executor's final handoff.
"One shot" constrains human involvement, not machine effort. Within the single run an executor may read files, consult documentation, use tools and subagents, run and observe its own artifact, test, find defects, repair, and iterate internally as much as it likes.
A run stops being one-shot if a human answers a mid-run question, edits the prompt after execution started, touches up the artifact afterward, or stitches multiple runs together by hand.
Compact core, separate showcase
The comparative core uses T0 atom, T1 compact, and T2 bounded tasks. T3 showcase challenges are richer stress tests on a separate board. A full browser game cannot be ranked honestly beside an eight-word SVG prompt merely because both started from one brief.
| Tier | Purpose | Typical budget |
|---|---|---|
| T0 atom | One directly observable capability | 5-15 minutes, one deliverable |
| T1 compact | One useful artifact, microgame, or narrow repair | 20-30 minutes, one or two deliverables |
| T2 bounded | Small multi-state workflow or fixture-backed repair | 45-60 minutes, up to four deliverables |
| T3 showcase | Rich vertical slice or sustained autonomous build | 90-180 minutes, separate board |
Prompt depth is recorded separately as brief, benchmark, or extensive. A rigorous evaluator does not require an oversized contestant prompt.
Do not mix run lanes
Raw completion is one frozen prompt and one generation, with no visual feedback, repair prompt, retry, or human edit.Agentic self-review still uses one frozen prompt and no human follow-up, but the executor may run, inspect, test, and repair its work inside a fixed tool and time budget.
These measure different systems. Results are comparable only inside the same lane, model identity, tool profile, and output envelope.
Verification tiers
| Tier | Meaning | Requirement |
|---|---|---|
| unverified | The submitter claims the checks passed. | Evidence attached for every passing claim; artifacts and checked-in evidence hash-verified. |
| verified | A listed maintainer attestation records that the claim was reviewed or reproduced. | Every criterion passes with required evidence, required artifacts are present and hashed, and a review attestation is recorded. |
| disputed | Review found problems with the claim or evidence. | Kept public with the dispute noted; never silently deleted. |
| withdrawn | The submitter retracted the result. | Manifest retained; claim marked withdrawn. |
Untrusted artifacts & sandbox policy
Result artifacts, including playable HTML, SVGs, and logs, begin as untrusted input. Repository review can make them eligible for display, but does not make their code safe. Concretely:
Submitted artifacts never execute automatically or with the site's origin. The gallery is fully static and renders only records accepted into the repository, not live submissions. Accepted HTML starts only after an explicit click inside an opaque-origin iframe whose sole sandbox grant is scripts. A restrictive document policy blocks data connections, external resources, forms, popups, downloads, and top-level navigation, and no privileged parent bridge is exposed. SVG is treated as code because it can carry script, so checked-in SVG previews are loaded only as static images, never inserted as inline document markup.
The sandbox makes accepted games playable on their result pages without granting them site privileges. It is a constrained demonstration, not verification evidence or a safety certificate. Exact checked-in bytes remain available for inspection, while result claims still come from preserved evaluator-controlled evidence.
Reproducibility
Every result declares one of four statuses:reproducible (independent reruns reproduced the claim),partially-reproducible (core outcome reproduced with documented variance), not-reproducible, oruntested. Only reproducible or partially-reproducible results can reach verified status, and validation enforces this mechanically rather than on the honor system.
Model drift is expected: model versions and snapshots are recorded per run, and identical outputs are never promised across time. Only honest attempts under the same frozen brief.
A maintainer ID in a result is an allowlisted identity label and shape check, not cryptographic proof of who typed the file. Repository authorship and code review remain the authentication boundary.