Lab 2 published runs
Published runs.
Not a leaderboard. A public ledger of what was asked, what ran, what survived verification, and what did not.
- Published
- 02
- Verified
- 00
- Disputed
- 02
- 01 / 02disputed
loopback / v1.0.0
DeepSeek V4 Flash
The model produced a recognizable offline puzzle game and reported all checks passing, but independent fresh-browser validation could not complete Level 1 at its stated par. Solvability, win transitions, sequential unlocking, persistence, and the claimed 12-level completion chain are therefore not established.
- Agent
- OpenCode 1.18.1
- Criteria
- 1 / 5 passed
- Evidence
- 13 checked files
- Reproducibility
- not-reproducible
- Executed
- 2026-08-21
- 02 / 02disputed
gravshift / v1.0.0
Qwen3.8 27B NVFP4
The attempt produced no challenge artifact. Its only model step reached OpenCode’s 32,000-token generation cap without a tool call or final handoff, leaving both required files absent; no gameplay, offline boot, deterministic rooms, or visual criteria could be exercised.
- Agent
- OpenCode 1.18.18
- Criteria
- 0 / 4 passed
- Evidence
- 7 checked files
- Reproducibility
- not-reproducible
- Executed
- 2026-08-21
What the labels mean
Verified means a listed maintainer reviewed the preserved evidence and reproduced the applicable checks. It does not mean the artifact is perfect.
Disputed means independent review contradicted a material claim. The artifact and evidence stay public so the failure remains inspectable.
Canonical identifies a maintainer-run reference attempt, not a passing result. Community and showcase runs face the same evidence contract.