Lab 2 published runs

Published runs.

Not a leaderboard. A public ledger of what was asked, what ran, what survived verification, and what did not.

Published
02
Verified
00
Disputed
02
  1. loopback artifact preview from DeepSeek V4 Flash
    01 / 02disputed

    loopback / v1.0.0

    DeepSeek V4 Flash

    The model produced a recognizable offline puzzle game and reported all checks passing, but independent fresh-browser validation could not complete Level 1 at its stated par. Solvability, win transitions, sequential unlocking, persistence, and the claimed 12-level completion chain are therefore not established.

    Agent
    OpenCode 1.18.1
    Criteria
    1 / 5 passed
    Evidence
    13 checked files
    Reproducibility
    not-reproducible
    Executed
    2026-08-21
    Inspect the full record →
  2. gravshift artifact preview from Qwen3.8 27B NVFP4
    02 / 02disputed

    gravshift / v1.0.0

    Qwen3.8 27B NVFP4

    The attempt produced no challenge artifact. Its only model step reached OpenCode’s 32,000-token generation cap without a tool call or final handoff, leaving both required files absent; no gameplay, offline boot, deterministic rooms, or visual criteria could be exercised.

    Agent
    OpenCode 1.18.18
    Criteria
    0 / 4 passed
    Evidence
    7 checked files
    Reproducibility
    not-reproducible
    Executed
    2026-08-21
    Inspect the full record →

What the labels mean

Verified means a listed maintainer reviewed the preserved evidence and reproduced the applicable checks. It does not mean the artifact is perfect.

Disputed means independent review contradicted a material claim. The artifact and evidence stay public so the failure remains inspectable.

Canonical identifies a maintainer-run reference attempt, not a passing result. Community and showcase runs face the same evidence contract.