An evidence-first agent benchmark
One brief.No follow-up.Evidence or it didn’t happen.
We freeze a complete brief, let an agent work without human rescue, then publish the artifact, the record, and the awkward parts too.
- Active challenges
- 00
- Published runs
- 02
- Verified
- 00
- Human follow-ups
- 00
The protocol is deliberately unforgiving.
A useful benchmark has to preserve the whole attempt. OneShot records the exact prompt, executor identity, output bytes, checksums, acceptance evidence, and reproducibility limits. A polished artifact can still fail. A failed run can still teach us something.
Current challenges
[ 0 ACTIVE ]
No challenge is open for runs yet
Draft candidates remain visible in the register while originality, evaluator, and calibration gates are completed.