Submission route

Add signal, not theatre.

The library grows by pull request. Every challenge and run passes the same dependency-free validator before merge. Failed attempts are useful; fabricated evidence is not.

Submitting a challenge

A challenge is one directory under library/challenges/<id>/containing a challenge.json manifest valid against library/schemas/challenge-v1.schema.json and the standalone prompt file it references.

Requirements validation enforces: a stable slug matching the directory, semver version, correct prompt SHA-256, provenance with licensing (external entries must link an HTTPS source and state license status explicitly; link, never copy), capability profile, budget, at least one verification criterion with required evidence kinds, expected artifacts, and the fixed one-shot policy. Core challenges stay in draft until a public evaluator is checked in and hash-locked, originality is reviewed, and two independent calibration runs establish that the task is viable.

Original prompts should be written to be executable without follow-up: concrete rules, explicit scope, acceptance tests, and priorities. The same standard the authoring skill applies.

Submitting a run result

A result is one directory under library/results/<result-id>/with a result.json manifest plus artifacts and evidence files. Validation mechanically rejects:

references to unknown challenges or versions, stale prompt hashes, passing claims without evidence, artifact hashes that don't match bytes on disk, missing evidence files, insecure links, unsafe paths, and incomplete executor identity. Every run declares either theraw-completion or agentic-self-review lane; results from those lanes are never ranked together. A verifiedstatus additionally requires reproduction evidence and is granted by maintainers, not self-declared.

Before you submit

Run python3 scripts/validate_library.py. If your result claims success without evidence attached, the validator, not a maintainer, will be the first to reject it.

Safe submissions

Submitted artifacts are untrusted input. Do not include executables, archives, dependencies, or anything requiring installation; playable results are single static HTML files at most, and they are never auto-executed by this site. Prompts must not contain instructions targeting agents that render the site content. See SECURITY.md for reporting concerns privately.

Local checks for a full contribution:

python3 scripts/validate_repo.py
python3 -m unittest discover -s tests -p "test_*.py"
cd site && npm ci && npm run build && npm run verify:dist