The submission workflow used to keep no copy of the source it evaluated. Public-repo submissions can be re-fetched from their upstream so long as that upstream exists, but private-repo submissions exist only in the submitter's account; if the submitter deletes the repo, rotates a tag, or otherwise rewrites history, the bytes the benchmark scored are gone. That defeats post-hoc auditability: a comparator regression, a soundness incident, or a research question about an older proof has no recoverable artifact to examine.
The archive exists to make every evaluated submission recoverable indefinitely, while keeping the source bytes inaccessible to anyone outside a small maintainer set.
One encrypted tarball plus one unencrypted JSON sidecar are pushed to
leanprover/lean-eval-audit
per submission, immediately after evaluation.
audit/
YYYY/
MM/
{submitter}-{issue}-{ref8}.tar.age # age-encrypted gzipped tar of source
{submitter}-{issue}-{ref8}.json # sidecar (issue, submitter, model, digests, verdict)
The tarball is the same source.tar.gz that fetch_submission.py
already produces — the same bytes the evaluator sees. Encryption uses
age with recipients listed in
.audit/recipients.txt. The sidecar
records the SHA-256 of both plaintext-tar and ciphertext so an
operator can verify integrity at decrypt time (against the plaintext
digest) and without decrypting (against the ciphertext digest).
For submissions made after publication disclosure was added, it also
preserves the submitter's solution_publication_status and optional
solution_publication_date snapshot.
Two new pieces live inside the existing submission.yml:
-
Size cap + encrypt, in the
evaluatejob, right after fetch.fetch_submission.pywritessource.tar.gzto/tmp/fetch-out/. The workflow rejects the submission ifsource.tar.gzexceeds 10 MiB (comments on the issue, closes asnot planned). Otherwise it shellsage --recipients-file .audit/recipients.txtover the tar and uploads only the ciphertext as an artifact. The plaintext is read for evaluation in this same job but never crosses the job boundary. -
Archive job, runs after
evaluate. Mints an installation token for thelean-eval-archiverGitHub App (scoped only tolean-eval-audit), downloads the ciphertext + sidecar artifact, merges in the per-problem evaluator verdict fromsummary.json["run_eval"]["problems"](the raw output oflake exe lean-eval run-eval --json), and uploads both objects via the GitHub Contents API.record(the leaderboard updater) is gated on this job succeeding: if archive fails, no leaderboard update happens.The push is idempotent on the source, which matters because a submission can be re-evaluated (e.g. after a benchmark toolchain bump) and neither the ciphertext nor the plaintext tar is reproducible —
agepicks a fresh file key per run, and gzip/tar packaging varies, so re-fetching the same source yields different bytes at the sameaudit/YYYY/MM/{submitter}-{issue}-{ref8}path. The script therefore decides on the recorded submission identity — the stable tuple(submitter, issue, submission_repo, submission_ref)in the sidecar.submission_refis a 40-char git SHA, which immutably pins the source tree content, so no digest is needed to identify the source (and including the non-reproduciblesha256_plaintext_tarwould misclassify a legitimate re-evaluation as a collision). Before uploading it reads any existing sidecar: if the identity matches, this exact source is already archived, the push is a no-op, and the original (immutable) ciphertext is left untouched. A path collision with a different identity is an operator-investigatable collision and fails hard, not something to silently resolve.The per-file Contents API call is a create-or-update: it fetches any existing file's Git blob SHA and supplies it on the PUT, so a rerun that finds an orphan ciphertext (uploaded by a prior run that crashed before the sidecar) updates it in place rather than failing the create with the API's
"sha" wasn't supplied422. That 422 is only treated as a benign already-exists conflict when the body says so; any other 422 (malformed path, oversize content, branch protection) is a real validation failure and fails fast.evaluateexposesaudit_ciphertext_readyas a job output, set to'true'only when both the encrypt step and the ciphertext-artifact upload step succeeded.archivegates on this output;notifybranches on it so an audit-encryption failure produces an "audit encryption failed" comment rather than the misleading "Submission.lean failed to compile" message that would otherwise fire from the generic evaluate-failure path.
The thing the design defends against is: the source bytes of any private submission leaking out of the maintainer set.
Concretely:
- Public-repo artifacts are downloadable by any authenticated user.
This is why the workflow already forbids
name: submission-source(seesubmission.yml's leading comment, and thetest_fetch_and_evaluate_share_one_jobinvariant). The ciphertext artifact is safe to upload because anyone who downloads it cannot decrypt without a recipient private key. - Runners that elaborate untrusted Lean can be compromised. The
archiver App's installation token must never appear in the env of
any job that runs untrusted Lean. It is therefore minted only in
the
archivejob, which runs on a separate runner that has never touched the submitted source. - App permission scoping.
lean-eval-archiverhasContents: writeonly onleanprover/lean-eval-audit, and is installed only on that repo. The pre-existinglean-eval-bot(which reads private submission repos) andlean-eval-recorder(which writes the leaderboard) keep their previous, narrower scopes — none of them is bumped to write to private submitter accounts or to audit.
- Recipient-private-key custody is the recipient's problem. Lose it and the corresponding archived entries become permanently undecryptable. The validate-recipients CI does not (and cannot) verify that someone holds each recipient's private key; that requires a manual decrypt drill.
- Pre-existing public submissions and their commits. Those exist
in the submitter's own public repo at a pinned SHA, recorded in
results/<login>.json. They are also archived going forward for symmetry and forensic completeness, but a deleted upstream remains primarily a problem for the submitter's own repo, not for the archive. - Anonymity. Submissions are not anonymous in the leaderboard, in the issue, or in the sidecar. The archive's privacy guarantee is about source bytes, not about whether a given submitter participated.
Open a PR that edits .audit/recipients.txt. The
validate-recipients.yml workflow lints each line by encrypting a
fixture to it; merging without the lint passing is impossible. Once
merged, every subsequent submission is encrypted to the new recipient
set. Pre-existing ciphertexts retain the recipient set they were
encrypted with — re-encrypting historical entries to a new recipient
is a manual operation requiring decryption first.
# 1. Install age.
brew install age # or: apt install age, or cargo install rage
# 2. Decrypt the ciphertext using one of the SSH private keys whose
# public half is in recipients.txt.
age -d -i ~/.ssh/id_rsa \
-o /tmp/source.tar.gz \
audit/2026/05/GanjinZero-73-52c6d202.tar.age
# 3. Verify the recovered bytes against the sidecar's plaintext SHA.
sha256sum /tmp/source.tar.gz # match sidecar.sha256_plaintext_tar
# 4. Extract.
mkdir /tmp/source && tar -xzf /tmp/source.tar.gz -C /tmp/sourceConduct a decrypt drill periodically (annually is plenty) so that recipient private keys are known to still exist on a reachable device.
The 10 MiB cap is on the compressed gzipped tar of the
post-.git-strip source tree, which is what the workflow uploads
and what du -h will show for a typical generated workspace plus a
modest Submission/ directory (well under 1 MiB for current
submissions). The cap exists because the archive is permanent and
public-repo-shaped (one file per submission, committed forever) and
because nothing in the current submission shape needs more space; if
a use case for >10 MiB submissions emerges, we bump the cap rather
than special-case some submissions.
A submission over the cap is rejected at the workflow level: the issue is commented and closed, no evaluation is run, no leaderboard update happens, and no audit entry is created.
A one-off script (scripts/backfill_audit.py) walks every entry in
results/*.json, re-fetches the submission via lean-eval-bot, and
archives it the same way the live workflow would have. Public-repo
submissions are best-effort: if the upstream has been deleted or
rewritten, the entry is logged and skipped. Private-repo submissions
where the bot has lost access fall into the same bucket. The
backfill is idempotent — entries that already exist in the audit
repo are not re-uploaded.