← Reports

Verification is all you need. Trust the gate, not the agent.

How Prover Labs verifies its own changes: a strict mechanical gate, then an independent reviewer that drives the app in a real browser, with a verdict that expires when the diff changes.

The bottleneck is belief, not code

Writing the feature

mostly solved

A base model already writes most of a feature correctly on the first try.

The bottleneck

Believing it is safe to merge

Does anyone, human or agent, actually believe the change is safe? At Prover, a company built on formally proving designs, that answer should not rest on instinct.

changethe diff Gate 1 make check local and CI, identical Gate 2 fresh reviewer agent drives a real browser Verdict: PASS .verify/<patch-id>.md git push · glab mr create what machines prove what only a running product shows bound to this exact diff
So we built the answer into the pipeline that ships Prover Labs itself: two gates, and a verdict filed under the diff's patch id that a local hook checks before a push or a merge request.

Attention Is All You Need

The paper the name echoes: one mechanism, applied consistently, replacing a pile of bespoke architecture.

Verification is all you need

Once a change is verified mechanically, then in a real browser, no trust ritual around the agent is left to perform.

Gate 1: what machines prove

make check runs identically on a developer's machine and in CI, the pipeline that runs on every merge request. Strict, and cheap.

296

backend tests, plus ruff and mypy, on every push

9

frontend tests, eslint, feature-map check, production build

2

container images built, not pushed, on every merge request

No exceptions0

skip flags: a failing check is a fail, not a TODO

Wrong endpoint

Gate 1 green

A migration applies cleanly, and a form still points at the wrong endpoint.

Unreachable button

Gate 1 green

A component passes every type check, and renders a sign-in button nobody can reach with the keyboard.

Ruff does not watch a browser. The second gate does what the first structurally cannot: it looks.

Gate 2: an independent witness, not a script

1 · Scope frontend/app/… backend/… unmapped path sign-in ask-prover run stops path-to-feature map · no skip flag · illustrative paths 2 · Hand-over request diff contracts Fresh reviewer agent did not write the change contract: promised behavior + boundary to distrust 3 · Drive and record Disposable instance own ports, database, build Chromium by ARIA role + accessible name Evidence per feature 1 snapshot + 1 screenshot + read-only second look
The contract names what was promised and where to be suspicious; the investigating is left to the agent. Steps are batched per feature to keep the token cost down, and the instance never touches a developer's running stack.

Why ARIA roles

They are the handles a screen-reader user relies on: a second reason to insist on them.

Why curated evidence

A click-by-click transcript would just be a new haystack.

The reviewer spends model time on judgment. Scripts handle scope, environment, browser mechanics and evidence storage.verification/README.md, Prover Labs

A verdict cannot outlive its patch

A PASS that survives the next commit is worse than no verdict. So the verdict is keyed to a stable git patch-id of the branch diff, not to a commit hash, a timestamp or someone's word.

git diff origin/main..HEAD | git patch-id --stable patch id of the diff 7e67022896… (new id after amend) .verify/7e67022896….md Verdict: PASS ids matchpush accepted no verdict for this idpush rejected Missing verdict → rejected · stale patch id → rejected · exact match → accepted. Cannot work out the patch id for any reason → refuses rather than guesses.
Amend one line after the verdict was written and the id changes under you: the old PASS applies to nothing, and the gate stays red until a fresh review covers the new diff. The id shown is truncated.

Local verdict hook

early warning

Only a developer's own Git can see it. GitLab cannot see an ignored file.

CI gate

protects main

What actually protects main.

The framework had to pass its own bar

✕

UI unit tests dropped

✕

Snapshot tests dropped

✕

Standing end-to-end suite dropped

Expensive in tokens to write and maintain, and for a product this size already covered by a real browser pass. What is left earns its keep every run.

The framework's own diet

Lines added by the merge request, lower is better

−1,695 net lines in the cleanup commit
Before
3,367
After
1,672
Verdict: PASS
Scope: all changed paths mapped; unmatched empty
Gate: missing verdict → rejected · stale patch id → rejected · exact match → accepted
Pipeline 51780: 4/4 jobs green on 6c64b85c

The cleanup touched only verification tooling, so the six-feature browser drive already on record still applied. Reusing a true verdict is not the same failure as reusing a stale one.

Trust is earned per diff, not per agent

Not this Agent writes→ Trust the agent more→ Merge
This Agent writes→ Gate 1 + Gate 2→ Verdict for this diff→ Merge

Cheap enough

to run on every change.

Strict enough

to mean something when it is green.

Expires

the moment the diff it describes does.

Not a bigger model: a narrower gap between "the change looks right" and "the change is right".

Prover Labs engineering note. Describes how the Prover Labs site itself is verified; figures are from the merge request that introduced the framework, September 2026 (AI/prover-lab!31).