Proof of Done

How it works

A merge request merging doesn’t mean the feature works for anyone. Proof of Done closes that gap: it checks every shipped issue in production, files the evidence where the team already works, fixes what it can, and keeps checking.

The loop

One human step starts it. Agents read and write; deterministic jobs decide and merge. Every arrow below leaves a record you can open in GitLab or on this site.

The Proof of Done loopA person merges; the pipeline deploys; the examiner agent plans checks; the runner photographs production; the clerk closes or reopens the issue; on failure the fixer opens a fix, the guard merges it within policy, and the re-check job replays the plan and closes the issue. Every night the light table replays all proven checks.passedplanframesfailedmergedheld: the issue closes with the new framesreplays through the runner; a regression reopens its issueA person mergesthe only human stepPipelinetests, scans, deploysExaminer agentcriteria into stepsRunnerphotographs productionClerk agentcloses or reopensFixer agenta fix merge requestGuard jobmerges only inside policyRe-check jobreplays the same planLight table, nightlyreplays every proven check
a persona GitLab Duo agenta deterministic job, no model
  1. A person merges a change that closes issues.
  2. The pipeline tests, scans, packages, releases and deploys it.
  3. When the pipeline passes, the GitLab Duo flow starts. The examiner agent reads each shipped issue’s acceptance criteria and writes steps.
  4. The runner walks those steps in production and photographs each one.
  5. The clerk agent closes each issue that held, with the frames, and reopens each one that didn’t.
  6. For each failure the fixer agent opens a small fix merge request.
  7. The guard job merges the fix only inside policy; the deploy that follows replays the same plan and closes the issue if it holds.
  8. Every night the light table replays every proven check and reopens anything that regressed, with each frame laid over the last one that held.
  9. If the regression came from the last deploy, and only from it, production goes back to the commit the issue last held on, within a written policy (see below). The issue stays open for a person to fix forward.

From a sentence to a frame

The issue says: “Add to calendar downloads an event that starts at the booked time, Pune time.” The examiner first asks the runner for the page’s outline (its roles, names and labels) so every step targets something real. Then it writes steps like these.

{
  "do": "remember",  "target": { "css": ".ticket .time" },
  "as": "time",      "note": "Note the booked time"
},
{ "do": "download", "target": { "role": "link", "name": "Add to calendar" }, "as": "ics" },
{ "do": "calendar_event", "file": "ics", "starts": "[[time]]", "time_zone": "Asia/Kolkata" }

The runner books a slot, notes the time on the confirmation, downloads the calendar file and reads its start time in the library’s time zone. Here it found 23:30 where the visitor had booked 18:00: the server ran in UTC while the tests ran in Pune time.

Every step is a frame on the contact sheet. A downloaded file is photographed too, with the comparison drawn large so it reads at thumbnail size.

Before any step, the runner checks that production reports the commit it was asked to prove. If the deploy hasn’t rolled out, nothing is checked and the roll says so.

Three outcomes, never two

A check that can’t be made isn’t a pass.

Held
Every step happened as written. The issue closes with the strip of frames and a link to the sheet; its plan joins the light table.
Failed
A step didn’t. The issue reopens with the failing frame and the runner’s sentence, and a fix merge request follows.
Needs a person
A browser can’t observe it (an email arriving, a phone alert). The reason is written on the issue; nothing is passed.

What runs where

Everything is in one repository, MIT licensed.

The flowThree GitLab Duo agents (examiner, clerk, fixer) defined in one flow, triggered by a passing pipeline.flows/proof-of-done.yml
The runnerNode and Playwright Chromium in an image the flow jobs run in. Validates plans, runs them, photographs, publishes.runner/
The pipelineEight CI stages, covering all nine lifecycle stages; deploys to Cloud Run with Google credentials from GitLab’s OIDC token, no keys stored..gitlab-ci.yml
The guardThe written policy for fixes that merge themselves, with its tests.board/lib/policy.ts
This siteNext.js on Cloud Run; frames in Cloud Storage; each roll hash-chained to the one before.board/
The app it checksCarrel, a study-room booking app for a public library in Pune, deployed to Cloud Run like any team’s product.demo/carrel/

When production rolls itself back

A regression found at 01:30 has two remedies: a fix merge request, which needs a person’s merge, or moving traffic back to the last revision that held, which doesn’t. The light table does the second only when all of this is true, and writes the table on the issue:

The regressed issues last held on one commitso there is one place to go back to
Production has changed sincean issue that fails on the very commit it held on is not a rollback: something else moved
The deploy is at most 7 days oldundoing an old deploy undoes too much else
No other deploy with proven features sits in betweenonly the last deploy is undone, never somebody else’s work

Outside any of these the issue says why a person is needed. The job that moves traffic holds no key: it authenticates to Google Cloud with GitLab’s OIDC token, and the board writes what happened on each issue, with the roll’s hash, and beside the roll’s record. The policy is code: board/lib/rollback.ts.

What changed, pixel by pixel

A replay photographs the same steps as the roll that last held. The board decodes both frames and marks every changed pixel in red, and the sheet says how much of the frame moved. It is a measurement beside the verdict, not a verdict: a date on a grid changes every day and that is fine; a page that is 40% different the night a check fails tells a person where to look first.

What it can't do yet

Replays of every proven check run nightly and after a fix, not yet after every ordinary deploy; a regression in a feature nobody touched is caught at 01:30, not at once. That is the next job to add, and it needs every saved check to be free of time-of-day words first.

Stated plainly, because a proof tool that overclaims is worse than none.

Questions

Why not just run end-to-end tests in CI?

CI tests run against a test environment with test configuration. The failures that reach users are the ones that only exist in production: a server in another time zone, a feature flag nobody turned on, a missing secret, a CDN rule. Proof of Done checks the thing users get, after it ships.

Can the model mark something as working?

No. Agents read the issue and write a plan in a fixed step vocabulary. The runner executes it with Playwright, and only its assertions decide. A plan with any other step is rejected before a browser starts.

What if a criterion can't be checked in a browser?

It's marked for a person, with the reason, and the issue gets a needs-person label. It is never passed silently.

What stops a bad fix from merging itself?

The guard job. A fix merges on its own only if it changes at most 3 files and 40 lines, leaves the pipeline, flow, runner and checks untouched, has no high or critical scan findings, and its pipeline is green. Otherwise it waits for a person, with the reason written on the merge request.

Who can publish a roll?

Only this project's own jobs, on the refs they belong on. The board asks GitLab who is calling: a job token must belong to a running job of this project, and the job must be a flow job, one of the named replay jobs on the default branch, or the guard in a merge request pipeline; a personal token is accepted only from the flow's own service account. No shared secret exists.

Does it cost a model call every night?

No. Replays, after a fix and every night, run the saved plan with no model involved. The agents run once per deploy.

Open the proof sheets or read the code.