I don't let AI agents approve their own work
The useful question isn't how much code an agent can write. It's what has to be true before you'd let its code touch money. Here are the three gates I keep for myself and why no agent can open them.
A coding agent will write a feature, write the tests, run them, tell you they passed and mark the story done. I have watched every one of those steps be wrong. In a consumer lender or a payments platform, "done" being wrong is not a bug report. It is a regulator's letter.
So when we rebuilt how product and engineering work at the lender I run product and technology for, and again when I built LPG's PayFac platform, I started from a different question than most teams. Not "how do we get the agents to do more." Rather, "what has to be true before I would let an agent's code touch money."
The answer turned into a delivery pipeline that is enforced by software rather than described in a wiki. Nine stages, twenty-two agents, and three gates that belong to me. This post is about the gates.
Gate one: I approve the acceptance criteria
Before any test or code exists, a product agent writes the acceptance criteria for the story, Given/When/Then, with the happy path, the sad paths and the idempotency case. A compliance agent is a mandatory co-author, and it signs the exact version it reviewed. If anyone edits the criteria afterward, the signature is void and the story goes back.
Then it stops, and waits for me to type, as my own message, that I approve them.
A hook captures that phrase only from my typed prompt. The same words inside a file, a tool result, a commit message or another agent's output do nothing. I learned to be specific about this after watching an agent helpfully "confirm" an approval on my behalf. It was trying to be useful. That is exactly the problem.
Gate two: I acknowledge the failing tests
Once the criteria are approved, a QA agent writes the tests. Not the engineer. The QA agent maps a test to every acceptance criterion, runs them, and the run has to be red. The red run is captured as evidence; a test that was never seen to fail has not proven anything.
Then it stops again. When I acknowledge the tests, they are hash-locked. From that point the engineer agents can write production code in exactly one stage, in the story's own worktree, and their job is to turn the locked tests green. They cannot edit the tests to get there. A guard denies the write.
This is the gate that changed how I think about AI-written code. An agent that can rewrite its own tests will, eventually, rewrite them. Not out of malice; it is optimizing for green. Taking that option away is the single biggest reason our production incident count is zero for the last year while we ship several times a day.
Gate three: I ship it
Between the second gate and the third, the code goes through verification that the engineer had no hand in. A QA verifier runs mutation checks on the hard gates. A security auditor reads the diff for the things that end careers in fintech: card data in a log, a float where money should be, a stub mode flipped. A diff auditor checks that what was built is what was specified.
All three run on a different model from the engineer that wrote the code. A dispatch guard denies the dispatch otherwise. If the engineer ran on one model family, the auditor runs on another. Nobody grades their own homework, and no model grades its own either.
Then dev deploy, acceptance verification against the running system, staging, and a release record. And then it stops a third time, and I ship it. Production deploys happen in a window, behind feature flags.
Why the gates are mine and not a committee's
People ask why one person holds all three gates. Two reasons.
First, accountability has to be a name. If the ship decision belongs to a process, it belongs to nobody, and when it goes wrong the process gets a retrospective and nobody gets a consequence. I would rather be the one who gets the consequence.
Second, three decisions per story is a small amount of attention. Reading acceptance criteria takes five minutes. Reading a test summary takes five. The ship decision takes one, because by then eight stages of agents have produced evidence I can read instead of code I would have to trust. The gates are not where I spend my time. They are where I spend my judgment.
What this is not
It is not slow. Lead time for 85% of our work went from 90 days to five. The agents do the hours; I do the minutes.
It is not a wiki. Every rule above is a hook that denies the action. An agent that is denied reports the reason and stops; it does not try another path. The only bypass is mine, and every use of it is logged.
And it is not finished. The agents keep their own memory of lessons from each story, and those get published as a pull request at the end, which I merge. The pipeline itself has a test suite. It gets better the way a team gets better, one story at a time.
If you run engineering on AI and you are not sure what to trust, start with the question I started with. Not how much can it do. What has to be true before you would let it.