The problem agentic engineering actually solves
An agent will write the code, write the tests, run them, call it green and tell you it's done. Every step can be wrong. In a regulated business that isn't a bug, it's a consent order. The problem to solve isn't getting agents to do more. It's deciding what has to be true before their work touches money.
Ask a coding agent to build a feature and it will do something remarkable. It will write the code, write the tests for the code, run the tests, report that they passed, and mark the work complete. The whole loop takes minutes. Then you look closer and discover that the tests assert nothing, the "passing" run used a mocked database, and the feature was marked done by the same process that wrote it.
None of that is a failure of the model. It is doing what it was asked. The failure is that nobody decided, in advance and in software, what has to be true before an agent's output counts.
Where it goes wrong in financial services
In most software companies that mistake costs a rollback. In a lender, a bank or a payments platform it costs more. A payout that goes to the wrong account, a disclosure that is missing a line, an underwriting rule that drifted by one decimal: these are not bug reports. They are regulator letters, sponsor-bank reviews, and sometimes consent orders. The institutions we work with cannot afford "done" to be wrong, and they know it. That is why so many of them have AI pilots and so few have AI in production.
The usual response is to add human review at the end. Every agent-written change gets a senior engineer reading the diff. That works for a week. Then the agents produce more than the seniors can read, the review becomes a skim, and you are back to trusting the agent, only now with a signature on it.
The question that reframes it
The question we set out to answer, first running product and engineering at a consumer lender and again building LPG's payments platform, was not "how do we get the agents to write more." It was: what has to be true before I would let an agent's code touch money?
That question has concrete answers, and none of them are "a human reads everything."
- The acceptance criteria have to exist before the code, and a named person has to have approved them.
- The tests that prove the criteria have to exist before the code too, and they have to be shown to fail first, so we know they test something.
- Once those tests are accepted they have to be locked, so the thing being built can't quietly redefine "done."
- Whoever audits the work for security and scope has to be a different mind from whoever wrote it. For agents, that means a different model.
- Every one of those facts has to be enforced by software, not by a wiki page that says "please."
- A person, not an agent, has to make the release decision, and the deploy underneath it has to be safe to reverse in seconds.
Write those down and you have an operating model. Enforce them with hooks and guards and you have the Iron Pipeline, which is what we call the engineering phase of it.
Why it is an operating model, not a tool
Every item on that list is a decision about who may do what, when, with what evidence. That is organizational design. The tools change every quarter; the model we use today is not the one we used a year ago, and the editor is different too. What has not changed is the set of things that must be true before the work counts.
That is the problem agentic engineering solves when it is done well. Not "how much code can an agent write," which is a solved problem and a boring one. Rather: how does a regulated company let agents do real work, at speed, while the accountability stays exactly where the regulator expects to find it, with a person whose name is on the decision.
The rest of this series walks through how we do it: the lifecycle from idea to production, the stages inside the pipeline, the rules that hold it together, the agents, greenfield versus brownfield, the tools, and the numbers it produced.