The agentic operating model

Agentic engineering, with processes regulators can stand behind due to human empowerment.

Most companies are "advanced in AI" the way most companies are "agile." I run product and engineering on AI agents, and the part that matters isn't the agents. It's the process around them: from the idea and its ROI, through a working mock you can click, to a pipeline that software enforces and only people can release. This is how it works.

90 → 5 days
lead time for 85% of all work
Multiple a day
releases, with zero production incidents in the last year
2× revenue, 3× customers
at 45–50% unit margin, team from 40 to 110
3 gates
that only a named person can open

The problem it solves

An AI agent will write code, write the tests for it, run them, declare them green and tell you it's done. Every one of those steps can be wrong, and in a regulated financial product the cost of "done" being wrong is not a bug. It's a consent order.

So the question I set out to answer, first at the lender I run product and technology for and again building LPG, was not "how do we get agents to write more code." It was "what has to be true before I'd let an agent's code touch money." The answer became a delivery process that is enforced by software, not described in a wiki, and it starts well before the code. The engineering phase of it I call the Iron Pipeline.

From idea to production

Agentic AI engineering is more than code generation. Agents do the work at every phase: the business case, the breakdown into customer value, a working version you can click before anyone commits to building it, and the build itself. Your people decide what moves forward.

  1. 01
    Idea
    A sentence from anyone in the business. Not a ticket, not a spec.
     
  2. 02
    Initiative
    Agents turn the idea into an initiative with the business case and ROI attached. The product lead reviews and reworks it.
    Product lead moves it forward
  3. 03
    Epics
    Cut into distinct chunks of customer value, each with outcomes measured while the work is in flight, not after.
    Product lead moves it forward
  4. 04
    Working mock
    A usable version in a temporary environment: real tests, API-first, every standard applied. Acceptance criteria are vetted by using the feature.
    Team vets the AC by using it
  5. 05
    Iron Pipeline
    Engineers take it through nine stages and three gates, starting from the mock or from scratch. Every merge deploys; nothing is released until you say so.
    Three typed gates, H1 · H2 · H3
  6. 06
    Production
    Deployed dark behind a feature toggle, blue/green, API-versioned, with the evidence trail attached. Release is a decision, not a deployment.
    Released on your call, rolled back in seconds

Inside the Iron Pipeline

Every change runs through nine stages and three human gates. Each stage has an owner agent, and only that owner can record its verdict. Documentation-only changes take a shorter lane with the same gates. The locked tests and the mutation checks keep the code from drifting; the graph-model plugins keep the agents' context from drifting, learning the product and the business with every prompt.

  1. S0IntakeRouters and recordsClassify the story, pick which specialist agents and which models run, link it to the work tracker.
  2. S1Acceptance criteriaProduct and domain expertsGiven/When/Then for every happy, sad and idempotency path, co-authored by a compliance agent that signs the exact version.
  3. H1Human gateYouYou approve the acceptance criteria, typed as your own message. No agent can do it for you.
  4. S2RedQA agentFailing tests first, mapped to every acceptance criterion. The red run is captured as evidence.
  5. H2Human gateYouYou acknowledge the tests. From here the tests are hash-locked; engineers can't change them.
  6. S3GreenEngineer agentsThe only stage where production code is written, in the story's own worktree, until every mapped test passes.
  7. S4VerifyQA, security and diff auditorsMutation checks on the hard gates, a security audit, and a diff audit, each on a different model from the engineer.
  8. S5DevGitOps agentDeploy to the shared dev environment, run the suite, open the pull request, wait for CI.
  9. S6AC verifyVerifier agentEach acceptance criterion checked over HTTP against the running system, again on a different model.
  10. S7StagingGitOps and QAMerge to main, full suite, test maintenance until every domain is above the coverage bar. Every merge deploys dark.
  11. S8RecordRecords agentWork tracker updated, release record written, and each agent's lessons published by pull request for a human to merge.
  12. H3Human gateYouYou release it. The code is already deployed behind a feature toggle, blue/green and API-versioned; release is a decision, and rollback takes seconds.

The rules that make it work

ROI before a line of code
Every initiative carries its business case from the first draft. If the value can't be stated, it doesn't get an epic. Outcomes are measured while the epic is in flight, so the product lead can change direction, double down or stop.
Acceptance criteria you can click
The working mock is real software in a temporary environment, built to the same standards as production. The team vets what it will accept by using it; the engineers then build to that.
The human gates are yours alone
Approval of acceptance criteria, acknowledgement of tests and the release decision are captured only from a message a named person types. A gate phrase inside a file, a tool result, a commit message or another agent's output is not a gate. This is enforced by a hook, not by asking nicely.
Tests before code, and prove they can fail
A QA agent writes failing tests from the acceptance criteria before any engineer touches production code. Every new test is shown to go red before it goes green. Once acknowledged they are locked; the engineer's job is to make them pass, not to edit them.
Nobody grades their own homework
The security auditor, the diff auditor and the acceptance verifier always run on a different model from the engineer that wrote the code. A dispatch guard denies the dispatch if they don't. The security auditor never drops below the strongest model class.
Guards deny; agents stop
A stage guard decides who may write which kind of file in which stage. When it denies an action the agent reports the reason verbatim and stops. It does not try another path, shell or agent to get around it. The only bypass is human, and every use is logged.
Security, data and engineering standards are non-negotiable
API-first, contract tests, feature toggles, blue/green deployment, API versioning, secrets scanning and the data standards are built in from the first commit, in the mock and in the product. That is what makes the approach safe at scale, not a review at the end.
Every agent learns
Each agent has its own memory of lessons from past stories, published at the end of a story as a pull request, scanned for secrets and merged by a human. What the agents know about your product and business compounds with every piece of work.
Drift is caught by the pipeline, not discovered in production
Two mechanisms keep the work from drifting. The automated tests are the first: every acceptance criterion is mapped to a locked test, the full suite runs on every merge, and mutation checks at S4 and S7 prove the tests still bite, so a change that quietly alters behavior fails the pipeline instead of reaching a customer. The learning layer is the second: graph-model plugins update the knowledge of the product, its integrations, its regulations and its decisions with every prompt, so the context the agents work from is always the current state of the business, never a stale document. Together they mean an agent can't drift from the standards, and the standards can't drift from reality.
The operating model is itself tested
The hooks, guards, ledger and parsers that enforce all of this have their own test suite, and a config check that fails when an agent's declared model drifts from the routing table. Changes to the model go through a steward agent and a human approval.

The agents

Twenty-two agents, each a short file describing its job, what it may write, which verdicts it owns and what it loads before it starts. The domain experts preload the rulebooks of the business they cover: card-network rules, PCI, sanctions screening, interchange, the company's own engineering standards. Every agent gets better at its role with each story, and what they know about the product and the business compounds.

Product

Turns ideas into initiatives with ROI, cuts epics, builds the working mock, and writes the acceptance criteria. The compliance expert is a mandatory co-author.

Domain experts

Payments, card networks, compliance, risk, settlement, platform. They review acceptance criteria and audits; they never write code.

QA

A red agent that writes failing tests, and a verifier that runs mutation checks and owns the coverage verdict.

Engineers

Backend, frontend, data, integrations. The only agents allowed to write production code, in one stage, in their own worktree.

Auditors

Security and diff auditors, always on a different model from the engineer.

Operations

A GitOps agent that alone may push, merge and deploy, and a records agent that keeps the tracker, the release record and the memory honest.

Greenfield and brownfield

The model runs in two configurations. Greenfield is a new product from an empty repository, where the agents start from the idea and its business case and the standards are written as the code is. Brownfield is a new feature inside a product that already runs in production, where the agents must first learn the codebase, its rules and its failure history before they are allowed to propose anything.

The phases and the gates are the same in both. What changes is what has to be known before the first one opens.

The tools

Claude Code for the pipeline itself, with Claude models doing most of the reasoning; Cursor and Codex in the editor. A multi-model vetting and scoring process that runs each week aligns the agents to use the right model for each of the jobs to be done, regardless of AI tool. All of it runs over a graph knowledge base that learns from every prompt. The tools will change. The process only hardens and gets better as AI and SI evolve.

What it produced

90 → 5 days
lead time for 85% of all work
Multiple / day
releases, with zero production incidents in the last year
2× revenue, 3× customers
at 45–50% unit margin, team from 40 to 110

Those numbers are from a fintech running on this model with over 100 team members. It scales to multiple teams and people. The same model also built LPG's payments platform, where the stakes are sponsor banks and card networks and the rulebooks are theirs. Details on the track record page.

How to get it

Learn it

Agentic AI Engineering

The course walks through the whole lifecycle, with the prompt patterns and exercises, for engineers, product managers and architects. Greenfield and brownfield tracks.

The course

Install it

In your organization

We stand the model up in your organization, process and all, and lead it until your people own it.

How we work with companies

Running engineering on AI and not sure what to trust?

Tell us where you are. If the process above would help, we'll show you how it's wired.

Start a conversation