Twenty-two agents, and what each one is allowed to touch
The agents in our operating model are not a swarm. Each is a short file that says what it does, what it may write, which verdicts it owns and what it loads before it starts. Here is the roster by group, why the domain experts never write code, and how the agents get better without being retrained.
When people hear "twenty-two agents" they picture a swarm: a cloud of bots passing work around until something comes out. Ours is the opposite. Each agent is a short, readable file with four things in it: its job, the kinds of files it may write, the verdicts it owns, and what it loads before it starts. The pipeline dispatches them by stage, and the guards refuse any agent that steps outside its file.
Here is the roster, by group.
Product
Turns ideas into initiatives with the business case attached, cuts epics with measurable outcomes, builds the working mock, and writes the acceptance criteria. The compliance expert is a mandatory co-author on every criterion and signs the exact version. The product agents produce the most important artifacts in the lifecycle and not one line of production code.
Domain experts
Payments, card networks, compliance, risk, settlement, platform. Before they start they preload the rulebooks of the business they cover: card-network operating rules, PCI, sanctions screening, interchange, the company's own engineering standards. They review acceptance criteria and audits. They never write code. An expert that can also implement is an expert with a conflict of interest, so we don't allow it.
QA
Two agents. A red agent writes failing tests from the acceptance criteria, one per criterion, and captures the red run as evidence. A verifier runs the mutation checks, owns the coverage verdict, and checks each acceptance criterion over HTTP against the running system. The verifier always runs on a different model from the engineers.
Engineers
Backend, frontend, data, integrations. These are the only agents allowed to write production code, in one stage of the pipeline, in the story's own worktree, against tests they cannot edit. Narrow on purpose. An engineer agent with a wide remit is an engineer agent that will eventually mark its own work done.
Auditors
A security auditor and a diff auditor. Always a different model from the engineer that wrote the code; the dispatch guard enforces it. The security auditor is pinned to the strongest model class and never falls below it, regardless of what the weekly model scoring says about cost.
Operations
A GitOps agent that alone may push, merge and deploy, and a records agent that keeps the tracker, the release record and the memory honest. If only one agent can deploy, then "how did that get to production" always has a one-word answer.
How they get better
None of these agents is retrained. Each has its own memory of lessons from past stories. At the end of a story the records agent publishes each agent's lessons as a pull request, scanned for secrets, and a human merges it. The next time that agent runs, it loads what it learned. Over a few months the engineer agents stop making the mistakes specific to this codebase, the domain experts stop missing the rule specific to this sponsor bank, and the whole roster behaves less like a model and more like a team that has been here a while.
Underneath all of them sits a graph knowledge base that learns from every prompt: the product, its integrations, the regulations, the decisions. That is what keeps the agents' context current rather than frozen at the date somebody last wrote a document.
Why the files are short
The agent files are short because they have to be reviewable. When the steward agent proposes a change to one of them, a person reads the diff and approves it. A two-thousand-line prompt cannot be reviewed; a sixty-line role definition can. The discipline of the model is partly a discipline of keeping each agent small enough that a human can still say what it is for.