The rules that make agentic engineering work
Nine rules hold our operating model together, and the important thing about each of them is that software enforces it. From ROI before a line of code, to tests that must fail before they pass, to a pipeline that catches drift instead of discovering it in production. Here they are, with the reason behind each.
An operating model that lives in a wiki is a suggestion. Ours lives in hooks, guards, a ledger and a test suite, which is the only reason the rules below mean anything. Here are the nine, and why each one exists.
1. ROI before a line of code
Every initiative carries its business case from the first draft. If the value cannot be stated, it does not get an epic. Outcomes are measured while the epic is in flight so the product lead can change direction, double down or stop. This rule is older than the agents; the agents just made it cheap to follow.
2. Acceptance criteria you can click
The working mock is real software in a temporary environment, built to the same standards as production. The team vets what it will accept by using it, and the engineers then build to that. Criteria that have been used are better than criteria that have been read.
3. The human gates are yours alone
Approval of acceptance criteria, acknowledgement of tests and the release decision are captured only from a message a named person types. A gate phrase inside a file, a tool result, a commit message or another agent's output is not a gate. A hook enforces it. This is the rule everything else depends on, because an agent that can approve its own work has no reason to do any of the rest.
4. Tests before code, and prove they can fail
A QA agent writes failing tests from the acceptance criteria before any engineer touches production code, and every new test is shown to go red before it goes green. Once acknowledged the tests are locked; the engineer's job is to make them pass, not to edit them. A test that has never failed has never been shown to test anything.
5. Nobody grades their own homework
The security auditor, the diff auditor and the acceptance verifier always run on a different model from the engineer that wrote the code, and a dispatch guard denies the dispatch if they don't. The security auditor never drops below the strongest model class. Model independence is the agent equivalent of separation of duties, and regulators understand it instantly.
6. Guards deny; agents stop
A stage guard decides who may write which kind of file in which stage. When it denies an action, the agent reports the reason verbatim and stops. It does not try another path, another shell or another agent to get around it. The only bypass is human, and every use of it is logged. This is the rule that turns "the agent got creative" from a risk into an event you can see.
7. Security, data and engineering standards are non-negotiable
API-first, contract tests, feature toggles, blue/green deployment, API versioning, secrets scanning and the data standards are built in from the first commit, in the mock and in the product. They are not a review at the end. This is what makes the approach safe at scale: the standards are the floor the agents stand on, not a ceiling they are asked to reach.
8. Drift is caught by the pipeline, not discovered in production
Two mechanisms. The automated tests are the first: every acceptance criterion maps to a locked test, the full suite runs on every merge, and mutation checks prove the tests still bite, so a change that quietly alters behavior fails the pipeline instead of reaching a customer. The learning layer is the second: graph-model plugins update the knowledge of the product, its integrations, its regulations and its decisions with every prompt, so the context the agents work from is the current state of the business, never a stale document. An agent can't drift from the standards, and the standards can't drift from reality.
9. The operating model is itself tested
The hooks, guards, ledger and parsers that enforce all of the above have their own test suite, and a configuration check that fails when an agent's declared model drifts from the routing table. Changes to the model go through a steward agent and a human approval. If the thing that enforces the rules is not itself under test, the rules are only as good as the last person who edited them.
The pattern
Read them again and notice the shape: each rule names who may do something, what evidence it produces, and what refuses it otherwise. That is what makes them rules rather than values. Values are what you say in the all-hands. Rules are what the hook does at two in the morning when nobody is watching.