The tools will change. The process only hardens.
Claude Code runs the pipeline, Cursor and Codex sit in the editor, and a weekly multi-model vetting and scoring process decides which model each agent uses for each job regardless of vendor. None of that is the model. The operating model is the thing that gets stronger every time the tools underneath it change.
People ask which tools we use before they ask how the model works, so here is the honest answer, followed by why it is the less important one.
What runs today
Claude Code runs the pipeline itself, and Claude models do most of the reasoning. Cursor and Codex sit in the editor for the humans. Every agent file declares which model class it runs on, and a configuration check fails the build if an agent's declared model drifts from the routing table.
Which model each agent gets is not a matter of taste. A multi-model vetting and scoring process runs each week: the same representative stories, the same acceptance criteria, run across the candidate models for each job to be done, scored on correctness, cost, latency and how often the auditors reject the output. The routing table is updated from the scores. If a new model is better at writing failing tests and worse at diff audits, the QA agent moves and the auditor stays. The security auditor is the one exception; it is pinned to the strongest class regardless of score, because that is not a job you optimize for cost.
The result is that the agents use the right model for each job to be done, regardless of which AI vendor made it, and the decision is re-made weekly with evidence rather than annually with a procurement deck.
The layer underneath
All of it runs over a graph knowledge base that learns from every prompt. The product, its integrations, the regulations it operates under, the decisions that were made and why: every interaction updates it, through plugins that write to the graph rather than to a document somebody has to remember to maintain. This is the second half of how we prevent drift. The tests keep the code honest; the graph keeps the context honest. An agent that starts a story loads the current state of the business, not the state as of the last time someone updated a wiki.
Why the tools are the least important part
Every tool named above will be replaced. The model we route most reasoning to today is not the one we used a year ago. The editor is different. The thing that has not changed, and that gets stronger each time the tools change, is the operating model: who may do what, in which stage, with what evidence, and who decides.
That is the inversion most AI programs get backwards. They adopt a tool, build process around it, and then the tool changes and the process goes with it. We built the process first, as hooks and guards and a ledger with their own tests, and we treat the tools as replaceable parts that plug into it. When a better model arrives, the weekly scoring moves the right agents onto it and nothing else changes. When a worse update ships, the scoring catches it and the agents move back.
As AI and SI evolve
Agentic AI is here. Superintelligence is coming behind it. The gap between them is where the operating model earns its keep, because the systems will keep getting more capable and the question "who approved this, what proves it, and who turned it on" will not change. A process that depends on the current limits of the model is a process with an expiry date. A process that is enforced by software and owned by named people only hardens as the models get better, because every improvement in the tools is an improvement in what runs inside the gates, never a change to the gates themselves.
The tools will change. The process only hardens and gets better as AI and SI evolve.