Operating Model
← All insights

A Fresh Pair of Agentic Eyes

If your organisation adopted AI in the last eighteen months, the chances are that everything in it runs on one model and nothing is positioned to tell you when it is wrong. What a fresh pair of eyes means for agents, and three questions to ask this week.

Operating Model

CEO, COO, CTO, CIO, PE Partner

Part of our guide: AI implementation and operating model

If your organisation adopted AI in the last eighteen months, the chances are that every assistant, agent and pipeline in it runs on one model, and that nothing in the arrangement is positioned to tell you when it is wrong. That is not a criticism of the model. It is the same problem any leader would see at once if it were people rather than software.

Every leader knows what a fresh pair of eyes does. The person who did not write the paper sees the gap in the argument, the figure that does not add up and the assumption the author stopped noticing three drafts ago. It is why review sits apart from authorship in any well-run function, why boards carry non-executives and why the best teams our founder has worked in were built from people who did not think alike. A team that agrees quickly tells its leader little they did not already know, and difference is what makes review worth doing.

The same is true of AI agents, and it is the part of the agentic conversation that is going missing. Most of the current effort is going into giving one model more tools, more autonomy and more agents to divide the work between. Almost none of it is going into the question a fresh pair of eyes depends on: whether the eyes are actually different.

Where this shows up in a business

A language model has one set of priors. It was trained one way, it favours certain constructions and certain conclusions, and it carries those consistently through everything it produces. Ask it to review its own output and it brings the same priors to the review that it brought to the writing. A second look is not a second view, and the same pair looking twice is not a fresh pair of eyes.

In practice that reaches a leadership team in three places. The first is the board paper or business case where a claim has been stated more confidently than the evidence supports, and nobody in the chain was positioned to notice because the drafting and the checking shared the same assumptions. The second is the customer-facing output, a letter, a decision or an explanation, where the same blind spot is repeated at scale rather than caught once. The third is the pipeline of agents that reviews its own work and reports that all is well.

In a regulated firm each of those lands on a named individual, and under Consumer Duty and the Senior Managers and Certification Regime the question of who was positioned to catch the error is not a rhetorical one.

What we have seen

Our founder runs his own work with two models from different providers, and the arrangement is wired into the workflow rather than performed by hand. Nothing is copied between two chat windows. One model leads the research, drafting and orchestration, and the second is invoked automatically when a piece of work is ready, arrives cold and is briefed to argue against the diagnosis rather than help with it. What comes back most weeks falls into three kinds: defects that would have shipped, claims stated more confidently than the evidence supports and a theory the author has grown attached to being taken apart with figures that were already to hand.

This is not a novel idea, and that is rather the point of it. It is what every engineering function already does with independent quality assurance, where the person who wrote the code does not sign off the testing. That control exists because authorship and review need different eyes rather than because anybody distrusts the author, and the reasoning does not stop applying when the author is a model. What is new is only that the reviewer has to be chosen deliberately, because a second instance of the same model looks like a control and is not one.

The finding that matters is the comparison. Given the same brief, the same cold context and the same instruction to argue, the lead model reviewing its own work is noticeably weaker than a different model reviewing it. It agrees more readily, it misses the assumptions it shares with the author and it polishes rather than challenges. That is experience in regular use rather than a controlled trial, and we would not present it as more, but it is exactly what thirty years of running mixed teams in regulated environments would predict.

Multi-agent is not multimodel

The two are easy to conflate and they are different axes. Multi-agent is division of labour: a researcher, a planner, a builder and a reviewer, each with a role and a hand-off. Multimodel is diversity of judgement: how many genuinely different sets of priors are in the room. The current trend runs strongly towards the first without the second, and a pipeline of five agents on one model is five boxes with one point of view.

This is not an argument against multi-agent design. The roles stay, and the division of labour is real and valuable. The argument is about who fills the roles, and putting a different model behind the review role is the least expensive change available to most organisations, because it changes nothing about the work and everything about what gets caught.

What good looks like

The arrangement that works is an operating model rather than a purchase, and it has four parts. Different models fill different roles, so that at least one set of judgement in the chain does not share the author's priors. At least one of them is briefed to argue rather than assist, and is given the whole context to do it with. Disagreements between them have a route to a named person who weighs them up and shapes the outcome. And that person is accountable for the output in the same way a leader is accountable for a team's output without writing every line of it.

That last point matters more than the technology. The work moves to the agents and the accountability does not, and an arrangement where nobody can say who owns the output is not ready to be relied on, however capable the models underneath it.

Three questions to ask this week

The state of an organisation's agents can usually be read from three answers.

How many models sit underneath the tooling your people use? If the answer is one, however many agents or assistants share it, the organisation has one set of priors at scale and no fresh pair of eyes anywhere in the chain.

Is anything in the arrangement briefed to disagree? A reviewer asked to help will help. If no agent or assistant has been told to find what is wrong, the review step is a second look rather than a second view.

Who is accountable for the output? If the answer is a named person who did not do the work but owns the result, the foundations are in place. If nobody can say, the arrangement is running on trust.

Where to start

Most organisations can answer those three questions in an afternoon, and the answers usually decide the next step. Where the picture is unclear, Breathe is a short discovery engagement that establishes where the organisation actually is and what the highest-priority opportunity looks like.

Where the need is to build the arrangement described above around a real use case, with a second model in the review role and governance and measurement designed in from the start, that is what Flow delivers in four to eight weeks. And where the missing piece is the accountable senior owner, someone who sits with the leadership team and answers for the output, Grow provides that on a fractional basis without a full-time hire.

We have argued elsewhere that the operating model matters more than the technology and that agents break the assumptions agile was built on. A fresh pair of agentic eyes is the same argument from a different angle. The value of a mixed team was never that it had more people. It was that they saw differently, and one of them was accountable for what they saw together. Diversity enriches, and that holds whether the team is people or models.

Work with us

Flow: Build something real

We take the highest-priority opportunity and build a working AI capability, with governance, measurement and an operating model designed from day one: a real, deployable tool (not a proof of concept).