Piling on more AI agents doesn't give you more control — it gives you the illusion of it. Agents trained the same way on the same data agree with each other, wrongly, and with more confidence.
- Anthropic research he cites: four models with partial info got it right ~17–36% of the time; one model with full info, nearly always.
- Two models doing the same job isn't segregation of duties. It's duplication.
- Worst pattern: agents chained together, each inheriting the previous one's error.
- Multi-agent works when the work is split, not when it's repeated.
- Build for disagreement: separate evidence, agents that challenge with proof, a human who approves the big stuff. Treat the one dissenting agent as the senior voice.
- The system should be able to stop — thin evidence, contradictions, low confidence, no human around.
Three AI agents made the same mistake.
The addition of agents creates a false sense of security over financial processes while removing any independent scrutiny of the work being performed.
For example:
You have an agent retrieving information from your ERP, an agent investigating all variances and a third agent checking over the management pack before it goes to your CFO. All three come to the same conclusion. What if they are all wrong?
This highlights the concern raised by Anthropic’s research team over the “hallucination” of different AI language models (and how they don’t always agree on the truth).
Anthropic investigated this further by testing how multiple models could work together, and how harmful it would be if they shared the same delusions about the world.
Their research shows that when presented with a series of hiring, investment and property purchase scenarios, groups of four models – given only partial information about the scenario – were able to find the correct answer 17-36% of the time. Whereas a single model, given complete information, could be right nearly 100% of the time.

(Group performance by type. The x-axis represents the percent of groups that selected the target answer; the y-axis represents the percent of single-model groups that selected the target answer, given full information.)
Individual models don’t perform badly, but problems arise with multiple models making the same, incorrect decision before addressing the most pertinent piece of information.
Like humans, these types of AI can be deluded, and can encourage delusion in others. The group of four is still a weakness, as the models agreed on an incorrect answer before tackling the most relevant piece of information.
For finance teams, the concept of independent review is nothing new. If there is someone preparing the reconciliation, there should also be someone reviewing it. Chances are, the person initiating the payment isn’t the same person authorizing it or the one who releases the funds for the accounting team to pay. Similarly, the person authorizing the forecast probably isn’t the only one who should be questioning the assumptions within it.
Likewise, if finance uses multiple models, the teams should be aware that any two models that have been trained on the same information using the same instructions and operating on the same assumptions would not offer true segregation of duties.
Put simply, there is no separation of responsibilities if two models replicate the work of one.
The addition of multiple models offers little protection in these cases and can give a false sense of security, as they are likely to deliver the same incorrect answer with greater confidence.
When can multiple models add value?
In circumstances where work has been segregated – for instance, if different models are reviewing different items on an invoice, categorizing different line items, investigating different possible reasons for variance or checking different elements of a policy, then yes, there is value in multiple models. Each can add speed and efficiency in certain areas.
The danger comes when models are used as a nexus, with one extracting the wrong contract date, another using that incorrect date to calculate revenue, a third writing up the explanation for the variance and a fourth formally approving the number because it looks good.
Consensus doesn’t equal control – especially not in the way finance teams might hope.
Better control might be achieved if there is different evidence for each model, and each is asked to consider different errors and issues before arriving at an answer.
I think it’s important that we look to systems where the agents actively challenge each other and offer evidence to back up their challenges, rather than seeking agreement, so that they can audit each other, not just rely on the work of another agent.
Audits are about evidence – not about people agreeing, believing or trusting each other.
They are about demonstrating that the numbers tie up, that the reconciliation has been completed correctly and that the figures actually reflect the real-world facts on the ground.
Having one agent challenging another is a start – but what if each agent audited the work of all the other agents? It might be a bit of a hurdle to get over at first, but here’s a suggestion for an effective multi-agent system for finance teams in the future: an evidence agent, a worker agent, an audit agent and a policy agent. Plus, a human owner.
Evidence agent
This agent would gather all the information and evidence relevant to a particular task – such as the ledger, the invoice, the contract, the approvals and any policies that apply.
Worker agent
This agent would perform the work needed to create the reconciliation, carry out the calculation, build the forecast or review the transaction, but would explicitly state its assumptions.
Audit agent
A challenger agent would seek to highlight discrepancies and issues with the recommendation by highlighting a lack of evidence, alternative explanations and contradictory facts.
Policy agent
This agent would explicitly check the recommendation against relevant policies, including limits, VAT, supplier rules and other finance policies.
Human owner
This person would review the evidence, challenges, assumptions and final recommendation before approving any high-value action.
I think it’s very healthy to be able to disagree – and to have the system stop while those concerns are addressed.
The best systems will create the gaps for people to step in, rather than trying to reach consensus.
If three agents approve a payment, but one highlighted issues with the supplier’s bank details that need to be resolved before any funds can be transferred, then the minority view in the agents should perhaps be treated as being from someone more senior.
Similarly, the system should be able to stop if there is a lack of evidence, if evidence contradicts other evidence, if confidence levels are low or if humans aren’t available to step in and review.
The questions facing finance teams
AI agents have the potential to free finance teams from a huge number of low-value tasks.
They can be used to review reconciliations, investigate variances, collate evidence for audits, review contracts, check policies and prepare presentations for managers.
The danger is that applying multiple agents only creates illusions of control while eliminating the independent review that finance teams need.
Before embracing such a system, finance leaders should ask these questions: Can the agents back up their evidence? Can one agent challenge another? Can the system identify contradictions? Can it escalate areas of uncertainty? Can a human understand the reasoning before approving actions? Can processing stop before money is transferred or tax filings are made?
The future of finance will almost certainly involve multiple agents working together.
The main question is whether those agents can flag the one piece of evidence that no one else saw.
Would you allow one agent to review the work of another as a form of second-level control? What safeguards would you want to see in place?

.avif)




