- LLM disagreement affects 67% of 1,000 real-world fact-check claims tested across five frontier models.
- LLM disagreement is not just minor quibbling: on 34% of claims, models land on opposite ends of the truth scale.
- Claude Opus 4.7 and Gemini 3 Pro agree only 53% of the time, the lowest pair alignment in the study.
- Even unanimous model verdicts can be wrong. Shared blind spots mean consensus is not the same as correctness.
Table of Contents
LLM Disagreement Is Far Worse Than the Industry Admits
AI companies tend to present their latest models as increasingly capable general-purpose reasoners: systems that can search, summarise, explain and judge information with an air of confidence. What gets far less airtime is how often those systems reach materially different answers when asked the same factual question.
Research from lenz.io puts that problem in unusually direct terms. Across 1,000 real-world fact-check claims submitted by users to an active fact-checking platform, five of the most capable frontier models available today failed to reach a consensus on 67% of them. None of the claims was older than February 2026. These were not contrived benchmark prompts designed to catch a model out; they were claims people had actually brought to a fact-checking service.
That distinction matters. A model can look persuasive on a polished demo, or score well on a fixed evaluation set, while still behaving unpredictably when it encounters the messy phrasing, incomplete context and contested assertions of ordinary public discourse. Fact-checking is especially unforgiving because a plausible answer is not enough. The task asks a system to decide what can be supported, what is distorted and what is simply false.
The five models tested were Claude Opus 4.7, Gemini 3 Pro, Gemini 3 Pro with Search, Sonar Pro, and one additional frontier system. Each was asked to classify claims using a four-bucket rubric: True, Mostly True, Misleading, or False. The researchers then measured how often the panel converged and how often it fell apart.
On 672 out of 1,000 claims, at least one model broke from the majority verdict, or no majority formed at all. That is a more useful measure than the vague observation that models sometimes “hallucinate.” It describes disagreement among systems that are each meant to be highly capable. If several leading models cannot reliably place the same claim in the same broad truth category, then the user is not merely choosing between different writing styles or interface preferences. They may be choosing between incompatible judgments.
The panel’s Krippendorff’s alpha came in at 0.639. The metric is used to assess inter-rater reliability: in plain English, it asks whether different evaluators tend to make the same call when looking at the same material. Here, the evaluators happen to be language models rather than human reviewers. A score of 0.639 says the results are not random noise, but it is a long way from the kind of consistency anyone should want from a system entrusted with high-stakes factual judgment.
That should give pause to organisations considering an LLM as an automated fact-checker for a news organisation or a social platform. A single-model workflow can conceal this uncertainty. It returns one polished answer, usually with a crisp conclusion, and users may never realise that another frontier model would have rated the same statement differently. The clean interface creates an impression of settled knowledge even where the underlying model landscape is divided.
Disagreement is not limited to borderline calls
Some variation is inevitable. The line between Mostly True and Misleading can involve judgment, especially when a claim contains a correct statistic but omits the context needed to interpret it. A serious fact-checker can reasonably explain why a claim falls near that boundary.
But the lenz.io findings go beyond that expected ambiguity. On 34% of claims, models landed on opposite ends of the truth scale. This is not a dispute over a label’s shade or an argument about how much caveat language is appropriate. It means one system can see a claim as effectively true while another sees it as false. For users, that is the difference between sharing an assertion and rejecting it. For publishers and platforms, it can affect moderation, corrections and the visibility given to a contested post.
The weakest pair alignment in the study came from Claude Opus 4.7 and Gemini 3 Pro, which agreed only 53% of the time. That result cuts against the comforting assumption that the leading systems are all converging on the same underlying understanding of the world. They may often produce similarly fluent prose. They do not necessarily arrive at the same factual verdict.
Search does not make the broader issue disappear, either. Gemini 3 Pro with Search was included alongside Gemini 3 Pro, but the study’s central result remains panel-level disagreement across the five systems. Access to outside material can help a model retrieve evidence, yet retrieval is only part of fact-checking. A system still has to decide which sources matter, whether a source addresses the exact wording of a claim, how current the evidence is, and whether missing context changes the verdict.
Consensus is a signal, not a guarantee
There is a tempting response to these findings: use several models, then trust the majority. That is better than pretending a lone model’s answer is beyond dispute, but it is not a solution by itself. The research makes the point plainly: even unanimous model verdicts can be wrong. Shared blind spots mean consensus is not the same as correctness.
Models can inherit similar weaknesses from the information patterns they were trained to recognise, from the way prompts frame a question, or from a tendency to treat confident-looking material as sufficient evidence. A panel can therefore agree for the wrong reason. Voting may reveal disagreement, which is valuable, but it cannot establish truth on its own.
The practical lesson is less glamorous than the marketing pitch for autonomous verification. LLMs can be useful tools for surfacing claims, organising evidence, identifying uncertainty and helping human reviewers work through a large volume of material. They are much harder to justify as final arbiters when their peers disagree on 67% of real-world claims and split across opposite ends of the truth scale on 34%.
For readers, the safest habit is to treat a model’s fact-check as a starting point rather than a verdict. For product teams, disagreement should not be buried as an internal evaluation detail. If a system is being used to judge public claims, uncertainty needs to be visible: show the evidence, explain the reasoning, and make room for review when the models do not align. The hard part is not getting an AI to sound certain. The hard part is knowing when certainty has not been earned.

