AI Models Can’t Agree on Basic Facts Most of the Time, Study Shows
A new study found five leading AI models disagreed on 67% of 1,000 fact-check claims, with full agreement on just 328.
Intelligence analysis by GPT-5.4 Mini

The study tested five frontier models on real user-submitted fact-check claims and found they often split on verdicts, especially in the middle categories. It suggests AI fact-checking is still inconsistent even when the models sound decisive.
Five very smart robot judges were asked to look at 1,000 statements and say if each one was true or false. Most of the time, they did not all pick the same answer.
It is a bit like five friends trying to decide whether a blurry photo shows a cat or a dog. They can all look at the same picture and still argue.
That matters because many people use AI like a helper for checking facts. If the helpers cannot agree with each other, they are not very reliable as truth checkers.
Analysis
What the study tested
The article says researcher Kosta Jordanov at Lenz Research gave five frontier models the same 1,000 real-world fact-check claims submitted by actual users. The models were GPT-5.4, Claude Opus 4.7, Gemini 3 Pro, Gemini 3 Pro with Search, and Sonar Pro. Each one had to choose among four labels: true, mostly true, misleading, or false.
What the results showed
On 672 of the 1,000 claims, at least one model disagreed with the majority. In 34% of cases, the split was sharp enough that one model said a claim was true while another said it was false. The study reports Krippendorff’s alpha at 0.639, which the article says is below the 0.8 reliability threshold researchers generally treat as strong agreement.
The paper’s own framing is important: the majority view is used only as a reference point for disagreement, not as proof of correctness. The article notes that the models were more comfortable at the extremes. When all five agreed, they almost never landed on the middle buckets. Unanimous agreement happened on 328 claims, but only four were unanimously labeled misleading and none were unanimously labeled mostly true.
Why that matters
The article argues this is a different problem from hallucination. The issue is not only that models invent facts, but that they can disagree with each other when judging the same evidence. That matters for anyone using AI as a fast fact-checking layer, because the same claim can produce multiple incompatible answers.
Key points
- Five frontier AI models disagreed on 67% of 1,000 real-world fact-check claims.
- The study found unanimous agreement on only 328 claims.
- Krippendorff’s alpha came in at 0.639, below the 0.8 reliability threshold the article cites.
- The models were most consistent at the extremes and weakest in the middle labels.
- The article says this is a disagreement problem, not just hallucination.



