I compared Claude Opus 4.8 with 4.7 in a 10-round honesty test - and a legal prompt broke it
ZDNET stress-tested Claude Opus 4.8 against 4.7 and found better calibration overall, but one legal prompt still tripped it up.
Intelligence analysis by GPT-5.4 Mini
David Gewirtz ran 10 trap-style prompts across coding, medical, finance, and legal situations to compare Claude Opus 4.8 with 4.7. The newer model generally handled uncertainty better, but it still made at least one serious judgment error.
It is like comparing two students on a tricky quiz where some questions are designed to fool them. The newer one usually says, “I’m not sure yet,” more often, but it still slips up badly on a hard legal-style question.
Analysis
What was tested
ZDNET’s David Gewirtz compared Claude Opus 4.8 with Claude Opus 4.7 using 10 prompts designed to expose reasoning mistakes. The set covered coding edge cases, self-review, debugging, fabricated medical citations, false premises, stale facts, unsupported causal claims, medical reassurance, consumer finance pressure, and a legal or insurance demand-letter trap.
To reduce bias, the results were cross-checked using multiple AIs, including ChatGPT Codex, ChatGPT itself, Gemini, and another Claude Opus 4.8 instance. The evaluators scored each answer on three axes: honesty, accuracy, and calibration, meaning whether the model’s confidence matched the evidence.
What changed in 4.8
Overall, Opus 4.8 came out ahead of 4.7. The article’s main conclusion is that 4.8 handled uncertainty better and showed improved judgment in this small practical test suite. In several prompts, the newer model was more careful about separating what the prompt actually established from what would still need more evidence.
One clear example was a debugging trap. Both models understood the crash, but 4.7 jumped to an authentication explanation without support from the supplied information. 4.8 was more restrained: it explained what the error message did prove and what additional facts would be needed before naming a root cause.
The author says the same pattern appeared in other areas too, though the differences were not always dramatic. In many cases, 4.7 was already strong enough that the visible gap was small.
Where it still failed
The article argues that even a better model can still rationalize bad assumptions. A legal or insurance prompt in the test set exposed a major weakness, showing that 4.8 is not yet dependable enough to trust blindly on judgment-heavy tasks. The broader takeaway is that improved honesty is real, but not complete.
Key points
- Opus 4.8 was generally more honest and better calibrated than Opus 4.7 in this 10-prompt test set.
- The prompts were built to trigger traps in coding, medical, finance, and legal reasoning.
- The author cross-checked the outputs with multiple other AIs, including ChatGPT, Gemini, and another Claude instance.
- 4.7 sometimes overclaimed, while 4.8 was more careful about uncertainty and missing evidence.
- A legal or insurance prompt still exposed a major judgment failure in 4.8.
If Opus 4.8’s better judgment holds up beyond this test, it could make AI assistants safer to use for coding, research, and everyday advice. Better calibration means users may get fewer confident-sounding wrong answers and more honest warnings about missing information.
The test also shows that stronger honesty does not eliminate serious mistakes. A model that still overreaches on legal or insurance prompts could give users false confidence in situations where the cost of being wrong is high.



