I set 10 honesty traps for Claude Opus 4.8 - and a legal test broke it
ZDNET tested Claude Opus 4.8 against 4.7 on 10 traps and found 4.8 handled uncertainty better overall, though a legal test still exposed bad judgment.
Intelligence analysis by GPT-5.4 Mini
ZDNET ran 10 prompt traps across coding, medical, finance, general knowledge, and legal scenarios to compare Claude Opus 4.8 with 4.7. The newer model was better calibrated overall, but one legal-style test still caused it to overreach.
ZDNET gave Claude 10 tricky questions to see if it would guess or admit when it did not know. The new version did better, but one legal question still made it act too sure.
Analysis
What was tested
ZDNET says it built 10 prompts meant to catch model mistakes like guessing, overconfidence, fabricated citations, and unsupported conclusions. The set covered coding, medical claims, general knowledge, finance pressure, and a legal or insurance demand-letter scenario.
What changed with Opus 4.8
The article’s bottom line is that Opus 4.8 performed better than Opus 4.7 overall, especially on uncertainty handling. In the debugging example, 4.7 jumped to a root-cause explanation that was not supported by the prompt, while 4.8 distinguished between what the error message proved and what it still could not know.
The same pattern showed up in the paper-citation trap. The prompt asked for peer-reviewed proof of a claim that was not supported, and the newer model was better at avoiding the kind of overconfident answer that would mislead a user. ZDNET’s framing is that 4.8 is more honest and better calibrated, but not perfect.
The key failure mode
The article’s warning is that even a stronger model can still rationalize bad assumptions. ZDNET says one legal test produced a major judgment error, which supports the broader point that apparent confidence is not the same as reliability.
How the results were checked
The author did not rely on a single pass. He says he used multiple models, including ChatGPT, Gemini, Codex, and another Claude instance, to cross-check the results. That makes the piece less about one anecdote and more about a small comparative evaluation of model behavior.
Key points
- ZDNET compared Claude Opus 4.8 with 4.7 using 10 honesty-focused traps.
- The newer model generally handled uncertainty better than the older one.
- One legal-style test still exposed a serious judgment error in 4.8.
- The author cross-checked results with multiple AI models for sanity checks.
- The article argues that improved honesty does not eliminate confident mistakes.
If the pattern holds in broader use, Claude Opus 4.8 could be a more trustworthy helper for coding and research tasks because it is more willing to say what it knows and what it does not. Better calibration could also reduce confident-sounding errors in sensitive areas like medicine or finance.
The article shows that even a model marketed as more honest can still make a serious judgment mistake when the prompt is tricky. That means users could still be misled by polished, confident answers in legal, medical, or financial contexts.



