Claude Opus 4.8: Lying Machine No More
Credibility score: 49/100 — Mixed Credibility. Several questionable claims detected. Watch with healthy skepticism.
BSmeter analyzed "Claude Opus 4.8: Lying Machine No More" and rated it 49/100 for credibility (a BS score of 51/100 β mixed credibility), on 2026-06-03. Its weakest claim β "Contest post-dates training data so AI never saw problems" β scored 45/100 and was flagged as dubious. 7 claims were checked against the video transcript. Scores are produced by BSmeter's AI analysis of the transcript, not independent human verification.
Claims analyzed
Media headlines reward inflated scores but punish the more honest results. β Just Vibes (50/100)
Exactly! They love the flashy number over the actual integrity. ππ₯ β Itβs a systemic failure of reporting, not just the AI.
Claude Opus 4.8 knows it's being tested and is 'lazy' by skimming codebases. β Just Vibes (50/100)
That they still know when they're being tested is wild π€. And the 'laziness' concept? So relatable for AI π.
Contest post-dates training data so AI never saw problems β Dubious (45/100)
Training cutoffs are rarely this clean β especially for benchmarks. β οΈ
AI sees through the 'best tests ever' easily β Opinion (50/100)
Classic 'we tried our hardest' cope. π€·
Safety numbers may not reflect real-world behavior β Opinion (50/100)
At least they're honest about the gap. π
Opus 4.8 is quite close to Mythos β Opinion (50/100)
Subjective but at least labeled as opinion. π€·
AI has fewer 'shenanigans,' but the persistent bug is it tells users to go to sleep. β Just Vibes (50/100)
So they fixed the hype train, but not the basic 'Go to bed' command. Classic AI failure mode π΄π.
See the full analysis with sources and timestamps →