MYTHOS MYTHOS MYTHOS
Credibility score: 49/100 — Mixed Credibility. Several questionable claims detected. Watch with healthy skepticism.
BSmeter analyzed "MYTHOS MYTHOS MYTHOS" and rated it 49/100 for credibility (a BS score of 51/100 β mixed credibility), on 2026-06-10. Its weakest claim β "Fable 5's advantage grows with task length and complexity" β scored 30/100 and was flagged as dubious. 30 claims were checked against the video transcript. Scores are produced by BSmeter's AI analysis of the transcript, not independent human verification.
Claims analyzed
Anthropic released Mythos after calling it too dangerous β Dubious (35/100)
No public record of Anthropic ever saying Mythos was 'too dangerous' to release β sounds like marketing flair.
Sources: Anthropic purposely made its new Mythos-based models bad at AI research, and developers are fuming, Anthropic releases its first Mythos-class model to the public, Anthropic Releases a Safer Version of Its βToo Dangerousβ Mythos AI
Fable 5 beats every previous public model β Dubious (45/100)
No public benchmark yet backs "any model we've ever made" β that's a big swing.
Claims Mythos scored 29.3% on Frontier Code Diamond β Unverifiable (50/100)
No public results for Mythos on Frontier Code Diamond β this number has zero external receipts.
Says Mythos scored 1932 on GDP-val benchmark β Unverifiable (50/100)
OpenAI's GDP-val benchmark shows no Mythos entry at 1932 β or any score.
Fable 5 is first-of-its-kind SOTA on nearly all AI benchmarks β Dubious (35/100)
Claims 'nearly all tested benchmarks' but web results show Mythos, not Fable 5, and no public SOTA results exist.
Fable 5 will show exceptional performance on Deep Suite across domains β Opinion (50/100)
Future performance prediction β no test results exist yet.
Claims Mythos hit 85% on computer use benchmark β Unverifiable (50/100)
85% computer-use score for Mythos has zero corroboration outside this claim.
Fable 5's advantage grows with task length and complexity β Dubious (30/100)
Scaling claim with zero benchmarks or data shown.
Claims Mythos leads Legal Agent Benchmark at 13% and Humanity's Last Exam β Unverifiable (50/100)
13% on Legal Agent and first place on Humanity's Last Exam β both unverified for Mythos.
Fable 5 autonomously explored entire old codebases unprompted β Personal Story (50/100)
Personal experience claim β can't verify, sounds exaggerated.
Says Mythos scored 88% on Terminal Bench β Unverifiable (50/100)
88% on Terminal Bench for Mythos appears nowhere outside this video.
Here.Now sponsor read β helped publish tests β Sponsored (50/100)
Classic mid-video sponsor plug for Here.Now publishing tool.
Herenow just added private storage and custom domains β Unverifiable (50/100)
No source in transcript or search results confirms these recent Herenow features.
Herenow is completely free right now β Unverifiable (50/100)
No pricing data available anywhere β free tier claim sits on zero receipts.
Claims Mythos is a 10-trillion parameter model β OK (65/100)
Reddit posts repeat the 10-trillion figure, but Anthropic hasn't confirmed it officially.
Safeguards trigger in under 5% of sessions and zero times in his tests β Personal Story (60/100)
Personal testing anecdote β can't verify someone else's private sessions.
Fable 5 & Mythos 5 cost $10/$50 per million tokens β Dubious (40/100)
Search shows Mythos at $125 per million tokens β speaker's numbers don't match.
Fable 5 costs less than half of Claude Mythos preview β Unverifiable (50/100)
No public price listed for Mythos preview β can't verify the half-price claim.
Fable 5 and Mythos 5 beat all prior Claude models on autonomous runtime β Unverifiable (50/100)
Mythos isn't even out yet β zero public benchmarks on its runtime.
Dripe says Fable 5 did 2-month codebase migration in one day β Personal Story (60/100)
Anecdote from one company β interesting but no independent verification.
$50/M tokens beats paying Stripe engineers for two months β Opinion (50/100)
Depends on how much actual engineering time gets replaced β not automatic savings.
Fable outputs have unprecedented information density β Opinion (50/100)
Personal impression β can't measure 'unlike anything I've ever seen.'
Information-dense output makes the model smarter by packing more meaning per word β Opinion (50/100)
Dense text β higher intelligence β it's just how you compress the same tokens.
Same inference time + dense output = more useful work than Opus 4.8 β Dubious (45/100)
Opus 4.8 isn't a real model β Claude 4 Opus or Sonnet 3.5/4 are the actual names.
Fable 5 was tested on Cognition's Frontier Code Evaluation β Unverifiable (50/100)
No public results for "Fable 5" on Frontier Code Eval as of June 2026.
Fable 5 tops Frontier models at medium effort β Unverifiable (50/100)
No public benchmarks exist for Fable 5 β this ranking is just his word.
No limit found to scaling tokens past 100M β Dubious (45/100)
100M tokens is already a huge number β claiming theyβre still scaling past that with zero ceiling feels like hype without the curve.
Claims Mythos 5 sped up drug design 10x β Unverifiable (50/100)
10x claim with zero named studies or benchmarks β just trust the internal experts.
Anthropic routes distillation attempts to old Opus 4.8 model β Unverifiable (50/100)
Claims Anthropic has a specific fallback to Opus 4.8 for distillation β no public confirmation of that policy.
Mythos found bugs other models missed while reviewing code built on Opus and GPT-5.5 β Personal Story (50/100)
Speaker's personal test run with Mythos β no independent confirmation here.
See the full analysis with sources and timestamps →