OpenAI Just Revealed Something More Dangerous Than AGI
Credibility score: 62/100 — Mostly Credible. Mixed credibility - some claims are solid, others need verification.
BSmeter analyzed "OpenAI Just Revealed Something More Dangerous Than AGI" and rated it 62/100 for credibility (a BS score of 38/100 — mostly credible), on 2026-09-17. Its weakest claim — "Dismissing quantity without acknowledging potential for emergent quality — a false dilemma." — scored 20/100 and was flagged as false dilemma. 27 claims were checked against the video transcript. Scores are produced by BSmeter's AI analysis of the transcript, not independent human verification.
Claims analyzed
OpenAI is chasing recursive self-improvement, which is more consequential than AGI. — Confidence Mismatch (45/100)
He says OpenAI is 'openly' chasing this, but the 'more dangerous' part is his own spin. They're developing, not declaring a threat. 😈
Claims recursive self-improvement is OpenAI's top focus and three teams have published systems for it. — No Frame (75/100)
He says it's the top focus and points to recent research. It's a straightforward statement of current AI development. 🔥
Admits no proof that recursive self-improvement is more dangerous than AGI. — No Frame (75/100)
A rare moment of honesty, admitting the danger isn't 'proven.' Acknowledges the lack of definitive evidence. 😈
Cites Reuters as treating recursive self-improvement as central to AI risk debate. — No Frame (75/100)
Citing Reuters for its coverage of AI risk. A direct, verifiable reference to a news source. 😈
Claims Brown says AI models are 100 times better than humans at some bug-finding tasks. — Anonymous Authority (45/100)
A specific number, '100 times better,' attributed to 'Brown' without context on how that metric was derived. Sounds impressive, but it's just a quote. 💀
Claiming human judgment is the 'last big advantage' — a bold, unproven assertion. — Confidence Mismatch (45/100)
He says 'might be' then declares it 'the last big advantage.' That's a leap, not a conclusion. 💀
Dismissing quantity without acknowledging potential for emergent quality — a false dilemma. — False Dilemma (20/100)
He sets up 'fast researchers' vs. 'knowing what's worth researching' as if they're mutually exclusive. Why not both? 😈
Introducing 'Dream RSI' and its creators — straightforward information. — No Frame (75/100)
Alright, they're actually naming the project and the institutions. A rare moment of clarity. 🔥
Describing current AI limitations where human-defined strategies are 'locked in place' — a factual observation. — No Frame (75/100)
He's just laying out how these systems currently work. No spin, just the setup. 😈
Explaining the 'Dream RSI' process — no tricks here, just setting the stage. — No Frame (75/100)
Just laying out how this 'dreaming' system supposedly works. Standard explanation. 😈
Claiming a 'built-in safety net' ensures no worse performance — a comforting thought with no real teeth. — Confidence Mismatch (45/100)
A 'safety net' that only guarantees 'not worse' isn't exactly a win, is it? That's a low bar for 'safety.' 💀
Highlighting 'striking results' with specific, impressive metrics — but only two examples. — Cherry-Picked (20/100)
They tested it on 'eight tasks' but only gave you the two best numbers. That's not 'striking results,' that's a highlight reel. 🍒
Presenting a 'most dramatic result' with a massive difference in calls — a stark comparison. — No Frame (75/100)
A 300 vs 51,000 call difference is genuinely significant. No trickery here, just a clear win. 🔥
Describing the system's adaptive strategy of 'holding back' then 'widening search' — a narrative of intelligent behavior. — Loaded Language (45/100)
It 'got more careful' and 'opened back up' — that's not just data, that's a story of intent. It's an algorithm, not a strategist. 😈
Describing a search strategy as a 'crude, narrow echo' of instinct. — No Frame (75/100)
Just setting the stage with a descriptive analogy. No trickery here, just a bit of poetic license. 😈
Claiming that giving AI specific search advice made results worse, boxing it in. — No Frame (75/100)
A straightforward observation about AI behavior. Sometimes less guidance is more. 🔥
Defining the 'harness' as the AI's operational machinery beyond the core model. — No Frame (75/100)
Clearly defining a technical term. This is just information, not a trick. 😈
Explaining how Modular RSI isolates and fixes faulty parts of the AI harness. — No Frame (75/100)
Describing a technical process. It's a clear explanation of how the system works. 🔥
Stating that harness parts evolve separately and then merge. — No Frame (75/100)
Just a continuation of the technical description. No hidden agenda here. 😈
Presenting specific accuracy improvements on Terminal Bench 2.0 and SWEBench Verified. — No Frame (75/100)
Giving specific numbers for performance improvements. The data is right there. 🔥
Claiming broad applicability for AI improvements based on specific benchmarks. — Confidence Mismatch (45/100)
They're showing specific percentage bumps and then leaping to 'it worked in different cars' and 'carried over' like it's a universal law. That's not how this works. 💀
Introducing 'Science Buddy' as a deeper step in recursive self-improvement. — No Frame (75/100)
Just introducing a new project and its claimed methodology. Nothing tricky here, for once. 😈
Describing Science Buddy's capabilities and its recursive learning process. — No Frame (75/100)
Detailing the features and the feedback loop of Science Buddy. It's a description of how the system is designed to work. 😈
Explaining the inner and outer loops of Science Buddy's self-improvement mechanism. — No Frame (75/100)
Breaking down the two-loop system for self-improvement. It's a technical explanation, not a trick. 😈
PHAI's evidence is only at the case study stage — a crucial caveat. — No Frame (75/100)
They're admitting the limitations of the evidence right up front. Rare honesty, I'll give them that. 😈
Not proof of intelligence explosion, but 'structurally' the same loop — a classic bait and switch. — Volume Game (45/100)
Dismisses 'intelligence explosion' then immediately pivots to 'structurally, this is the kind of loop.' They give with one hand and take with the other. 😈
Admits no AI redesigns itself without humans, then immediately pivots to 'Brown's comment lands differently.' — Volume Game (45/100)
They tell you 'nobody has shown' autonomous AI, then immediately use that to say 'Brown's comment lands differently.' It's a setup for manufactured urgency. 💀
See the full analysis with sources and timestamps →