There has been a situation in AI
Credibility score: 43/100 — Mixed Credibility. Several questionable claims detected. Watch with healthy skepticism.
BSmeter analyzed "There has been a situation in AI" and rated it 43/100 for credibility (a BS score of 57/100 — mixed credibility), on 2026-06-19. Its weakest claim — "Anthropic's Fable model is a "crazy hacky" cybersecurity risk – setting up a straw man." — scored 20/100 and was flagged as straw man. 21 claims were checked against the video transcript. Scores are produced by BSmeter's AI analysis of the transcript, not independent human verification.
Of 21 claims analyzed: 3 scored under 40, 17 between 40 and 69, and 1 at 70 or above.
Claims analyzed
Opening with vague 'series of situations' to build anticipation. — Loaded Language (45/100)
At 0:00
Starts with 'series of situations' without specifics. Classic hype-building, keeps you guessing 🙄.
Why this score: The speaker uses vague, dramatic language like 'a situation' and 'a series of situations' without immediately providing context or specific details. This rhetorical device is used to create intrigue and draw the audience in, making them curious about what the 'situations' are, rather than directly stating the topic.
Original quote: “What is going on everybody? There has been a situation. I mean, there's really been a series of situations. Uh, and I just uh I just want to sit down and kind of give my thoughts and talk about it a little bit here.”
Downplaying 'Claude Fable stuff' as a 'tiny piece' while immediately focusing on it. — Volume Game (45/100)
At 0:15
Calls it a 'tiny piece' then immediately dives deep into it. Classic volume game: downplay the main topic, then make it the main topic 🤡.
Why this score: The speaker states that the 'Claude Fable stuff' is 'only but a tiny piece of the situation,' yet immediately proceeds to explain it in detail. This is a 'volume game' where a topic is initially minimized to manage expectations or create a sense of a larger, more complex narrative, but then it becomes the immediate focus, indicating its actual importance to the speaker's discussion.
Original quote: “I'm going to open with the Claude Fable stuff, but I'm actually that's that's only but a tiny piece of the situation here. So if if you don't know the TLDDR on the Claude Fable stuff is, you know, Anthropic's been running around talking about this Fable model and how it's this”
Anthropic's Fable model is a "crazy hacky" cybersecurity risk – setting up a straw man. — Straw Man (20/100)
At 0:30
He's describing Anthropic's marketing with loaded terms like 'crazy hacky' and 'cybersecurity risk' to set up a straw man he can then knock down. 🤡
Why this score: The speaker is characterizing Anthropic's claims about their Fable model in an exaggerated, negative light ('crazy hacky', 'cyber security risk') before he even evaluates it. This sets up a straw man argument, where he can then dismiss a more extreme version of their claims rather than their actual, nuanced statements. It's a classic move to make his own critique seem more impactful.
Original quote: “Anthropic's been running around talking about this Fable model and how it's this crazy hacky like it can hack anything uh type model and it's a cyber security risk all this.”
Claiming Claude Fable 5 is export restricted like weapons of war. 🚩 — Loaded Language (45/100)
At 2:30
Comparing an AI model to 'weapons of war' for export restrictions is pure emotional loading. It's not a direct equivalence. 🙄
Why this score: The speaker uses highly charged language ('insane,' 'weapons of war') to describe the export restriction on Claude Fable 5. While the restriction itself might be true, framing it as equivalent to firearms is an appeal to emotion, exaggerating the severity and implications without providing a nuanced comparison of the actual risks or regulatory frameworks. It's designed to provoke a strong reaction rather than inform.
Original quote: “there's been like Claude Fable 5 has been export restricted. Uh, which is insane. That that's what you do to like firearms stuff like weapons of war get export restricted. Um and somehow Claude Fable 5 has found itself in the category of um export restrictions and um the government of the United…”
Anthropic tried to 'do to us' — vague threat framing. — Loaded Language (45/100)
At 4:30
Uses 'do to us' to imply malicious intent without specifying what 'it' is. Classic fear-mongering. 🚩
Why this score: The phrase 'what Anthropic tried to do to us' is emotionally charged and vague. It suggests a negative, potentially harmful action without providing any specific details or evidence of what that action was, leaving the audience to infer the worst. This is a common tactic to create an emotional response and align the audience against a perceived antagonist.
Original quote: “let's not lose sight of what Anthropic tried to do to us.”
Doubts M3's benchmark access, then pivots to personal experience as proof. — Confidence Mismatch (45/100)
At 6:30
Starts with 'maybe I'm mistaken' then ends with 'in all my experience no you cannot.' That's a quick jump from doubt to certainty 🚩.
Why this score: The speaker initially expresses uncertainty about M3's access to benchmark data, then quickly dismisses the possibility based on personal experience, presenting it as definitive proof. This is a confidence mismatch where personal anecdote overrides initial doubt without new evidence.
Original quote: “This benchmark came out, I believe this came out be like too soon to the release of M3 for M3 to um to have possibly had access to the information. So, let me see if I can find Yes. Okay. So, if I go to V1, that almost does concern me that may maybe I'm mistaken. I don't think I am. I don't think…”
Calling a company's actions 'evil' and 'insane' – pure emotional button pushing. — Emotional Button (45/100)
At 8:30
Uses 'evil' and 'insane' to describe a company's actions. That's not a fact, that's a feeling. 😭
Why this score: The speaker uses highly charged emotional language ('evil,' 'insane') to describe a company's decision. This framing aims to elicit a strong emotional reaction from the audience rather than presenting a neutral, fact-based critique of the decision itself. It's designed to make the audience feel outrage or disgust without necessarily providing objective evidence for the 'evil' or 'insane' nature of the action.
Original quote: “I still it breaks my brain that a company purposely did that. Um it's that's insane to me.”
Setting up the GLM 52 trial as a casual, almost reluctant, discovery. — No Frame (75/100)
At 10:30
Just laying out the backstory of how he got the key, no real spin here. Straightforward setup. 🤷♂️
Why this score: The speaker is simply recounting the circumstances of receiving access to the GLM 52 model, emphasizing his initial busyness and focus on other projects. This sets the stage for his later positive reaction, making it seem like an unexpected discovery rather than a pre-arranged endorsement. It's a common narrative technique to build anticipation and make the eventual praise more impactful, but it's not manipulative framing.
Original quote: “But I get this DM from ZI and they're like, "Hey, uh, we're about to release 52 and, um, would you like to try it out?" And I was like, "Okay, sure." You know, you give me a key. I was busy at the time. I'm like busy trying to make Miniax M3 actually work for, you know, local. And I know GLM GLM…”
Claims GLM52 is a "true frontier model" comparable to Opus48 and GPT55. — Confidence Mismatch (45/100)
At 12:30
Calling GLM52 a "true frontier model" and equating it to Opus48/GPT55 after 10 minutes of use? That's a bold claim with zero evidence. 🚩
Why this score: The speaker makes a very strong assertion that GLM52 is not 'just another open source model' but is on par with 'Opus48 GPT55,' calling it the 'first true Frontier model.' This claim is based on a very short personal experience ('took me like 10 minutes of like working with it to realize'). While personal experience can be valid, equating a new model to established, high-performing proprietary models like GPT-5.5 (which isn't even publicly released yet, making the comparison speculative) and Opus48 after such limited interaction is a significant leap of confidence without presenting any…
Original quote: “trying to use Miniax with um uh the Hermes agent, I just couldn't like it just wasn't it would never get out of reasoning. And then so I'm like, man, I really need to get this out of reasoning and I just couldn't figure it out. I knew it had something to do with these tokens, but then the tokens…”
Claiming GLM52 is equivalent to Opus48 and GPT55. — Confidence Mismatch (45/100)
At 12:30
Calling GLM52 'Opus48 GPT55' after '10 minutes of working with it' is a bold claim with zero evidence. That's a vibe, not a benchmark. 🤡
Why this score: The speaker makes a very strong claim about GLM52's performance, equating it to top-tier models like Opus48 and GPT55, based on only '10 minutes of like working with it.' This is a subjective assessment presented with high confidence, lacking any objective benchmarks or comparative data to support such a significant equivalence. It's an emotional reaction, not a technical evaluation.
Original quote: “trying to use Miniax with um uh the Hermes agent, I just couldn't like it just wasn't it would never get out of reasoning. And then so I'm like, man, I really need to get this out of reasoning and I just couldn't figure it out. I knew it had something to do with these tokens, but then the tokens…”
Claims 'holy grail' then immediately walks it back. Classic Volume Game. — Volume Game (45/100)
At 14:30
Loudly calls it 'holy grail' then whispers 'maybe not' in the same breath. Classic hype-and-retreat. 🙄
Why this score: The speaker uses an emotionally charged term ('holy grail') to describe the MIT license, then immediately qualifies it with 'maybe it's not the holy holy grail.' This is a volume game, where a strong, attention-grabbing statement is made, followed by a quiet, less impactful retraction or qualification, allowing the initial strong impression to linger.
Original quote: “I'm not a lawyer, but you know, read the MIT license, but it's the holy grail.”
Nvidia and Neotron are the pinnacle of open-source AI, a bold claim. — Confidence Mismatch (45/100)
At 16:30
Calling them the 'pinnacle' is a huge claim, especially in such a fast-moving field. Where's the data for that? 📈
Why this score: The speaker asserts Nvidia and Neotron are the 'pinnacle' of open-source AI without providing any specific metrics, benchmarks, or comparative analysis to support such a definitive statement. It's a strong, subjective opinion presented as a widely accepted fact.
Original quote: “Neimotron is kind of the they are the the pinnacle of like Nvidia and Neotron is the pinnacle of actual open source AI.”
Claiming a model is 'better' than established ones based on personal 'impression' — Confidence Mismatch (45/100)
At 18:30
Saying a model is 'better' than Opus 48 and GBD55 based on 'impression' but admitting benchmarks don't show it. That's a vibe, not a fact. 🤷♀️
Why this score: The speaker asserts their personal 'impression' and 'usage' lead them to believe a model is superior to well-known ones (Opus 48, GBD55), despite acknowledging that 'benchmarks don't reveal that.' This is a classic confidence mismatch where subjective feeling overrides objective data, presented as a bold, insider take.
Original quote: “felt like it it was a better model than”
Presents Anthropic's claims as extreme, setting up a straw man. 🤡 — Straw Man (20/100)
At 20:30
Exaggerates Anthropic's position to 'hack the planet' and 'take everyone's jobs' — setting up a straw man to knock down. Classic move. 🙄
Why this score: The speaker characterizes Anthropic's claims about AI as overly dramatic and fear-mongering ('all the jobs are going to be replaced,' 'hack the planet'). While AI companies do discuss job displacement and advanced capabilities, framing it in such extreme terms creates a straw man argument that is easier to refute, rather than engaging with the nuanced discussions these companies actually have about AI's impact and risks.
Original quote: “so what this changes though is it changes the entire calculus because up until this point we've had these companies like Anthropic where they're saying all the jobs are going to be replaced. We have we our model is so smart it's going to take everyone's jobs. our model is so smart it's going to…”
Declaring a 'moat' gone with absolute certainty, no evidence. 💀 — Confidence Mismatch (45/100)
At 22:30
Absolute certainty that a 'moat' is 'gone' — zero data to back that up. Just vibes. 🚩
Why this score: The speaker states definitively that a competitive advantage ('moat') is 'gone' without providing any specific metrics, market analysis, or competitive data to support such a strong claim. It's a confident assertion without evidence.
Original quote: “and the situation now is they simply don't they just don't anymore. It's gone. It's it's just gone.”
Pre-emptively shutting down criticism about cable management. Straw Man 🙄 — Straw Man (20/100)
At 24:30
He's attacking a hypothetical comment about cable management before anyone even says it. Classic straw man. 🛡️
Why this score: The speaker anticipates and dismisses criticism about his cable management by telling 'the comments' to 'shut up' and challenging them to do better. This is a straw man argument because he's creating a hypothetical critic and then refuting them, rather than addressing actual feedback or simply acknowledging the issue without defensiveness. It's a preemptive strike against an imagined opponent.
Original quote: “Um cable management is not my strong
suit. But also, you know what? Shut up
in the comments because you try to make
this pretty because you you have to
understand like even like the these like
little adapters. So like this this
stuff, right? Each one of these is a
PCIe um you know to to your GPU.…”
Guessing performance at 2 tokens/second for 8-bit precision. It's a guess, not a hard stat. 🤷♂️ — Confidence Mismatch (45/100)
At 26:30
Says 'I'm guessing' but presents 'two tokens a second' like it's a known benchmark. It's a vibe, not data. 🤡
Why this score: The speaker explicitly states 'I'm guessing' regarding the performance (two tokens a second). While it's a reasonable estimate given the memory requirements, it's still an unverified guess presented with a degree of specificity that could be mistaken for an actual benchmark. It's a personal estimate, not a tested fact.
Original quote: “Good luck. So, uh, if you have that in RAM, awesome. I do, uh, on one of my machines, but it's probably, you know, if you try to serve this over RAM, it's I'm guessing like you'd get like two tokens a second at that, uh, precision.”
Claims 97.5% of model is in Q4bit, using a tiny, unreadable graph. 📈 — Confidence Mismatch (45/100)
At 28:30
Cites a specific number (97.5%) from a graph he admits is 'hard to read.' The confidence doesn't match the visual evidence. 💀
Why this score: The speaker is presenting a precise statistic (97.5%) while simultaneously acknowledging that the visual evidence (the graph labels) is 'so tiny' and 'really hard to read.' This creates a mismatch between the certainty of the claim and the verifiable source being presented. It's hard to trust a number when the source is illegible.
Original quote: “This is 8 bit, right? Uh, and this is 4bit and this number here is 97.5. So, theoretically, you know, it's like this is like a 97.5% of the model is in Q4bit.”
Framing IPO rush as 'locking in at AI bubble price' 🚩 — Loaded Language (45/100)
At 30:30
Calling current valuations an 'AI bubble price' is pure speculation, not a market analysis. 📈
Why this score: The speaker uses 'AI bubble price' to imply current valuations are inflated and unsustainable, which is a common sentiment but presented as fact rather than opinion or a potential scenario. It's loaded language designed to evoke caution without providing concrete evidence of a bubble's imminent burst.
Original quote: “especially as all these companies are
scrambling to like IPO in the United
States and like lock in at the AI bubble
price.”
Explaining performance dips with 'getting hammered' due to popularity. — Loaded Language (45/100)
At 32:30
Uses 'getting freaking hammered' to explain performance, which is a vibe, not a technical reason. 🔨
Why this score: The speaker attributes lower token speeds to providers 'getting hammered' due to popularity. While popularity can impact performance, 'getting hammered' is an emotional, non-technical explanation that lacks specific details about infrastructure or load management issues. It's a colorful way to say 'overloaded' without providing any data.
Original quote: “anyone who's not doing 60, the reason they're not doing 60 is because they're getting freaking hammered because this model is very popular right now.”
Predicting government backstopping of shares due to public reaction to AI models like ZI. — Confidence Mismatch (45/100)
At 34:30
Predicting the government 'would have to' backstop shares and the public 'will see' AI models like ZI. That's a lot of certainty for a future event. 🔮
Why this score: The speaker is making a confident prediction about future government action ('would have to backstop') and public reaction ('public will see something like ZI'). While it's a plausible scenario in a discussion about AI's impact, the phrasing implies a level of certainty that isn't backed by current policy or public sentiment data. It's a strong assertion about future events without specific evidence to support the inevitability.
Original quote: “And so these two companies would have to do something similar and then the government would have to backs stop a large amount of those shares because the public will see something like ZI and there's just going to be more models. It's going to keep happening.”
See the full analysis with timestamps →