They Spent Billions on Moats. This Free Model Just Destroyed Them.

Credibility score: 46/100 — Mixed Credibility. Several questionable claims detected. Watch with healthy skepticism.

BSmeter analyzed "They Spent Billions on Moats. This Free Model Just Destroyed Them." and rated it 46/100 for credibility (a BS score of 54/100 — mixed credibility), on 2026-06-23. Its weakest claim — "GLM 2.5 is a "massive attack" on large corporations' business models, presented with high confidence." — scored 45/100 and was flagged as confidence mismatch. 25 claims were checked against the video transcript. Scores are produced by BSmeter's AI analysis of the transcript, not independent human verification.

Of 25 claims analyzed: 0 scored under 40, 24 between 40 and 69, and 1 at 70 or above.

Claims analyzed

GLM 2.5 is a "massive attack" on large corporations' business models, presented with high confidence. — Confidence Mismatch (45/100)

At 0:00

Calling it a "massive attack" right out the gate with zero evidence. That's a bold claim, chief, where's the beef? 🥩

Why this score: The speaker makes a very strong, definitive statement about the impact of GLM 2.5 without providing any immediate supporting evidence or explanation for why it constitutes a 'massive attack.' This is a classic confidence mismatch, where the certainty of the claim far outstrips the presented data.

Original quote: “GLM 2.5 is out and this is huge. Not just because it's a incredibly good model, but also because it's a massive attack to those business models for those large corporations. Let me explain because there is a lot to unpack here.”

GLM 2.5 has a "1 million contest window" and an MIT open-source license. — Missing Context (45/100)

At 0:16

Mentions a "1 million contest window" but doesn't say what kind of contest or what the '1 million' refers to. Is it dollars? Participants? Unicorns? 🦄

Why this score: The speaker mentions a '1 million contest window' without specifying the currency, the nature of the contest, or any other details that would make this claim understandable. It's a piece of information dropped without the necessary context for the audience to grasp its significance, especially in relation to the 'massive attack' claim.

Original quote: “First of all, this is of course a 1 million contest window, okay? And so we have all the requirement that we want, but plus there is a MIT open-source license, okay? And this is”

Claiming a new model is "sometimes better than GPT 5.5" with no proof. Big talk, no receipts. 🤡 — Confidence Mismatch (45/100)

At 0:30

Comparing it to GPT 5.5 and Opus 4.8 like it's a known fact, but then immediately saying "we need to test it." Which is it, champ? 🤷‍♀️

Why this score: The speaker makes a bold claim about the model's performance being 'sometimes better than GPT 5.5' and 'comparable to Opus 4.8' without providing any specific benchmarks or evidence. This is a classic confidence mismatch, hyping it up before admitting it needs testing.

Original quote: “no restriction, no limits, no borders. You can do anything that you want with this. But a lot of people have been talking about this model in the last couple of days because this is comparable to Opus 4.8 and sometimes better than GPT 5.5. Okay, so if you see what's going on here, it's pretty…”

Compares a new model to Opus 4.8 and GPT 5.5, calling it 'insane' without showing the benchmarks. — Confidence Mismatch (45/100)

At 0:35

Claims it's 'comparable' and 'sometimes better' than top models, then calls it 'insane' without showing the actual benchmarks. Just vibes and hype, baby! 🚀

Why this score: The speaker makes a strong claim about a new model's performance relative to established, high-performing models (Opus 4.8, GPT 5.5) and uses emotionally charged language ('insane'). However, they don't display or cite the specific benchmarks that would support this comparison, leaving the audience to trust the speaker's assessment without evidence.

Original quote: “But a lot of people have been talking about this model in the last couple of days because this is comparable to Opus 4.8 and sometimes better than GPT 5.5. Okay, so if you see what's going on here, it's pretty insane.”

Immediately contradicts the previous claim about benchmarks, saying we 'can't trust' them. — Volume Game (45/100)

At 0:51

Just hyped up benchmarks, then immediately said 'we can't trust what a benchmark is saying.' The pivot was faster than a startup's burn rate. 🤡

Why this score: After using the implied performance from benchmarks to hype up the new model, the speaker immediately undermines the credibility of benchmarks themselves. This is a classic 'volume game' where a strong, attention-grabbing claim is made, then quietly walked back or contradicted in the same breath, allowing the initial impression to linger while providing an out.

Original quote: “Now, of course, we need to test it because we can't trust what a benchmark is saying, but”

Dismissing benchmarks after just using them to hype the model. The convenient skepticism switch. 🙄 — Volume Game (45/100)

At 0:51

Just hyped the model by comparing it to top-tier AI, then immediately said "we can't trust what a benchmark is saying." That's a quick pivot, chief. 🎭

Why this score: The speaker just used implied benchmarks (comparable to Opus 4.8, better than GPT 5.5) to build excitement, then immediately undermines the credibility of benchmarks by stating 'we can't trust what a benchmark is saying.' This is a volume game, making a loud claim and then quietly retracting the basis for it.

Original quote: “Now, of course, we need to test it because we can't trust what a benchmark is saying, but”

Advises against yearly plans based on personal experience with Minimax models changing. — Personal Story (70/100)

At 1:08

Sharing personal experience with yearly plans and model changes. It's a valid take based on their own wallet. 💸

Why this score: The speaker is giving advice based on their personal experience with a specific AI model (Minimax) and its performance fluctuations over time. This is a subjective opinion and anecdote, not a universal fact, but it's presented as a personal lesson learned, which is a legitimate way to share insights.

Original quote: “subscribe if you want to. Here, they have price for monthly, quarterly, and yearly. I personally don't suggest to use a yearly plan. I have the Minimax yearly plan and I realized that their old model are just going up and down in the sense they are you know, nowadays GLM is better than Minimax.…”

Using the "1.7 billion unbanked" stat to push crypto payments for AI. The emotional button for a tech solution. 😭 — Emotional Button (45/100)

At 2:02

Bringing up 1.7 billion unbanked people to justify crypto for AI access. That's a heavy emotional lift for a payment method. 💔

Why this score: The speaker uses the statistic of '1.7 billion people in this planet without bank account' to highlight the importance of cryptocurrency payments for AI access. While the statistic itself might be accurate, it's used as an emotional button to frame cryptocurrency as a necessary solution, potentially overstating its immediate impact or necessity for AI access specifically, rather than just payment in general.

Original quote: “pay the the token with a cryptocurrency. Now, why this is important, okay? I There is 1.7 billion people in this planet without bank account. And if you want to use anything on internet, okay? So, those whatever frontier model, whatever API key from another company, you need to be able to pay with…”

Highlights cryptocurrency payment as crucial for 1.7 billion unbanked people to access internet services. — Emotional Button (45/100)

At 2:06

Connects crypto payments to '1.7 billion unbanked people' accessing internet services. A bit of a leap, but it tugs at the heartstrings. ❤️‍🩹

Why this score: The speaker uses the statistic of '1.7 billion people without bank accounts' to frame cryptocurrency payments as a crucial solution for internet access. While financial inclusion is a real issue, directly equating crypto payments with universal access to 'whatever frontier model' or 'API key' for this entire demographic is a broad generalization. It leverages an emotional appeal (helping the unbanked) to emphasize the importance of crypto, potentially overstating its immediate impact for this specific use case.

Original quote: “Now, why this is important, okay? I There is 1.7 billion people in this planet without bank account. And if you want to use anything on internet, okay? So, those whatever frontier model, whatever API key from another company, you need to be able to pay with your bank account. Okay? Here there is uh…”

Anyone on the planet can use AI because of crypto payments. Bold claim. — Confidence Mismatch (45/100)

At 2:30

From 'can pay with crypto' to 'anyone on the planet can use AI' is a leap. Access isn't just about payment methods, chief. 🌍💸

Why this score: The speaker implies that simply enabling cryptocurrency payments makes AI accessible to 'anyone on this planet.' This ignores significant barriers like internet access, digital literacy, device availability, and regulatory restrictions in various countries. Payment is one hurdle, not the only one.

Original quote: “Okay, so now anyone on this planet can actually use uh AI because of this.”

Claiming universal AI access due to crypto payments. — Confidence Mismatch (45/100)

At 2:30

Crypto payments don't magically give 'anyone on this planet' AI access. That's a leap of faith, not logic. 🌍💸

Why this score: The ability to pay with cryptocurrency, while expanding options, doesn't inherently solve issues like internet access, digital literacy, or the cost of AI services for 'anyone on this planet.' It's a confident overstatement of impact.

Original quote: “Okay, so now anyone on this planet can actually use uh AI because of this.”

Claims Cloud Fable is 'a little bit better' than GPT 5.5 — Missing Context (45/100)

At 4:30

Says 'a little bit better' and 'five points better' without defining what those points mean or what metric is being used. What are we even measuring here? 🤷‍♀️

Why this score: The speaker states Cloud Fable is 'a little bit better' and 'five points better' than GPT 5.5. However, they don't specify the metric or benchmark used for this comparison (e.g., accuracy, speed, specific task performance). Without this context, 'five points better' is vague and hard to evaluate.

Original quote: “we can see that Cloud Fable it's just a little bit better than GPT 5.5.”

Suggests using a leaderboard for token usage to pick models, then immediately caveats it's not about 'smartness' — Volume Game (45/100)

At 4:30

Tells you to use a leaderboard for model choice, then immediately says it's not about intelligence. The ol' give-and-take. 🤷‍♀️

Why this score: The speaker presents a method for choosing models (leaderboard based on token usage) but then quickly qualifies it by stating it doesn't reflect the model's 'smartness.' This is a volume game, presenting a suggestion and then retracting its core utility in the same breath.

Original quote: “also a really good way if you don't know what to you should use, have a look about this and already give you kind of an idea of what is worth using it. But remember, this is not how smart it is, but the number of token that are being used. So, because on the top, you will see models that are…”

Noting a 'contradiction' where a cheaper model (Pro) is less popular than MiniMax. — Missing Context (45/100)

At 6:30

Calling it a 'contradiction' that people prefer MiniMax over a cheaper 'Pro' model. Price isn't the only factor, chief. 🤷‍♂️

Why this score: The speaker identifies a 'contradiction' based solely on cost, implying that a cheaper model should automatically be more popular. This misses the context that 'love' or preference for an AI model can be driven by many factors beyond just price, such as performance, features, ease of use, community support, or specific use cases. It's not necessarily a contradiction, just a more nuanced market reality.

Original quote: “now this is a sort of contradiction because I see that it's even cheaper the the Pro than MiniMax, but people love MiniMax.”

Estimates for Cloud Fable 5 parameters are all over the place, no solid number 🤷‍♀️ — Confidence Mismatch (45/100)

At 8:30

Went from 'should be double' to 'some people say 10 billion' to 'could be 6-8' real quick. That's a lot of wiggle room for a 'should be' statement 🤡

Why this score: The speaker gives multiple, widely varying estimates for Cloud Fable 5's parameters, starting with a confident 'should be double' and then immediately walking it back with 'some people say' and 'could be around'. This shows a lack of a definitive source or clear understanding, despite the initial confident tone.

Original quote: “Cloud Fable 5, it should be around at least double. There's some people that say it's in between the 10 billion parameters. Um but it could be around 6-8, something like this, okay?”

Parameter count for Cloud Fable 5 is a wild guess, not a hard number. — Confidence Mismatch (45/100)

At 8:30

Went from "at least double" to "some people say 10 billion" to "could be 6-8." That's not a number, that's a vibe check. 🤷‍♂️

Why this score: The speaker presents a range of parameter counts for Cloud Fable 5, starting with a confident 'at least double' but quickly devolving into vague estimates ('some people say,' 'could be around 6-8'). This shows a lack of concrete data, yet the delivery implies a level of insight that isn't backed by specific figures.

Original quote: “at least double. There's some people that say it's in between the 10 billion parameters. Um but it could be around 6-8, something like this, okay?”

Chinese models are smaller but achieve 'incredible results' — a bold claim with no specifics. — Confidence Mismatch (45/100)

At 10:30

Calling results 'incredible' without a single metric or comparison is just vibes, not data. Show the receipts! 📈

Why this score: The speaker uses highly subjective language ('incredible results') to describe the performance of Chinese models without providing any objective benchmarks, comparative data, or specific examples. This makes the claim difficult to verify and relies on the audience's trust rather than evidence.

Original quote: “we have those Chinese model actually much smaller achieving incredible results.”

Chinese models are smaller and achieve 'incredible results' — a bit vague on the 'incredible' part. — Loaded Language (45/100)

At 10:30

Calling results 'incredible' without a single metric to back it up. That's just vibes, not data, chief. 💅

Why this score: The speaker uses 'incredible results' as a descriptor, which is subjective and emotionally charged. It lacks specific, quantifiable evidence or examples to support the claim, making it a loaded statement rather than a factual one. What's 'incredible' to one person might be 'meh' to another.

Original quote: “we have those Chinese model actually much smaller achieving incredible results.”

Cites 50-60 tokens/second for 35B model, but only 10-15 for 27B model. — Confidence Mismatch (45/100)

At 12:30

Wait, the BIGGER model (35B) is doing 50-60 tokens/sec, but the SMALLER one (27B) is only 10-15? That math ain't mathing, chief. 🤨

Why this score: The speaker presents a performance comparison where the larger 35 billion parameter model (Gwen 3.6) is stated to run significantly faster (50-60 tokens/second) than the smaller 27 billion parameter model (Gwen 3.6) at only 10-15 tokens/second. This is counter-intuitive, as generally, larger models require more computational resources and would typically run slower or at a similar speed given the same hardware and optimization. This discrepancy suggests either a misunderstanding, a specific optimization detail that isn't explained, or an error in the stated numbers, leading to a mismatch…

Original quote: “50 60 even more sometimes I talk a second. This one instead the 27 billion only around 10 15 token per second.”

Suggests Nvidia for speed, then gives a vague $3000+ price tag for hardware. — Missing Context (45/100)

At 12:30

Mentions Nvidia for speed, then throws out a $3000+ price without specifying what that gets you. What kind of card, chief? 💸

Why this score: The speaker suggests Nvidia cards for speed but then gives a broad price range ($3000+) without detailing what specific hardware or configuration that price point refers to. This lacks crucial context for the viewer to understand the actual cost-performance trade-off.

Original quote: “But, if you want to have something that go fast, you need to have maybe an Nvidia card and and so on. So, the price that you're going to spend it's around probably from the 3,000 up foot. Prices are going higher and higher.”

Claims RTX 5090 runs 60-70 tokens/sec, but the card doesn't exist yet 🤡 — Confidence Mismatch (45/100)

At 14:30

Talking about an RTX 5090 like it's real and available. Bro, that card is still in the metaverse 💀

Why this score: The speaker confidently cites performance numbers for an 'RTX 5090,' but as of June 2026, the RTX 50 series (Blackwell architecture) has not been released. The 5090 is a future product, making any specific performance claims speculative at best. It's a classic confidence mismatch, talking about unreleased tech as if it's current.

Original quote: “RTX I think it's called RTX 5090 can run at around 60 70 token per second which”

Claims most people find $200 subscriptions extremely expensive, using an emotional appeal. — Emotional Button (45/100)

At 16:30

Says 'most people' find $200 'extremely expensive' — a broad generalization to hit the emotional button. No data, just vibes. 💸

Why this score: The speaker uses a sweeping statement about 'most people' and an emotionally charged term ('extremely expensive') without any supporting data or context. This is designed to evoke empathy and agreement rather than present a verifiable fact. It's a common tactic to frame a product or service as inaccessible or unfair.

Original quote: “Most of the people in this planet they find the $200 subscription extremely expensive.”

OpenAI did a decent job in 8 minutes and 15 seconds, but the 'insane speed' is subjective. — Confidence Mismatch (45/100)

At 18:30

Calling 8 minutes and 15 seconds for a game 'insane speed' is a bit much — it's fast, but 'insane' is doing heavy lifting here. 💨

Why this score: The speaker uses 'insane' to describe the speed, which is an emotional exaggeration. While 8 minutes and 15 seconds is quick for building an HTML game, the term 'insane' injects a level of hype that isn't strictly objective, especially without a clear benchmark for what would be considered 'sane' speed.

Original quote: “OpenAI working at medium doing a job, a decent job, but at a speed is insane. It took 8 minutes and 15 okay?”

Comparing Minimax time to another model's time, but the context is missing. — Missing Context (45/100)

At 20:30

Comparing times without saying what they're measuring or for what task. It's like comparing race times without mentioning the distance 🏃‍♂️💨

Why this score: The speaker presents two time durations (19 minutes 7 seconds vs. 18 minutes 52 seconds) for 'Minimax' and another unnamed entity, implying a direct comparison. However, the specific task or metric being measured is not clearly stated, making it difficult to understand the significance or fairness of the comparison. What exactly took these times? What are we comparing here?

Original quote: “and uh 7 seconds. And this is Minimax that took 18 minutes and 52 seconds.”

Claiming MiniMax isn't a huge difference, despite being smaller. Confidence Mismatch. — Confidence Mismatch (45/100)

At 22:30

He's saying 'not a huge difference' for a model that's 'much smaller.' That's like saying a mini-fridge isn't that different from a full-size one because they both keep food cold. The size IS the difference! 🤏

Why this score: The speaker downplays the difference in performance between MiniMax and other models, despite explicitly stating MiniMax is 'much smaller.' The implication is that a smaller model performing comparably is a significant achievement, but the speaker's phrasing minimizes this, creating a mismatch between the implied context and the stated conclusion. It's like he's trying to be humble but accidentally underselling his own point.

Original quote: “So, if you see now it's not a huge difference even with MiniMax on which is much smaller it's not a huge difference with the other models.”

See the full analysis with timestamps →