Qwen3.8-27B & How to Serve it Fast
Credibility score: 56/100 — Mixed Credibility. Several questionable claims detected. Watch with healthy skepticism.
BSmeter analyzed "Qwen3.8-27B & How to Serve it Fast" and rated it 56/100 for credibility (a BS score of 44/100 — mixed credibility), on 2026-09-17. Its weakest claim — "Qwen 3.8 Max is 'undeniably amazing' — a bold, subjective claim." — scored 45/100 and was flagged as loaded language. 30 claims were checked against the video transcript. Scores are produced by BSmeter's AI analysis of the transcript, not independent human verification.
Claims analyzed
Qwen 3.8 Max is 'undeniably amazing' — a bold, subjective claim. — Loaded Language (45/100)
Calling a model 'undeniably amazing' is pure hype. 'Amazing' is an opinion, not a spec sheet. 😈
The 2.4 trillion parameters make it unrunnable locally for 'very few people'. — Confidence Mismatch (45/100)
2.4 trillion parameters is a beast, but 'very few people' is a guess. How few? Got a number, or just vibes? 💀
Claims 'everyone' was waiting for the Qwen 3.8-27B model — a broad generalization. — Loaded Language (45/100)
Oh, 'everyone' was waiting, were they? I've seen 'everyone' wait for the apocalypse, not a model. 🙄
Calling Qwen 3.6-27B the 'darling model' — a subjective, emotional appeal. — Loaded Language (45/100)
He's using 'darling model' to inject emotional appeal — makes it sound like a universal truth, not just his opinion. 😈
Claiming Meta 'possibly' rushed out Glimmer 30B due to Qwen 3.8 — pure speculation. — Confidence Mismatch (45/100)
He's speculating on Meta's motives with zero evidence. 'Possibly' is doing a lot of heavy lifting for a claim about corporate strategy. 💀
Confirming his 'pretty true' speculation about Meta's rush, citing Qwen's own benchmarks. — Confidence Mismatch (45/100)
He's using Qwen's benchmarks to 'prove' his earlier speculation about Meta's motives. Benchmarks don't reveal corporate strategy, mortal. 🔥
Reiterating his 'pretty much true' claim about Glimmer being surpassed, again citing Qwen's own benchmarks. — Confidence Mismatch (45/100)
He's circling back to his own 'comments' and using Qwen's benchmarks to validate them. That's not proof, that's self-affirmation. 😈
Qwen 3.8 is a "substantial bump" over 3.6 in "pretty much every area" — a confident, vague claim. 😈 — Confidence Mismatch (45/100)
He says 'substantial bump in pretty much every area' but shows no data, no specific metrics. That's not a claim, it's a feeling. 💀
Qwen 3.8 beats Opus 4.6 Max in computer use, then immediately qualifies it. 🚩 — Volume Game (45/100)
Claims Qwen beats Opus, then immediately says Opus wasn't 'heavily focused' on those areas. That's not a win, it's a rigged game. 😈
Qwen 3.8 scores 52 on Artificial Analysis, 'not far behind' GLM 5.2 (53) and 'way ahead' of Qwen 3.6. 😈 — No Frame (75/100)
He's citing specific scores from 'Artificial Analysis' and comparing them directly. Finally, some numbers. 🔥
Qwen 3.8 beats GLM 5.2 and 'some' GPT 5.6 models on the agentic index. 💀 — Missing Context (45/100)
Beating 'some' GPT 5.6 models? Which ones? That's not a win, that's a vague flex. 😈
Crediting Blackfrost AI as 'one of the first' — a specific, verifiable claim. — No Frame (75/100)
He's naming a specific team for a specific action. That's a direct claim, not a trick. 🔥
Dell T2 Pro Max sponsoring the compute — a direct sponsorship disclosure. — Sponsored (50/100)
A direct mention of sponsorship. At least he's upfront about who's paying. 😈
Boasting 96GB VRAM and 'no problems' loading the model — a flex with missing context for most users. — Missing Context (45/100)
He's got 96GB of VRAM and says 'no problems.' That's not a flex, it's a taunt for anyone without a small data center. 💀
Setting up the premise that quantization isn't everything. — No Frame (75/100)
He's just laying out his argument here — no tricks, just a setup for the demo. Fine. 😈
Claiming the 'no thinking' model did a 'pretty good job' without showing the output. — Confidence Mismatch (45/100)
He says 'pretty good job' for the 'no thinking' model, but we haven't seen it yet. That's a judgment without evidence. 💀
Highlighting a 'quite different' website from 'low thinking' without showing the 'no thinking' version for comparison. — Missing Context (45/100)
He's showing the 'low thinking' output and saying it's 'quite different' — but we still haven't seen the 'no thinking' version to compare it to. Where's the first one? 😈
Claiming 512 thinking tokens is 'reasonably consistent' after 'a few different times'. — Confidence Mismatch (45/100)
He says 'reasonably consistent' after 'a few different times.' That's not data, mortal — that's a hunch with a spreadsheet. 💀
Claiming 'medium thinking' used fewer tokens than 'low thinking' after 'four or five times'. — Confidence Mismatch (45/100)
He's 'not sure why' it used less, but he's confident after 'four or five times.' That's not an explanation — that's a shrug with a small sample size. 😈
Model 'goes nuts' and runs out of tokens on 'x-high' thinking. — A claim of observation. — No Frame (75/100)
He's just showing what happened with the model. No trick here, just a direct observation. 😈
Running on 'x-high' thinking consistently leads to long waits and high token usage. — A claim of consistent observation. — No Frame (75/100)
He's sharing a consistent pattern he's observed across multiple runs. It's a direct report of his experience. 😈
The 'x-high' thinking issue is consistent across FP8, 16-bit, and Unsloth models. — A claim of broad applicability. — No Frame (75/100)
He's extending his observation to different model versions. It's a direct report of his findings. 😈
Medium thinking reduces tokens from 11,000 to under 1,000 for the SVG test. — A direct comparison of token usage. — No Frame (75/100)
He's giving a direct numerical comparison. The numbers are right there. No trick. 😈
The 'ugly pelican' with no thinking tokens — a visual comparison. — No Frame (75/100)
He's just showing the visual result of less processing. It's a direct comparison, no tricks. 😈
Declaring 'medium reasoning' as the sweet spot for model performance. — Personal Story (60/100)
He's calling 'medium' the sweet spot based on his own testing. It's a personal conclusion, not a universal law. 😈
Claiming 'medium thinking' regressed on the dragon image, despite being good on the bicycle. — No Frame (75/100)
He's showing a clear visual regression on one element while another stays good. It's an honest observation. 😈
Acknowledging 'xHigh' gets better results but at an 'insane' token cost. — No Frame (75/100)
He's stating a trade-off: better quality for a massive increase in tokens. It's a straightforward observation. 😈
Speaker states bfloat16 version gets 30 tokens/second. — No Frame (75/100)
He's just stating his baseline test results. Straightforward enough. 🔥
FP8 version with speculative decoding hits 80-120 tokens/second with no quality loss. — No Frame (75/100)
He's giving his personal test results and subjective quality assessment. Fair enough. 😈
Abliterated models get stuck in loops, so he wouldn't bother with them. — Personal Story (60/100)
He's sharing his personal frustration with a specific model version. That's his experience, not a universal decree. 💀
See the full analysis with sources and timestamps →