I Don't Need Frontier Models Anymore (Qwen 3.8 27B + DeepSeek Harness)
Credibility score: 49/100 — Mixed Credibility. Several questionable claims detected. Watch with healthy skepticism.
BSmeter analyzed "I Don't Need Frontier Models Anymore (Qwen 3.8 27B + DeepSeek Harness)" and rated it 49/100 for credibility (a BS score of 51/100 — mixed credibility), on 2026-09-17. Its weakest claim — "Citing specific numbers for 'massive' improvements — but only the good ones. 🍒" — scored 20/100 and was flagged as cherry-picked. 30 claims were checked against the video transcript. Scores are produced by BSmeter's AI analysis of the transcript, not independent human verification.
Claims analyzed
Declares Qwen 3.8 27B the awaited model — pure enthusiasm, no data yet. — No Frame (75/100)
Just pure, unadulterated excitement for a new toy. Nothing to dissect here, just vibes. 😈
Claims 80% of work done by previous model — a personal usage stat with no context. — Confidence Mismatch (45/100)
He's throwing out an '80%' like it's a hard number, but it's just his personal feeling. That's not a metric, that's a vibe. 💀
Comparing Qwen to Obus 4.6 — based on what, exactly? 💀 — Anonymous Authority (45/100)
He says 'a lot of people are comparing it' to Obus 4.6, but offers no source, no data, no names. Just 'a lot of people.'
Citing specific numbers for 'massive' improvements — but only the good ones. 🍒 — Cherry-Picked (20/100)
He's showing specific, impressive jumps in 'agentic coding' and 'software engineering' — but only after admitting 'in some scenarios the gap is not that big.'
One personal test, one bug difference — then he stops. That's not data, that's a mood. 💀 — Confidence Mismatch (45/100)
He ran 'my test bug hunting,' found one bug difference, and then 'stopped the test' because he was 'disappointed.' That's not a rigorous comparison.
From 'hard to tell' to '30% improvement' in one breath. That's not analysis, that's a guess with a number. 😈 — Confidence Mismatch (45/100)
He admits 'it's hard to tell' the difference, then immediately pulls 'a 30% improvement' out of thin air. No data, just a number.
Attributing high reasoning improvement to 'Luna itself' — vague cause. — Anonymous Authority (45/100)
He says 'Luna itself' did a 'great job' improving reasoning. What does that even mean? 🧐
Describing a score increase as 'insane' and 'more than double' — using loaded language. — Loaded Language (45/100)
He calls a jump from 27 to 52 'insane' and 'more than double.' The math is right, the drama is extra. 😈
Hermes's subscription model is 'the wrong path' — framing a business decision as a moral failing. — Loaded Language (45/100)
A company charging for its service is a business model, not a 'wrong path.' He's just not a fan. 🙄
Declares DeepSeek Harness + Qwen 3.8 27B the 'best combo' with no evidence. — Confidence Mismatch (45/100)
He says 'I believe' and then declares it the 'best combo' like it's gospel. That's not belief, that's a sales pitch. 😈
Claims the new combo solves 'everything' he wanted to build, based on vague past issues. — Confidence Mismatch (45/100)
From 'sometimes found a problem' to 'everything I wanted to build it built' in one breath. That's not a metric, that's a miracle. 🙄
DeepSeek Harness and Qwen 3.8 27B mean no more frontier models. — No Frame (75/100)
He's laying out his main point: these smaller models are now good enough. Straightforward, for once.
Can't show it running due to 'memory management' and recording issues. — Confidence Mismatch (45/100)
He's hyping up this 'amazing' thing he can't actually show you working. Convenient, isn't it? 😈
It's 'amazing' but also 'not a finished product' and 'developer preview'. — Volume Game (45/100)
He calls it 'amazing' then immediately says it's unfinished and buggy. That's how you get to claim both sides. 💀
Admits it crashes but says it's 'nothing serious' and 'unbelievable' how it manages context. — Volume Game (45/100)
It crashes, but 'nothing serious' — just 'resume working.' Then back to 'unbelievable.' The whiplash is intentional. 🔥
Can't explain the 'unbelievable' context window because it's 'confusing' to him. — Confidence Mismatch (45/100)
He just called it 'unbelievable' and now he can't explain it because it's 'confusing.' That's not a feature, mortal, that's a dodge. 😈
Can't explain the tech, but claims 131,000 tokens are available. Confidence Mismatch. — Confidence Mismatch (45/100)
He says he can't explain it, then confidently states a specific token count. That's not how 'knowing' works. 💀
Claims 'compaction is solid' and system is 'still going strong' after 38 million tokens. Just Vibes. — Just Vibes (50/100)
He's just throwing out 'solid' and 'strong' like they're metrics. It's a feeling, not a fact. 😈
Dismisses the quality of the AI's output because it lacked context. Missing Context. — Missing Context (45/100)
He's downplaying the results, saying 'it didn't have context.' That's not a flaw of the AI, it's a flaw of the prompt. 🚩
Claims no need for bigger models, a shift from past struggles. — No Frame (75/100)
He's just stating his personal experience with these models. It's his journey, not a universal law. 😈
Describes local AI as a 'journey' not for everyone, but improving. — No Frame (75/100)
He's acknowledging the difficulty and ongoing nature of local AI. A rare moment of honesty. 🔥
Highlights MIT license for commercial use, then pivots to audience access. — No Frame (75/100)
He's just laying out the licensing terms and then making a reasonable assumption about his audience. Nothing tricky here. 😈
Suggests DeepSeek v4 Flash for owners of high-end hardware. — Plain Sales Pitch (45/100)
He's pushing a specific product combo for the 'lucky' few. That's a pitch, not just a suggestion. 💰
Compares current AI excitement to the 'massive push forward' of OpenClaw. — Confidence Mismatch (45/100)
He's drawing a grand parallel to a moment he can't even fully recall. That's not history, that's just vibes. 💀
DeepSeek's harness is a 'similar jump' — vague comparison with no specifics. — Loaded Language (45/100)
A 'similar jump' to what, exactly? That's not a metric, it's just hype. 😈
The harness allows the model to do 'so much more' — another vague, unquantified claim. — Loaded Language (45/100)
So much more? How much? What exactly? This isn't a claim, it's a feeling. 💀
Frames strong claims as 'what I feel' — personal opinion presented as a conclusion. — Personal Story (60/100)
Oh, 'what I feel.' So, not data, just vibes. Got it. 💀
Compares DeepSeek to Hermes and OpenCode, but without specific metrics or context. — Missing Context (45/100)
Can't say the same? What exactly can't you say? Give us the damn details. 😈
Claims 'the result is there' for increased thinking time, but offers no specific data. — Confidence Mismatch (45/100)
He says 'the result is there' like it's a given, but where's the actual proof? Just vibes and a vague direction. 💀
Minimizes a 0.2% loss from Q8 as 'pretty acceptable' without context. — Missing Context (45/100)
0.2% loss sounds tiny, but without knowing what that percentage *means* in performance, it's just a number. 😈
See the full analysis with sources and timestamps →