Ternary Bonsai 27B benchmarked and tested vs Qwen 27B - 16GB Local LLM setup
Credibility score: 51/100 — Mixed Credibility. Several questionable claims detected. Watch with healthy skepticism.
BSmeter analyzed "Ternary Bonsai 27B benchmarked and tested vs Qwen 27B - 16GB Local LLM setup" and rated it 51/100 for credibility (a BS score of 49/100 — mixed credibility), on 2026-07-20. Its weakest claim — "Drafters work like MTP but sit outside the model" — scored 45/100 and was flagged as missing context. 30 claims were checked against the video transcript. Scores are produced by BSmeter's AI analysis of the transcript, not independent human verification.
Claims analyzed
Ternary weights (-1/0/+1) make inference much faster than floats — No Frame (75/100)
Straight technical description of the actual mechanism — no trick.
Drafters work like MTP but sit outside the model — Missing Context (45/100)
Compares drafters to MTP then immediately notes the key difference — the comparison is loose.
Exact model sizes presented as fact — no frame — No Frame (75/100)
Just reads the file sizes straight — no trick, no hype.
Lists exact model sizes like they're settled facts. — No Frame (75/100)
Straight numbers pulled from files. No framing tricks here.
Calls Q4KM 'standard' precision for their model — Missing Context (45/100)
Labels their choice 'standard' — but Q4KM is one of several common options, not the default.
Equates MTP and drafter as 'kind of similar' — Confidence Mismatch (45/100)
Calls two different techniques 'kind of similar' with zero technical comparison shown.
Lists benchmarks without explaining what 'various depths' means — Missing Context (45/100)
'Various depths' sounds rigorous — but no definition or methodology given yet.
Calls agency benchmark 'a good test' without evidence — Confidence Mismatch (45/100)
Labels it 'a good test' — confidence with zero validation shown.
Calls dungeon crawler 'really good' for algorithm testing — Confidence Mismatch (45/100)
'Really good evaluation' — asserted, never justified.
Introduces new Blender test via MCP — no prior validation — Missing Context (45/100)
New test announced as if it's ready — but it's literally the first run.
Chains Godot success to Blender success — conditional test — Missing Context (45/100)
Test only happens if Blender works — so failure is already baked in.
Ends with 16 GB VRAM system spec — straightforward — No Frame (75/100)
Just states the hardware — no exaggeration.
Claims results generalize to 'your' hardware — zero shared data — Missing Context (45/100)
Promises performance on 'your' rig while admitting zero shared benchmarks — classic extrapolation without evidence.
Q4 drafter 22 t/s vs BF16 drafter 1 t/s — drafter shouldn't change output — Missing Context (45/100)
Calls the 1 t/s BF16 result 'interesting' while the drafter comment shows why it shouldn't happen at all.
BF16 'one token a second' vs Q4 '22' — no mention of why the gap exists — Missing Context (45/100)
Presents 22× speed difference as model choice when memory offloading is the real culprit.
Calls 22 t/s 'quite significant' while skipping that base model is offloaded — Missing Context (45/100)
Hails the drafter's 22 t/s as a big win — quietly compares it to a base model that's already crippled by RAM swapping.
22 t/s is double my normal 11 t/s on 27B — Missing Context (45/100)
Calls it double the speed — never says what model or setup gave the 11 t/s baseline.
22 t/s on 27B is 'quite significant' — confidence with no baseline given — Confidence Mismatch (45/100)
Calls 22 t/s 'quite significant' without saying what normal 27B speeds look like on the same rig.
Drafter "helped" at 256K — admits results mixed, no conclusion — Confidence Mismatch (45/100)
Calls it "definitely helped" then immediately says "mixed" and "not sure what conclusion." Classic pivot without the receipts.
BF16 drafter suddenly makes Bonsai find everything across 256K — Confidence Mismatch (45/100)
Calls it 'the story changes' like the drafter fixed the depth problem — zero data on why it worked or if it will again.
Base model 'almost perfect' except on #13 where 'most models fail' — Missing Context (45/100)
Calls base 'almost perfect' while admitting it failed the same question as everyone else — the bar is doing all the work.
Bonsai beats base on both speed and score — No Frame (75/100)
Straight comparison with numbers attached. No hidden framing here.
Claims Bonsai faster and better score than base model — Missing Context (45/100)
Calls it 'better score' after saying base got almost perfect and only failed on #13 — the exact same spot most models fail.
77% pass rate, 90% on answered questions, 85% answer rate — No Frame (75/100)
Breaks out the three different metrics clearly. Numbers, not narrative.
Presents 90% on answered questions as the real number while burying the 24 skips — Volume Game (45/100)
Loudly drops the 90% figure, then quietly adds it skipped 24 questions. The 77% pass rate gets mentioned once and never revisited.
Bonsai edges out base on answered-pass rate and answer rate — No Frame (75/100)
Direct side-by-side without inflating the margin. Clean.
Blames context burning for skips, implying Bonsai is smarter for answering more — Missing Context (45/100)
Attributes Bonsai's fewer skips to better thinking instead of the drafter setup that might be forcing shorter responses.
Bonsai answers twice as many questions as base Qwen — No Frame (75/100)
Clean head-to-head on refusal count. No spin added.
Claims Bonsai 2h faster + smarter than base on same test — Missing Context (45/100)
Speed win is real — but the 'better intelligence' part is just declared, not measured.
Blames model for file write failure without testing harness first — Missing Context (45/100)
Assumes model fault before checking if the harness itself blocks writes — classic premature diagnosis.
See the full analysis with sources and timestamps →