Ternary Bonsai 27B benchmarked and tested vs Qwen 27B - 16GB Local LLM setup
Credibility score: 51/100 — Mixed Credibility. Several questionable claims detected. Watch with healthy skepticism.
Claims analyzed
Ternary weights (-1/0/+1) make inference much faster than floats — No Frame (75/100)
Straight technical description of the actual mechanism — no trick.
Drafters work like MTP but sit outside the model — Missing Context (45/100)
Compares drafters to MTP then immediately notes the key difference — the comparison is loose.
Exact model sizes presented as fact — no frame — No Frame (75/100)
Just reads the file sizes straight — no trick, no hype.
Lists exact model sizes like they're settled facts. — No Frame (75/100)
Straight numbers pulled from files. No framing tricks here.
Calls Q4KM 'standard' precision for their model — Missing Context (45/100)
Labels their choice 'standard' — but Q4KM is one of several common options, not the default.
Equates MTP and drafter as 'kind of similar' — Confidence Mismatch (45/100)
Calls two different techniques 'kind of similar' with zero technical comparison shown.
Lists benchmarks without explaining what 'various depths' means — Missing Context (45/100)
'Various depths' sounds rigorous — but no definition or methodology given yet.
Calls agency benchmark 'a good test' without evidence — Confidence Mismatch (45/100)
Labels it 'a good test' — confidence with zero validation shown.
Calls dungeon crawler 'really good' for algorithm testing — Confidence Mismatch (45/100)
'Really good evaluation' — asserted, never justified.
Introduces new Blender test via MCP — no prior validation — Missing Context (45/100)
New test announced as if it's ready — but it's literally the first run.
Chains Godot success to Blender success — conditional test — Missing Context (45/100)
Test only happens if Blender works — so failure is already baked in.
Ends with 16 GB VRAM system spec — straightforward — No Frame (75/100)
Just states the hardware — no exaggeration.
Claims results generalize to 'your' hardware — zero shared data — Missing Context (45/100)
Promises performance on 'your' rig while admitting zero shared benchmarks — classic extrapolation without evidence.
Q4 drafter 22 t/s vs BF16 drafter 1 t/s — drafter shouldn't change output — Missing Context (45/100)
Calls the 1 t/s BF16 result 'interesting' while the drafter comment shows why it shouldn't happen at all.
BF16 'one token a second' vs Q4 '22' — no mention of why the gap exists — Missing Context (45/100)
Presents 22× speed difference as model choice when memory offloading is the real culprit.
Calls 22 t/s 'quite significant' while skipping that base model is offloaded — Missing Context (45/100)
Hails the drafter's 22 t/s as a big win — quietly compares it to a base model that's already crippled by RAM swapping.
22 t/s is double my normal 11 t/s on 27B — Missing Context (45/100)
Calls it double the speed — never says what model or setup gave the 11 t/s baseline.
22 t/s on 27B is 'quite significant' — confidence with no baseline given — Confidence Mismatch (45/100)
Calls 22 t/s 'quite significant' without saying what normal 27B speeds look like on the same rig.
Drafter "helped" at 256K — admits results mixed, no conclusion — Confidence Mismatch (45/100)
Calls it "definitely helped" then immediately says "mixed" and "not sure what conclusion." Classic pivot without the receipts.
BF16 drafter suddenly makes Bonsai find everything across 256K — Confidence Mismatch (45/100)
Calls it 'the story changes' like the drafter fixed the depth problem — zero data on why it worked or if it will again.
Claims Bonsai faster and better score than base model — Missing Context (45/100)
Calls it 'better score' after saying base got almost perfect and only failed on #13 — the exact same spot most models fail.
Base model 'almost perfect' except on #13 where 'most models fail' — Missing Context (45/100)
Calls base 'almost perfect' while admitting it failed the same question as everyone else — the bar is doing all the work.
Bonsai beats base on both speed and score — No Frame (75/100)
Straight comparison with numbers attached. No hidden framing here.
Presents 90% on answered questions as the real number while burying the 24 skips — Volume Game (45/100)
Loudly drops the 90% figure, then quietly adds it skipped 24 questions. The 77% pass rate gets mentioned once and never revisited.
77% pass rate, 90% on answered questions, 85% answer rate — No Frame (75/100)
Breaks out the three different metrics clearly. Numbers, not narrative.
Blames context burning for skips, implying Bonsai is smarter for answering more — Missing Context (45/100)
Attributes Bonsai's fewer skips to better thinking instead of the drafter setup that might be forcing shorter responses.
Bonsai edges out base on answered-pass rate and answer rate — No Frame (75/100)
Direct side-by-side without inflating the margin. Clean.
Bonsai answers twice as many questions as base Qwen — No Frame (75/100)
Clean head-to-head on refusal count. No spin added.
Claims Bonsai 2h faster + smarter than base on same test — Missing Context (45/100)
Speed win is real — but the 'better intelligence' part is just declared, not measured.
Blames model when harness fails — false dilemma — False Dilemma (20/100)
Presents only two options — harness or model — while ignoring the third: the setup itself.
See the full analysis with sources and timestamps →