Nvidia, You’re Late. World’s First 128GB LLM Mini Is Here!
Credibility score: 73/100 — Mostly Credible. Mixed credibility - some claims are solid, others need verification.
BSmeter analyzed "Nvidia, You’re Late. World’s First 128GB LLM Mini Is Here!" and rated it 73/100 for credibility (a BS score of 27/100 — mostly credible), on 2026-04-27. Its weakest claim — "ROCm not available yet for this setup" — scored 45/100 and was flagged as dubious. 30 claims were checked against the video transcript. Scores are produced by BSmeter's AI analysis of the transcript, not independent human verification.
Claims analyzed
Nvidia's DJX Spark is still weeks away, always 'coming soon' — Opinion (50/100)
Sassy shade on the eternal 'coming soon' — fair roast on delays 😏
GMKtec Evo X2 has 128GB unified RAM for LLMs, half price of DJX Spark — Solid (85/100)
👌💯
Dual PCIe Gen4 SSDs up to 16TB, Wi-Fi 7, shareable memory via static partitioning — Solid (75/100)
Checks out 👌
Testing Asus ROG Flow Z13 with 128GB chip — Solid (80/100)
👌
LM Studio offloads all 64/64 layers of Q4 Quinn 32B to GPU — Solid (75/100)
✅
Static partitioning: copies from system to GPU memory, unlike Apple's unified — Verified (90/100)
Spot on 👌
Quinn 32B Q4 is 18.4GB, gets 10.58 tokens/sec — OK (65/100)
Plausible benchmark ⚠️
Offloading LLM layers to CPU makes it very slow — Solid (85/100)
👌✅
M4 Mac Mini at 31W, 19.94 tokens/sec — OK (60/100)
Specific benchmark shown — no contradiction but no public match either ⚠️
Apple M4 chip has 120 GB/s memory bandwidth — Verified (95/100)
Spot on number ✓
Chat LLM Teams sponsors with model routing — Sponsored (50/100)
Full-on ad read for Chat LLM Teams — $10/mo dashboard 🚀
Measured 96 GB/s triad on M4 STREAM benchmark — Personal Story (70/100)
Plausible real-world measurement ⚠️
DeepSeek R1-Distill-Qwen-7B Q4 is ~4 GB — Solid (80/100)
👌
Chat LLM features PDF chat, humanizer, agents — Sponsored (50/100)
Continuing the pitch — agents, humanizer, all-in-one AI suite 💼
Llama CPP compiles for ROCm and Vulkan — Solid (85/100)
✅
ROCm not available yet for this setup — Dubious (45/100)
ROCm 'not available'? Tell that to AMD's 2026 docs 🚩
Sources: Compatibility matrix - ROCm Documentation - AMD, Limitations and recommended settings — Use ROCm on Radeon GPUs, Installation prerequisites — ROCm installation (Linux) - ROCm Documentation - AMD
Vulkan Llama.cpp: 935 PP, 46 TG tokens/sec — OK (65/100)
Plausible benchmark numbers for Vulkan on beefy AMD hardware ⚠️
M4 Mac Mini gets 228 PP512, 21 tokens/sec on Llama Bench — Personal Story (70/100)
Dude's own benchmarks — fair play 👌
M4: 120GB/s, M4 Pro: 273GB/s, M4 Max: 540GB/s, M3 Ultra: 819GB/s memory bandwidth — Verified (100/100)
✅
Nvidia DGX Spark has 279 GB/s memory bandwidth — Solid (80/100)
👌
This machine's bandwidth at 279 GB/s, higher than M4 Pro — Solid (80/100)
👌
M4: 96/120 GB/s Stream; M4 Pro: 209/273 GB/s — Solid (85/100)
✅
This machine's Stream benchmark: 120 GB/s vs expected 256 — OK (65/100)
Low real-world vs spec — common but underwhelming ⚠️
Gemma3-1B: ~160 t/s; Llama3.2-1B: 130 t/s on this machine — OK (60/100)
Self-reported benchmarks — impressive if real 📈
Gemma 3 12B: 23.9 t/s, Qwen 2.5 32B: 10.8 t/s, Llama 3 70B Q4: 5 t/s — OK (65/100)
Benchmark numbers look real but Llama 3 '370B' is sus — it's 70B 😂⚠️
Knuckbox Evo X2 desktop beats Flow Z13 laptop in most LLM benchmarks — Dubious (45/100)
'Knuckbox' or 'Nugbox'? And no hits on that exact desktop model 🚩🤔
M4 Max: 49 t/s Gemma 3 12B; Nvidia Spark Blackwell has 278 bandwidth — Solid (80/100)
👌 Numbers align with unified memory advantages
Static memory partitioning is stable but wastes memory and lacks flexibility for modern PCs — Solid (85/100)
✅
AMD at Computex promised better ROCm support for LLM chips like this one — Solid (80/100)
👌
Michael Larabel got ROCm working on this chip but unstable with seg faults — OK (65/100)
Specific tester named but sounds like fragile hack job ⚠️
See the full analysis with sources and timestamps →