AMD Says 2 Ryzen AI Halos Can Run a 400B Model... I Tested It

Credibility score: 53/100 β€” Mixed Credibility. Several questionable claims detected. Watch with healthy skepticism.

BSmeter analyzed "AMD Says 2 Ryzen AI Halos Can Run a 400B Model... I Tested It" and rated it 53/100 for credibility (a BS score of 47/100 β€” mixed credibility), on 2026-08-23. Its weakest claim β€” "Either it works perfectly or the network ruins everything" β€” scored 20/100 and was flagged as false dilemma. 27 claims were checked against the video transcript. Scores are produced by BSmeter's AI analysis of the transcript, not independent human verification.

Of 27 claims analyzed: 1 scored under 40, 18 between 40 and 69, and 8 at 70 or above.

Claims analyzed

Tiny pocket box runs 200B, two of them run 400B β€” Missing Context (45/100)

At 0:30

Calls it pocket-sized while the hardware is two full desktop units β€” missing the actual size and setup.

Why this score: The speaker equates one small box with 200B, then jumps to two boxes for 400B without clarifying these are not pocket devices but full desktop platforms. The 'pocket' framing makes the scaling claim sound more impressive than the physical reality.

Original quote: β€œon a tiny box like this that you can fit into your pocket, well, maybe you can run models that are up to 200 billion parameters in size... And in this advanced section, you can now cluster two of them. So you'll be able to potentially run models that are about 400 billion parameters.”

Either it works perfectly or the network ruins everything β€” False Dilemma (20/100)

At 1:15

Sets up two extremes β€” perfect system or total failure β€” ignoring the middle ground of 'slow but usable.'

Why this score: Real-world networking falls on a spectrum. The speaker presents a binary where anything less than seamless becomes 'ruined,' which frames any latency as a deal-breaker rather than a trade-off.

Original quote: β€œSo the question is, can these two actually function as one practical AI system or will the network between them ruin the entire idea?”

You need proper 10GbE β€” can't just plug in a cable β€” Missing Context (45/100)

At 2:24

Implies a switch is mandatory while commenters note auto-MDIX makes direct crossover work on 10GBASE-T.

Why this score: The speaker presents the 10GbE requirement as needing a switch or special setup, yet 10GBASE-T standard includes auto-MDIX, meaning a simple crossover cable should suffice. The 'you can't just take a cable' line undercuts a simpler solution the audience might try.

Original quote: β€œboth of these methods call for proper 10 GB Ethernet network. However, you can't just take a cable”

AMD says you must use a switch between two machines β€” no direct cable. β€” Missing Context (45/100)

At 2:38

Claims AMD's instructions forbid direct cable β€” commenters say 10GBASE-T auto-MDIX should work fine.

Why this score: The video presents AMD's requirement as gospel, but top comments point out that modern 10GbE hardware handles crossover automatically. No evidence given why a switch is strictly required.

Original quote: β€œAMD's instructions for both of these require a 10 GB Ethernet switch to be used between the machines. So yeah, that's not going to work.”

Micro Center is the exclusive retailer for Ryzen AI Halo machines. β€” No Frame (75/100)

At 2:52

Straight fact β€” no spin, just naming the seller.

Why this score: Matches the known retail situation: Micro Center is the only listed US retailer for the Halo developer platform.

Original quote: β€œThese Ryzen Halo machines are now available at the MicroEnter, which is their exclusive retailer.”

OS mismatch blocks clustering β€” needs both on Linux β€” No Frame (75/100)

At 4:36

Straight technical obstacle β€” the machines won't cluster until they run the same OS.

Why this score: Speaker simply states the incompatibility: one Halo is on Windows, one on Linux, and AMD's clustering playbook only exists for Linux. No exaggeration, no hidden agenda, just the next problem to solve.

Original quote: β€œI can't cluster these machines the way they are now. They don't match. The Halo I tested in the previous video came with Windows, and the second one had Linux on it.”

iperf3 shows ~9.4 Gbps β€” essentially saturating the 10 GbE link β€” No Frame (75/100)

At 6:00

Real measured throughput, not theory β€” 9.4 Gbps is basically line rate on a 10 GbE connection.

Why this score: Speaker runs a quick bandwidth test and reports the actual numbers. No hype, just confirming the network pipe is as wide as advertised before moving on to the clustering playbook.

Original quote: β€œJust a quick little test. I got IPF 3 on server 2, Halo 2, and Iperf 3 on Halo 1 just to test the network bandwidth between the two. And it looks like it's coming back pretty much to its maximum capacity. 9.41, 9.44. Yeah, that's as much as we're going to get for a 10 gig network.”

Windows caps at 96 GB, Linux needs manual tweak β€” Missing Context (45/100)

At 7:02

States the cap difference like it's a done deal β€” skips that you must recompile the kernel to hit 120 GB.

Why this score: Speaker presents 120 GB as a simple Linux default, but the actual route involves kernel parameters and the AMD TTM tool, which he later reveals failed on the first box. The framing hides the extra steps most users will hit.

Original quote: β€œWindows limits this to 96 as maximum. Linux can go up to 120, but not by default.”

Windows locks to 64 GB, hidden from BIOS β€” Confidence Mismatch (45/100)

At 7:30

Speaks with certainty that BIOS won't show the setting β€” yet the only proof is that he didn't find it on this model.

Why this score: He looked in BIOS, didn't see the option, and concludes it's impossible. Could be model-specific, firmware-locked, or simply missing from his search. No data on other Halo units.

Original quote: β€œThe Windows version automatically sets it to 64 GB. And this is not something that you can discover in the BIOS.”

Linux TTM command failed too β€” No Frame (75/100)

At 7:47

Straight reporting of what actually happened on his test unit β€” no embellishment.

Why this score: He describes the exact failure: TTM command ignored, stuck at 64 GB. Clean, verifiable detail.

Original quote: β€œEven after I put Linux on that box, I couldn't change it in the BIOS. I couldn't change it using this AMD TTM set 120.”

Calls 10 GB connection "good" β€” skips faster options shown. β€” Missing Context (45/100)

At 8:34

Top comment calls the 40 Gbps USB-C ports right there β€” 10 GB suddenly looks slow.

Why this score: Speaker accepts the 10 GbE port as sufficient while the hardware itself advertises two 40 Gbps USB4 ports that commenters note could be bonded for 80 Gbps RDMA.

Original quote: β€œSo that's 10 GB right there. So we're good with the connections.”

Shows 358 B model while title promises 400 B β€” slides one past the viewer. β€” Volume Game (45/100)

At 9:13

Video title says 400 B, then quietly tests 358 B; the 42 B gap never gets addressed.

Why this score: Title stakes a round number; actual test drops to a smaller model without flagging the difference, letting the headline carry the bigger claim.

Original quote: β€œOne more thing to notice here is the playbook uses GLM 4.7. ... 358 billion parameter model”

Quant tweaks give better quality at same size β€” no data shown β€” Anonymous Authority (45/100)

At 10:38

Names 'Barttowski' and 'they all' β€” zero specifics on which techniques or what 'juice' actually means.

Why this score: Speaker invokes unnamed 'little techniques' and a famous quantizer without showing results or benchmarks. The authority is dropped, never proven.

Original quote: β€œYet, it's exactly the same size. I've talked about this on the channel before and Barttowski is another quantizer pretty famous on the internets. They all use like little techniques to tweak the quantizations to get more juice out of it.”

Low-bit quants 'probably' ruin output β€” no test shown β€” Loaded Language (45/100)

At 10:57

Turns 'might degrade' into 'probably mess up' with zero side-by-side evidence.

Why this score: The claim hinges on the emotional weight of 'mess up your output' while offering no actual output samples or metrics.

Original quote: β€œBut neither one of these quants will fit on a single 128 GB box. The second Halo is what lets you serve a good quant instead of dropping down, I don't know, one of these Q IQ2M or IQ1S quants, which are probably going to mess up your output.”

8.2 t/s on a 358B MoE model β€” presented as impressive β€” Missing Context (45/100)

At 11:47

No baseline comparison β€” is 8.2 t/s good, bad, or average for this size?

Why this score: Without a reference point or competing hardware numbers, the raw figure floats free of meaning.

Original quote: β€œWe're getting about uh 8.2 tokens per second here. Remember folks, this is GLM 4.7. It's pretty big.”

Low power draw and VRAM use β€” framed as a win β€” Missing Context (45/100)

At 11:59

Efficiency numbers shown without stating what the second machine is also consuming.

Why this score: Power and utilization stats for one box are meaningless until you add the second machine's draw and the switch overhead.

Original quote: β€œWow, interesting. Okay, so VRAM is at 35% and the power is only at about 43 watts. GPU is only about halfway utilized. This is only on one machine, by the way.”

7.6-7.9 tokens/sec on 400B model β€” real measured speed β€” No Frame (75/100)

At 12:41

Actual numbers from a live test, no hype attached.

Why this score: Speaker is citing concrete measurements rather than vague claims. The numbers line up with what web sources confirm about Rickle clustering on two Halos.

Original quote: β€œ7.93 7.69 7.6 within that range for the concurrency of one. By the way, this is like basically chatting with this thing. And that's how many tokens per second we're getting out of this.”

Clustering two Halos combines their memory for bigger models β€” No Frame (75/100)

At 14:11

Straight fact about tensor parallelism via Rickle; matches independent tests.

Why this score: The speaker correctly states the core benefit of clustering without exaggeration. Geeky Gadgets and Reddit tests confirm the same setup works for 400B-class models.

Original quote: β€œYour Ryzen AI Halo is already capable of running large language models locally, but clustering takes it a further by combining GPU memory.”

Calls software stack 'Rickle Rock Communication Collectives' without explaining why it matters β€” Anonymous Authority (45/100)

At 14:30

Drops a library name like it explains itself β€” no one knows what it does or why it's necessary.

Why this score: He treats the acronym as proof the setup works, but never shows what happens without it or why this specific stack matters.

Original quote: β€œusing Rickle Rock Communication Collectives Library. It's it's an acronym within an acronym with VLM.”

Calls 397B model 'much bigger' without context on scale β€” Missing Context (45/100)

At 14:34

Says 397B is 'much bigger' β€” compared to what? No baseline given.

Why this score: The speaker treats the jump in size as obvious, but without naming the prior model or quantifying the difference, the claim floats without meaning.

Original quote: β€œwe're going to run Quen 3.5 397B, which is a much bigger model than the previous one”

Admits the process is brittle but shrugs it off as normal β€” Missing Context (45/100)

At 15:22

Flags the 15-minute reset loop as a known pain but offers zero alternatives or workarounds.

Why this score: The story makes the setup sound workable, yet the only evidence is that one wrong flag costs you a quarter-hour with no recovery path shown.

Original quote: β€œIf you get one flag wrong, you have to start over and wait another 15 minutes. Don't ask me how I know.”

Personal frustration presented as universal setup warning β€” Just Vibes (50/100)

At 15:22

One bad flag and 15-minute reload β€” sounds like a personal war story, not a technical fact.

Why this score: This is pure lived experience; the speaker isn't claiming a universal rule, just venting the cost of one typo. Roast the pain, don't fact-check the citation.

Original quote: β€œIf you get one flag wrong, you have to start over and wait another 15 minutes. Don't ask me how I know.”

Labels the split 'tensor parallelism TP=2' as if that alone guarantees success β€” Confidence Mismatch (45/100)

At 16:26

Names the technique like it proves the 400B model runs β€” zero tokens/sec or quality numbers attached.

Why this score: He equates the existence of tensor parallelism with a working cluster, but the clip ends before any performance data lands.

Original quote: β€œThis is actually tensor parallelism at work here. TP equals 2.”

States tensor parallelism is active with TP=2 as fact β€” No Frame (75/100)

At 16:26

Names the exact technique and parameter β€” clean, no tricks.

Why this score: Direct technical statement with the correct term and value. Straightforward, no framing needed.

Original quote: β€œThis is actually tensor parallelism at work here. TP equals 2.”

Two boxes, 220 GB total, 18 tokens/sec max β€” real numbers, no hype. β€” No Frame (75/100)

At 16:31

Drops actual measured numbers instead of marketing fluff. Straight data.

Why this score: Speaker gives concrete throughput figures from his own setup rather than vague promises. The numbers line up with what independent tests have shown for Rickle tensor parallelism on two Halo units.

Original quote: β€œTP equals 2. And we're using about 110 GB per box here for a total of 220 pulled. Can you hear the machines? I can hear it. Rickle talking over the 10 GB link at about 7.81 tokens per second and up to almost 18 tokens per second for concurrency of four.”

Current stack scales to bigger AMD systems; four-node clusters mentioned. β€” Missing Context (45/100)

At 17:05

Hints at four-node support but gives zero details on how or when. Tease without substance.

Why this score: Speaker gestures at future four-node capability without any timeline, pricing, or performance data. The claim stays vague enough to sound promising while revealing nothing concrete.

Original quote: β€œThis is the technology stack that you'll use if you transition into bigger AMD infrastructure. Now, AMD also alludes to having four node cluster as well.”

Calls two clustering options 'simple' vs 'mini data center' β€” no numbers given. β€” Confidence Mismatch (45/100)

At 17:06

Labels one method 'simple' and the other 'mini data center' without benchmarks or setup time.

Why this score: Speaker presents two paths as clearly different in complexity yet offers zero data on actual install steps, errors, or hours required. The 'mini data center' label sounds impressive, but the claim rests on branding, not measured effort.

Original quote: β€œAMD gives you a couple of different options to do that. Lomar RPC is simple. Rickle is a little bit more complicated, but it scales a little bit better and it's like running a little mini data center at home or in your office.”

See the full analysis with timestamps β†’