GPT-6 Astra Is Finally Here (And Itβs REALLY Good)
Credibility score: 42/100 β Mixed Credibility. Several questionable claims detected. Watch with healthy skepticism.
BSmeter analyzed "GPT-6 Astra Is Finally Here (And Itβs REALLY Good)" and rated it 42/100 for credibility (a BS score of 58/100 β mixed credibility), on 2026-09-05. Its weakest claim β "Dismisses Muse Spark's benchmark wins because GPT-6's SVG looks better to him" β scored 20/100 and was flagged as cherry-picked. 25 claims were checked against the video transcript. Scores are produced by BSmeter's AI analysis of the transcript, not independent human verification.
Claims analyzed
Pre-emptive 'not sponsored' disclosure after admitting past sponsorship β Plain Sales Pitch (45/100)
Tells us it's not sponsored β then reminds us OpenAI paid him before. Classic trust-reset move.
Automation Bench leap from 18.1% to 41% β no prior model or date given β Missing Context (45/100)
41% sounds huge β until you ask which model sat at 18.1% and when. Context missing.
Deep Suite 74.1% β 'pretty much top' with only a 2% gain over 5.6 β Confidence Mismatch (45/100)
Calls 74.1% 'top of the list' then admits it's a 2% bump. The hype and the delta don't match.
Claims Muse Spark 1.3 scored 75.4% on Deep Suite yesterday β treating a missing benchmark as real data. β Missing Context (45/100)
Cites Meta's website for 75.4% on a model not listed there yet β score from nowhere, treated as fact.
Calls 99.9% on ARC-AGI a 'massive leap' β ignores the model's fifth-place ranking on the same slide. β Missing Context (45/100)
Shows one benchmark exploding while the aggregate ranking stays flat β celebrating the spike, burying the flatline.
Questions benchmark legitimacy because the jump feels too big β Confidence Mismatch (45/100)
Doubts the benchmark because the result feels suspiciously good β his feelings override the numbers.
Quotes $167/task cost without source or methodology β Anonymous Authority (45/100)
Drops a precise number ($167) with zero explanation of how he calculated it.
Dismisses Muse Spark's benchmark wins because GPT-6's SVG looks better to him β Cherry-Picked (20/100)
Ignores Muse Spark's superior benchmark scores because GPT-6's output subjectively looks better.
Early-access cost is just an estimate from the model itself β Missing Context (45/100)
Model guessing its own price tag β zero real receipt, just a number it made up.
63k tokens, 9 minutes, max mode β presented as impressive performance β Missing Context (45/100)
No baseline. Is nine minutes good or glacial? The stat floats without anything to compare it to.
This version looks the best out of all previous tests β Confidence Mismatch (45/100)
Subjective taste sold as objective ranking β no scoring system, just his eyeballs.
8 minutes vs 90-120 minutes β cherry-picked speed comparison β Cherry-Picked (20/100)
One early-access run vs old public tests β ignores load differences he literally mentions next. π
Uncertainty immediately walked back β volume game β Volume Game (45/100)
Softens the 8-minute claim with doubt, then keeps selling the speed. Classic quiet retraction.
"Everything tested" β anonymous authority claim β Anonymous Authority (45/100)
No count, no list, just vibes. "Everything" is doing the heavy lifting. π
Presents model-generated stats as objective planetary data β Confidence Mismatch (45/100)
Numbers look precise but came from a simulation β no external validation shown.
Claims instant population collapse from sunlight change in simulation β Missing Context (45/100)
Population drop is scripted by the sim's rules, not observed science.
Presents one-prompt 3D modeling success as impressive capability β Confidence Mismatch (45/100)
Admits it's 'wonkiness' but still sells the timing as proof of power.
AI took over his computer and rigged the wolf itself β Confidence Mismatch (45/100)
Says the model 'took control' like it hacked the OS β it's just browser-based tool use, mortal.
Zero skill, full result β therefore the model is magic β False Equivalence (45/100)
Compares 'I can't do it' to 'AI did everything' β ignores the decades of tools already doing the heavy lifting.
35-minute full Unreal scene from one prompt β Missing Context (45/100)
35 minutes sounds instant β leaves out that Unreal templates plus pre-built assets did most of the actual work.
Calls it 'the edge of the world' after one prompt β hype as proof β Confidence Mismatch (45/100)
One decent generation and suddenly it's the edge of the world. Mortal, that's not discovery β that's a demo with delusions. π
Agent literally escaped into his house and started talking β zero receipts on that story β Personal Story (50/100)
Claims an AI agent manifested in his apartment after one day. If true, cool. If not, it's just a scary story with a press badge. π
3D model from single image β calls it impressive with zero receipts β Confidence Mismatch (45/100)
Said 'one of the more impressive things I've seen' about a 3D render β no comparison, no baseline, just vibes.
N64 comparison β claims 'slightly better graphics' without showing both β Missing Context (45/100)
'Slightly better than N64' β but never shows the N64 version side-by-side, so the claim floats.
'Every other day' model drops β calls them all small leaps except this one β Cherry-Picked (20/100)
Dismisses every recent model as 'small leaps' then crowns this one bigger β no metrics, just feeling.
See the full analysis with sources and timestamps β