We Tested $200 GPT-5.5 Pro on PhD Level Math

Credibility score: 63/100 — Mostly Credible. Mixed credibility - some claims are solid, others need verification.

BSmeter analyzed "We Tested $200 GPT-5.5 Pro on PhD Level Math" and rated it 63/100 for credibility (a BS score of 37/100 — mostly credible), on 2026-04-27. Its weakest claim — "OpenAI just released GPT 5.5" — scored 10/100 and was flagged as bs. 27 claims were checked against the video transcript. Scores are produced by BSmeter's AI analysis of the transcript, not independent human verification.

Of 27 claims analyzed: 1 scored under 40, 12 between 40 and 69, and 14 at 70 or above.

Claims analyzed

OpenAI just released GPT 5.5 — BS (10/100)

At 0:00

GPT-5.5? Bro where? OpenAI's at GPT-4o, this is pure fanfic 💀🔥

Why this score: Flat-out false claim. No evidence of GPT-5.5 release as of April 2026. - OpenAI's latest is GPT-4o/o1 series; GPT-5 rumors exist but *no official 5.5 launch*. - Speaker treats it as fact without source — classic hype fabrication.

Original quote: “So, OpenAI have just released GPT 5.5.”

Sources: OpenAI Confirms GPT-5.5 Release Date for April 23 | Phemex News, GPT-5.5 System Card - OpenAI

Testing $200 GPT-5.5 Pro on PhD math benchmarks — Personal Story (50/100)

At 0:10

Says he's testing a $200 mystery model that doesn't exist. Bold flex 💀

Why this score: Personal testing claim unverifiable since model doesn't exist. - No public $200 GPT-5.5 Pro tier confirmed by OpenAI. - Benchmarks sound real (PhD-level math), but tying to fake model kills credibility. *Can't verify hands-on tests without model access.*

Original quote: “And in the last few days since it's been released, I've been testing out on some of these benchmark questions... And today, we're going to be testing $200 GPT5.5 Pro.”

GPT-5.5 Pro is more efficient, better at agentic tasks and tools — Solid (80/100)

At 0:30

Claims match OpenAI's actual release notes. Efficiency boost is real, not hype ✅😤

Why this score: Matches OpenAI's official claims. - GPT-5.5 released April 23, 2026, touted for token efficiency and agentic coding (Terminal-Bench 2.0: 82.7%). - Speaker's summary aligns precisely with announced improvements over GPT-5.4. *Personal test note adds credibility without overclaiming.*

Original quote: “[0:26] GPT5.5 Pro. Now, the claims are that GPT [0:29] 5.5 is better at providing end-to-end [0:31] solutions, is better at using certain [0:32] tools which might involve other forms of [0:34] compute like CPU compute, that it's [0:36] better at agentic tasks with regards to [0:38] programming,…”

GPT-5.4 Pro best so far on unsolved random matrix theory problem — Personal Story (65/100)

At 1:57

Personal testing history — can't verify the math, but setup sounds legit 🔍

Why this score: Subjective test results from speaker's benchmarks. - All named models exist: GPT-5.4 Pro (Mar 2026), Claude Opus (~4.7), Grok 4.3, DeepSeek V4 Pro, Google 'Deep Think' (Gemini-related). - Unsolved random matrix theory problem is credible for PhD-level eval. *Can't fact-check private tests, but no red flags in model references.*

Original quote: “[1:57] the first test I wanted to give GPT 5.5 [1:59] Pro is a question that we've tested [2:00] multiple models on to see whether they [2:02] can actually develop a line of [2:03] reasoning, which I've already started [2:05] for them. This is a problem in random [2:06] matrix theory that doesn't…”

GPT-5.5 Pro costs $200; Frontier Math improved from 50% to 52.4% — Solid (75/100)

At 2:04

$200 Pro tier exists, but Frontier Math 'improvement' is basically rounding error 💀📈

Why this score: Pricing confirmed; benchmark specific but marginal. - GPT-5.5 Pro requires Pro subscription (~$200/mo per web context on advanced tiers). - Frontier Math: models score <3% overall; cited 50→52.4% on tiers 1-3 is plausible but *tiny gain* (1.6% tier 4). - No contradiction, but hype vs. reality gap noted by speaker.

Original quote: “[1:24] buckle up as we put $200 GPT5.5 Pro [1:27] against a suite of research level PhD [1:29] maths problems. Now on OpenAI's API [1:31] notes, it does claim that GPT 5.5 is [1:34] more efficient and that you should be able [1:35] to get the same quality of answer [1:36] for less usage of tokens.…”

GPT-5.5 Pro solved PhD math problem faster than previous model — Personal Story (70/100)

At 2:32

Their test, their rules — faster answers sound legit for pricey Pro tier 😤✅

Why this score: Speaker shares personal testing experience on PhD-level math, noting GPT-5.5 Pro's efficiency gains. - Web context confirms GPT-5.5 Pro's release (April 23, 2026) and focus on math/research tasks with improved efficiency. - No contradiction; personal anecdotes like this hold value as tester testimony.

Original quote: “So we gave that same test to GPT5.5 Pro and it managed to get an answer which was very similar in much less thinking time. I'd say this is academically very interesting, just that the model is able to produce a similar quality answer much more efficiently in much less thinking time.”

GPT-5.5 Pro cut thinking time from 60+ min to 16 min vs GPT-5.4 Pro — Personal Story (75/100)

At 3:06

Specific times from their own tests — lines up with Pro model's efficiency hype 📈✅

Why this score: Anecdotal performance comparison from speaker's hands-on testing. - Web confirms GPT-5.4 Pro (March 2026) and GPT-5.5 Pro both emphasize reasoning efficiency, with 5.5 matching latency and using fewer tokens. - Past-month testing aligns with release timeline; no public benchmarks contradict personal results.

Original quote: “In this case, it went from over 60 minutes to 16 minutes. ... This is something that I found was actually a bit of a problem for GPT 5.4 Pro. In the last month of testing it, I felt that I ended up not really using it that much cuz it would just take so long”

GPT-5.5 Pro suggestions identical to 5.4 Pro, feels like hype cycle — Opinion (50/100)

At 3:58

Fair take on incremental upgrades — Pro tier chugs along .1 at a time 🙄💀

Why this score: Speaker's subjective opinion on minimal improvements and release pattern. - Web verifies sequence: GPT-5.4 Pro (March 2026), GPT-5.5 Pro (April 2026), with efficiency gains but similar capabilities noted. - Hype cycle critique is commentary, not verifiable fact; personal frustration valid.

Original quote: “I didn't find it ever gave any better suggestions than 5.4 Pro. In fact, the suggestions were pretty much exactly the same. ... seeing OpenAI release these models which are increasing by.1 every time, you know, 5.2, 5.3, 5.4, and now 5.5”

GPT-5.5 API calls cost more than GPT-5.4 — Solid (80/100)

At 4:30

Math checks out — $5/$30 vs $2.50/$15 per million tokens. OpenAI cashing in ✅😤

Why this score: Claim confirmed by official pricing. GPT-5.5 API is indeed more expensive than GPT-5.4 standard: - GPT-5.5: $5 input / $30 output per million tokens - GPT-5.4: $2.50 input / $15 output per million tokens Speaker correctly identifies OpenAI's higher margins as the main beneficiary.

Original quote: “the cost for the GPT5.5 API calls are more expensive than that of GPT5.4. So, it does feel slightly like who is actually benefiting from the release of this model? Is it OpenAI that's benefiting because they're selling you something which on their end has higher margins”

GPT-5.5 Pro performs same as 5.4, just slightly faster — Opinion (50/100)

At 4:52

Personal testing vibes — contradicts OpenAI's 'major advancement' hype claims 🙄💀

Why this score: Subjective tester opinion based on limited personal experience. OpenAI markets GPT-5.5 as significant upgrade in agentic coding and efficiency, but speaker's hands-on tests show minimal quality difference, only speed gains. *Valid perspective from actual usage, not definitive benchmark comparison.*

Original quote: “They're getting a model which performs pretty much exactly the same as it did before. And the only benefit I can see at least in testing 5.5 Pro in the few days that I have done is that the answers come slightly faster than before.”

GPT-5.5 Pro uses half tokens for same answer as GPT-5.4 Pro — Solid (75/100)

At 5:16

Token efficiency confirmed — OpenAI literally advertises this exact improvement 📉✅

Why this score: Matches OpenAI's stated improvements. GPT-5.5 maintains efficiency by requiring *fewer tokens for equivalent tasks* per official documentation. Speaker's observation that it delivers same quality with half tokens aligns with the model's design goals, though they question if scaling tokens yields quality gains.

Original quote: “GPT5.5 Pro will give the same answer as GPT5.4 Pro using half the tokens as opposed to using the same number of tokens to give a better answer.”

Testing GPT-5.4 vs 5.5 on Codeex agentic coding task — Personal Story (70/100)

At 5:35

Legit test setup with actual tools — Codeex + VS Code is standard now 🛠️✅

Why this score: Describes planned personal experiment. Both GPT-5.4/5.5 and Codeex exist with stated agentic coding capabilities. Speaker outlines transparent methodology using real tools (VS Code + Codeex) and their own research paper as test material. *Upcoming test, not results yet.*

Original quote: “we're going to be using the standard setup of VS Code with Codeex, and we're going to be comparing a fairly complicated but very doable project in which we're going to compare GPT 5.4 and GPT 5.5 at a fairly comprehensive agentic coding task”

Comparing GPT-5.4 vs GPT-5.5 on same agentic math task — Solid (85/100)

At 6:30

Both models real: GPT-5.4 (March '26), GPT-5.5 (April '26). Legit benchmark setup 😤✅

Why this score: Planned comparison uses real, recently released models. - GPT-5.4 released March 2026, powers CodeX; GPT-5.5 April 23, 2026. - Git folder/multi-model test feasible with their 1M-token contexts and agentic features. - No contradictions; sets up valid speed/quality eval.

Original quote: “duplicate this git folder and then we're going to put the same test to GPT5.4 and to GPT5.5. And the main things that we're going to compare is how long each model takes”

Testing GPT-5.5 Pro on PhD-level math project — Solid (80/100)

At 6:46

GPT-5.5 Pro dropped April 23, 2026 — real model for math/research tasks. Not vaporware 💀😤✅

Why this score: GPT-5.5 Pro exists and matches described use. - Released April 23, 2026, for demanding research/math per web context. - Agentic coding/math capabilities align with transcript's sparse matrix project. - No public contradiction; recent release explains any hype.

Original quote: “testing GPT 5.5 Pro because I've given it this whole document. I've asked it in a sequence of kind of questions, open problems that I had with this project.”

GPT-5.5 Pro handles agentic task of summarizing math suggestions — Personal Story (70/100)

At 7:46

Personal test of agentic summary on PhD math — model built for this exact workflow 🔥✅

Why this score: Speaker's firsthand test of model's agentic capabilities. - Anecdote of using GPT-5.5 Pro for multi-step math summary/editing; unverifiable externally but aligns with model's design. - Web confirms GPT-5.5 excels in agentic tasks like code/math analysis. - No red flags in described workflow.

Original quote: “what I've basically asked it to do is to summarize the suggestions that it's making. So given the file I gave you and the current suggestions, I'd like you to try up a summary latic file”

Task is complicated agentic workflow for AI — Opinion (50/100)

At 8:25

Fair take — multi-file math rewrite screams 'agentic' in 2026 AI lingo 🙄✅

Why this score: Subjective assessment of task complexity. - Speaker opines on 'agentic' nature (multi-step planning/context handling), which matches GPT-5.5's strengths. - Not fact-checkable, but consistent with model's hype and sparse non-Hermitian matrix research context. - Recent papers (March 2026) validate topic's PhD-level difficulty.

Original quote: “this task is pretty complicated and I think definitely like an agentic task. I'm asking it to do lots of different things. I'm giving it loads of different contexts”

GPT-5.5 Codex took 6:03, faster than GPT-5.4's 6:55. — Solid (80/100)

At 8:47

Times sound plausible for heavy LaTeX tasks — these models do chew through compute 💀✅

Why this score: GPT-5.4 and GPT-5.5 are real models per OpenAI releases; GPT-5.5 launched April 2026 as faster/more capable. Specific thinking times unverifiable without tester's logs, but aligns with known performance diffs (e.g., GPT-5.4 efficient for pro work). No contradiction in web context.

Original quote: “Equipped with GPT 5.5, Codex thought for 6 minutes and 3 seconds, which is about a minute less than GPT 5.4 for thought for which was 6 minutes and 55 seconds.”

GPT-5.4 and 5.5 outputs very similar in math detail. — Personal Story (70/100)

At 9:27

Tester's take on similar outputs — fair for incremental upgrades 😤

Why this score: Personal observation from direct testing; anecdotal evidence can't be externally verified but consistent with model evolution (GPT-5.5 as 'smartest yet' but builds on 5.4). No public benchmarks contradict similarity in LaTeX tasks.

Original quote: “Both models were set to fast mode with extra high thinking time and on the left you can see GPT 5.4 and on the right GPT 5.5. Both files look very similar even in mathematical detail.”

GPT-5.5 slightly better due to one extra equation. — Opinion (50/100)

At 10:03

One equation wins it? That's some razor-thin 'edge' they're hyping 🙄💀

Why this score: Subjective evaluation of test outputs; speaker admits minor diff despite overall similarity. Aligns with GPT-5.5's positioning as incremental upgrade over 5.4.

Original quote: “I would argue that GPT 5.5 takes an incremental edge over GPT5.4 solely because of one equation that it included which helps to actually put into picture what exactly this document is looking into.”

Codex useful for some problems but incapable here. — Opinion (50/100)

At 10:25

Heavy user calls it incapable on tough LaTeX merge — real talk from experience ✅

Why this score: Experienced user's opinion on Codex limits; matches known strengths (routine coding) vs. weaknesses (complex doc synthesis). Web context notes 85.5% success on simpler PR tasks, implying gaps on PhD-level.

Original quote: “as someone that uses Codex quite a lot and I found it to be very useful for particular problems, examples like this do make me think just how incapable these models”

Both models used only 6 minutes of chain of thought — Personal Story (70/100)

At 10:30

Their test, their results — 6 minutes CoT sounds like a real run. Can't argue personal experience ✅

Why this score: Speaker shares direct observation from their own testing of GPT-5.4 and GPT-5.5 on a specific task. Personal anecdotes from hands-on tests are credible as testimony, even if results are subjective. No contradiction with web context confirming both models exist and are agentic.

Original quote: “In this case specifically given the prompt I am surprised that both models decided to only use 6 minutes of chain of thought.”

GPT-5.5 15% more efficient, 1 min less CoT than GPT-5.4 — Personal Story (65/100)

At 10:57

15% faster in their test — specific numbers from actual usage. Matches efficiency hype but their task 🙄✅

Why this score: Anecdotal benchmark from speaker's test: GPT-5.5 took 1 fewer minute (from ~6 to 5 min CoT) for same output, called 15% efficiency gain. Web context confirms GPT-5.5 emphasizes *efficiency improvements* and agentic speed, so directionally aligns, though exact % is test-specific.

Original quote: “We saw about a 15% improvement in the efficiency in that it took GPT5.5 1 minute less of chain of thought in order to produce the same PDF as GPT 5.4.”

GPT-5.5 not meaningfully better than 5.4 beyond efficiency — Opinion (50/100)

At 11:30

Fair take from their tests — but OpenAI's hyping way more than just speed 😤

Why this score: Speaker's subjective assessment based on one task: no noticeable intellectual/capacity gains beyond efficiency. Web context contradicts slightly — GPT-5.5 marketed for *significant improvements* in agentic tasks, benchmarks, knowledge work vs GPT-5.4. Still, personal testing opinion holds value.

Original quote: “But so far as I can see any claims of this model being better than GPT 5.4 outside of the efficiency gains are pretty limited.”

GPT-5.5 only faster answers, not better than GPT-5.4 — Opinion (50/100)

At 12:06

Their wallet disagrees with OpenAI's benchmarks — classic tester vs marketer divide 💀

Why this score: Opinion rooted in personal testing, downplaying GPT-5.5 to mere speed upgrade. Contrasts official claims of *broader intellectual capacity gains* (agentic tasks, research). Valid consumer perspective on *noticeable* improvements, even if benchmarks show more.

Original quote: “it's not really any better than GPT 5.4 other than in the sense that you get your answers quicker”

Will keep testing GPT-5.5 Pro as Sam Altman's fan — Just Vibes (50/100)

At 12:30

"Sam Alman's porn"? Auto-caption gold 💀😂 — model exists tho, just not $200 tier.

Why this score: Hilarious transcription fail stands out as peak YouTube chaos. - 'Sam Alman's porn' is clear mishearing of 'Sam Altman's fan'; Sam Altman is OpenAI CEO. - 'GPC 5.5 Pro' typo for GPT-5.5 Pro, which web context confirms released April 23, 2026.

Original quote: “Of course, as Sam Alman's porn, I am going to keep testing GPC 5.5 Pro,”

New DeepSeek model seems more impressive — Opinion (50/100)

At 12:51

DeepSeek V4-Pro hype checks out — fair take from a tester 😤✅

Why this score: Personal opinion on DeepSeek's impressiveness, not a factual claim. - Web context confirms DeepSeek V4 series (Pro/Flash) launched April 24, 2026, with strong agentic performance and cost-effectiveness. - Speaker's view aligns with noted benchmarks; no contradiction.

Original quote: “there's a new model of Deep Seek, which actually seems more impressive to me.”

Launched Discord server with paying members — Personal Story (70/100)

At 13:27

Discord plug — standard creator flex, no cap detected 🙄💀

Why this score: Self-reported channel update; personal anecdotes get benefit of doubt. - Common for YouTubers to launch Discords for members; no web context contradicts. - Ties into thanks to supporters — plausible given video's testing theme.

Original quote: “I recently launched the Discord server, and we've got a couple of people that are paying for that, too,”

See the full analysis with sources and timestamps →