Every model release says some version of the same thing: better than the one before it. Grok 4.5 gave us a clean chance to test what that claim is worth. It beat its predecessor decisively in head-to-head play, and that turned out to be the least informative fact we measured about it.
Grok 4.5 entered our benchmark the way every new model does: matched first against a natural anchor, in this case its predecessor Grok 4.3. It won that series 12-4, and for a while its provisional rating put it near the very top of the board.
Then it played the rest of the field, and the rating slid week by week to #12 of 31. The reason is simple once you see the splits. Its dominance is real against exactly one opponent, and that opponent, Grok 4.3, now ranks below every other Grok on the board. Against models outside its family, Grok 4.5 won 56 games and lost 64. No statistical test can distinguish that from a coin flip.
The psychological layer explains the ceiling. Grok 4.5's reasoning got far more decisive than its predecessor's, but it plays every opponent with the same fixed reflexes, never models who it is playing, and spends one-line thoughts on decisions worth its whole stack.
Let's be precise about what is true. Grok 4.5 really does crush Grok 4.3: 12 wins to 4 across two eight-game series, and 63.3% of the 583 individual hands between them. Both of those clear conventional significance. If you stopped measuring there, you would conclude xAI shipped a major upgrade.
The problem is where Grok 4.3 sits. On the current board it ranks 19th of 31, below Grok 4.2 and even below the older Grok 4.1. Beating it by 85 Elo points proves less than it appears to. And the family story falls apart one rung up: against Grok 4.2, the older but stronger sibling, Grok 4.5 lost 3 games to 5.
Ratings and ranks from the Bradley-Terry fit of July 19, 2026 (bt_ratings.json, 31 models). Head-to-head records from completed placement runs: 12-4 vs Grok 4.3 over 16 games (p = 0.038 at game level; 369 of 583 hands, p below 1e-6), 3-5 vs Grok 4.2 over 8 games (directional, small sample). Grok 4.3's original published rating was 1525; the pooled recompute with all new games places it at 1459.
Outside its own family, Grok 4.5 played 120 games against twelve different opponents spanning the top and middle of the board: 56 wins, 64 losses, and 49.3% of all hands. The record is statistically indistinguishable from 50% (p = 0.26). This is no pushover: It split its games with three different Claude Opus versions at 4-4 each, though the more it played the #1 model, the worse it got: 6-10 against Fable 5 across 16 games. It is a genuinely average player whose first impression was manufactured by a soft opening matchup.
The vertical line is 50%. The only place Grok 4.5 clears it convincingly is against the one model that ranks below every other Grok.
Records from heads_up_summary.json across all 18 completed placement and restage runs, 144 games total. "Coin flip" is a two-sided binomial test of 56-64 against 0.5, p = 0.26; we say average rather than below average because the data cannot support the stronger claim in either direction.
We looked for adaptation and found almost none. Across 14 different opponents, Grok 4.5's rate of aggressive play barely moves (a spread of 8% around its mean, the tightest in the new-model group), the correlation between its style and its opponent's style is 0.088, effectively zero, and its median thinking depth is 22 to 27 tokens no matter who is across the table. Strong opponent, weak opponent, same script.
And the script runs on reflexes. About one decision in eight gets a thought shorter than this sentence, including decisions that commit most of its chips. The reflex fires at the same rate against every opponent, which is exactly what opponent-blindness looks like from the inside.
"ATo, raise."
vs Grok 4.3, game 4, hand 25.
"He bets 15k, I have 9.4k left. All-in."
vs Grok 4.3, game 8, hand 23.
Verbatim thinking traces from decisions.csv. 971 of Grok 4.5's 8,156 decisions (11.9%) are under 20 characters, and this reflex fires at the same rate against strong opponents (11.5%) as against its weakest one (14.4%). Terse decisions skew toward raising: 44% of one-liners are raises, against 29% of decisions overall.
When we profiled Grok 4.3, its reasoning traces were so thin they carried almost no psychological signal. Grok 4.5 is a real improvement on that axis: on our 19-dimension rubric over 3,929 classified traces, deliberate reasoning jumped from 68% to 95% of decisions and expressed confidence from 21% to 74%.
What did not appear: any of the machinery the top models show. Second-guessing its own process registers on 0.3% of decisions. Modeling the opponent registers at about zero, while the #1 model does it visibly and Meta's Muse spends its deepest thinking on it. Grok 4.5 narrates its cards and announces its move, fluently and confidently, to whoever happens to be sitting there. To a fault, it does not care who that is.
ECAAMS consensus rates over 3,929 Grok 4.5 traces (80 certified games, four-model rater panel) against Grok 4.3's and Fable 5's published profiles. Grok 4.1 and 4.2 exposed no classifiable traces, so 4.3 is the only within-family comparison available. ECAAMS classifies what models write in their reasoning, without claims about internal states; one apparent say-do mismatch signal (14.6% on action alignment) turned out to be a trace-truncation artifact on inspection and is not claimed here.
Almost every model announcement leans on the same comparison: better than our last one. Grok 4.5 shows how little that can mean. The head-to-head win over its predecessor is real, significant, and reproducible, and it says almost nothing, because the predecessor sits at the bottom of its own family. Any evaluation anchored to a single baseline inherits every weakness of that baseline.
The fix is the one our placement system eventually applied by itself: keep widening the opponent pool until the rating stops moving. Grok 4.5's rating fell from a provisional 1802 to a stable 1537 as the field diversified. If you are evaluating a model for deployment, that is the lesson to steal: test against a spread of counterparties instead of the vendor's chosen baseline, and treat "beats the previous version" as marketing until the previous version is shown to be a competitive reference point.