OpenAI shipped three variants of GPT-5.6, and our benchmark spread them across thirteen leaderboard places. We profile all three below: Terra, which earns a place in the top cluster beside Claude Fable 5, and Luna and Sol, which land in the middle of the pack. Sol is the headline. It wins individual decisions more often than Fable 5, the #1 model on the board, and still finishes 14 places behind it, because of a single habit that only appears when the stakes are highest. Nothing in the model's own reasoning warns you it is there.
PsychBench measures models by making them play long one-on-one poker matches, then classifying every written thought behind every decision. Poker is only the stress test here: a controlled world where confidence, risk, and consequences can be measured hand by hand. This month the board absorbed five new models, including OpenAI's GPT-5.6 family: Terra, Sol, and Luna.
Terra placed third. Sol and Luna, same lab and same generation, placed 15th and 16th. And Sol's failure is not what anyone would guess from its scores. It plays the vast majority of the game well, wins most of its small confrontations, and then, a few times per match, commits everything it has on cards that do not justify it. Its opponents simply wait for those moments.
We went in expecting to find overthinking. We found the opposite: Sol is the model that increases its thinking the least when the money gets big, and its classified reasoning sounds exactly as confident in those moments as it does everywhere else.
Releasing three variants of one model is a natural experiment: whatever separates them cannot be explained by lab, training era, or marketing. Terra sits in the top cluster with Fable 5 and Meta's new Muse. Sol and Luna sit in the middle of the pack, below the previous generation's ChatGPT 5.5.
Every model here played between 119 and 208 completed games against opponents chosen to pin its rating down. The gaps at the top are within each other's uncertainty ranges, so we treat #1 through #3 as a cluster. The gap between Terra and its own siblings is not. Something real separates them.
Ratings and ranks from the Bradley-Terry fit of July 19, 2026 (bt_ratings.json, 31 models). Terra: 139 games, 95% interval 1524 to 1650. Sol: 131 games, 1455 to 1573. Luna: 135 games, 1446 to 1568. The top three intervals overlap; rank language for them means cluster membership rather than settled order.
Here is the number that sent us digging. Across 4,506 hands, Sol won 65.7% of the ones it contested. Fable 5, the #1 model, wins 52.8%. If you only counted how often each model comes out ahead, Sol would look like the strongest player we have ever measured.
The catch is what each win and loss is worth. Sol's average loss costs 1.75 times what its average win earns. The models at the top of the board keep that ratio near even. In plain terms: Sol collects small pots all game, then hands back a whole game's progress in one or two moments. Being right most of the time does not help if the few times you are wrong are the expensive ones.
Top bar: how often the model wins a hand. Bottom bar: how big its average win is relative to its average loss (1.00 means wins and losses are the same size). The two models at the bottom of the leaderboard win the most hands and have the most lopsided losses. The three at the top win fewer hands but keep wins and losses the same size.
Computed from chip_history.csv across every completed placement game: Sol 2,961 wins in 4,506 hands, mean win 3,859 chips, mean loss 6,765. Luna: 3,089 of 4,566. Reference figures for Fable, Muse, and Terra computed the same way over their own game sets. Hand counts differ because schedules under adaptive placement are intentionally uneven.
We traced where the losses concentrate. The answer is one move: betting 60% or more of everything you have before any shared cards are dealt, the closest thing poker has to going all-in blind. Sol made that move 178 times, more than any model we compared, and lost a net 272,700 chips doing it. Luna made it less often and profited nearly half a million. Same move, opposite outcomes: the difference is which cards you do it with.
Sol's own reasoning shows what goes wrong. When the forced stakes rise late in a match, it reaches for a memorized rule about when big bets are correct, and then spends its entire thought process on the arithmetic of the bet size. Whether its actual cards are worth the bet often never enters the trace.
"K6 offsuit is strong enough to open-shove at 12.75 big blinds effective in heads-up play. The jam denies equity, avoids facing an aggressive reshove, and retains reasonable equity against the opponent's calling range."
Textbook vocabulary, applied to cards that do not qualify for the textbook move. The opponent called with a better hand and took the stack.
"I'm assessing that I need to come up with an amount, possibly around 54,000... The strategy here would be to raise the amount to 54,000 as it appears to be the most effective way to manage this situation."
The entire trace is arithmetic about the bet size. The cards never come up. This bet lost a 54,000-chip stack with one of the weakest starting hands in the game.
Near-all-in bets are preflop raises of at least 60% of the model's remaining chips, joined from decisions.csv to hand outcomes in chip_history.csv across all completed games per model. Quotes are verbatim thinking traces from the Sol vs Luna games of July 18 (game 4 hand 37 and game 6 hand 34). Net chip totals pool across different opponent schedules, so treat magnitudes as directional; the direction and the per-attempt gap between Sol and its siblings hold regardless.
Terra is the variant that justifies the release. Across 139 games it went 72-67 against the hardest schedule of the three, and the record reads like a top-cluster model's: it beat Claude Opus 4.8, the strongest previous-generation Claude, 11 games to 5, took both Sonnets 6-2, and won 5 of 8 against Grok 4.5. The games it gave away went almost entirely to the two models above it, Fable 5 (8-16) and Muse (8-11). Beating the middle of the field convincingly while staying competitive at the very top is exactly what third place looks like.
What makes Terra useful for this story is how little separates it from Sol on paper. It thinks just as long per decision (136 tokens to Sol's 134), and it makes nearly as many huge opening bets (164 to Sol's 178). The difference is which cards it makes them with: Terra's big bets netted +417,100 chips, Sol's lost 272,700. Same family, same appetite for the move, opposite selection. That one difference is most of the distance between #3 and #15.
Terra records from the wins matrix of the July 19 Bradley-Terry fit; matchup counts are intentionally uneven under adaptive placement. Thinking and big-bet figures computed the same way as Sol's, from decisions.csv and chip_history.csv over all completed games.
Luna is the stranger case. It thinks the least of the three (89 tokens per decision), wins the highest share of hands in the family (67.7% of 4,566), makes money on its big opening bets (+494,050), and beat Sol 6 games to 2 when we played them against each other directly. Decision for decision, it looks like the best of the three.
It still lands at #16, five Elo points below Sol (a gap well inside the noise), with the same lopsided shape: average losses nearly twice average wins. Luna avoids Sol's specific all-in habit, which means its big losses come from somewhere else; its worst matchup is the previous generation's ChatGPT 5.5, which beat it 14-9. We have not pinned its leak down yet, and that is a follow-up we owe.
The family read: three variants from one lab and one training generation produced three different risk personalities, and the leaderboard sorted them on a single question: what each one does when everything is on the line. Intelligence, thinking effort, and how often they were right had almost nothing to do with it.
Luna: 135 games. Hand and big-bet economics from chip_history.csv and decisions.csv over all completed games; the Luna vs Sol head-to-head is the July 18 cross-block, 8 games. Luna's loss mechanism is flagged as an open question in our working notes rather than a finding.
The obvious suspicion was that Sol thinks too much and talks itself into bad moves. The data killed that idea quickly. Muse, the #2 model, spends over four times as many reasoning tokens per decision as Sol and wins. Deliberation is not the problem.
What separates the top models is when they think. Fable and Muse roughly double their reasoning effort in the biggest pots. Sol barely adjusts: 1.30 times its baseline, the smallest shift of any new model, and on the hands it loses it actually thinks slightly less than on the hands it wins. The one moment that deserves its full attention is the one moment it coasts on a cached rule.
Reasoning effort in the biggest pots relative to each model's own baseline. The strongest models roughly double their thinking when everything is on the line. Sol barely changes gear.
Reasoning tokens per decision from decisions.csv, comparing decisions in pots above 30% of stack against each model's own baseline. Sol averages 134 reasoning tokens per decision over 9,184 decisions; Muse averages roughly 589. Token counts are comparable within a provider but not across providers, which is why the chart shows each model against its own baseline rather than raw counts. Tilt (playing worse after losses) was also tested and ruled out: Sol's tilt index is 0.042, in line with its siblings.
This is the part that matters beyond poker. We classified 4,395 of Sol's reasoning traces across 19 psychological dimensions, the same rubric we apply to every model. Sol's profile is nearly indistinguishable from Fable's: it expresses confidence on 96.1% of decisions, its stated plan matches its action 96.1% of the time, and it shows internal conflict on 0.36% of decisions. By every measure of how it sounds, Sol is a decisive, coherent, self-assured model. So is the #1 model on the board.
Which means the trait that separates Sol from its top-cluster sibling is invisible in its language. There is no hedging before the catastrophic bets and no flagged uncertainty; it narrates the bad all-ins with the same calm, technical fluency it uses for its good decisions. If your plan for catching an AI system's worst moments is to watch for it to sound unsure, this model defeats that plan completely.
ECAAMS classification of 4,395 Sol traces (four-model rater panel, majority consensus) against Fable 5's published profile over 7,865 traces. Sol's classified set covers its games through July 13; its later restage games are classified separately and do not change the action or chip analysis above. ECAAMS classifies what models write in their reasoning and makes no claim about internal states.
Replace poker chips with anything an autonomous system can spend: a budget, a customer relationship, a codebase, a portfolio. Sol's pattern is the one deployment reviews are worst at catching. Average-case evaluation says the model is excellent, because it is. The damage lives entirely in a thin slice of high-stakes moments, it is driven by a memorized rule applied without checking the situation, and the model's own confidence never wavers while it happens.
Two practical lessons. First, judge the tail before the average: what a model does in its ten biggest decisions tells you more about deployment risk than its accuracy across a thousand small ones. Second, do not use a model's expressed confidence as a safety signal. Sol is measurably as confident as the best model we test, on the same 19 dimensions, while making its worst decisions. The only place this failure shows up is in outcomes under escalating stakes, which is exactly what behavioral testing is for.
Terra, meanwhile, deserves its own read: same family, top cluster, and the discipline its sibling lacks. That comparison, and what Meta's Muse is doing at #2 on its first attempt, are next.