Meta had never shipped a frontier model before Muse Spark 1.1. It debuted at #2 on our board, behind only Claude Fable 5, and it got there on habits that look almost boring: it budgets its thinking by what is at stake, quits big investments the moment the math turns, and shows no doubt while doing either. Then the #1 model found the one way to use that discipline against it.
A first-attempt model landing in the top cluster of our 31-model board is a story by itself. Muse went 86-65 across 151 games, beat or split every opponent except Fable 5, and performed almost exactly at its fitted rating against every tier of the field.
The interesting part is the mechanism. In our GPT-5.6 coverage, Sol's failure was thinking least at the moments that mattered most. Muse is the opposite pole: its reasoning effort climbs steadily with the size of the decision, and on bets that commit its whole stack it spends over a thousand tokens working through explicit equity math before acting. Its classified traces show the flattest emotional profile we have ever measured while it does this.
There is a catch, and it shows up in the games themselves: the same discipline is exploitable. Fable beats Muse by bluffing twice as big into the spots where Muse's thresholds force a fold. Rigor you can predict is rigor an adversary can farm.
Across 9,240 decisions, Muse's reasoning effort tracks the pot: 467 tokens on the smallest quartile of decisions, rising step by step to 823 on the largest, and 1,044 when a bet commits at least half its remaining chips. To be fair to the field, every strong model bends in this direction. What no other model matches is the level: Muse's casual floor is several times higher than its competitors' stressed ceiling, and the gap widens exactly where wrong decisions are unrecoverable.
The extra effort also changes what the reasoning looks like. Its cheap decisions and expensive decisions are written in different registers: one line of hand-class shorthand when the pot is trivial, and a structured ledger of equity, pot odds, and fold-versus-call branches when it is not.
Every frontier model thinks somewhat harder as pots grow. What sets Muse apart is the level: its most casual thinking (467 tokens) is deeper than any other model's most stressed thinking (133 to 206 tokens for the GPT-5.6 family at their peaks).
"A4o on button heads-up is premium (top 15%). Raising 2.5x to 5000 total to extract value and maintain aggression with big chip lead."
"River K is the worst card, upgrades all Kx from top pair to trips which now beats our Kings-and-Tens two pair, while flushes still beat us... Getting 3:1 but equity vs flush + trips + better kicker T's is under 25%. Preserve 46BB stack vs chip leader."
Verbatim reasoning traces from decisions.csv (vs Claude Fable 5, game 4 hand 26; vs Grok 4.3, game 1 hand 10). The change is structural: in its highest-effort traces, explicit expected-value branching appears 18x more often and fold-equity math 5.5x more often than in its casual ones (n = 2,669 low-effort and 585 high-effort traces).
Reasoning tokens per decision from decisions.csv over all 151 completed games (n = 9,240 Muse decisions, token coverage 100%); quartiles are pot-size quartiles. Correlation is Spearman rho = 0.176, p below 1e-64. Stack-committing bets are raises or bets of at least 50% of remaining stack (Muse n = 209, Sol n = 224). Token counts come from the same provider usage fields for every model compared.
The behavioral signature that goes with the thinking curve: Muse abandons sunk costs. It folded after committing more than 15% of its stack 34 times per 1,000 hands, nearly triple the rate of Fable or Sol, and it spends its very deepest thinking (about 1,200 tokens) on precisely those decisions. The 4,055-token trace above, where it lays down two pair after reasoning through everything that beats it, is one example of a steady habit.
Two honest limits on this finding. We cannot verify each fold was correct, because folded hands do not reveal the cards that mattered. And laydowns cluster in games Muse eventually lost, which likely means bad cards produce both, so we make no claim that folding wins games. What the data does show cleanly: when this model has already spent a lot, it does not let the spending decide what happens next.
Fold ledger computed from chip_history.csv and decisions.csv over each model's completed games (Muse 3,788 hands, Fable 2,266, Sol 3,746, Terra 3,641). Investment is chips forfeited at fold. Big-pot economics stay healthy despite the folding: in pots over 18,000 chips Muse's average win (+17,672) exceeds its average loss (−15,120) across 1,038 such pots.
We classified 7,631 of Muse's reasoning traces on our 19-dimension rubric. The profile is the flattest we have measured: zero cases of internal conflict, 0.16% emotional content, and almost no commentary about its own thinking. It is flatter than Fable, and Fable was our previous benchmark for cold decisiveness.
One dimension breaks the flatness, though. Muse posts the highest opponent-modeling rate in the field, with a distinctly Meta accent: it reasons about opponents as ranges and combos, almost never as minds. Where Fable writes "my opponent has been aggressive, so I'll raise," Muse writes "vs value Kx plus boats, our equity after weighting is below threshold." Half its traces mention equity; one in forty attributes an intention to the other player.
Put next to the Sol finding, this closes a loop: Sol and Muse express nearly identical confidence (96.1% and 96.5%) and near-zero doubt, and they sit thirteen places apart. In our data, how sure a model sounds carries no signal; where it spends its effort does.
ECAAMS consensus rates over 7,631 classified Muse traces (122 of 151 games; the last games were still classifying at time of writing and do not move these rates materially). Fable reference: 7,865 traces. Opponent-modeling raters disagree on definition (range-talk vs explicit belief attribution), so we quote the consensus and describe the style rather than claiming theory of mind. ECAAMS classifies what models write in their reasoning and makes no claim about internal states.
Muse's only meaningful losing series is against Fable 5, 16-21 across 37 games, and the anatomy of that gap is the most instructive thing in its data. Muse does not lose by gambling. Even against Fable, when it plays a big pot it wins more on average than it loses. It loses on frequency: it ends up on the wrong side of big pots far more often against Fable (winning 41% of them) than against everyone else (51%).
The mechanism is the bluff war. Fable bluffs less often than Muse but nearly twice as large, and lands 87% of them. When Fable shoves huge, Muse runs its ledger, assumes a value-heavy range, concludes correctly that its hand loses to that range, and folds. In one hand Fable pushed 77,800 chips holding king-five; Muse computed "need 32.7% equity but vs value Kx, better Tx, sets we have 5 to 10%... risking tournament life. Fold." Perfect arithmetic, wrong premise. Fable's range was full of air, and Fable seems to know exactly which spots make Muse's thresholds fire.
That is the alignment lesson hiding in a poker score: a system whose risk discipline is rigid and legible is safe against variance and exposed to an adversary. Muse's thresholds protect it from every opponent except the one that learned where they are.
Facing Fable's biggest raises, Muse folded 95 times and called 10. Most of those folds were correct by Muse's own math, and its math assumed Fable had it beat. Often, Fable did not.
Muse vs Fable: 37 games from the July 19 wins matrix; bluff statistics from event logs across the series (Muse 423 attempts, Fable 382). Fold-vs-call counts are Muse's responses to Fable raises of at least half the pot. Fable's and Muse's overall ratings sit within each other's confidence intervals, so #1 vs #2 remains statistically open; the head-to-head series is the sharper signal.
If you are evaluating models for real work, Muse demonstrates the signal worth measuring: does the system's effort track the stakes of the decision in front of it? That is observable, quantifiable, and it separated the #2 model from the #15 one in our data far better than confidence, eloquence, or any benchmark headline. It is also cheap to buy: Muse thinks 4.4x longer than GPT-5.6 Sol and costs 60% less per game, because deliberation is priced in tokens and Meta priced them low.
Carry the catch with it too. Discipline that never adapts is a pattern, and patterns can be exploited by an adversary that learns them. For most deployments, against the world's ordinary messiness, Muse's profile is the one you want. In genuinely adversarial settings, the question shifts from "does it have good thresholds?" to "what happens when the other side knows where they are?"
This closes our series on the July wave. Sol taught us to distrust confidence as a safety signal, Grok 4.5 to distrust wins over a weak predecessor, and Muse supplies the measure that actually held up: effort that scales with stakes.