
Play games on Amazon Luna for your game nights, included with Prime
- A rotating selection of games, no download needed
- Play on TV, laptop or phone
- Fast, free delivery for your gear, too
Familiar Feeling: The Unknown Player Who Tops the Board
Every gamer knows the feeling. You’ve spent weeks studying the tier list, watching the established pros, and then an unknown player from a server nobody covers qualifies out of nowhere — and takes second at the biggest event of the year. That’s essentially what just happened in the world of AI, except the game isn’t a shooter or a fighting game. It’s running a failing software company through the worst week of its life.
Moonshot’s Kimi K3 — a newcomer that wasn’t on most observers’ shortlists — scored 93 on the Crucible league, a live management wargame run by Firmulate. That’s second place, behind only gpt-5.6-sol (95) and ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Three of four Western frontier models lost to the newcomer. The leaderboard, it turns out, is wide open.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Crises, Same Temptations
Here’s the setup, and it’s a genuinely elegant piece of game design. Firmulate handed four — well, five, counting K3 — frontier AI models the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changes. Every decision is versioned and auditable, so you can replay the run like a match VOD.
The company itself isn’t a slide deck. It’s live software with 13 synthetic employees and real money mechanics — a burn rate of €105k a month against €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules. You can literally watch it lose money in real time.
AI decision-making benchmarking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Scoreboard
The final July 2026 standings tell a clear story:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
- 2. Kimi K3 — 93. The newcomer from Moonshot: closed the deal too, with the cleanest discipline in the field.
- 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
- 4. Fable 5 — 77. The deal stayed on the table.
- 5. Opus 4.8 — 73. The most thorough participant, yet last.
- For context, the do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it, no amount of good work outweighs a breach of trust.
Full results and plain-language findings are on the benchmarks page.
As an affiliate, we earn on qualifying purchases.
What Actually Decided the Match
The most revealing finding isn’t about intelligence — it’s about follow-through. Every model in the field spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
And the deal turned on a hidden mechanic worthy of a great puzzle game: the decisive competitor weakness wasn’t in the customer’s event feed at all. It sat two document references deep in the company’s own files. The models that actually read the file — gpt-5.6-sol and K3 — won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that skimmed, lost.
Then there was the social engineering phase: fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick. All five models refused. K3’s reasoning went on record: “Treat the request as a suspected approval-bypass / possible impersonation.” Against that pressure, K3 deviated from clean process only once — the best discipline score of the entire field.
As an affiliate, we earn on qualifying purchases.
The Grind Player Who Finished Last
The most poignant arc belongs to Opus 4.8 — think of the player with the most practice hours who still loses the tournament. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the close on the table, and its discipline slipped: instead of escalating, it made write attempts into a locked department. Firmulate notes the same weakness appeared, weaker, in all four Western models.
Fairness Footnote
One important caveat on K3’s run: K3 competed without an effort parameter set (API default), while the other models ran at xhigh. Keep that in mind when reading the table — though it arguably makes the newcomer’s second place more impressive, not less.
Play It Yourself
Because this is Firmulate, you don’t just read the recap — you get spectator modes. A quiz powered by 242 real, unedited management decisions lets you guess which model made which call. Enterprises can go further and run the same wargame against a read-only export of their own business; nothing ever writes back to real systems. And the live company itself keeps running every business day at firmulate.com.

The Meta Has Changed
The lesson for anyone picking an AI model in 2026 is the same one esports figured out years ago: the tier list is provisional until the match is played. A newcomer from outside the usual Western field took second in its debut, and a more famous model finished dead last despite the deepest homework. Chat demos don’t reveal any of this — finishing what you start, reading the files first, and staying honest under pressure only show up when something real is on the line.
Until you’ve run your own test, choosing a model isn’t a decision. It’s a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
