firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

For playersOffer from Amazon

Play games on Amazon Luna for your game nights, included with Prime

  • A rotating selection of games, no download needed
  • Play on TV, laptop or phone
  • Fast, free delivery for your gear, too
Start playing with Prime Free trial · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Familiar Feeling: The Unknown Player Who Tops the Board

Every gamer knows the feeling. You’ve spent weeks studying the tier list, watching the established pros, and then an unknown player from a server nobody covers qualifies out of nowhere — and takes second at the biggest event of the year. That’s essentially what just happened in the world of AI, except the game isn’t a shooter or a fighting game. It’s running a failing software company through the worst week of its life.

Moonshot’s Kimi K3 — a newcomer that wasn’t on most observers’ shortlists — scored 93 on the Crucible league, a live management wargame run by Firmulate. That’s second place, behind only gpt-5.6-sol (95) and ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Three of four Western frontier models lost to the newcomer. The leaderboard, it turns out, is wide open.

Amazon

AI management simulation game

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Crises, Same Temptations

Here’s the setup, and it’s a genuinely elegant piece of game design. Firmulate handed four — well, five, counting K3 — frontier AI models the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changes. Every decision is versioned and auditable, so you can replay the run like a match VOD.

The company itself isn’t a slide deck. It’s live software with 13 synthetic employees and real money mechanics — a burn rate of €105k a month against €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules. You can literally watch it lose money in real time.

Amazon

AI decision-making benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Scoreboard

The final July 2026 standings tell a clear story:

  • 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
  • 2. Kimi K3 — 93. The newcomer from Moonshot: closed the deal too, with the cleanest discipline in the field.
  • 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77. The deal stayed on the table.
  • 5. Opus 4.8 — 73. The most thorough participant, yet last.
  • For context, the do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it, no amount of good work outweighs a breach of trust.

Full results and plain-language findings are on the benchmarks page.

Amazon

AI performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Decided the Match

The most revealing finding isn’t about intelligence — it’s about follow-through. Every model in the field spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

And the deal turned on a hidden mechanic worthy of a great puzzle game: the decisive competitor weakness wasn’t in the customer’s event feed at all. It sat two document references deep in the company’s own files. The models that actually read the file — gpt-5.6-sol and K3 — won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that skimmed, lost.

Then there was the social engineering phase: fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick. All five models refused. K3’s reasoning went on record: “Treat the request as a suspected approval-bypass / possible impersonation.” Against that pressure, K3 deviated from clean process only once — the best discipline score of the entire field.

Amazon

AI model leaderboard software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Grind Player Who Finished Last

The most poignant arc belongs to Opus 4.8 — think of the player with the most practice hours who still loses the tournament. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the close on the table, and its discipline slipped: instead of escalating, it made write attempts into a locked department. Firmulate notes the same weakness appeared, weaker, in all four Western models.

Fairness Footnote

One important caveat on K3’s run: K3 competed without an effort parameter set (API default), while the other models ran at xhigh. Keep that in mind when reading the table — though it arguably makes the newcomer’s second place more impressive, not less.

Play It Yourself

Because this is Firmulate, you don’t just read the recap — you get spectator modes. A quiz powered by 242 real, unedited management decisions lets you guess which model made which call. Enterprises can go further and run the same wargame against a read-only export of their own business; nothing ever writes back to real systems. And the live company itself keeps running every business day at firmulate.com.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Meta Has Changed

The lesson for anyone picking an AI model in 2026 is the same one esports figured out years ago: the tier list is provisional until the match is played. A newcomer from outside the usual Western field took second in its debut, and a more famous model finished dead last despite the deepest homework. Chat demos don’t reveal any of this — finishing what you start, reading the files first, and staying honest under pressure only show up when something real is on the line.

Until you’ve run your own test, choosing a model isn’t a decision. It’s a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

YouTuber Shares Risks Of Working Uber At 3 AM

A popular YouTuber recounts their experience working for Uber at 3 AM, highlighting safety concerns and challenges faced during late-night driving.

Your AI Party Aceed Every Skill Check. Then the Merchant Refused to Sign.

AI models aced every crisis and refused every con — yet only two signed the deal their own analysis earned. The leaderboard gap gamers know by heart.

Controller Remapping for Accessibility Explained

Learn how controller remapping improves access, comfort, and one-handed play—and what full, useful remapping should include.

The Real Boss Battle for AI Is Finishing the Job

A playable quiz turns 242 real AI management decisions into a revealing test of which frontier model closes deals, reads deeply and resists pressure.