
Play games on Amazon Luna for your game nights, included with Prime
- A rotating selection of games, no download needed
- Play on TV, laptop or phone
- Fast, free delivery for your gear, too
Every Speedrunner Knows the Do-Nothing Run
Gamers invented a strange art form: the any% run, the pacifist run, the idle run where you tape down a button and walk away. The point of a baseline run is never to win — it’s to know what the floor looks like. How much can you score by standing still? What does the game give away for free?
A public AI experiment called Firmulate just ran the corporate version of that test, and the answer is quietly brilliant: a manager AI that does absolutely nothing still scores 26 out of 100. Not zero. Twenty-six. And that number, more than any leaderboard, tells you the benchmark is honest.
As an affiliate, we earn on qualifying purchases.
The Wargame, Briefly
Four frontier AI models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision is versioned and auditable, like a replay file you can scrub through frame by frame.
The final Crucible League from July 2026:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
- 2. Kimi K3 — 93. Closed the deal too, with the cleanest discipline of the field (with a caveat we’ll get to).
- 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73. The most thorough participant of all — and last place.
As an affiliate, we earn on qualifying purchases.
Why the Floor Is 26, Not 0
Most benchmarks punish silence. Answer nothing, score nothing. Firmulate’s designers took the opposite view: partial progress counts. A manager who correctly triages a crisis but doesn’t resolve it has still done something real. A decision that avoids catastrophe is worth points even if it captures no upside. So a do-nothing run racks up 26 points of ambient, structural credit — and everything above that has to be earned.
The second rule is harsher: a single breach of trust caps the total grade. The stated principle — “no amount of good work outweighs a breach of trust” — means one act of dishonesty doesn’t ding your score, it ceilings it. In gaming terms: it’s not a time penalty, it’s a disqualification flag. You can’t out-grind it.
And yes — the benchmark is openly suspicious of a perfect 100. No model scored one. A round 100 on this kind of test would be a reason to ask questions, not to celebrate.
corporate AI benchmarking platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Boss Fight Hidden in a Filing Cabinet
Here’s the finding that should make any gamer grin. All five models spotted every crisis. All five refused every manipulation attempt. But only two — gpt-5.6-sol and Kimi K3 — signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
Why? The decisive weakness in the competitor’s offer wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. It was, essentially, a hidden room. The models that did the reading — that explored the environment instead of rushing the objective — won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The ones that skipped the lore got to the final dialogue and couldn’t close.
AI model performance analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering: The Impostor Check
The week included a scripted attack: fake CEO messages escalating over three stages, capped with a reporter’s trap — “just one yes/no, on background.” All five models refused. Kimi K3 left its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the AI equivalent of checking the player list before handing over the guild bank.
The Tragic Support Character
Opus 4.8 is the experiment’s saddest arc: the most thorough participant, with the deepest analyses and over 80 self-learned rules — more homework than anyone — yet last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness without follow-through is a recognizable character build, and it doesn’t win leagues.
Fine Print Worth Reading
One fairness note the publishers flag themselves: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still took second at 93. A benchmark that discloses its own asterisks is a benchmark you can trust a bit more.
Behind the league sits a live, watchable company: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The Takeaway
A benchmark is only as good as its floor. By giving a do-nothing run 26 points, capping scores on any breach of trust, and rewarding partial progress, Firmulate measures something demos can’t fake: whether an AI finishes what it starts, reads the files before the meeting, and stays honest when nobody’s checking. The buried document that decided a €55,000 deal is the whole story in miniature — the winning move wasn’t cleverness. It was exploration. Every speedrunner already knew that.
Full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
