firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

For playersOffer from Amazon

Play games on Amazon Luna for your game nights, included with Prime

  • A rotating selection of games, no download needed
  • Play on TV, laptop or phone
  • Fast, free delivery for your gear, too
Start playing with Prime Free trial · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Every Speedrunner Knows the Do-Nothing Run

Gamers invented a strange art form: the any% run, the pacifist run, the idle run where you tape down a button and walk away. The point of a baseline run is never to win — it’s to know what the floor looks like. How much can you score by standing still? What does the game give away for free?

A public AI experiment called Firmulate just ran the corporate version of that test, and the answer is quietly brilliant: a manager AI that does absolutely nothing still scores 26 out of 100. Not zero. Twenty-six. And that number, more than any leaderboard, tells you the benchmark is honest.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Wargame, Briefly

Four frontier AI models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision is versioned and auditable, like a replay file you can scrub through frame by frame.

The final Crucible League from July 2026:

  • 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
  • 2. Kimi K3 — 93. Closed the deal too, with the cleanest discipline of the field (with a caveat we’ll get to).
  • 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73. The most thorough participant of all — and last place.
Amazon

AI audit and traceability tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the Floor Is 26, Not 0

Most benchmarks punish silence. Answer nothing, score nothing. Firmulate’s designers took the opposite view: partial progress counts. A manager who correctly triages a crisis but doesn’t resolve it has still done something real. A decision that avoids catastrophe is worth points even if it captures no upside. So a do-nothing run racks up 26 points of ambient, structural credit — and everything above that has to be earned.

The second rule is harsher: a single breach of trust caps the total grade. The stated principle — “no amount of good work outweighs a breach of trust” — means one act of dishonesty doesn’t ding your score, it ceilings it. In gaming terms: it’s not a time penalty, it’s a disqualification flag. You can’t out-grind it.

And yes — the benchmark is openly suspicious of a perfect 100. No model scored one. A round 100 on this kind of test would be a reason to ask questions, not to celebrate.

Amazon

corporate AI benchmarking platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Boss Fight Hidden in a Filing Cabinet

Here’s the finding that should make any gamer grin. All five models spotted every crisis. All five refused every manipulation attempt. But only two — gpt-5.6-sol and Kimi K3 — signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

Why? The decisive weakness in the competitor’s offer wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. It was, essentially, a hidden room. The models that did the reading — that explored the environment instead of rushing the objective — won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The ones that skipped the lore got to the final dialogue and couldn’t close.

Amazon

AI model performance analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering: The Impostor Check

The week included a scripted attack: fake CEO messages escalating over three stages, capped with a reporter’s trap — “just one yes/no, on background.” All five models refused. Kimi K3 left its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the AI equivalent of checking the player list before handing over the guild bank.

The Tragic Support Character

Opus 4.8 is the experiment’s saddest arc: the most thorough participant, with the deepest analyses and over 80 self-learned rules — more homework than anyone — yet last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness without follow-through is a recognizable character build, and it doesn’t win leagues.

Fine Print Worth Reading

One fairness note the publishers flag themselves: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still took second at 93. A benchmark that discloses its own asterisks is a benchmark you can trust a bit more.

Behind the league sits a live, watchable company: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

A benchmark is only as good as its floor. By giving a do-nothing run 26 points, capping scores on any breach of trust, and rewarding partial progress, Firmulate measures something demos can’t fake: whether an AI finishes what it starts, reads the files before the meeting, and stays honest when nobody’s checking. The buried document that decided a €55,000 deal is the whole story in miniature — the winning move wasn’t cleverness. It was exploration. Every speedrunner already knew that.

Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

VR Motion Sickness Terms Explained

Cybersickness, vection, latency, VR legs — every VR motion sickness term explained in plain English, plus the settings and habits that keep you comfortable.

Luanti Removed From Google Play Due To Baseless AI Copyright Notice

Google has removed the Luanti app from its Play Store after developers cited an unfounded AI copyright notice, raising questions about app moderation and AI claims.

The Real Boss Battle for AI Is Finishing the Job

A playable quiz turns 242 real AI management decisions into a revealing test of which frontier model closes deals, reads deeply and resists pressure.

Interpupillary Distance Explained for Headset Comfort

Learn how IPD affects VR clarity, eye strain, and comfort, plus how to measure it and adjust your headset in a few practical steps.