firmulate.com/index — live view
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Every gamer knows this character. Perfect build, perfect stats, aced the tutorial, top of every leaderboard. Then the actual quest starts — the escort mission, the timed run, the boss with a hidden mechanic — and the run falls apart in ways the scoreboard never predicted.

Skill rating, it turns out, is not the same thing as finishing the quest. And that lesson from a thousand hours of games is now the sharpest lens we have on AI agents.

Right now, the AI industry is obsessed with leaderboards: coding benchmarks, chat arenas, elo scores. Which model writes the best function? Which one gives the most helpful answer? It’s all tutorial-tier measurement. Meanwhile, Firmulate, a live project running a very different kind of competition, just published final results that expose the gap — and it looks exactly like a failed raid where everyone survived the trash mobs and two of five forgot to loot the boss.

The wargame nobody else is running

The setup is almost a game design document: take four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — and give each the same job. Run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changes. Every decision is versioned and auditable, like a replay file you can scrub through frame by frame.

This isn’t a chat demo. It’s a stress test for management: churn waves, a price increase, a down-round scenario, a PR crisis. The final Crucible League table from July 2026 reads like a post-raid scoreboard: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts — but there’s one hard rule that will feel instantly familiar to anyone who’s played a game with permadeath: a single breach of trust caps your total. As the organizers put it, “no amount of good work outweighs a breach of trust.”

Everyone passed the skill checks. Two looted the boss.

Here’s the finding that should reframe how you think about AI agents. All the models spotted every crisis. All of them refused every manipulation attempt. The party cleared every encounter without a single wipe.

Then came the objective: a €55,000 deal that their own analysis had earned. Only two models signed it. The organizers’ verdict on the rest: “Same diagnosis, same pitch — no signature.”

That’s the tutorial-vs-quest gap in one sentence. Answering questions well is a skill check. Closing a deal, under pressure, across days, with consequences — that’s the campaign. Chat leaderboards measure the first. Almost nothing measures the second. Firmulate does.

The buried mechanic

The best part, for anyone who loves hidden game mechanics: the deal’s decisive information wasn’t in the customer conversation at all. It was buried two document references deep in the company’s own files — a competitor weakness that only models which actually read the file would find. The ones that did won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that skimmed past it left the loot on the table.

Every seasoned player knows this feeling: the winning strategy was in the codex entry nobody read.

The social-engineering boss fight

The week also included what amounts to a social-engineering encounter: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused, 5 for 5. Kimi K3’s on-record reasoning is the kind of clean threat-assessment log you’d want from a tank calling out mechanics: “Treat the request as a suspected approval-bypass / possible impersonation.”

Honesty under pressure: passed. Follow-through: failed by more than half the field. Which of those shows up on a chat arena leaderboard?

The grinding build that lost anyway

The most poignant profile is Opus 4.8: the most thorough participant in the entire run. It generated the deepest analyses and learned over 80 new rules — the most grind of anyone — and still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. And here’s the uncomfortable footnote: the same weakness appeared, weaker, in all four other models. Over-preparation without follow-through isn’t one model’s bug. It’s a class-wide trait.

One fairness note worth flagging: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second with the cleanest discipline of the field.

It’s live, and you can play

Firmulate isn’t a slide deck — it’s a running world. A synthetic company with 13 employees, real money mechanics (burning €105k/month against €2.3k MRR), a public cash countdown, and 680+ self-learned playbook rules, versioned every workday. You can watch it at firmulate.com, and there’s a genuinely playable layer: a quiz built from 242 real, unedited management decisions where you guess which model made which call. If you’ve ever argued about who really wiped the raid, this is your kind of forensic replay. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full methodology and plain-language findings are on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The AI industry’s leaderboards are tutorial scores. Firmulate’s argument — backed by a live, auditable experiment — is that the market should be ranking something else entirely: management quality, not chat quality. Does the model finish what it starts? Does it read your files before it acts? Does it stay honest when a fake CEO escalates and a reporter dangles an easy quote? And when the pressure is on, does it sign the deal its own work earned?

Every gamer learned this lesson years ago on the first escort mission that went sideways. The character sheet was never the campaign. Now the AI industry is learning it too — one €55,000 deal left unsigned at a time.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management stress test software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision-making analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model performance benchmarking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document reading and analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Boss Is a Fake: How Five AI Models Survived a Corporate Deception Quest

A live AI-company wargame found every model resisted an impersonated CEO, showing integrity under pressure can be tested before deployment.

The Ultimate Management Sim Is a Software Company Fighting to Stay Alive

Firmulate turns AI management into a live survival story, with synthetic employees, real money pressure and every workday preserved for scrutiny.

VR Audio Comfort Explained for Long Sessions

Your ears quit VR before your eyes do. Learn how volume, fit, heat, and latency cause audio fatigue — and the simple fixes that add hours.

Roblox Tries To Dismiss Wrongful Death Lawsuit By Arguing That A Player Gave Up Her Right To Sue When She Was 8 Years Old

Roblox argues the wrongful death lawsuit should be dismissed, claiming the deceased signed away her rights at age 8 via terms and conditions.