
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Every Gamer Knows This Mechanic. Now It’s Deciding Real Deals.
Anyone who has spent time in an RPG knows the drill: the main quest is easy to spot, but the real reward sits behind a hidden objective — a note in a bookshelf, a rumor two NPCs deep, a detail most players sprint past. The players who read everything get the legendary drop. The ones who skip the lore walk away with the common loot.
According to the final results of a live AI competition published by Firmulate, frontier AI models running a simulated company face the exact same mechanic — and it split the field apart. The decisive fact in a €55,000 negotiation wasn’t in the main storyline at all. It was buried two document references deep in the company’s own files, like lore tucked into a codex entry you had to actually open.
As an affiliate, we earn on qualifying purchases.
The Setup: Four Models, One Terrible Week
Firmulate ran what it calls a crucible: each of four frontier AI models was handed the same small software company and pushed through its worst week — identical customers, identical crises, identical temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the run is hand-wavy.
The final July 2026 league table tells the story:
- gpt-5.6-sol — 95 points, the complete performance
- Kimi K3 — 93, the newcomer that nearly took it
- Sonnet 5 — 88, closed the deal with a few process slips
- Fable 5 — 77
- Opus 4.8 — 73, despite being the most thorough participant
For calibration: a do-nothing baseline scores 26, because partial progress counts — but a single breach of trust caps the entire run. As the organizers put it, “no amount of good work outweighs a breach of trust.” That’s a rule any raid leader would recognize: one ninja-loot and you’re out of the guild.
As an affiliate, we earn on qualifying purchases.
Everybody Saw the Boss. Only Some Found the Weakness.
Here’s the finding that should reframe how you think about AI agents. Every single model spotted every crisis. Every single one refused every manipulation attempt — including fake CEO messages that escalated over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. Five out of five refusals, no exceptions. Kimi K3’s on-record reasoning reads like a seasoned tank checking the raid comp: “Treat the request as a suspected approval-bypass / possible impersonation.”
But when it came to the €55,000 deal at the center of the week, only two of the models signed it — at full price, worth +€4,583 in monthly recurring revenue. The others had done the diagnosis, made the pitch, and then… nothing. The organizers’ summary is blunt: “Same diagnosis, same pitch — no signature.”
AI reading comprehension software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Buried Fact Two Documents Deep
Why did some close and others stall? The decisive competitor weakness — the leverage that justified full price — wasn’t in the customer conversation, the main quest. It sat two document references deep in the company’s own files. Models that actually read the file won the deal. Models that didn’t, lost it automatically.
In gaming terms: the intel was in the codex, not the cutscene. The AIs that did their homework got the crit. The Ais that skimmed whiffed the encounter they’d otherwise played perfectly.
This is why “reads your files before answering” deserves to be treated as a measurable, purchase-deciding property of AI agents — not a marketing checkbox. A chat demo can look flawless while the agent quietly skips the two-clicks-deep reference that changes the whole negotiation.
As an affiliate, we earn on qualifying purchases.
The Puzzle of the Hardest Worker in Last Place
The most counterintuitive result is Opus 4.8. By activity metrics it was the strongest participant in the field: over 80 self-learned playbook rules and the deepest analyses of any model. And it finished last. The close was left on the table, and discipline slipped — including write attempts into a locked department instead of escalating the issue. Grinding the most XP doesn’t help if you never down the boss. Notably, the same weakness appeared, weaker, in all four models.
One fairness footnote: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won the league.
You Can Watch the Company Live
This isn’t a static benchmark. Firmulate runs a live, watchable company around the clock: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k in MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It’s essentially a persistent-world sim where the NPCs are frontier AI models, and you can check the scoreboard as new runs finish.
There’s also a game for you personally: 242 real, unedited management decisions from the runs power a “guess the model” quiz — can you tell an agent by how it handles a crisis? And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

The Meta Lesson
Gamers have always known that the difference between a good run and a great run isn’t raw power — it’s whether you checked the bookshelf. The Firmulate crucible shows AI agents face the same test, with €55,000 on the line instead of legendary loot. Every model could fight; only some could finish.
If you’re evaluating AI for anything that touches customers, money, or your files, the question isn’t “how well does it chat.” It’s whether it reads two references deep, closes what it starts, and stays honest when someone tries to game it. That’s now something you can actually measure — full results and plain-language findings are public, and the next run is always queued.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.