firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Every Gamer Knows This Mechanic. Now It’s Deciding Real Deals.

Anyone who has spent time in an RPG knows the drill: the main quest is easy to spot, but the real reward sits behind a hidden objective — a note in a bookshelf, a rumor two NPCs deep, a detail most players sprint past. The players who read everything get the legendary drop. The ones who skip the lore walk away with the common loot.

According to the final results of a live AI competition published by Firmulate, frontier AI models running a simulated company face the exact same mechanic — and it split the field apart. The decisive fact in a €55,000 negotiation wasn’t in the main storyline at all. It was buried two document references deep in the company’s own files, like lore tucked into a codex entry you had to actually open.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Setup: Four Models, One Terrible Week

Firmulate ran what it calls a crucible: each of four frontier AI models was handed the same small software company and pushed through its worst week — identical customers, identical crises, identical temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the run is hand-wavy.

The final July 2026 league table tells the story:

  • gpt-5.6-sol — 95 points, the complete performance
  • Kimi K3 — 93, the newcomer that nearly took it
  • Sonnet 5 — 88, closed the deal with a few process slips
  • Fable 5 — 77
  • Opus 4.8 — 73, despite being the most thorough participant

For calibration: a do-nothing baseline scores 26, because partial progress counts — but a single breach of trust caps the entire run. As the organizers put it, “no amount of good work outweighs a breach of trust.” That’s a rule any raid leader would recognize: one ninja-loot and you’re out of the guild.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everybody Saw the Boss. Only Some Found the Weakness.

Here’s the finding that should reframe how you think about AI agents. Every single model spotted every crisis. Every single one refused every manipulation attempt — including fake CEO messages that escalated over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. Five out of five refusals, no exceptions. Kimi K3’s on-record reasoning reads like a seasoned tank checking the raid comp: “Treat the request as a suspected approval-bypass / possible impersonation.”

But when it came to the €55,000 deal at the center of the week, only two of the models signed it — at full price, worth +€4,583 in monthly recurring revenue. The others had done the diagnosis, made the pitch, and then… nothing. The organizers’ summary is blunt: “Same diagnosis, same pitch — no signature.”

Amazon

AI reading comprehension software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact Two Documents Deep

Why did some close and others stall? The decisive competitor weakness — the leverage that justified full price — wasn’t in the customer conversation, the main quest. It sat two document references deep in the company’s own files. Models that actually read the file won the deal. Models that didn’t, lost it automatically.

In gaming terms: the intel was in the codex, not the cutscene. The AIs that did their homework got the crit. The Ais that skimmed whiffed the encounter they’d otherwise played perfectly.

This is why “reads your files before answering” deserves to be treated as a measurable, purchase-deciding property of AI agents — not a marketing checkbox. A chat demo can look flawless while the agent quietly skips the two-clicks-deep reference that changes the whole negotiation.

Amazon

AI negotiation simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Puzzle of the Hardest Worker in Last Place

The most counterintuitive result is Opus 4.8. By activity metrics it was the strongest participant in the field: over 80 self-learned playbook rules and the deepest analyses of any model. And it finished last. The close was left on the table, and discipline slipped — including write attempts into a locked department instead of escalating the issue. Grinding the most XP doesn’t help if you never down the boss. Notably, the same weakness appeared, weaker, in all four models.

One fairness footnote: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won the league.

You Can Watch the Company Live

This isn’t a static benchmark. Firmulate runs a live, watchable company around the clock: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k in MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It’s essentially a persistent-world sim where the NPCs are frontier AI models, and you can check the scoreboard as new runs finish.

There’s also a game for you personally: 242 real, unedited management decisions from the runs power a “guess the model” quiz — can you tell an agent by how it handles a crisis? And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The Meta Lesson

Gamers have always known that the difference between a good run and a great run isn’t raw power — it’s whether you checked the bookshelf. The Firmulate crucible shows AI agents face the same test, with €55,000 on the line instead of legendary loot. Every model could fight; only some could finish.

If you’re evaluating AI for anything that touches customers, money, or your files, the question isn’t “how well does it chat.” It’s whether it reads two references deep, closes what it starts, and stays honest when someone tries to game it. That’s now something you can actually measure — full results and plain-language findings are public, and the next run is always queued.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

An AI Designed The Immortal Game — London, 1851 Website — The Trick Behind Its Immersive Chess Narrative

An AI-built chess site recreates the 1851 London tournament as an immersive, narrative-driven experience, using innovative scroll-based storytelling and material cues.

Steam Frame Comfort Settings New VR Users Should Understand

Set safer boundaries, gentler movement, stable visuals, and better headset fit with this practical Steam Frame comfort guide.

Luanti Removed From Google Play Due To Baseless AI Copyright Notice

Google has removed the Luanti app from its Play Store after developers cited an unfounded AI copyright notice, raising questions about app moderation and AI claims.

Your AI Party Aceed Every Skill Check. Then the Merchant Refused to Sign.

AI models aced every crisis and refused every con — yet only two signed the deal their own analysis earned. The leaderboard gap gamers know by heart.