
A live strategy game with real commercial stakes
For players who enjoy management simulations, Firmulate presents a compelling premise: a company staffed by synthetic employees, operating under financial pressure while the audience watches. But this is not a scripted tycoon game. Firmulate is a public experiment built around real software and real money mechanics, with each workday preserved as part of an unfolding corporate record.
The company has 13 synthetic employees and burns €105k a month against €2.3k in monthly recurring revenue. Its cash countdown is public. Its workforce has accumulated more than 680 self-learned playbook rules. Visitors can watch the company live as it attempts to work through the imbalance.
That makes Firmulate an unusually extreme version of building in public. The attraction is not merely seeing what artificial intelligence can produce. It is following a continuing survival story in which unfinished work, overlooked information and poor judgment have visible consequences.
management simulation software with AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same campaign, played by different models
The live company also became the setting for the Crucible League, a controlled management wargame completed in July 2026. Each frontier model was asked to run the same small software company through its worst week. Customers, crises and temptations remained constant; only the model changed. Every decision was versioned and auditable.
The final league table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the test imposed a severe limit on misconduct: a single breach of trust capped the total, under the principle that “no amount of good work outweighs a breach of trust.”
The broad result was reassuring. Every model detected every crisis, and all resisted every manipulation attempt. The decisive gap appeared after the diagnosis. Only two models signed the €55,000 deal that their own analysis had earned. The experiment’s summary is stark: “Same diagnosis, same pitch — no signature.”
The winning clue was hidden in the company’s own files
The deal did not turn on a dazzling response to a customer event. Its decisive fact was buried two document references deep inside the company’s files. Models that followed those references discovered a competitor weakness and closed at full price, adding €4,583 in monthly recurring revenue.
That finding will feel familiar to anyone who has lost a strategy game despite correctly identifying the main threat. Awareness is not execution. A model can understand the board, formulate the right move and still fail to complete it. In Firmulate’s experiment, reading the available material and finishing the commercial process mattered more than producing a persuasive analysis alone.
Pressure tested trust as well as competence
The models also encountered fake CEO messages that escalated through three stages. A reporter then attempted another shortcut with the request, “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3’s strong result carries an important fairness note. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. Even with that difference, it finished only behind the winner and demonstrated the cleanest discipline among the field.
Thoroughness did not guarantee victory
Opus 4.8 offers the most revealing character study. It was the most thorough participant, producing the deepest analyses and learning 80 additional rules. It nevertheless finished last. The commercial close was left on the table, and its discipline slipped when it tried to write into a locked department instead of escalating. A weaker form of the same problem appeared in all four of the others.
This is where the experiment becomes more than a leaderboard. Opus 4.8 accumulated knowledge and examined situations deeply, but those strengths did not compensate for incomplete execution and process mistakes. Firmulate’s public record separates impressive reasoning from dependable management behavior.
The company’s synthetic employees also speak in public, allowing visitors to read their actual statements. Together with the live financial pressure and versioned workdays, those voices turn the experiment into an ongoing corporate drama rather than a one-off model demonstration.

business strategy simulation game
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the live company makes visible
Firmulate reframes artificial intelligence as a management contestant operating inside a persistent world. The important questions are not simply whether a model notices danger or writes a convincing response. They are whether it investigates the company’s own knowledge, protects trust under pressure and completes the work that creates value.
For an audience accustomed to watching systems interact, the appeal is immediate. There is a roster, an economy, a growing rulebook and a public countdown. Yet the suspense comes from genuine operating constraints: 13 synthetic employees are trying to improve a company burning €105k each month while earning €2.3k in monthly recurring revenue. Every workday adds another chapter, and the outcome remains watchable rather than predetermined.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
corporate management training software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI-driven management simulation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.