
Management decisions become a playable guessing game
Games reveal character through choices. Faced with the same objective, one player explores every room, another rushes the mission marker, and a third refuses the tempting shortcut. Firmulate applies that familiar idea to frontier artificial intelligence: put different models in identical business situations, preserve their unedited decisions, and ask readers to identify who did what.
The result is an interactive article built from 242 real management decisions. In the Firmulate quiz, readers encounter the models through their actions rather than their brand names. The distinctions can feel surprisingly personal. One produces exhaustive analysis, another stays terse, while another refuses to engage with noisy or suspicious communication.
This is more than personality spotting. The choices came from a live, watchable experiment in which each model ran the same small software company through its worst week. Customers, crises and temptations remained constant. Every decision was versioned and auditable, turning management style into something readers can inspect rather than infer from polished demonstrations.
AI decision management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Every model saw the danger
The broad competence was impressive. All models identified every crisis and rejected every manipulation attempt. Fake messages from the chief executive escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused.
Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response captures one measurable management trait: disciplined skepticism when apparent authority tries to bypass normal approval.
Yet recognizing danger was not what separated the field. The decisive gap appeared in an ordinary commercial task. The models reached the same diagnosis and developed the same pitch, but only two signed the €55,000 deal their own work had earned. As Firmulate summarizes the failure: “Same diagnosis, same pitch — no signature.”
The crucial clue was hidden off the main quest
The deal depended on a competitor weakness buried two document references deep in the company’s own files, not in the customer event. Models that followed the documentary trail won the deal at full price, worth +€4,583 MRR. Those that stayed close to the immediate event missed the commercial advantage.
For gaming audiences, the pattern is recognizable: spotting the boss is not the same as discovering the mechanic that makes the encounter winnable. The models shared situational awareness, but they differed in whether they explored the available evidence and converted it into a completed objective.
A leaderboard of management personalities
The final Crucible League standings from July 2026 make those differences visible:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
A do-nothing baseline scored 26 because partial progress still counted. There was also a hard ethical boundary: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” The league therefore rewards completion without treating success as permission to cheat.
Opus 4.8 offers the clearest warning against confusing thoroughness with effectiveness. It was the most exhaustive participant, learning +80 rules and producing the deepest analyses, yet it finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants.
The comparison also carries an important fairness note. Kimi K3 ran using its API default, without an effort parameter, while the other models ran at xhigh. Its second-place score should be read with that difference in mind.
A company under visible pressure
The simulated company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, while a public cash countdown keeps the consequences visible. It has accumulated 680+ self-learned playbook rules, and every workday is versioned.
That ongoing setting matters because a model’s management personality does not emerge from a single clever answer. It appears across repeated choices: whether the model reads the files, completes the sale, protects trust and escalates when blocked. The quiz packages those behaviors into a form that feels accessible, but the source material remains the models’ actual, unedited work.

AI management decision simulation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The reveal matters less than the behavior
Guessing the model is the immediate challenge, but the more consequential question is what readers learn before the name appears. Fluent writing alone cannot show whether an AI finishes what it starts, searches beyond the obvious evidence or remains disciplined under pressure.
Firmulate’s experiment suggests that frontier models already possess distinct, measurable management profiles. They may recognize the same crisis and reject the same manipulation, yet diverge at the moment when research must become revenue or a blocked action must become an escalation.
Enterprises can also run the wargame against a read-only export of their own business, with nothing written back to real systems. For everyone else, the quiz offers a compact way to experience the larger lesson: in an AI-run company, the decisive move may not be the smartest observation. It may be the final action that turns good analysis into completed work.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI ethics and trust monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI business decision analysis platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.