
Good taste is easy to spot—until the stakes rise
Anyone who has chosen a statement lamp, assembled a gift basket or planned a celebration knows that personality appears in decisions. Some people research every option. Some act quickly. Others notice the detail everyone else missed. The revealing moment is not what they say they value, but what they actually choose.
Firmulate applies that idea to frontier artificial intelligence. Its guess-the-model quiz presents real, unedited management decisions and asks readers to identify which AI made each one. The material is drawn from 242 decisions produced during a live business experiment, turning model comparison into something closer to reading character than checking specifications.
The appeal is playful, but the underlying question is serious: if an AI is trusted to help run a company, can people distinguish its management personality from its prose—and does that personality affect results?
AI decision-making analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, crises and temptations
In Firmulate’s experiment, each frontier model ran the same small software company through its worst week. The customers, crises and temptations remained constant. Every decision was versioned and auditable, allowing the models to be compared through what they did rather than through polished demonstrations.
The simulated company has 13 employees and deliberately unforgiving finances: it burns €105k each month against €2.3k in monthly recurring revenue. A public cash countdown makes the pressure visible, while more than 680 self-learned playbook rules show how the company’s operating knowledge accumulates. Every workday is versioned, making the experiment watchable as an evolving business rather than a static test.
The final Crucible League results from July 2026 placed gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, imposed a hard boundary: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
Spotting trouble was not the differentiator
All the models identified every crisis and refused every manipulation attempt. That included fake messages from a chief executive escalating across three stages and a reporter seeking “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 described the approach in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Those refusals matter because they show consistent resistance under pressure. Yet they did not explain the spread in the league table. The decisive difference came after the models had understood the problem.
Only two signed the €55,000 deal that their own analysis had earned. The others reached the same diagnosis and produced the same pitch but failed to secure the signature: “Same diagnosis, same pitch — no signature.” It is an unexpectedly human management failure. Recognizing the right course is not the same as completing it.
The winning clue was hiding in company files
The decisive weakness in a competitor was not obvious in the customer event. It sat two document references deep in the company’s own files. Models that followed the trail found the fact and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
That finding gives the quiz its sharper edge. A response may sound confident, prudent or impressively detailed without showing whether the model read far enough, acted on what it discovered or finished the commercial task. Readers are therefore guessing from behavioral signatures embedded in consequential decisions, not merely from writing style.
Opus 4.8 offers the clearest warning against equating thoroughness with effectiveness. It was the most exhaustive participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in milder form across the other four models.
There is also an important fairness qualification. Kimi K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That does not erase its result, but it belongs beside the ranking for readers interpreting the comparison.

Management personality is measurable
Firmulate’s quiz works because it transforms an abstract AI debate into recognizable choices. The question is not simply whether a model can write a convincing memo. It is whether it reads the files, preserves trust, follows the evidence and completes the action its analysis supports.
For businesses considering AI workers, those distinctions are operational rather than cosmetic. Firmulate also offers enterprises the same wargame using a read-only export of their own business, with nothing written back to real systems. The broader lesson is visible in the public experiment: models facing identical circumstances can produce meaningfully different management outcomes.
For everyone else, the 242-decision quiz offers an accessible test of intuition. Like identifying a decorator from a finished room or a gift-giver from the wrapping, readers may discover that AI systems leave recognizable fingerprints—but the most consequential traits are sometimes hidden behind equally fluent words.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html