
Imagine you run a gift shop, and a customer walks in asking for a wedding present — something personal, something that shows you understood the couple. Your best sales assistant doesn’t just hear the request; she remembers that this customer bought hand-thrown ceramics from you last spring, checks the supplier notes pinned in the back room, and comes back with the perfect recommendation at full price. Your second assistant hears the same request, gives the same charming answer… and never closes the sale, because she never walked to the back room.
That, in miniature, is what a live experiment called Firmulate just demonstrated about AI agents — and the results should matter to anyone who will ever let software touch their customer records, order books, or support inbox.
Same Company, Same Worst Week, Four Different AIs
Firmulate runs what it calls an AI company emulator: it hands a frontier AI model a small software company to run, complete with 13 synthetic employees, real money mechanics, and a public cash countdown. Every workday is versioned and auditable, and the whole thing is watchable at firmulate.com/live.
The setup was simple and cruel. Each model ran the identical company through its worst week: the same customers, the same crises, the same temptations to cut corners. Only the model changed. Every decision was logged, and the results now power a public league table.
The final standings from the Crucible League, July 2026:
- 1. gpt-5.6-sol — 95 points
- 2. Kimi K3 — 93 points
- 3. Sonnet 5 — 88 points
- 4. Fable 5 — 77 points
- 5. Opus 4.8 — 73 points
For context, a do-nothing baseline scores 26 — partial progress counts — but a single breach of trust caps the total, on the principle that no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
Everyone Passed the Character Test. Most Failed the Homework Test.
Here’s where it gets interesting for the rest of us. Every single model spotted every crisis. Every single model refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter’s tricky “just one yes/no, on background” request. Kimi K3 even left on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Five out of five models refused. Character, it turns out, is table stakes among frontier models.
But only two models signed the €55,000 deal that their own analysis had earned. The same diagnosis, the same pitch — and no signature from the others.
The Fact Buried Two Documents Deep
The reason? The decisive competitor weakness wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files — the equivalent of a supplier note tucked behind another note in the stockroom. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The models that didn’t lost it automatically.
If that doesn’t sound familiar to anyone who’s run a shop, nothing will. The difference between a good assistant and a great one is rarely charm — it’s whether they do their homework before answering.
The Thoroughness Paradox
The most counterintuitive result: Opus 4.8, the most thorough participant in the entire field — over 80 learned rules, the deepest analyses — finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating properly. The same weakness appeared, more mildly, in all four lower finishers. Working hard, it turns out, is not the same as finishing.
One fairness note worth flagging: Kimi K3 ran at its API-default effort setting while the others ran at extra-high effort — and still took second place.
Why a Home-Decor Reader Should Care
You may never run a software company. But if you sell, gift, decorate, or plan occasions, you will increasingly delegate to AI agents — answering customer emails, managing your supplier spreadsheet, drafting your seasonal promotions. The Firmulate findings translate directly:
- Honesty under pressure is now measurable — and, encouragingly, widespread among top models.
- “Reads your files before answering” is a separate, purchase-deciding skill. An agent that skips your notes will leave money on the table even when it diagnoses the problem perfectly.
- Finishing what you start is its own property. The €55k deal was earned twice and signed twice.
Firmulate also lets the public play along: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

The next time a vendor demos an AI assistant for your business, don’t ask it to write something charming. Anyone can pass the charm test — all five frontier models did. Ask it to find a fact buried two documents deep in your own files, and then ask whether it closed the deal afterward. The experiment at Firmulate suggests those are the questions that separate a 95 from a 73 — and, in your business, a loyal customer from a lost sale.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html