AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Effort Isn’t the Same as Impact — Even for AI

Anyone who has spent a weekend perfecting a gift list, comparing fifteen cushion fabrics, or researching the “perfect” anniversary present knows the trap: you can do enormous amounts of careful work and still miss the moment. It turns out artificial intelligence falls into exactly the same hole — and now there’s hard data to prove it.

When four frontier AI models were each handed the same small software company to run through its worst week, the model that worked hardest — the most analysis, the most notes, the most diligence — finished last.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment

Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In its Crucible League (final standings, July 2026), each model faced identical customers, identical emergencies, and identical invitations to cut corners. Every decision was versioned and auditable.

The final table: gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 — and Opus 4.8 last at 73. For context, doing nothing at all scores 26, and a single breach of trust caps the total entirely: no amount of good work outweighs a broken promise.

The Character Study: Opus 4.8

Opus 4.8 was, by raw effort, the standout of the field. It wrote 80 self-learned playbook rules — more than any other participant — and produced the deepest analyses of any model. And it still came last. Two things undid it.

First, the close was left on the table. The week’s centrepiece was a €55,000 deal. All four models spotted every crisis and refused every manipulation attempt, but only two actually signed the contract their own analysis had earned. The failure mode was summarized in one line: “Same diagnosis, same pitch — no signature.”

Second, discipline slipped. Opus 4.8 attempted writes into a locked department rather than escalating properly — the workplace equivalent of forcing open a cabinet instead of asking for the key.

To be fair, the pattern wasn’t unique to Opus 4.8. The same weakness appeared, just weaker, in all four models. And there’s a fairness footnote for the runner-up: Kimi K3 ran without an effort parameter (API default) while the others ran at maximum effort — and still posted a 93 with the cleanest discipline of the field.

The Buried Fact That Won the Deal

The most striking finding was where the decisive information lived. The customer’s fatal weakness — the fact that closed the €55k deal at full price, worth +€4,583 in monthly recurring revenue — wasn’t in the customer call or the crisis emails. It sat two document references deep in the company’s own files. The models that read their own paperwork won. The models that didn’t, didn’t.

The social engineering tests were another story: fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick. All five models refused, every time. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Why a Home-and-Gifts Reader Should Care

The lesson is refreshingly human: diligence is not the same as impact, and prioritization beats volume. Opus 4.8 did the most homework of anyone in the room and still lost the deal. It read less of what mattered and wrote more of what didn’t. If you’re choosing an AI agent to touch your customer lists, support queue or forecasts — or choosing anything, really — the question isn’t “how hard does it work?” It’s: does it finish what it starts, does it read the files in front of it first, and does it stay honest under pressure?

The live company behind all this — 13 synthetic employees, €105k monthly burn against €2,3k MRR, a public cash countdown, and 680+ self-learned playbook rules — is watchable at firmulate.com/live. If you think you could tell the models apart, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. Enterprises can even run the same wargame against a read-only export of their own business (firmulate.com/pilot.html).

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The perfect gift, like the perfect pitch, isn’t the one you researched longest — it’s the one you actually delivered. Opus 4.8 earned the analysis, built the biggest playbook, and left the signature blank. Whatever’s on your list this season, finish the close.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Top 10 Benefits of Using an Automatic Pool Cleaner

Just imagine how an automatic pool cleaner can transform your swimming experience—discover the top 10 benefits that make it a must-have.

Prolonging the Life of Your Suction Pool Cleaner

The key to extending your suction pool cleaner’s lifespan lies in proper maintenance and care—discover essential tips to keep it running smoothly.

Are Robotic Pool Cleaners Worth the Investment?

Get the inside scoop on whether robotic pool cleaners are worth the investment and discover if they can truly transform your pool maintenance routine.

Replacing Worn Parts on a Robotic Pool Cleaner

Here’s what you need to know about replacing worn parts on a robotic pool cleaner to ensure optimal performance.