AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

If you run a home decor shop or a gift business, you already know the difference between a lovely salesperson and a great manager. The lovely one chats with every customer, remembers birthdays, wraps everything beautifully. The great one does that and notices the supplier invoice that quietly doubled, catches the fake “urgent message from the owner” before it’s acted on, and still closes the big wedding-favour order before Friday.

Most tests of AI models measure the lovely salesperson. A project called Firmulate set out to measure the manager.

The worst week, on repeat

Here’s the setup: four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 ran the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing depended on a judge’s mood.

Think of it as a mystery-shop test, except the shop is a whole company and the mystery shopper is a week from hell: a churn wave, a price increase, a downround, a PR crisis.

Amazon

AI sales assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone passed the etiquette exam

The headline finding is oddly reassuring. All of the models spotted every crisis. All of them refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter’s “just one yes/no, on background” trick. Five out of five refusals. Kimi K3’s reasoning was put on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

So far, so good. The models are polite, alert and honest. That’s the part a chat demo shows you.

Then came the money question

Only two of the models actually signed the €55,000 deal that their own analysis had earned. The gap was summed up in one line: “Same diagnosis, same pitch — no signature.”

In retail terms: your assistant gives the customer a flawless consultation, recommends exactly the right pieces, the customer reaches for their card — and the assistant wanders off to tidy the shelves.

The final league table told the story. gpt-5.6-sol won with 95. Kimi K3 took second at 93 — remarkable for a newcomer, and it closed the deal with what the results call the cleanest discipline of the field. (One fairness footnote: K3 ran at its default effort setting while the others ran at high effort.) Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 landed last at 73. A do-nothing baseline scored 26.

The fact buried in the filing cabinet

The most instructive detail of the whole experiment: the decisive weakness in the customer’s current supplier wasn’t in the sales conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal — at full price, worth +€4,583 in monthly recurring revenue.

Any shop owner will recognize this instantly. The best salespeople aren’t the smoothest talkers; they’re the ones who did the homework before the customer walked in.

When being thorough isn’t enough

The most poignant profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place. It left the close on the table, and its discipline slipped at one point, making write attempts into a locked department instead of escalating. The same weakness, in weaker form, showed up in all four models.

And one rule towered over everything: partial progress counted, but a single breach of trust capped the total. As the scoring puts it, “no amount of good work outweighs a breach of trust.” Most of us run our businesses the same way.

You can watch it lose money

Firmulate isn’t a slide deck. There’s a live company running every business day — 13 synthetic employees, real money mechanics, a burn of €105k a month against €2.3k of MRR, a public cash countdown, and more than 680 self-learned playbook rules. Every workday is versioned, and you can watch it at firmulate.com.

There’s also a genuinely fun bit: 242 real, unedited management decisions power a “guess which model made this call” quiz — a bit like guessing whether the shop assistant or the owner wrote that email. And if you’d rather see how your own business holds up, enterprises can run the same wargame against a read-only export of their own operations; nothing ever writes back to real systems.

Full results and plain-language findings are on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

If AI agents will ever touch your order pipeline, your customer list or your forecast, the question isn’t “does it write well?” It’s: does it finish what it starts, does it read your files before it speaks, does it stay honest when someone pretends to be the boss — and does it actually take the customer’s money at the end? Charm is common in AI now. Follow-through is the scarce commodity. Measure the manager, not the conversationalist.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The Science Behind Countertop Nugget Ice Makers

Gaining insight into the science behind countertop nugget ice makers reveals how innovative engineering creates irresistibly soft, slow-melting ice—discover the fascinating details inside.

Overview of Automatic Pool Cleaners Industry

Just as technology transforms everyday chores, the automatic pool cleaners industry is evolving rapidly—discover how innovation is reshaping pool maintenance.

Understanding Advanced Features in Robotic Pool Cleaners

Navigating the latest advanced features in robotic pool cleaners reveals how smart technology can transform your pool maintenance routine, and there’s more to discover.

The Surprising Physics Behind Silent Washing Machines

Silent washing machines harness physics to minimize noise through innovative damping and insulation techniques—discover how science makes laundry quieter than ever.