AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Anyone who has hired a contractor to renovate a living room knows the type: flawless mood boards, eloquent consultations, perfect paint samples — and then the project stalls at 90 percent and the final invoice never quite gets signed off. The vision was right. The diagnosis was right. The finish never came. It turns out frontier AI models behave exactly the same way when you hand them a business to run — and a live experiment called Firmulate has built a way to catch it before you hire.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get decor and gifts delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The worst week in business, on repeat

Firmulate’s setup is elegantly simple. Four frontier AI models were each given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision is versioned and auditable, so nothing rests on a chat transcript someone happened to screenshot.

The crucible league results from July 2026 tell the story. gpt-5.6-sol took first place with 95 points, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a do-nothing manager scores 26, not zero

Here’s the detail that makes this benchmark honest rather than theatrical. Before grading any AI, Firmulate runs a baseline: a manager who does essentially nothing. That baseline scores 26 points — not zero. Why? Because partial progress counts. A company left alone still keeps some of its lights on; a manager who at least avoids catastrophes has done something measurably better than one who creates them. A benchmark that starts every participant at zero flatters everyone. One that starts at 26 tells you exactly how much value the AI added above the empty chair.

Then there’s the ceiling rule, and it’s the one most benchmark designers would shy away from: a single breach of trust caps the total grade. As the experiment’s own framing puts it, “no amount of good work outweighs a breach of trust.” That’s the same standard you’d apply to a house painter who did beautiful work in three rooms and sold your furniture from the fourth. Brilliance doesn’t average out dishonesty.

And yes — the designers are openly suspicious of round 100s. A perfect score in a live, messy, judgment-heavy simulation should raise eyebrows, not applause.

Everyone diagnosed. Only some finished.

The headline finding sounds almost reassuring until you read it twice. All the models spotted every crisis. All of them refused every manipulation attempt, including a three-stage fake-CEO escalation and a reporter’s disarming “just one yes/no, on background” trick — five out of five refusals, with Kimi K3 leaving its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. It’s the AI equivalent of the designer who nails the client brief, presents beautifully, and never sends the contract.

The buried fact

Why did the others stall? The decisive competitive weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read what was already in the filing cabinet won the deal at full price — worth an additional €4,583 in monthly recurring revenue. The lesson generalizes far beyond software: the answer is often already in your own records, and the skill being tested is whether the agent bothers to look.

Thoroughness isn’t the same as finishing

The most poignant profile belongs to Opus 4.8 — the most thorough participant in the entire field, with over 80 learned rules and the deepest analyses, yet last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating the way a manager should. The same weakness appeared, weaker, in all four models. Diligence without follow-through is a familiar character in any industry, from interior design to enterprise software.

One fairness note worth flagging: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

You can watch the company run

This isn’t a one-off lab report. Firmulate operates a live company with 13 synthetic employees and real money mechanics: it burns €105,000 a month against just €2,300 in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live — the site rebuilds itself twice a day as new benchmark runs finish and the league table grows automatically. For readers who want to test their own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html.

Enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems (firmulate.com/pilot.html).

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Furnish-effect analogy holds: a beautiful room isn’t judged by the swatches, and an AI employee shouldn’t be judged by its chat. Firmulate grades what happens after the pleasant conversation — whether the agent reads your files, finishes what it starts, stays honest under pressure, and knows that one breach of trust outweighs any amount of good work. That’s a standard most human hiring processes haven’t even articulated. The fact that a do-nothing manager scores 26 — and nobody scores an unexamined 100 — is exactly what makes it worth trusting.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Pressure Pool Cleaners Market Trends 2025

Market trends for pressure pool cleaners in 2025 reveal innovative features that could transform your pool maintenance routine—discover what’s next.

How Programmable Slow Cookers Maintain Safe Temperatures

How programmable slow cookers maintain safe temperatures through advanced sensors and controls, ensuring your food stays perfectly cooked—and here’s how they do it.

How Long Do Automatic Pool Cleaners Last?

Pool cleaner lifespan varies, but understanding factors that influence longevity can help you maximize its performance and decide when to replace it.

Robotic Pool Cleaner Vs Hiring a Pool Service

Getting the right pool maintenance solution depends on your needs—discover whether a robotic cleaner or professional service suits you best.