AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get decor and gifts delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

You Wouldn’t Hire a Contractor From a Brochure

Anyone who has ever redone a kitchen knows the drill. Every contractor interviews beautifully. Every portfolio looks flawless. And yet some of them leave the job half-finished, some miss the wiring note buried on page two of the building plans, and a rare few actually close up the walls, clean the site, and hand over the keys on price.

Choosing an AI model for your business turns out to work the same way — and a live, public experiment just proved it in a way no chat demo ever could.

In the final July 2026 standings of the Crucible league, run by the AI company emulator Firmulate, a newcomer — Moonshot’s Kimi K3 — scored 93 and beat three of four Western frontier models at the most mundane, most human job there is: running a small software company through its worst week.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Crisis, Same Temptations

The setup is elegantly unfair. Firmulate handed each frontier model the identical small software firm — same customers, same crises, same temptations to cut corners — and let it manage. Every decision was versioned and auditable, so nothing could be hand-waved after the fact.

The final league table: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing nothing at all scores 26 — and a single breach of trust caps the total, because, as the experiment’s own rule puts it, “no amount of good work outweighs a breach of trust.”

The Finding That Should Worry Every Buyer

Here is the part that chat demos never reveal. All five models spotted every crisis. All five refused every manipulation attempt. But only two — gpt-5.6-sol and Kimi K3 — actually signed the €55,000 deal their own analysis had earned. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.”

The deal hinged on a buried fact: the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer conversation. The models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It is the AI equivalent of the contractor who never opened the building plans.

The Newcomer’s Clean Week

Kimi K3’s week reads like the tradesperson you recommend to neighbors. It found the buried security needle, won the €55,000 deal, saved the churning customer, and resisted all three bait attempts — fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick — with only a single deviation, the cleanest discipline in the field. Its on-record reasoning for refusing the fake CEO: “Treat the request as a suspected approval-bypass / possible impersonation.”

The most instructive profile, though, belongs to Opus 4.8. It was the most thorough participant — over 80 learned rules, the deepest analyses — and still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. Effort, in other words, is not the same as judgment. Firmulate notes the same weakness appeared, weaker, in all four models.

This Isn’t a Simulation on a Slide

The company the models manage is live software: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k in MRR — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it in real time at firmulate.com. There is even a “guess the model” quiz built from 242 real, unedited management decisions.

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — a caveat worth keeping in mind when comparing scores.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The League Is Open

The comfortable assumption that a handful of familiar Western frontier models sit in a class of their own just took a hit. A newcomer, running at default effort, outmanaged three of them. For anyone deploying AI against a CRM, a support queue, or a forecast, the lesson is the same one home-decor veterans learned long ago: references beat brochures. Picking a model without testing it on your own worst week is now a bet, not a decision.

Enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Until then, the leaderboard is public, the company is losing money in public, and the race is very much on.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Heat‑Pump Dryers vs. Ventless Models: What Science Says About Savings

Gaining insights into heat-pump versus ventless dryers reveals surprising savings benefits that could change your laundry choices—discover what science says.

UV‑C in Air Purifiers: What It Can Do (and What It Can’t)

Beneath its powerful surface, UV-C in air purifiers can neutralize pathogens, but understanding its true capabilities and limitations is essential for safe, effective use.

Troubleshooting Common Robotic Pool Cleaner Issues

Keep your robotic pool cleaner running smoothly by troubleshooting common issues—discover simple solutions that can save you time and frustration.

AI and Smart Tech in Robotic Pool Cleaners

Many robotic pool cleaners now feature AI and smart technology, revolutionizing pool maintenance—discover how these innovations can improve your cleaning routine.