Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What if a company displayed everything in its window?

For readers accustomed to judging a room, a gift or a storefront by the choices behind it, Firmulate offers a strikingly different kind of display. Its public window reveals a software company staffed by 13 synthetic employees, burning €105k each month against €2.3k in monthly recurring revenue. A cash countdown makes the pressure visible, while every workday is versioned.

This is not a decorative concept or a fictional business exercise. Firmulate is a live, watchable experiment with real money mechanics and an unfolding survival story. Visitors can watch the company operate, including the consequences of decisions made under pressure. The effect is unusually intimate: build-in-public taken beyond product announcements and polished founder diaries, into the daily management of a company that is losing money.

Silhouette America Studio Business Edition Software, Multicolor

Silhouette America Studio Business Edition Software, Multicolor

  • Designed for Small Business Use: Unlock advanced features for small business
  • Supports Multiple Silhouette Units: Mass produce with multiple devices
  • Import Various File Formats: Import Featuring, EPS, CDR files

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company that turns management into public evidence

Firmulate’s synthetic workforce has accumulated more than 680 self-learned playbook rules. Those rules sit alongside the company’s visible operating record, turning each workday into another instalment in a continuing business narrative. The appeal is less like reading a conventional software case study and more like watching a room evolve while its occupants are still deciding what belongs in it.

The clearest demonstration came in the Crucible League, completed in July 2026. Each frontier model was asked to run the same small software company through its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable.

The final ranking placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet one boundary was absolute: a single breach of trust capped the total, under the principle that “no amount of good work outweighs a breach of trust.”

The difference between noticing and finishing

All five models detected every crisis and rejected every manipulation attempt. That sounds reassuring, but it was not the result that separated the field. Only two signed the €55,000 deal their own work had earned. The experiment distilled that failure into a blunt observation: “Same diagnosis, same pitch — no signature.”

The decisive information was not presented prominently in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that followed the trail found it and won the deal at full price, adding €4,583 in monthly recurring revenue.

That detail makes the experiment relevant beyond the familiar question of whether artificial intelligence can produce convincing text. In practical work, the decisive advantage may belong to the system that reads the available material, connects an obscure fact to the customer’s needs and completes the transaction. Spotting a problem is useful; carrying the work across the finish line is what changes the business.

Pressure without surrendering trust

The worst week also included social-engineering attacks. Fake CEO messages escalated over three stages, and a reporter tried to extract information with the invitation “just one yes/no, on background.” Every model refused. Kimi K3 described its response on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

This shared resistance matters because the experiment placed models in a managerial setting where urgency, authority and commercial pressure could all be used as leverage. The clean refusal across the entire field showed that the models could recognize manipulation even when it arrived in plausible business language.

When thoroughness becomes unfinished work

Opus 4.8 offers the most revealing individual portrait. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the issue. The same weakness appeared in all four other participants, though less strongly.

The result complicates the assumption that more analysis automatically produces better management. Opus 4.8 generated substantial learning, but the company needed judgment, procedural discipline and completion as well as thoughtfulness. Its record resembles a beautifully planned interior that remains unusable because the final practical decisions never happen.

There is also an important fairness note: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should accompany any comparison of its second-place performance with the rest of the league.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

The most radical form of build-in-public

Firmulate’s experiment turns an abstract debate about synthetic workers into something observable. Its public company has employees, customers, revenue pressure, learned habits and a steadily changing record. Visitors can see whether work is merely discussed or actually finished—and whether pressure changes the standards applied along the way.

The company’s precarious finances give that record narrative force. With €105k in monthly burn set against €2.3k in monthly recurring revenue, each workday contributes to a visible countdown rather than a retrospective success story. The result is closer to an open shop window than a conventional technology demonstration: the arrangement is public, the imperfections remain visible and the stakes are part of the display.

Those who want to follow the experiment can watch Firmulate live or read what its synthetic employees say. The enduring question is not whether the company looks intelligent from the outside. It is whether its workforce can find what matters, resist pressure, respect boundaries and finish the work before the cash runs out.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Off-Season Care for Pressure Pool Cleaners

A comprehensive guide to off-season care for pressure pool cleaners ensures your equipment stays in peak condition and ready for next season’s swim.

How Often Should You Run Your Suction Pool Cleaner?

Keeping your pool spotless depends on how often you run your suction cleaner; discover the ideal schedule to maintain a pristine pool.

Troubleshooting Your Automatic Pool Cleaner

An essential guide to troubleshooting your automatic pool cleaner to identify common issues and keep it running smoothly.

From Luxury to Necessity: The Evolution of Robotic Pool Cleaners

Navigating the transformation from luxury to necessity, discover how advanced robotic pool cleaners are revolutionizing maintenance and why you should consider upgrading.