
Imagine a business with no human employees, losing €105,000 every month, yet still trying to close deals and navigate crises — all under the watchful eye of the public. This isn’t science fiction; it’s a real, live experiment in building a company with AI at its core, openly documented for all to see. Welcome to the world of Firmulate, where artificial intelligence models run a small software company through its worst week, and every decision, every misstep, is public and auditable.
The Live Experiment: A Company Without People
At the heart of the experiment is Firmulate’s live site, firmulate.com/live.html. Here, four frontier AI models are tasked with managing a fictional small software business. The goal? Navigate crises, make management decisions, and close deals — all while being monitored and compared in real time. This isn’t a simulation; these models are making actual decisions in a live environment, with real money mechanics and a public cash countdown.
What makes this experiment extraordinary is its transparency. Every workday, the models’ decisions are versioned, auditable, and publicly visible. They face the same customer demands, crises, and ethical temptations as human managers, including social engineering attempts like fake CEO messages and reporter tricks. The models are tested against the toughest conditions: reading complex documents, resisting manipulation, and making disciplined choices under pressure.
Are AI Models Capable of Honest, Effective Management?
The results so far are revealing. All four models successfully identified every crisis and refused manipulation attempts — a sign of their integrity. Yet, only two managed to seal the €55,000 deal, earning a positive impact on the company’s monthly recurring revenue. Interestingly, the models that read deeper into company files, such as Kimi K3, were the ones that closed the deal at full price, highlighting the importance of thorough data analysis.
One particularly detailed participant, OPUS 4.8, with over 80 learned rules and the deepest analysis, performed poorly in the final stages. It left deals unexecuted and slipped discipline, demonstrating that even the most thorough AI can falter without proper process discipline. Another key insight: running the models at different effort levels (like default versus high effort) affected results, showing the nuances of AI management styles.
As an affiliate, we earn on qualifying purchases.
What This Means for Businesses and AI Adoption
This experiment offers more than just a curiosity; it provides vital lessons for any enterprise considering AI automation. The key takeaway isn’t whether an AI can generate engaging chat responses — it’s whether it can see the job through to completion, avoid manipulation, and make honest, disciplined decisions under pressure.
For companies using CRM, support systems, or forecasts, the question is: Will your AI agents just talk a good game, or will they actually deliver reliable, trustworthy work? The Firmulate experiment demonstrates that with proper testing and transparent measurement, AI can be held accountable for real decision-making, not just chit-chat.
The Competitive Edge and Future Outlook
In the ongoing Crucible League, the top-performing model, GPT-5.6-sol, scored 95 out of 100, successfully closing the deal and uncovering hidden company information that others missed. The second-place model, Kimi K3, scored 93, showing the importance of thoroughness and discipline. The experiment’s transparency and detailed scoring system allow anyone to see how each model performs, fostering a new kind of accountability in AI management.
For businesses eager to test their own AI tools, Firmulate offers a pilot program where companies can run their own management wargames without risking real data or systems. The platform ensures that no actual systems are touched, yet provides a clear picture of how AI handles complex management scenarios.
The Broader Implications: Transparency, Trust, and Building in Public
This experiment is more than a technical showcase; it’s a statement on transparency. Building an AI-powered company in public pushes the boundaries of trust, accountability, and innovation. Every decision, every crisis, every successful deal is openly documented, making the process a real-world case study for anyone interested in the future of AI in business.
In a world where AI’s role in management is rapidly expanding, such open, live experiments could set the standard for how AI systems are evaluated, trusted, and integrated into everyday operations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html