
Imagine observing a business in real time — not a staged demo, but a genuine experiment where AI models navigate crises, make decisions, and even attempt to cheat, all while losing money every day. This is the reality of a live company experiment that’s reshaping how we think about AI in management.
The Public Show of an AI Business in Crisis
At first glance, it looks like a typical startup struggle — but this is no ordinary company. It operates with 13 synthetic employees and faces the harsh realities of cash burn, with €105,000 spent monthly against a mere €2,300 in monthly recurring revenue. This real-time experiment is accessible publicly at firmulate.com/live.html, where anyone can watch how AI models perform under pressure.
AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Experiment Works
The core idea? Four advanced AI models are tasked with running a small software company through its worst week. They face the same customers, crises, and temptations — but their decisions are always documented and can be reviewed later. Each version is fully auditable, providing transparency into how and why decisions are made.
The Tested AI Models and Their Performance
- gpt-5.6-sol: Scored the highest with a 95 out of 100, successfully uncovering a hidden fact in company files that led to closing a €55,000 deal, adding +€4,583 in monthly recurring revenue (MRR).
- Kimi K3: Slightly behind at 93, but the newcomer proved disciplined, also closing the same deal but without uncovering the hidden information.
- Sonnet 5: With an 88, it closed the deal but made some process slips.
- Fable 5: Scored 77 and left the deal unexecuted despite identifying the opportunity, showing strong rule discipline but weaker decision execution.
business AI management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Insights Beyond the Scores
Interestingly, the models’ ability to detect and act on critical information in company documents was decisive. The winning models read and understood files deep enough to find a buried fact that was not apparent in the customer interactions — a rare but crucial edge in AI decision-making.
AI data analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Refusing Manipulation and Social Engineering
In addition to managing crises, the models faced social engineering attempts, such as fake CEO messages escalating through multiple stages and a reporter trick requiring a simple yes/no. All five models refused to be manipulated, with Kimi K3 explicitly treating such requests as impersonation risks. This highlights AI’s potential for integrity under pressure, a vital trait for management tools.
AI cybersecurity social engineering protection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Business Model Under Fire
Despite the impressive capabilities, the live company is deeply unprofitable — burning €105,000 each month with only €2,300 in income. Its cash countdown is visible to the public, making this a high-stakes experiment in survival and AI performance.
What Does This Mean for Your Business?
While the experiment may seem abstract, it’s highly relevant. If AI tools are going to interact with your customer support, sales, or forecasting systems, the critical questions aren’t about language quality or conversation flow. Instead, they focus on whether AI can follow through, read crucial data, remain honest, and deliver tangible results — especially under stress.
The Results of the League
These models are ranked based on their ability to succeed in this brutal week:
- gpt-5.6-sol: Top scorer, completed the deal, uncovered hidden info.
- Kimi K3: Closely behind, disciplined, but didn’t find the hidden fact.
- Sonnet 5: Slightly lower, with minor slips.
- Fable 5: Focused on rule discipline but failed to act on the opportunity.
These results are publicly available for review — transparency is key in understanding AI’s real-world competencies, not just its conversational skills.
Watch the Action Live
Curious to see this in action? Visit firmulate.com/live.html to watch the ongoing experiment. Every day, the company’s decisions are versioned, recorded, and made accessible, providing a rare window into how AI manages genuine business challenges.

This experiment reveals that AI’s potential in management isn’t just about chat quality — it’s about decision integrity under pressure. Watching this live company fight for survival exposes core strengths and weaknesses every business should consider before deploying AI in real operations.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html