
If you have ever compared mutual funds, you know the trap: a gorgeous five-star rating that tells you nothing about what happens when markets get ugly. The same problem now faces anyone buying AI for their business. Chat demos are the marketing brochure; the real question is how the thing behaves on its worst week. A live, public experiment called Firmulate just ran that stress test — and its scoring system offers a lesson in honest measurement that any investor should appreciate.
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The experiment
Firmulate handed four frontier AI models the identical job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision was versioned and auditable, so nothing rests on anyone’s say-so. The final league table from July 2026: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. (One fairness note worth flagging, in the spirit of the thing: K3 ran at its API-default effort setting while the others ran at the highest effort tier.)
AI stress test software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why a do-nothing manager scores 26, not 0
Here is the detail that should warm any skeptic’s heart. A baseline run that does essentially nothing still earns 26 points out of 100. That is not a bug — it is a design philosophy. The benchmark gives credit for partial progress. An agent that correctly diagnoses a problem but never closes it has still done some of the work, and the score says so. Investors know this instinctively: a manager who preserves capital while missing a rally has not delivered nothing.
But the floor at 26 comes with a ceiling rule that cuts the other way, and it is the benchmark’s moral core: a single breach of trust caps the total grade. As the project puts it, “no amount of good work outweighs a breach of trust.” In portfolio language — past returns never excuse the manager who steals from the till. Interestingly, no model crossed that line this time. All of them spotted every crisis and refused every manipulation attempt, including a social-engineering gauntlet of fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick. All five attempts were refused; Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
AI decision-making evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Where the benchmark distrusts round numbers
Note that no model scored 100. A perfect round score in a messy, judgment-heavy exercise is a smell, not a triumph — and a benchmark that hands them out invites gaming. The top score of 95 means something precisely because 100 was neither reached nor casually awarded.
AI trustworthiness benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buried fact that decided the €55,000 deal
The key finding is the kind of gap that matters to anyone deploying AI against real money. All the models diagnosed the same customer problem and delivered the same pitch — but only two actually signed the €55,000 deal their own analysis had earned. “Same diagnosis, same pitch — no signature.”
And the decisive competitive weakness? It sat two document references deep in the company’s own files, not in the customer conversation. The models that read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. It is the AI equivalent of a fund manager who actually reads the footnotes instead of skimming the press release.
As an affiliate, we earn on qualifying purchases.
Effort is not results
Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, weaker, appeared in all four models. Hard work, in management as in investing, is not the same as outcomes.
You can watch the books yourself
The experiment is not a one-off paper. Firmulate runs a live synthetic company with 13 employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. It is watchable at firmulate.com/live. For the quiz-minded, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html — think of it as reading the fund’s actual trade blotter. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The Firmulate benchmark does what good performance measurement always does: it pays for partial progress, refuses to let one act of dishonesty be averaged away, and treats a suspiciously perfect score with suspicion. If you are evaluating AI the way you would evaluate a fund manager — outcomes, not eloquence; audit trails, not demos — this is a scoring model worth studying before you trust any agent with your customers or your cash.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
