
Every serious investor knows the rule: never buy what you can’t stress-test. You wouldn’t put your savings into a fund based on a glossy pitch deck — you’d want to know how it performs in a crash. Yet enterprises are doing exactly that with AI agents, buying on chat demos and slideware promises, with no idea how the system behaves on its worst week.
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A public experiment called Firmulate just ran that stress test — and the results read like a cautionary note from an auditor. The lessons matter to anyone whose money sits in a company about to hand its decisions to software.
The wargame, not the demo
Firmulate ran four frontier AI models through the same job: operate an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, like a trading log you can replay line by line.
In the final July 2026 Crucible League standings, the spread was wide: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it, no amount of good work outweighs a breach of trust.
What the audit found
The headline findings cut both ways. On the plus side, all models spotted every crisis and refused every manipulation attempt — including a three-stage fake-CEO escalation and a reporter’s disarming “just one yes/no, on background” trick. Five out of five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
But here’s the gap that should worry any shareholder: only two of the models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That failure is invisible in a chat demo. It only shows up when you run the full week and read the ledger.
And the buried fact is better still: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. Diligence paid, literally.
Effort isn’t everything
The most instructive profile is Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more faintly, in all four models. Thoroughness without follow-through is a cost, not a feature — the same reason you don’t want an analyst who writes brilliant memos and never places the trade.
One fairness note the league itself discloses: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still finished second.
It’s live, and it’s watchable
This isn’t a one-off paper. Firmulate operates a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned, and anyone can watch at firmulate.com. There’s even a quiz built on 242 real, unedited management decisions where you guess which model made which call — a humbling exercise.

The investment frame is simple: chat quality is the pitch; management quality is the track record. Firmulate measures the latter — outcomes, crisis triage, and integrity — because that’s what compounds, or doesn’t, when an agent touches your CRM, support queue, or forecast.
For enterprises, the natural next step from watching to acting is a pilot: run the same wargame against a read-only export of your own business — your customers, your pipeline, your rules — through churn waves, price increases, competitor attacks, and social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks. Critically, nothing ever writes back to your real systems: it’s a flight simulator, not an autopilot.
If your company — or one you’re invested in — is about to put AI in the decision loop, run the stress test first. Start a pilot at firmulate.com/pilot.html, or reach out to contact@firmulate.com. Cheaper to find the breach of trust in a simulation than in your P&L.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
