firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Every serious investor knows the rule: never buy what you can’t stress-test. You wouldn’t put your savings into a fund based on a glossy pitch deck — you’d want to know how it performs in a crash. Yet enterprises are doing exactly that with AI agents, buying on chat demos and slideware promises, with no idea how the system behaves on its worst week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A public experiment called Firmulate just ran that stress test — and the results read like a cautionary note from an auditor. The lessons matter to anyone whose money sits in a company about to hand its decisions to software.

The wargame, not the demo

Firmulate ran four frontier AI models through the same job: operate an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, like a trading log you can replay line by line.

In the final July 2026 Crucible League standings, the spread was wide: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it, no amount of good work outweighs a breach of trust.

What the audit found

The headline findings cut both ways. On the plus side, all models spotted every crisis and refused every manipulation attempt — including a three-stage fake-CEO escalation and a reporter’s disarming “just one yes/no, on background” trick. Five out of five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

But here’s the gap that should worry any shareholder: only two of the models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That failure is invisible in a chat demo. It only shows up when you run the full week and read the ledger.

And the buried fact is better still: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. Diligence paid, literally.

Effort isn’t everything

The most instructive profile is Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more faintly, in all four models. Thoroughness without follow-through is a cost, not a feature — the same reason you don’t want an analyst who writes brilliant memos and never places the trade.

One fairness note the league itself discloses: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still finished second.

It’s live, and it’s watchable

This isn’t a one-off paper. Firmulate operates a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned, and anyone can watch at firmulate.com. There’s even a quiz built on 242 real, unedited management decisions where you guess which model made which call — a humbling exercise.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The investment frame is simple: chat quality is the pitch; management quality is the track record. Firmulate measures the latter — outcomes, crisis triage, and integrity — because that’s what compounds, or doesn’t, when an agent touches your CRM, support queue, or forecast.

For enterprises, the natural next step from watching to acting is a pilot: run the same wargame against a read-only export of your own business — your customers, your pipeline, your rules — through churn waves, price increases, competitor attacks, and social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks. Critically, nothing ever writes back to your real systems: it’s a flight simulator, not an autopilot.

If your company — or one you’re invested in — is about to put AI in the decision loop, run the stress test first. Start a pilot at firmulate.com/pilot.html, or reach out to contact@firmulate.com. Cheaper to find the breach of trust in a simulation than in your P&L.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Stablecoin Rule Changes Nobody Expected and What They Mean for Investors

New stablecoin rule changes could reshape your investments—discover what these unexpected shifts mean for your financial future.

Top Cryptocurrencies of 2025: Which Projects Emerged Strongest?

Learn which cryptocurrencies are dominating in 2025 and discover the surprising projects that are reshaping the landscape of digital finance.

Decoding Signal Peak 2026: Microsoft’s AI Innovation With Anthropic’s Models

Microsoft prepares to launch Project Perception, an AI security tool integrating Anthropic’s models, signaling a shift toward model routing for enterprise AI.

The Switch: You Never Owned the AI You Depend On

Recent events reveal that AI reliance is vulnerable to instant shutdowns by governments or companies, exposing dependency without ownership.