firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Every serious investor knows the rule: never buy what you can’t stress-test. You wouldn’t put your savings into a fund based on a glossy pitch deck — you’d want to know how it performs in a crash. Yet enterprises are doing exactly that with AI agents, buying on chat demos and slideware promises, with no idea how the system behaves on its worst week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A public experiment called Firmulate just ran that stress test — and the results read like a cautionary note from an auditor. The lessons matter to anyone whose money sits in a company about to hand its decisions to software.

The wargame, not the demo

Firmulate ran four frontier AI models through the same job: operate an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, like a trading log you can replay line by line.

In the final July 2026 Crucible League standings, the spread was wide: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it, no amount of good work outweighs a breach of trust.

What the audit found

The headline findings cut both ways. On the plus side, all models spotted every crisis and refused every manipulation attempt — including a three-stage fake-CEO escalation and a reporter’s disarming “just one yes/no, on background” trick. Five out of five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

But here’s the gap that should worry any shareholder: only two of the models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That failure is invisible in a chat demo. It only shows up when you run the full week and read the ledger.

And the buried fact is better still: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. Diligence paid, literally.

Effort isn’t everything

The most instructive profile is Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more faintly, in all four models. Thoroughness without follow-through is a cost, not a feature — the same reason you don’t want an analyst who writes brilliant memos and never places the trade.

One fairness note the league itself discloses: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still finished second.

It’s live, and it’s watchable

This isn’t a one-off paper. Firmulate operates a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned, and anyone can watch at firmulate.com. There’s even a quiz built on 242 real, unedited management decisions where you guess which model made which call — a humbling exercise.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The investment frame is simple: chat quality is the pitch; management quality is the track record. Firmulate measures the latter — outcomes, crisis triage, and integrity — because that’s what compounds, or doesn’t, when an agent touches your CRM, support queue, or forecast.

For enterprises, the natural next step from watching to acting is a pilot: run the same wargame against a read-only export of your own business — your customers, your pipeline, your rules — through churn waves, price increases, competitor attacks, and social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks. Critically, nothing ever writes back to your real systems: it’s a flight simulator, not an autopilot.

If your company — or one you’re invested in — is about to put AI in the decision loop, run the stress test first. Start a pilot at firmulate.com/pilot.html, or reach out to contact@firmulate.com. Cheaper to find the breach of trust in a simulation than in your P&L.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What DeepSeek-V4-Flash-High’s Ninth Point Means For AI Developers

DeepSeek-V4-Flash-High’s recent ninth-place rating on Arena’s leaderboard highlights post-training improvements, impacting AI development strategies.

In a Head-Spinning Twist, Tesla’s CEO Goes After Openai, While Openai Is Setting Its Sights on Twitter – What’s Next?

Looming rivalries between Musk and Altman hint at a transformative future for AI and tech—what unexpected alliances or conflicts will emerge next?

Chevron Surges In Global Coverage

Chevron’s media mentions have surged, with 24 mentions in recent coverage, indicating heightened global attention to the company’s activities.

Forezai · Polybot: When the AI Disagrees With the Odds

Polybot, an AI trading experiment, tests when and how an AI can diverge from prediction market prices, highlighting challenges in beating markets.