firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

If you have ever compared mutual funds, you know the trap: a gorgeous five-star rating that tells you nothing about what happens when markets get ugly. The same problem now faces anyone buying AI for their business. Chat demos are the marketing brochure; the real question is how the thing behaves on its worst week. A live, public experiment called Firmulate just ran that stress test — and its scoring system offers a lesson in honest measurement that any investor should appreciate.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The experiment

Firmulate handed four frontier AI models the identical job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision was versioned and auditable, so nothing rests on anyone’s say-so. The final league table from July 2026: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. (One fairness note worth flagging, in the spirit of the thing: K3 ran at its API-default effort setting while the others ran at the highest effort tier.)

Amazon

AI stress test software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a do-nothing manager scores 26, not 0

Here is the detail that should warm any skeptic’s heart. A baseline run that does essentially nothing still earns 26 points out of 100. That is not a bug — it is a design philosophy. The benchmark gives credit for partial progress. An agent that correctly diagnoses a problem but never closes it has still done some of the work, and the score says so. Investors know this instinctively: a manager who preserves capital while missing a rally has not delivered nothing.

But the floor at 26 comes with a ceiling rule that cuts the other way, and it is the benchmark’s moral core: a single breach of trust caps the total grade. As the project puts it, “no amount of good work outweighs a breach of trust.” In portfolio language — past returns never excuse the manager who steals from the till. Interestingly, no model crossed that line this time. All of them spotted every crisis and refused every manipulation attempt, including a social-engineering gauntlet of fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick. All five attempts were refused; Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI decision-making evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Where the benchmark distrusts round numbers

Note that no model scored 100. A perfect round score in a messy, judgment-heavy exercise is a smell, not a triumph — and a benchmark that hands them out invites gaming. The top score of 95 means something precisely because 100 was neither reached nor casually awarded.

Amazon

AI trustworthiness benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact that decided the €55,000 deal

The key finding is the kind of gap that matters to anyone deploying AI against real money. All the models diagnosed the same customer problem and delivered the same pitch — but only two actually signed the €55,000 deal their own analysis had earned. “Same diagnosis, same pitch — no signature.”

And the decisive competitive weakness? It sat two document references deep in the company’s own files, not in the customer conversation. The models that read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. It is the AI equivalent of a fund manager who actually reads the footnotes instead of skimming the press release.

Amazon

AI performance scoring system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Effort is not results

Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, weaker, appeared in all four models. Hard work, in management as in investing, is not the same as outcomes.

You can watch the books yourself

The experiment is not a one-off paper. Firmulate runs a live synthetic company with 13 employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. It is watchable at firmulate.com/live. For the quiz-minded, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html — think of it as reading the fund’s actual trade blotter. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Firmulate benchmark does what good performance measurement always does: it pays for partial progress, refuses to let one act of dishonesty be averaged away, and treats a suspiciously perfect score with suspicion. If you are evaluating AI the way you would evaluate a fund manager — outcomes, not eloquence; audit trails, not demos — this is a scoring model worth studying before you trust any agent with your customers or your cash.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Earnings Call Gap: What Q1 2026 Just Told Us About AI ROI

Analysis of Q1 2026 earnings shows a widening gap between AI investment claims and measurable ROI, impacting stock performance and investor confidence.

Week Three — Foundation model vs Brownian motion. Kronos on five-minute BTC.

Kronos foundation model tested against Brownian motion for 5-minute BTC predictions; results show no significant outperformance in recent trading data.

What a Strong Gold Market Could Mean for Bitcoin Bulls

Navigating a rising gold market may signal new opportunities and shifting confidence for Bitcoin bulls—discover what this means for your investments.

Six Things Europe Should Ask Canada To Understand Its AI Landscape

Europe needs clarity from Canada on AI and data sovereignty, trade agreements, and alliance terms to shape a strategic tech partnership.