firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

If you have ever compared mutual funds, you know the trap: a gorgeous five-star rating that tells you nothing about what happens when markets get ugly. The same problem now faces anyone buying AI for their business. Chat demos are the marketing brochure; the real question is how the thing behaves on its worst week. A live, public experiment called Firmulate just ran that stress test — and its scoring system offers a lesson in honest measurement that any investor should appreciate.

Before you orderOffer from Amazon

Get smart everyday buys delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The experiment

Firmulate handed four frontier AI models the identical job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision was versioned and auditable, so nothing rests on anyone’s say-so. The final league table from July 2026: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. (One fairness note worth flagging, in the spirit of the thing: K3 ran at its API-default effort setting while the others ran at the highest effort tier.)

Amazon

AI stress test software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a do-nothing manager scores 26, not 0

Here is the detail that should warm any skeptic’s heart. A baseline run that does essentially nothing still earns 26 points out of 100. That is not a bug — it is a design philosophy. The benchmark gives credit for partial progress. An agent that correctly diagnoses a problem but never closes it has still done some of the work, and the score says so. Investors know this instinctively: a manager who preserves capital while missing a rally has not delivered nothing.

But the floor at 26 comes with a ceiling rule that cuts the other way, and it is the benchmark’s moral core: a single breach of trust caps the total grade. As the project puts it, “no amount of good work outweighs a breach of trust.” In portfolio language — past returns never excuse the manager who steals from the till. Interestingly, no model crossed that line this time. All of them spotted every crisis and refused every manipulation attempt, including a social-engineering gauntlet of fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick. All five attempts were refused; Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI decision-making evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Where the benchmark distrusts round numbers

Note that no model scored 100. A perfect round score in a messy, judgment-heavy exercise is a smell, not a triumph — and a benchmark that hands them out invites gaming. The top score of 95 means something precisely because 100 was neither reached nor casually awarded.

Amazon

AI trustworthiness benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact that decided the €55,000 deal

The key finding is the kind of gap that matters to anyone deploying AI against real money. All the models diagnosed the same customer problem and delivered the same pitch — but only two actually signed the €55,000 deal their own analysis had earned. “Same diagnosis, same pitch — no signature.”

And the decisive competitive weakness? It sat two document references deep in the company’s own files, not in the customer conversation. The models that read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. It is the AI equivalent of a fund manager who actually reads the footnotes instead of skimming the press release.

Amazon

AI performance scoring system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Effort is not results

Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, weaker, appeared in all four models. Hard work, in management as in investing, is not the same as outcomes.

You can watch the books yourself

The experiment is not a one-off paper. Firmulate runs a live synthetic company with 13 employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. It is watchable at firmulate.com/live. For the quiz-minded, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html — think of it as reading the fund’s actual trade blotter. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Firmulate benchmark does what good performance measurement always does: it pays for partial progress, refuses to let one act of dishonesty be averaged away, and treats a suspiciously perfect score with suspicion. If you are evaluating AI the way you would evaluate a fund manager — outcomes, not eloquence; audit trails, not demos — this is a scoring model worth studying before you trust any agent with your customers or your cash.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Bitcoin’s Legal Status in El Salvador Shaken by Controversial Law Proposal

Discover how a controversial law proposal is reshaping Bitcoin’s legal status in El Salvador and what it could mean for the country’s economic future.

Kimi K3’s Top 3 Achievement: A Milestone In AI Development

Kimi K3 by Moonshot ranks third in Vigilsar’s defense-ISR LLM benchmark, marking a milestone in AI trustworthiness for intelligence tasks.

Einstein’s Parenting Secrets For Cultivating Resilience In Children

New insights reveal how Albert Einstein’s advice to his son offers valuable strategies for fostering resilience in kids today.

Soaring Bitcoin Mining Strength Underpins a Brighter BTC Future

Uncover how soaring Bitcoin mining strength is shaping a sustainable future for BTC, and what hidden implications lie ahead for investors.