firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Your next business risk may be an AI hire

Investors weigh management, cash flow and customer loyalty when judging a company. As businesses hand more work to AI agents, there is another question to ask: can the system make sound decisions when money and trust are on the line? Firmulate has put several leading models through the same rough week at a small software company. The results suggest polished answers are only part of the picture.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate’s live experiment gives each model the same customers, crises and temptations, then versions and audits every decision. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. It is a watchable experiment, not a slide deck; the company runs every business day at Firmulate.

In the final July 2026 Crucible league, gpt-5.6-sol led with 95 points. Moonshot’s Kimi K3 came second with 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The difference between seeing an opportunity and taking it

All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The key clue was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The experiment’s blunt summary: “Same diagnosis, same pitch — no signature.”

Kimi K3 found the buried security weakness, won the deal, saved the churning customer and resisted all three baits. It had one deviation, the cleanest discipline in the field. When fake CEO messages escalated over three stages and a reporter asked for “just one yes/no, on background,” all five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 makes a different point about performance. It was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet finished last. It left the close on the table and attempted work in a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four participants. The lesson for a business buyer is practical: diligence and detailed analysis matter, but so does carrying a decision through properly.

What an investor should take from the test

For readers weighing the promise of AI against operational and financial risk, a leaderboard is a starting point, not a buying decision. A model can identify a crisis and resist pressure, yet still fail to complete a commercially important task. The ranking also comes with a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Firmulate says enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. Its benchmark page presents the results and plain-language findings at firmulate.com/benchmarks.html.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI risk management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the manager, not just the pitch

AI promises efficiency, but businesses and investors should care about whether an agent reads the relevant files, protects trust and finishes the work it recommends. Firmulate’s experiment shows why picking a model on reputation alone can be a bet: test it against the decisions your business actually needs it to make.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Amazon

enterprise AI security solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business decision analysis

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Forezai · Polybot: When the AI Disagrees With the Odds

Polybot, an AI trading experiment, tests when and how an AI can diverge from prediction market prices, highlighting challenges in beating markets.

Bitcoin, XRP, and DOGE Lose Steam as Trade War Fears Grip Crypto Markets

Gripped by trade war fears, Bitcoin, XRP, and DOGE face significant declines—what could this mean for the future of these cryptocurrencies?

Why Crypto Correlation With Tech Stocks Keeps Changing

Just as market conditions shift, the changing correlation between crypto and tech stocks reveals complex dynamics that every investor should understand.

SPY (SPY) Up Or Down On July 24?

Investors are watching SPY closely on July 24, with recent trading data showing significant activity. Here’s what is confirmed and what remains uncertain.