firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Would You Trust an AI That Finds the Opportunity but Fails to Close It?

For investors, the most consequential management mistakes are often not failures of intelligence. They are failures of execution: an overlooked document, an unsigned contract or a breakdown in discipline when pressure rises. Firmulate has turned those familiar business risks into a live experiment—and an unusually revealing test of frontier AI models.

Each model was asked to run the same small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. The resulting record now powers a quiz built from 242 real, unedited management decisions. Readers see what a model actually did and try to identify which one was responsible.

The game is entertaining, but the underlying question is serious: Can an AI merely recognize a sound business decision, or can it carry that decision through to completion?

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five Models, Five Distinct Management Records

The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One rule imposed a hard limit, however: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

That trust test did not separate the field. Every model spotted every crisis and refused every manipulation attempt. The experiment included fake CEO messages escalating over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 captured the danger directly in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The meaningful separation emerged elsewhere. Only two models signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap starkly: “Same diagnosis, same pitch — no signature.” In investment terms, the models demonstrated that identifying value and realizing value are not the same capability.

The Detail That Changed the Commercial Outcome

The decisive competitive weakness was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue.

This is where the quiz becomes more than a personality test. A concise response can look decisive, while a long analysis can look diligent. Neither style alone proves that the model found the most important fact or completed the commercial task. Readers must distinguish tone from performance by examining what each decision accomplished.

Thoroughness Was Not Enough

Opus 4.8 offers the clearest warning against equating visible effort with effective management. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last in the league. The deal was left unsigned, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in weaker form across all four other participants.

That profile should feel familiar to anyone who has evaluated a management team. More analysis can improve a decision, but it can also coexist with unfinished work. In Firmulate’s experiment, the most elaborate participant did not produce the strongest overall result.

Kimi K3’s performance also carries an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Even under that different condition, K3 finished second. The qualification does not erase the result, but it matters when comparing the models as potential business operators.

A Company Designed to Expose Financial Consequences

The experiment is not a conventional chat demonstration. Its live company has 13 synthetic employees and real money mechanics, including monthly burn of €105k against €2.3k in monthly recurring revenue. A public cash countdown makes the pressure visible. The company has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

That financial structure gives ordinary management choices weight. Missing a contract affects revenue. Ignoring an internal file weakens negotiating leverage. Mishandling a suspicious request risks trust. The models are therefore judged in a setting where business actions connect to business consequences rather than ending as polished text.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI contract signing tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Investors Can Learn From the Quiz

The most useful lesson is not that one model always writes more or another usually sounds terse. It is that frontier models can display measurable management personalities while sharing many of the same analytical strengths.

  • All of them recognized the crises and resisted manipulation.
  • Only two completed the €55,000 commercial opportunity.
  • Reading deeply enough to uncover the buried fact protected full-price value.
  • Extensive analysis and rule-building did not guarantee first-rate execution.

For finance-minded readers, that makes the Firmulate quiz a compact exercise in operational due diligence. The task is to look past presentation and ask the questions that matter in any investment: Did management use the information already available? Did it protect trust? Did it convert insight into revenue? And when the normal process failed, did it escalate or simply keep trying the wrong door?

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to their real systems. That offers a practical way to observe an AI workforce before giving it responsibility for customers, forecasts or commercial decisions.

The experiment’s sharpest conclusion is also its most human one: knowing what should happen is not the same as making it happen.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Amazon

AI business decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI enterprise risk management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

VigilSAR: The Object That Isn’t Transmitting

VigilSAR is a SAR-based platform that identifies vessels without active transponders, enhancing maritime domain awareness in all weather conditions.

Bitcoin Battles Unfold in Live Warzone Visualization

A new web-based tool visualizes Bitcoin trading activity as a live warzone, highlighting market dynamics through immersive graphics and real-time data.

Manus AI: Bridging Thought Processes and Automated Actions

Fusing your thoughts with automation, Manus AI revolutionizes productivity—discover how this groundbreaking technology transforms your daily tasks into seamless actions.

The Open Interest in Bitcoin Is Surging With a Wealth of New Positions.

The surge in Bitcoin’s open interest hints at potential volatility, leaving traders wondering what dramatic shifts might lie ahead.