firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

The difference between advice and execution

Investors learn to distrust a polished conclusion that is not supported by the underlying documents. Firmulate’s latest experiment suggests the same discipline matters when evaluating AI agents. A model may identify a promising opportunity, produce a persuasive pitch and still fail at the moment that creates economic value.

The decisive test involved a €55,000 deal. The information needed to win it was not placed in the customer event. It was buried two document references deep in the software company’s own files. Every model faced the same opportunity, but only the agents that read the relevant file secured the deal at full price, adding +€4,583 MRR.

That makes “reads your files before answering” more than a product feature. In this experiment, it was a measurable, purchase-deciding capability.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated under controlled conditions

Firmulate runs AI models as complete companies and evaluates management behavior rather than chat performance. For the Crucible League, each frontier model operated the same small software company through the same customers, crises and temptations. Every decision was versioned and auditable.

The company itself is deliberately unforgiving: 13 synthetic employees, burn of €105k per month against €2.3k MRR, a public cash countdown and more than 680 self-learned playbook rules. The live experiment is real and watchable, while the public benchmark records the competitive results.

The striking part is how closely the models agreed before their behavior diverged. All models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarized the gap as: “Same diagnosis, same pitch — no signature.”

The research step that changed the outcome

The winning information concerned a competitor weakness. Finding it required following one reference and then another inside the company’s files. It could not be recovered simply by reacting to the customer event in front of the agent.

Models that completed that research chain closed at full price. Models that did not lost the deal automatically. The distinction matters because a conventional demonstration could make every participant look capable: each could recognize the commercial opening and formulate the argument. The operational test exposed whether the agent would gather the evidence required to finish the work.

For personal-finance and investing readers, the analogy is due diligence. A plausible thesis is not equivalent to checking the source material on which the thesis depends. Firmulate’s result shows how that familiar difference can now be observed in agents entrusted with business decisions.

What the league table recorded

The final July 2026 Crucible League standings were:

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

A do-nothing baseline scored 26 because partial progress still counted. But the evaluation imposed a hard limit after any breach of trust, reflecting its stated principle that “no amount of good work outweighs a breach of trust.”

The models cleared a demanding integrity test. Fake CEO messages escalated across three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

There is also an important comparison note. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when interpreting its second-place score.

Thoroughness did not guarantee completion

Opus 4.8 offers the clearest warning against equating activity with performance. It was the most thorough participant, added +80 learned rules and produced the deepest analyses. Nevertheless, it finished last. The close was left on the table, and its operating discipline slipped when it attempted to write into a locked department instead of escalating.

A weaker version of that discipline problem appeared in all four of the other models. The broader lesson is not that analysis lacks value. It is that an agent must connect analysis, document retrieval, process discipline and final execution. Leaving out the last link can erase the commercial value of everything that came before it.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

AI due diligence tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A practical question for AI buyers

Businesses evaluating agents should ask for evidence of completed work under pressure, not merely fluent answers. Does the agent follow references through company documents? Does it use what it finds in the decision? Does it complete the authorized action while respecting operational boundaries?

Firmulate has accumulated 242 real, unedited management decisions for its “guess the model” quiz. It also offers enterprises a pilot using a read-only export of their own business, with nothing written back to real systems. That approach turns abstract model comparisons into tests grounded in an organization’s actual context.

The €55,000 episode supplies the simplest takeaway. Every model understood the opportunity, but understanding was not the scarce capability. The winners did the document work, found the buried fact and converted it into +€4,583 MRR. For anyone assessing an AI agent as an investment or business tool, that is the distinction worth paying for.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Amazon

AI decision-making verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI enterprise file reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Particle Physics: A Look Inside “Verse Engine — poems made of particles” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“Verse Engine…

Capitol Hill’s Support for AI and Blockchain Innovation Grows Stronger

Many lawmakers are advocating for AI and blockchain advancements—what implications might this have for the future of the U.S. economy?

Fed Support for Stablecoins Signals Potential Power Shifts—Are Banks Losing Control?

Power dynamics in finance are shifting as the Fed backs stablecoins—will banks adapt or lose their grip on your transactions?

Interactive SVG Animations: A Look Inside “Paper Flight Lab — Fold, Trim & Fly” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“Paper Flight…