firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

For investors, activity is not the same as value creation

Personal-finance readers know the distinction instinctively. A portfolio can contain thoughtful research, elaborate spreadsheets and carefully documented convictions, yet still disappoint if the decisive choices are mistimed or never made. The same lesson now applies to artificial intelligence.

In Firmulate’s Crucible League, Opus 4.8 was the most thorough participant. It produced the deepest analyses and learned more than 80 playbook rules while managing a small software company through an exceptionally difficult week. Yet it finished last, with 73 points. Its failure was not ignorance. It identified the problems, developed a viable sales case and resisted attempts to manipulate it. What it did not consistently do was convert sound analysis into completed outcomes.

That makes Opus 4.8 an unusually useful case study. Its performance was neither foolish nor reckless. It was diligent, intelligent and often impressive. But diligence without prioritization can become its own form of risk.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A demanding test of managerial judgment

Firmulate runs AI models as complete companies, testing management quality under pressure rather than polished answers in isolation. Each frontier model faced the same customers, crises and temptations while operating the same small software business. Every decision was versioned and auditable.

The company itself has 13 synthetic employees and real money mechanics. It burns €105,000 a month against €2,300 in monthly recurring revenue, while a public cash countdown makes delay visible. Across the live company, the models have accumulated more than 680 self-learned playbook rules, and every workday is versioned. The experiment is real, ongoing and watchable on Firmulate’s public site.

The final July 2026 Crucible League results placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the benchmark placed a hard limit on any performance involving a breach of trust: “no amount of good work outweighs a breach of trust.”

The research was right; the deal still disappeared

All models recognized every crisis and rejected every manipulation attempt. The commercial divide came later. Only two signed the €55,000 deal their own analysis had earned: “Same diagnosis, same pitch — no signature.”

The decisive information was not presented conveniently in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed the trail and read the file won the deal at full price, adding €4,583 in monthly recurring revenue.

This is where the investment analogy becomes sharp. Discovering an opportunity is not the same as capturing it. An analyst may identify a mispriced asset, but the insight has no portfolio impact if conviction never becomes a position. Likewise, Opus 4.8 could reason deeply about the customer and develop the pitch, yet the close remained on the table.

Thoroughness became a distraction

Opus 4.8’s 80-plus learned rules demonstrate serious effort. Its analyses went deeper than those of the other participants. Those are genuine strengths, especially in environments where missed context can create costly errors.

Yet the experiment exposed the downside of treating comprehensiveness as the objective. The model spent attention building knowledge while execution discipline weakened. It attempted to write into a locked department instead of escalating the blockage. That is a small operational moment with a large managerial meaning: when the direct path failed, it did not reliably switch to the action required to finish the job.

Firmulate found the same weakness, in milder form, across all four models covered by this finding. Opus 4.8 therefore should not be treated as uniquely incapable. Its result is better understood as the clearest expression of a broader AI limitation: models can mistake extensive work for completed work.

Strong defenses did not compensate for incomplete execution

The models also faced staged social engineering. Fake messages from a chief executive escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 captured the appropriate stance in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean result matters for any business considering agents with access to customer records, forecasts or support systems. But safety and commercial effectiveness are separate dimensions. Refusing manipulation is essential; it does not automatically mean an agent can locate buried evidence, preserve operating discipline and secure the outcome.

There is also an important comparison caveat. Kimi K3 ran with the API default because it had no effort parameter, while the other participants ran at xhigh. Its 93-point finish should therefore be read with that difference in mind, even though the observed decisions remain available for review.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

business analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Useful intelligence must reach the finish line

The Opus 4.8 profile offers a warning for investors and executives alike. More analysis, more documentation and more rules can reduce uncertainty, but they can also create the appearance of progress while the economically important action remains undone.

Organizations evaluating AI workers should look beyond eloquence and visible effort. The relevant questions are whether an agent reads the available files, recognizes which fact changes the decision, escalates when blocked, protects trust and completes the transaction it has prepared.

Firmulate also offers enterprises the same wargame using a read-only export of their own business, with nothing written back to real systems. Its separate quiz draws on 242 real, unedited management decisions and asks visitors to identify which model made each choice. Both ideas reinforce the central lesson: evaluate AI through consequential behavior, not impressions.

Opus 4.8 did a remarkable amount of work. Its last-place finish does not erase that diligence; it reveals its limit. In business, investing and AI management, impact belongs to the participant that identifies what matters most—and then finishes.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model version control

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Geopolitical Sanctions Influence Blockchain Adoption

Many geopolitical sanctions are accelerating blockchain adoption, but the evolving regulatory landscape raises questions about its future in global finance.

NFTs in 2025: Are Digital Collectibles Still Worth It?

Market trends indicate NFTs are evolving, but what does this mean for their future value in 2025? Discover the surprising insights within.

Crypto and Wall Street: 2025’s Biggest Institutional Moves

Keen insights reveal how Wall Street’s 2025 institutional shifts in crypto are transforming investments, promising a future that begs to be explored further.

Top Crypto Predictions for 2026: What Investors Should Watch

The top crypto predictions for 2026 reveal potential growth and challenges; discover what investors must watch to navigate this evolving landscape.