Why AI Managers Never Drop To Zero In This Persistent Benchmark
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why AI Managers Never Drop To Zero In This Persistent Benchmark on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A new AI management benchmark shows scores never drop to zero, highlighting the importance of trust, partial progress, and integrity in AI-driven business processes. The final standings reveal key strengths and weaknesses of current models.

The final results of a groundbreaking AI management benchmark for July 2026 confirm that no AI model scored zero, even in worst-week scenarios, as detailed in the original analysis. This reveals that, in real business contexts, partial progress and trustworthiness are valued over perfect performance, a principle embedded in the benchmark’s design. The findings highlight why AI managers are unlikely to drop to zero, emphasizing the importance of trust and incremental work in AI-driven management systems. For more context, see the detailed review at this analysis.

The benchmark, developed by Firmulate, tested four frontier AI models managing a simulated company facing seven days of crises, customer manipulations, and trust challenges. The top scorer, gpt-5.6-sol, achieved 95 out of 100, while the lowest, Opus 4.8, scored 73. The do-nothing baseline, intentionally designed as a minimal effort manager, scored 26, illustrating that even minimal management has measurable value. Notably, no model reached a perfect score of 100, which the designers consider a red flag indicating unmeasured or idealized performance.

The scoring system reflects a core principle: “no amount of good work outweighs a breach of trust.” This concept is explored in depth in the original report. A model that breaches trust even once is disqualified from achieving top scores, underscoring the importance of integrity over partial competence. The results also reveal that models capable of reading and referencing internal documentation performed better, especially in closing deals, highlighting the critical role of thoroughness and follow-through. For example, two models identified key documents deep in the company’s files, securing a €55,000 deal, whereas others failed to do so.

During the week, models faced social engineering attacks, such as fake CEO messages and staged reporter inquiries. All five models refused to comply, with Kimi K3 explicitly treating such requests as potential impersonation attempts. Despite high levels of analysis and rule adherence, some models faltered in follow-through, demonstrating that thoroughness does not necessarily translate into disciplined execution. The benchmark’s design emphasizes that trust, integrity, and the ability to complete tasks are more critical than superficial performance or complexity.

At a glance
reportWhen: announced July 2026
The developmentThe final July 2026 results of a persistent AI management benchmark demonstrate that no model scores zero, reflecting the importance of trust and partial progress in AI management.
Why AI Managers Never Drop To Zero In This Persistent Benchmark
26
AI MANAGEMENT BENCHMARK — JULY 2026

Why AI Managers Never Drop to Zero in This Persistent Benchmark

Firmulate’s persistent management benchmark simulated a company’s worst week. The result: no model scored zero, no model scored 100 — and trust outranked raw talent.

0
Models scoring zero — even in worst-week scenarios
0 / 5
Perfect scores — a 100 is treated as a red flag
€55K
Deal closed by models that read internal docs
4
Frontier AI models tested
7
Days of crises per run
95/100
Top score — gpt-5.6-sol
26/100
Do-nothing baseline floor
Final Standings

Every Score Lives Between 26 and 100

The league is intentionally bounded: a minimal-effort manager still earns 26 points, while a perfect 100 is structurally blocked. Partial progress always counts.

gpt-5.6-sol
95
Kimi K3
88*
Opus 4.8
73
Do-nothing baseline
26
No amount of good work outweighs a breach of trust.
Core scoring principle — one breach disqualifies any top score
How Scoring Works

The Path From Crisis to Score

Each model manages an identical simulated week of crises, customer manipulation, and trust challenges — with every decision auditable and traceable.

1
Simulated Worst Week

Seven days of crises, manipulation attempts, and trust tests at a small software company.

2
Auditable Decisions

Identical scenarios for each model; every action is fully traceable for scoring.

3
Trust Screen

A single breach of trust caps the score — integrity outranks partial competence.

4
Bounded Score

Floor of 26 for minimal effort; ceiling just below 100 to block idealized results.

Capability Analysis

What Separated the Leaders From the Rest

Thoroughness in documentation and disciplined follow-through proved more decisive than analysis quality or rule adherence alone.

STRENGTH — DOCUMENTATION

Deep File Readers Win Deals

Two models dug deep into internal company files, found key documents, and closed a €55,000 deal. Others never found them.

STRENGTH — SECURITY

All Models Refused Manipulation

Faced with fake CEO messages and staged reporter inquiries, all five models refused. Kimi K3 explicitly flagged requests as impersonation attempts.

WEAKNESS — EXECUTION

Follow-Through Falters

Strong analysis and rule adherence did not reliably translate into disciplined execution — thoroughness isn’t the same as finishing the job.

Capabilitygpt-5.6-solKimi K3Opus 4.8
Crisis handling✓ Strong✓ Strong~ Mixed
Internal doc referencing✓ Deep search✓ Found key docs✗ Missed docs
Resisted social engineering✓ Refused✓ Flagged impersonation✓ Refused
Disciplined follow-through~ Occasional gaps~ Occasional gaps✗ Weak
Trust integrity maintained✓ No breach✓ No breach✓ No breach
Implications for Business

Integrity Beats Perfection

The results challenge the perception that AI must be perfect to be useful. For support, sales, and operations, reliability and trust are the differentiators.

FOR ENTERPRISES

Value Incremental Progress

Some management is better than none — partial progress earns points by design, mirroring real-world value.

FOR EVALUATION

A Perfect Score Is a Red Flag

Designers cap scores below 100: an unmeasured or idealized result signals a problem, not a triumph.

FOR HIGH STAKES

Trust Breaches Are Costly

In sensitive environments, a single breach outweighs accumulated good work — evaluation must reflect that.

Key Questions

Frequently Asked

Why do AI models never score zero?

The benchmark values partial progress and trustworthiness — even minimal management effort earns points, reflecting that some management beats none.

What does a high score indicate?

The model can handle crises, follow rules, reference internal documentation, and maintain trust under pressure — key traits for real-world management.

How does the system prevent perfect scores?

Scores are intentionally capped below 100 to avoid overestimating AI capability and to signal that trust and integrity are non-negotiable.

Do results predict real business performance?

They are indicative of management qualities, but based on simulation. Real-world reliability under varied, unpredictable conditions remains under evaluation.

Traceability Chain

What Matters Most in AI Management

🔒 Trust ⚖️ Integrity 📄 Documentation 🔁 Follow-Through 📊 Reliable Management

Implications for AI-Driven Business Management

This benchmark underscores that in real-world applications, AI managers are valued for their ability to maintain trust, follow through on commitments, and handle pressure ethically, rather than simply generating high-quality outputs. The fact that scores never reach zero, even in worst-case scenarios, highlights that partial progress and trustworthiness are fundamental to effective AI management. For businesses deploying AI in customer support, sales, or operations, these findings suggest that focusing on integrity and reliability is more crucial than raw competence or sophistication.

The results challenge the common perception that AI models must be perfect to be useful. Instead, they demonstrate that incremental progress, transparency, and adherence to rules are what differentiate successful AI management systems. This has direct implications for how companies evaluate and trust AI tools, especially in sensitive or high-stakes environments where breaches of trust can be costly.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of the Persistent Management Benchmark

The benchmark, launched by Firmulate, simulates a small software company’s worst week, testing AI models’ ability to manage crises, customer relations, and trust under pressure. The setup involves identical scenarios for each model, with decisions fully auditable and traceable, ensuring transparency in performance evaluation. The league’s design intentionally includes a minimum score of 26 for minimal management effort, and a maximum of just below 100, to reflect real-world limitations and the importance of trust.

Previous benchmarks primarily measured language or problem-solving ability, but this one emphasizes management skills, ethical decision-making, and trustworthiness. The July 2026 results build on earlier phases, confirming that partial progress and integrity are more valued than perfection. The benchmark also incorporates social engineering tests, which models successfully resisted, indicating robustness against manipulation.

While models like gpt-5.6-sol and Kimi K3 performed well, the results reveal persistent weaknesses in follow-through and discipline, even among highly analyzed systems. The benchmark is regarded as a significant step toward assessing AI’s readiness for real-world management tasks, especially in high-pressure scenarios.

“The design intentionally prevents perfect scores, reinforcing that integrity and follow-through matter more than superficial performance.”

— Thorsten Meyer

Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About AI Trust and Performance

It remains unclear how these results will translate to real-world business environments outside the simulated benchmark. While models demonstrated resistance to social engineering attacks, the long-term reliability and trustworthiness under varied, unpredictable conditions are still being evaluated. Additionally, the impact of different training methods, rule sets, and operational parameters on follow-through and integrity needs further investigation. The benchmark does not measure all aspects of management, such as strategic decision-making or emotional intelligence, leaving gaps in understanding how AI will perform in complex, nuanced scenarios.

Amazon

trustworthy AI management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps in AI Management Benchmarking

Following these results, developers and organizations will likely focus on improving models’ discipline, follow-through, and internal referencing capabilities. Future benchmarks may incorporate more diverse scenarios, longer timeframes, and additional trust challenges to better simulate real-world management. Researchers will also explore how different training regimes impact trustworthiness and whether models can be designed to prioritize integrity without sacrificing performance. Meanwhile, enterprises are encouraged to test AI tools in controlled environments, emphasizing transparency and trust, as the industry moves toward more responsible AI deployment.

Expect ongoing updates from Firmulate, including more detailed analyses of model weaknesses and strengths, as well as new features for testing AI’s ability to manage complex, high-stakes business processes under pressure.

Amazon

AI document referencing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do AI models never score zero in this benchmark?

Because the benchmark values partial progress and trustworthiness, even minimal management efforts earn points. The scoring system is designed to reflect real-world scenarios where some management is better than none, and trust breaches disqualify top scores.

What does a high score in this benchmark indicate about an AI model?

A high score suggests the model can handle crises, follow rules, reference internal documentation, and maintain trust under pressure, making it more suitable for real-world management tasks.

How does the scoring system prevent perfect scores?

The designers intentionally cap scores below 100 to prevent overestimation of AI capabilities and to highlight that trust and integrity are non-negotiable in management scenarios.

Can these results predict AI performance in actual businesses?

While indicative of certain management qualities, the results are based on simulated scenarios. Real-world environments may introduce additional complexities, so further testing is necessary to confirm applicability.

What are the main weaknesses revealed by the benchmark?

Models often struggle with follow-through, referencing internal documentation, and disciplined execution, despite high levels of analysis and rule adherence.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Bank Of America Advises Hedging Portfolios Ahead Of Potential Q3 S&P 500 Pullback, Warns Of ‘Three-Wave Correction’

Bank of America advises investors to hedge portfolios ahead of a potential Q3 decline in the S&P 500, citing a ‘three-wave correction’ forecast.

CNN Newsroom Wary of Bari Weiss’ Potential Oversight Amid Paramount-Warner Bros. Merger

CNN Newsroom is cautious about Bari Weiss’s potential oversight responsibilities amid the ongoing merger of Paramount and Warner Bros., raising questions about editorial independence.

SEALSQ Corp Reports Preliminary H1 2026 Results; Revenue Up 120%, FY 2026 Guidance Reaffirmed

SEALSQ Corp announces preliminary H1 2026 results with a 120% revenue increase; reaffirms full-year guidance amid strong growth signals.

SK Telecom Pursues 15GW AI Data Center Buildout, Aiming To Become Asia’s AI Infrastructure Hub

SK Telecom announces plans to build a 15GW AI data center network, aiming to become Asia’s leading AI infrastructure provider. Details are emerging.