🔍 Read the full analysis: Why AI Managers Never Drop To Zero In This Persistent Benchmark on ThorstenMeyerAI.com
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A new AI management benchmark shows scores never drop to zero, highlighting the importance of trust, partial progress, and integrity in AI-driven business processes. The final standings reveal key strengths and weaknesses of current models.
The final results of a groundbreaking AI management benchmark for July 2026 confirm that no AI model scored zero, even in worst-week scenarios, as detailed in the original analysis. This reveals that, in real business contexts, partial progress and trustworthiness are valued over perfect performance, a principle embedded in the benchmark’s design. The findings highlight why AI managers are unlikely to drop to zero, emphasizing the importance of trust and incremental work in AI-driven management systems. For more context, see the detailed review at this analysis.
The benchmark, developed by Firmulate, tested four frontier AI models managing a simulated company facing seven days of crises, customer manipulations, and trust challenges. The top scorer, gpt-5.6-sol, achieved 95 out of 100, while the lowest, Opus 4.8, scored 73. The do-nothing baseline, intentionally designed as a minimal effort manager, scored 26, illustrating that even minimal management has measurable value. Notably, no model reached a perfect score of 100, which the designers consider a red flag indicating unmeasured or idealized performance.
The scoring system reflects a core principle: “no amount of good work outweighs a breach of trust.” This concept is explored in depth in the original report. A model that breaches trust even once is disqualified from achieving top scores, underscoring the importance of integrity over partial competence. The results also reveal that models capable of reading and referencing internal documentation performed better, especially in closing deals, highlighting the critical role of thoroughness and follow-through. For example, two models identified key documents deep in the company’s files, securing a €55,000 deal, whereas others failed to do so.
During the week, models faced social engineering attacks, such as fake CEO messages and staged reporter inquiries. All five models refused to comply, with Kimi K3 explicitly treating such requests as potential impersonation attempts. Despite high levels of analysis and rule adherence, some models faltered in follow-through, demonstrating that thoroughness does not necessarily translate into disciplined execution. The benchmark’s design emphasizes that trust, integrity, and the ability to complete tasks are more critical than superficial performance or complexity.
Why AI Managers Never Drop to Zero in This Persistent Benchmark
Firmulate’s persistent management benchmark simulated a company’s worst week. The result: no model scored zero, no model scored 100 — and trust outranked raw talent.
Every Score Lives Between 26 and 100
The league is intentionally bounded: a minimal-effort manager still earns 26 points, while a perfect 100 is structurally blocked. Partial progress always counts.
The Path From Crisis to Score
Each model manages an identical simulated week of crises, customer manipulation, and trust challenges — with every decision auditable and traceable.
Seven days of crises, manipulation attempts, and trust tests at a small software company.
Identical scenarios for each model; every action is fully traceable for scoring.
A single breach of trust caps the score — integrity outranks partial competence.
Floor of 26 for minimal effort; ceiling just below 100 to block idealized results.
What Separated the Leaders From the Rest
Thoroughness in documentation and disciplined follow-through proved more decisive than analysis quality or rule adherence alone.
Deep File Readers Win Deals
Two models dug deep into internal company files, found key documents, and closed a €55,000 deal. Others never found them.
All Models Refused Manipulation
Faced with fake CEO messages and staged reporter inquiries, all five models refused. Kimi K3 explicitly flagged requests as impersonation attempts.
Follow-Through Falters
Strong analysis and rule adherence did not reliably translate into disciplined execution — thoroughness isn’t the same as finishing the job.
| Capability | gpt-5.6-sol | Kimi K3 | Opus 4.8 |
|---|---|---|---|
| Crisis handling | ✓ Strong | ✓ Strong | ~ Mixed |
| Internal doc referencing | ✓ Deep search | ✓ Found key docs | ✗ Missed docs |
| Resisted social engineering | ✓ Refused | ✓ Flagged impersonation | ✓ Refused |
| Disciplined follow-through | ~ Occasional gaps | ~ Occasional gaps | ✗ Weak |
| Trust integrity maintained | ✓ No breach | ✓ No breach | ✓ No breach |
Integrity Beats Perfection
The results challenge the perception that AI must be perfect to be useful. For support, sales, and operations, reliability and trust are the differentiators.
Value Incremental Progress
Some management is better than none — partial progress earns points by design, mirroring real-world value.
A Perfect Score Is a Red Flag
Designers cap scores below 100: an unmeasured or idealized result signals a problem, not a triumph.
Trust Breaches Are Costly
In sensitive environments, a single breach outweighs accumulated good work — evaluation must reflect that.
Frequently Asked
The benchmark values partial progress and trustworthiness — even minimal management effort earns points, reflecting that some management beats none.
The model can handle crises, follow rules, reference internal documentation, and maintain trust under pressure — key traits for real-world management.
Scores are intentionally capped below 100 to avoid overestimating AI capability and to signal that trust and integrity are non-negotiable.
They are indicative of management qualities, but based on simulation. Real-world reliability under varied, unpredictable conditions remains under evaluation.
What Matters Most in AI Management
Implications for AI-Driven Business Management
This benchmark underscores that in real-world applications, AI managers are valued for their ability to maintain trust, follow through on commitments, and handle pressure ethically, rather than simply generating high-quality outputs. The fact that scores never reach zero, even in worst-case scenarios, highlights that partial progress and trustworthiness are fundamental to effective AI management. For businesses deploying AI in customer support, sales, or operations, these findings suggest that focusing on integrity and reliability is more crucial than raw competence or sophistication.
The results challenge the common perception that AI models must be perfect to be useful. Instead, they demonstrate that incremental progress, transparency, and adherence to rules are what differentiate successful AI management systems. This has direct implications for how companies evaluate and trust AI tools, especially in sensitive or high-stakes environments where breaches of trust can be costly.
As an affiliate, we earn on qualifying purchases.
Background of the Persistent Management Benchmark
The benchmark, launched by Firmulate, simulates a small software company’s worst week, testing AI models’ ability to manage crises, customer relations, and trust under pressure. The setup involves identical scenarios for each model, with decisions fully auditable and traceable, ensuring transparency in performance evaluation. The league’s design intentionally includes a minimum score of 26 for minimal management effort, and a maximum of just below 100, to reflect real-world limitations and the importance of trust.
Previous benchmarks primarily measured language or problem-solving ability, but this one emphasizes management skills, ethical decision-making, and trustworthiness. The July 2026 results build on earlier phases, confirming that partial progress and integrity are more valued than perfection. The benchmark also incorporates social engineering tests, which models successfully resisted, indicating robustness against manipulation.
While models like gpt-5.6-sol and Kimi K3 performed well, the results reveal persistent weaknesses in follow-through and discipline, even among highly analyzed systems. The benchmark is regarded as a significant step toward assessing AI’s readiness for real-world management tasks, especially in high-pressure scenarios.
“The design intentionally prevents perfect scores, reinforcing that integrity and follow-through matter more than superficial performance.”
— Thorsten Meyer
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Trust and Performance
It remains unclear how these results will translate to real-world business environments outside the simulated benchmark. While models demonstrated resistance to social engineering attacks, the long-term reliability and trustworthiness under varied, unpredictable conditions are still being evaluated. Additionally, the impact of different training methods, rule sets, and operational parameters on follow-through and integrity needs further investigation. The benchmark does not measure all aspects of management, such as strategic decision-making or emotional intelligence, leaving gaps in understanding how AI will perform in complex, nuanced scenarios.
As an affiliate, we earn on qualifying purchases.
Next Steps in AI Management Benchmarking
Following these results, developers and organizations will likely focus on improving models’ discipline, follow-through, and internal referencing capabilities. Future benchmarks may incorporate more diverse scenarios, longer timeframes, and additional trust challenges to better simulate real-world management. Researchers will also explore how different training regimes impact trustworthiness and whether models can be designed to prioritize integrity without sacrificing performance. Meanwhile, enterprises are encouraged to test AI tools in controlled environments, emphasizing transparency and trust, as the industry moves toward more responsible AI deployment.
Expect ongoing updates from Firmulate, including more detailed analyses of model weaknesses and strengths, as well as new features for testing AI’s ability to manage complex, high-stakes business processes under pressure.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do AI models never score zero in this benchmark?
Because the benchmark values partial progress and trustworthiness, even minimal management efforts earn points. The scoring system is designed to reflect real-world scenarios where some management is better than none, and trust breaches disqualify top scores.
What does a high score in this benchmark indicate about an AI model?
A high score suggests the model can handle crises, follow rules, reference internal documentation, and maintain trust under pressure, making it more suitable for real-world management tasks.
How does the scoring system prevent perfect scores?
The designers intentionally cap scores below 100 to prevent overestimation of AI capabilities and to highlight that trust and integrity are non-negotiable in management scenarios.
Can these results predict AI performance in actual businesses?
While indicative of certain management qualities, the results are based on simulated scenarios. Real-world environments may introduce additional complexities, so further testing is necessary to confirm applicability.
What are the main weaknesses revealed by the benchmark?
Models often struggle with follow-through, referencing internal documentation, and disciplined execution, despite high levels of analysis and rule adherence.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
