firmulate.com/quotes.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

A security result investors should notice

For anyone evaluating a company’s finances, artificial intelligence creates an uncomfortable new category of risk. An AI worker may draft polished emails and analyze customer data, but what happens when someone claiming to be the chief executive demands that it ignore procedure and disclose sensitive information?

Firmulate put that question to five frontier models in a live, watchable business experiment. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All five models refused every manipulation attempt.

That clean sweep is an encouraging result. More importantly, it suggests that integrity under pressure can be tested before an AI system enters production—not discovered later in an incident report, after financial or reputational damage has already occurred.

Amazon

AI security testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week shared by every model

Firmulate gave each model the same assignment: run the same small software company through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable.

The simulated company makes the financial stakes visible. It has 13 synthetic employees and burns €105,000 a month against €2,300 in monthly recurring revenue. A public cash countdown turns weak execution into something more concrete than a disappointing chatbot response. The company has also accumulated more than 680 self-learned playbook rules, with every workday versioned.

Across the experiment, all models detected every crisis and rejected every manipulation attempt. In the social-engineering sequence, the attacker used urgency and supposed executive authority to push for a customer-list disclosure without normal process. The requests became more forceful over three stages, and the separate reporter trick tried to make disclosure sound harmless and informal.

Kimi K3 captured the correct stance in a concise on-record assessment: “Treat the request as a suspected approval-bypass / possible impersonation.” More model statements can be explored on Firmulate’s public quotes page.

Refusal was necessary, but not sufficient

The broader results complicate the reassuring security story. Although every model identified the problems and resisted manipulation, only two signed the €55,000 deal their own analysis had earned. The experiment’s summary is blunt: “Same diagnosis, same pitch — no signature.”

The difference came from evidence hidden inside the business itself. A decisive competitor weakness sat two document references deep in the company’s files rather than in the customer event. Models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

That finding matters to investors and business owners because safe inaction can still be expensive. A model may avoid an obvious breach yet fail to complete valuable work. Conversely, commercial ambition without disciplined handling of permissions and sensitive information can create an unacceptable downside. The experiment tests both sides of that tension.

The final league table

The July 2026 Crucible League placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, while a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.” The complete standings and plain-language findings appear on the Firmulate benchmarks page.

K3’s result carries a fairness qualification: it ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should accompany any comparison of its second-place finish.

Opus 4.8 offers another useful warning against judging capability by diligence alone. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four of the other models.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI integrity verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the behavior before funding the rollout

Firmulate’s result does not say that every AI system will resist every social-engineering attack. It shows something narrower and more actionable: identical pressure can be applied across models, their decisions can be audited, and differences in judgment can become visible before deployment.

For finance-minded readers, that turns AI selection into a due-diligence question. The relevant issue is not merely whether a model produces convincing language. It is whether the model protects entrusted information, investigates the company’s own evidence, respects operational boundaries and completes work that creates value.

The live company also makes those trade-offs observable rather than hypothetical. Its 242 real, unedited management decisions power a “guess the model” quiz, while enterprises can run the same wargame against a read-only export of their own business. Nothing writes back to real systems.

The fake CEO failed against every participant. The harder lesson is that trustworthiness and execution must be evaluated together: resisting the wrong instruction protects the downside, while disciplined follow-through determines whether the promised return ever arrives.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Amazon

AI model manipulation resistance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethical hacking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Ethereum at $3,000? Short Squeeze Potential, Say Analysts

Surging interest in Ethereum hints at a potential rise to $3,000, but will a short squeeze change everything? Discover the implications.

Trump’s Reserve Announcement Hits Bitcoin—Is the Market Off Track?

Market reactions to Trump’s Bitcoin reserve announcement reveal uncertainty—will this reshape government crypto policies or lead to more instability? Discover the implications.

Grok-3: Musk’S Xai Takes AI Evolution to the Next Level

Shattering limits in AI, Grok-3’s groundbreaking technology promises to redefine our daily interactions with artificial intelligence—what changes can we expect?

Are Polymarket Trading Bots Actually Profitable? The Math Behind 2026’s Prediction-Market Arbitrage Industry

An on-chain study reveals only 0.51% of wallets profit over $1,000 on Polymarket in 2024-2025; most retail bots lose money or break even in 2026.