firmulate.com/quotes.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Urgency is where trust usually gets tested

Music and creator businesses run on compressed timelines: releases move, audiences react and decisions that felt distant suddenly become immediate. That makes an urgent message from a senior executive especially potent. The more convincing the pressure, the easier it can become to treat ordinary safeguards as obstacles.

Firmulate put that problem in front of frontier AI models during a live, auditable company experiment. Fake CEO messages escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” The result was unusually clear: 5 of 5 models refused every manipulation attempt.

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A security test disguised as a working week

Firmulate gave each model the same job: run the same small software company through its worst week. The customers, crises and temptations remained constant, while every decision was versioned and auditable. This was not a conversational test about what an AI claimed it would do. The models had to make management decisions while the company continued operating.

The social-engineering campaign tested whether apparent authority and mounting urgency could override judgment. All the models spotted every crisis and refused every manipulation attempt. Kimi K3 summarized the correct posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That sentence matters because it does not depend on proving who sent the message. K3 treated the attempt according to its risk: someone appeared to be asking the company to bypass approval. For businesses entrusting AI with customer communications, commercial information or operational decisions, that is the useful behavior. The model preserved the boundary while the identity question remained unresolved.

Integrity did not automatically mean effectiveness

The encouraging security result came with a harder business lesson. Although every model detected the crises and resisted manipulation, only two signed the €55,000 deal their own analysis had earned. Firmulate’s summary captures the gap: “Same diagnosis, same pitch — no signature.”

The decisive commercial fact was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that read that material won the deal at full price, worth +€4,583 MRR. The distinction is important: refusing a dangerous instruction protects the company, but completing legitimate work still requires persistence, research and follow-through.

The final July 2026 Crucible League ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress counts, but a breach of trust caps the total. Firmulate states the principle plainly: “no amount of good work outweighs a breach of trust.”

The most thorough model still finished last

Opus 4.8 offers the most revealing counterexample to the idea that more analysis always produces a better operator. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it placed last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, though less strongly, in all four other participants.

K3’s result also needs a fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference does not change the recorded refusals, but it belongs beside any comparison of the final standings.

A company built to expose operational behavior

The live company contains 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, and the experiment can be watched at firmulate.com/live.

Readers can also inspect the human texture behind the league table. A “guess the model” quiz at firmulate.com/quiz.html is powered by 242 real, unedited management decisions. For enterprises, Firmulate offers the same wargame against a read-only export of their own business; nothing writes back to real systems.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Rethinking AI Reasoning: From Prompts to Governed Thinking

Rethinking AI Reasoning: From Prompts to Governed Thinking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the refusal before the emergency

The central finding is not that AI can recognize an obviously suspicious prompt in a demo. It is that integrity under pressure can be observed while models are responsible for an operating company, facing identical distractions and commercial demands.

For creator-tech companies considering AI access to consequential work, this points to a practical standard: test whether a model preserves trust when authority appears urgent, then test whether it can still finish the legitimate job. Firmulate’s results show that those are separate capabilities. In this field, every model resisted the fake executives and the reporter. Only part of the field converted sound analysis into the deal.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


The AI Edge for Solo Entrepreneurs: An Ethical Guide to Leveraging AI for Competitive Advantage

The AI Edge for Solo Entrepreneurs: An Ethical Guide to Leveraging AI for Competitive Advantage

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI Model Validation & Testing: Ensuring Reliable AI Systems — Bias Testing, Robustness Evaluation & Regulatory Compliance (AI Compliance Toolkit)

AI Model Validation & Testing: Ensuring Reliable AI Systems — Bias Testing, Robustness Evaluation & Regulatory Compliance (AI Compliance Toolkit)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Simple Way to Understand Soundstage Without Overcomplicating It

Ongoing exploration reveals how simple adjustments can transform your audio experience into a more immersive and realistic soundstage.

How to Clean Up Harsh Treble Without Killing Detail

Discover effective ways to tame harsh treble without sacrificing detail and learn how subtle adjustments can transform your listening experience.

How to Think About Headphone Amplifier Power Without Overbuying

Great headphone amplifier choices depend on matching power to impedance, but understanding your needs ensures optimal sound—so, here’s what you need to know.