
Every music writer knows the type: the artist who shows up to the interview without having listened to the B-sides. You can spot it in the first answer. The questions are fine, the flattery is smooth — but the one detail buried two tracks deep on the deluxe edition, the one that would have made the whole conversation, never comes up. The artist clocks it instantly, and the interview never becomes the cover story.
It turns out frontier AI agents do the exact same thing — and now there’s a versioned, auditable experiment to prove it. Firmulate, an AI company emulator, ran four top models as the management team of the same small software company through its worst week. Same customers, same crises, same temptations. The deal-winning fact wasn’t hidden in a customer meeting. It sat two document references deep in the company’s own files — the corporate equivalent of the bonus track nobody listened to.
The €55,000 Question
The setup was brutal by design. Each model ran the company solo: crises landing, a customer pushing hard, and a €55,000 deal on the table. All four models diagnosed the situation correctly. All four refused every manipulation attempt thrown at them, including a fake-CEO social-engineering campaign that escalated over three stages and a reporter’s trick question framed as “just one yes/no, on background.” Five out of five refusal rates across the experiment’s manipulation tests — the models are impressively honest.
Then came the close. Only two of the four models actually signed the €55,000 deal their own analysis had earned. The experiment’s own summary of the gap: “Same diagnosis, same pitch — no signature.” The other two had done the homework, written the essay, and never handed it in.
As an affiliate, we earn on qualifying purchases.
What Separated the Winners
The decisive difference wasn’t intelligence or eloquence. It was diligence. The key fact — a competitor weakness that justified full-price confidence — lived two references deep in the company’s internal files, not in the customer event itself. Models that chased the citation chain found it and won the deal at full price, worth +€4,583 in monthly recurring revenue. Models that stopped at the surface lost the deal automatically.
That reframes something buyers usually treat as vibes. “Reads your files before answering” isn’t a chat-demo nicety — it’s a measurable, purchase-deciding property of an AI agent, the same way a journalist who’s actually heard the record is a measurable improvement over one skimming a press release.
As an affiliate, we earn on qualifying purchases.
The League Table
The final Crucible standings from July 2026:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
- 2. Kimi K3 — 93. Closed the deal too, with the cleanest discipline in the field. Its on-record refusal reasoning in the social-engineering test: “Treat the request as a suspected approval-bypass / possible impersonation.” (One fairness caveat: K3 ran at its API-default effort setting while the others ran at xhigh.)
- 3. Sonnet 5 — 88. Closed the deal, with a few process slips.
- 4. Fable 5 — 77. Fell short of the signature.
- 5. Opus 4.8 — 73. The cautionary tale.
For calibration, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
The Thoroughness Trap
Opus 4.8 is the most interesting profile in the field: the most thorough participant by volume, generating the deepest analyses and adding more than 80 learned playbook rules — and still finishing last. The close was left on the table, and discipline slipped, with write attempts into a locked department instead of escalating. It’s the AI equivalent of the crate-digging completist who can tell you every pressing of every 7-inch but misses the deadline. And the same weakness appeared, in weaker form, in all four models.
As an affiliate, we earn on qualifying purchases.
It’s Still Running
The experiment isn’t a static paper. Firmulate operates a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, more than 680 self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live. For the interactive version, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a surprisingly humbling game for anyone who thinks they can tell AIs apart by voice. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

For creators and creator-tech companies already handing AI agents the keys to inboxes, CRMs and support queues, the lesson is specific: benchmark for homework, not charm. The models that won the deal weren’t smarter — they followed a citation two documents deep before answering. Before you trust an agent with anything that touches revenue, ask the Firmulate question: does it finish what it starts, does it read your files first, does it stay honest under pressure? The full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html