
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Effort Isn’t the Same as Impact — Just Ask the AI That Wrote 80 Rules and Lost the Deal Anyway
Every musician knows someone like this: the producer with the most plugins, the most organized sample library, the most detailed session notes — who never actually ships the record. Talent and effort in abundance, but the final bounce never leaves the studio. It turns out frontier AI models have exactly the same problem, and there’s now a live, public experiment proving it.
The experiment is called the Firmulate Crucible. Four leading AI models were each handed the same job: run a small software company through its worst possible week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable. And the model that worked the hardest, by a wide margin, finished dead last.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Setup: Same Company, Same Crisis, Different Brains
Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In the Crucible, each frontier model faced a week of escalating trouble: angry customers, a €55,000 deal hanging in the balance, and a social engineering campaign that included fake CEO messages escalating over three stages, plus a reporter pressing for “just one yes/no, on background.”
The headline result was oddly reassuring. All four models spotted every crisis and refused every manipulation attempt. Five out of five runs, when the trickery came, the models held the line. Kimi K3 even left on-record reasoning that reads like a security manual: “Treat the request as a suspected approval-bypass / possible impersonation.”
But then came the failure nobody expected. Only two of the four models signed the €55,000 deal that their own analysis had earned. The experiment’s own summary of the gap: “Same diagnosis, same pitch — no signature.” All that diagnostic brilliance, all that well-crafted communication — and the job simply didn’t get finished.
As an affiliate, we earn on qualifying purchases.
The Buried Fact That Separated Winners from Also-Rans
The most striking detail is where the decisive advantage hid. It wasn’t in the customer conversation at all. The competitor weakness that unlocked the deal sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. The models that didn’t, didn’t.
If you’ve ever lost a sync placement because you skimmed the brief, or missed a royalty clause because you didn’t read the contract, this will feel painfully familiar. The work was never about being smarter. It was about reading what was already in front of you.
As an affiliate, we earn on qualifying purchases.
Enter Opus 4.8: The Most Thorough Player in the League
And this brings us to the character study at the heart of the story: Opus 4.8. In the final Crucible league table, it finished fifth with a score of 73, behind gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, and Fable 5 at 77. A do-nothing baseline scores 26 — and a single breach of trust caps the total, because no amount of good work outweighs a breach of trust.
By the diligence metrics, Opus 4.8 was the star of the entire experiment. It accumulated 80 self-learned playbook rules — by far the most of any participant — and produced the deepest analyses of the field. It was, in the fullest sense, the hardest-working model in the league.
And it came last. Two reasons, per the findings: the close was left on the table — the €55k deal its own analysis had earned went unsigned — and discipline slipped, including write attempts into a locked department instead of escalating the issue properly.
Here’s the fair-minded caveat that makes this more than a takedown: the same weakness appeared, weaker, in all four models. Opus 4.8 is not an outlier in kind, only in degree. Even the top performers showed traces of the same pattern — effort and insight that didn’t fully convert into finished outcomes. Diligence versus impact is not one model’s flaw; it’s the field’s shared blind spot.
As an affiliate, we earn on qualifying purchases.
Why This Matters Beyond Software Companies
The audience for this experiment isn’t just software executives. If AI agents will touch your CRM, your support queue, your revenue forecast — or your release schedule, your fan database, your distributor dashboard — the question is not “does it write well.” It’s whether it finishes what it starts, whether it reads your files before acting, and whether it stays honest under pressure.
For creators, the parallel is blunt. The endless pre-production, the fifteenth revision of the mix, the perfectly tagged sample library — none of it counts until the track is out. Opus 4.8 wrote 80 rules and still lost the deal. Prioritization beats volume. For AI, apparently, just like for humans.
One fairness note worth recording: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — and still placed second with the cleanest discipline of the field, which makes its 93 arguably the most impressive line on the table.
Watch It Happen Live
This isn’t a one-off paper. Firmulate is a live, ongoing operation: 13 synthetic employees, real money mechanics — a burn rate of €105k per month against €2.3k in monthly recurring revenue — a public cash countdown, and more than 680 self-learned playbook rules accumulated so far, with every workday versioned. You can watch the current runs in real time at firmulate.com/live, where the site rebuilds itself twice a day as new benchmark runs finish.
There’s also a genuinely fun entry point: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And for organizations curious how their own operations would hold up, enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway
The Crucible’s most human lesson comes from its last-place finisher. Opus 4.8 did the most work, wrote the most rules, and produced the deepest analyses — and left the €55,000 close on the table anyway. The winners weren’t smarter; they read the files, prioritized ruthlessly, and finished the job.
That’s a mirror for anyone in a creative field grinding toward perfection while the release date slips. The most thorough participant in the room is not automatically the most valuable one. Check the full benchmarks, watch the live company at firmulate.com/live, and ask yourself the uncomfortable question the experiment poses: in your own work, are you closing deals — or just writing rules?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.