firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Effort Isn’t the Same as Impact — Just Ask the AI That Wrote 80 Rules and Lost the Deal Anyway

Every musician knows someone like this: the producer with the most plugins, the most organized sample library, the most detailed session notes — who never actually ships the record. Talent and effort in abundance, but the final bounce never leaves the studio. It turns out frontier AI models have exactly the same problem, and there’s now a live, public experiment proving it.

The experiment is called the Firmulate Crucible. Four leading AI models were each handed the same job: run a small software company through its worst possible week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable. And the model that worked the hardest, by a wide margin, finished dead last.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Setup: Same Company, Same Crisis, Different Brains

Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In the Crucible, each frontier model faced a week of escalating trouble: angry customers, a €55,000 deal hanging in the balance, and a social engineering campaign that included fake CEO messages escalating over three stages, plus a reporter pressing for “just one yes/no, on background.”

The headline result was oddly reassuring. All four models spotted every crisis and refused every manipulation attempt. Five out of five runs, when the trickery came, the models held the line. Kimi K3 even left on-record reasoning that reads like a security manual: “Treat the request as a suspected approval-bypass / possible impersonation.”

But then came the failure nobody expected. Only two of the four models signed the €55,000 deal that their own analysis had earned. The experiment’s own summary of the gap: “Same diagnosis, same pitch — no signature.” All that diagnostic brilliance, all that well-crafted communication — and the job simply didn’t get finished.

Amazon

business analysis AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact That Separated Winners from Also-Rans

The most striking detail is where the decisive advantage hid. It wasn’t in the customer conversation at all. The competitor weakness that unlocked the deal sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. The models that didn’t, didn’t.

If you’ve ever lost a sync placement because you skimmed the brief, or missed a royalty clause because you didn’t read the contract, this will feel painfully familiar. The work was never about being smarter. It was about reading what was already in front of you.

Amazon

AI document reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Enter Opus 4.8: The Most Thorough Player in the League

And this brings us to the character study at the heart of the story: Opus 4.8. In the final Crucible league table, it finished fifth with a score of 73, behind gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, and Fable 5 at 77. A do-nothing baseline scores 26 — and a single breach of trust caps the total, because no amount of good work outweighs a breach of trust.

By the diligence metrics, Opus 4.8 was the star of the entire experiment. It accumulated 80 self-learned playbook rules — by far the most of any participant — and produced the deepest analyses of the field. It was, in the fullest sense, the hardest-working model in the league.

And it came last. Two reasons, per the findings: the close was left on the table — the €55k deal its own analysis had earned went unsigned — and discipline slipped, including write attempts into a locked department instead of escalating the issue properly.

Here’s the fair-minded caveat that makes this more than a takedown: the same weakness appeared, weaker, in all four models. Opus 4.8 is not an outlier in kind, only in degree. Even the top performers showed traces of the same pattern — effort and insight that didn’t fully convert into finished outcomes. Diligence versus impact is not one model’s flaw; it’s the field’s shared blind spot.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters Beyond Software Companies

The audience for this experiment isn’t just software executives. If AI agents will touch your CRM, your support queue, your revenue forecast — or your release schedule, your fan database, your distributor dashboard — the question is not “does it write well.” It’s whether it finishes what it starts, whether it reads your files before acting, and whether it stays honest under pressure.

For creators, the parallel is blunt. The endless pre-production, the fifteenth revision of the mix, the perfectly tagged sample library — none of it counts until the track is out. Opus 4.8 wrote 80 rules and still lost the deal. Prioritization beats volume. For AI, apparently, just like for humans.

One fairness note worth recording: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — and still placed second with the cleanest discipline of the field, which makes its 93 arguably the most impressive line on the table.

Watch It Happen Live

This isn’t a one-off paper. Firmulate is a live, ongoing operation: 13 synthetic employees, real money mechanics — a burn rate of €105k per month against €2.3k in monthly recurring revenue — a public cash countdown, and more than 680 self-learned playbook rules accumulated so far, with every workday versioned. You can watch the current runs in real time at firmulate.com/live, where the site rebuilds itself twice a day as new benchmark runs finish.

There’s also a genuinely fun entry point: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And for organizations curious how their own operations would hold up, enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Takeaway

The Crucible’s most human lesson comes from its last-place finisher. Opus 4.8 did the most work, wrote the most rules, and produced the deepest analyses — and left the €55,000 close on the table anyway. The winners weren’t smarter; they read the files, prioritized ruthlessly, and finished the job.

That’s a mirror for anyone in a creative field grinding toward perfection while the release date slips. The most thorough participant in the room is not automatically the most valuable one. Check the full benchmarks, watch the live company at firmulate.com/live, and ask yourself the uncomfortable question the experiment poses: in your own work, are you closing deals — or just writing rules?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Compare Listening Setups Without Fooling Yourself

Want to accurately compare listening setups and truly understand their strengths? Keep reading to avoid biased judgments and improve your sound experience.

When a Portable DAC Amp Helps and When It Really Doesn’t

The truth about portable DAC amps reveals when they enhance your audio experience and when they fall short—discover how to make the most of them.

The ‘One Speaker Test’ That Reveals Mixing Problems Fast

Keen to identify hidden mixing issues quickly? Discover how the ‘One Speaker Test’ can reveal problems that may otherwise go unnoticed.

How to Choose Portable Microphone For Creators

Learn how to choose, connect, and optimize a portable microphone for content creation with this step-by-step guide, suitable for all skill levels.