firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get audio and creator gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When Doing Nothing Still Earns a Score

Any producer knows the type: the session player who shows up, plugs in, plays nothing wrong — and nothing at all. They didn’t wreck the take, but they didn’t make the record either. How do you grade that?

That question sits at the heart of one of the more interesting AI experiments running right now. Firmulate, a public project that runs AI models as complete companies through a staged worst week of crises and temptations, hands out a striking number in its final league table: a do-nothing baseline — an AI manager that simply does nothing — scores 26 out of 100. Not zero. Twenty-six.

For a benchmark that crowns winners at 95 and 93, deliberately publishing the floor is an unusual act of honesty. And it says a lot about how the people behind it think management should be measured.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Same Company, Same Worst Week

Firmulate ran four frontier AI models through an identical scenario: run the same small software company through its worst week. Same customers, same crises, same temptations — only the model changes. Every decision is versioned and auditable, and the whole thing is watchable at firmulate.com/benchmarks.html.

The final Crucible League standings, concluded in July 2026:

  • 1. gpt-5.6-sol — 95 (“Found the buried fact, closed the deal — the complete performance”)
  • 2. Kimi K3 — 93 (closed the deal too, cleanest discipline of the field)
  • 3. Sonnet 5 — 88

    4. Fable 5 — 77

    5. Opus 4.8 — 73

One fairness note worth flagging: Kimi K3 ran without an effort parameter while the others ran at xhigh — and still nearly won.

Amazon

AI benchmark testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the Do-Nothing Baseline Gets 26, Not 0

The reasoning is refreshingly practical. In a real company, a manager who does nothing still avoids making things worse. Crises they ignore may partially resolve; customers they neglect may stay anyway; partial progress on existing work still counts. So a zero would actually overstate how bad inaction is.

The 26-point floor means every score above it measures what a model added, not merely what it failed to break. It’s the same logic a studio uses when judging a mix engineer: the baseline is the rough mix that already exists, not silence.

Amazon

AI decision-making analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Nobody Scores a Suspicious 100

Just as telling: no model hit a perfect score, and the benchmark’s design seems built to distrust one. The cap comes from a single rule the project states plainly: “no amount of good work outweighs a breach of trust.” One breach — one attempt to slip something past the customer, the team, or the rules — and the total grade is capped, however brilliant everything else was.

That principle showed up in the social-engineering leg of the test. Fake CEO messages escalated over three stages, plus a reporter offering a quick out: “just one yes/no, on background.” All five models refused every manipulation attempt. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI trust and ethics evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Deal That Separated the Winners

The key finding: all models spotted every crisis and refused every manipulation — yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature.

The buried fact explains it: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The others left the close on the table.

The cautionary tale is Opus 4.8: the most thorough participant, with 80+ learned rules and the deepest analyses, yet last place. The close went unfinished, and discipline slipped — write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four.

It’s Live, and You Can Play

The company is real in the ways that matter: 13 synthetic employees, burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, every workday versioned. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

If AI agents will touch your CRM, support queue, or forecast, the question isn’t whether they write well. It’s whether they finish what they start, read your files first, and stay honest under pressure. Firmulate’s benchmark design — a floor at 26 for doing nothing, partial credit for real progress, and a hard cap for any breach of trust — is what measuring that honestly looks like. A perfect 100 isn’t on the table, and that’s exactly the point.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Portable Party Speaker Buying Basics for Tailgates and Camps

Jumpstart your tailgate or camp party with essential tips to choose the perfect portable speaker that keeps the fun going all day.

How to Build a Listening Routine That Makes Upgrades Smarter Later

A thoughtful listening routine can unlock smarter upgrades—discover how small changes can transform your understanding and growth.

Mono vs Stereo for Live Music: When Each One Wins

Join us as we explore whether mono or stereo sound elevates your live performance—discover which setup truly wins for your next show.

Planar vs Dynamic Headphones Explained Without the Hype

What sets planar and dynamic headphones apart, and which one truly suits your listening style? Keep reading to find out the honest differences.