firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get audio and creator gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Every Studio Has a Hidden Champ. So Does Every AI League Table.

If you work in music or creator tech, you know the feeling: the plugin or DAW everyone ignores turns out to outperform the big-name flagship. This month, the same thing happened in AI — not in a chat demo, but in something far stranger: a live, watchable experiment where frontier models were handed the same struggling software company and told to run it through its worst week.

The result? Kimi K3, a newcomer from Moonshot, scored 93 in the Crucible league — second place, just behind gpt-5.6-sol (95), and ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). A model many Western buyers hadn’t shortlisted beat three of the four Western frontier models at actual management. The league is open — and picking a model without testing it yourself is now a bet, not a decision.

Amazon

AI decision-making software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Company, One Terrible Week, Five Models

The experiment, run by Firmulate and watchable live, gave each frontier model the identical job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing rests on anyone’s impressions.

For context, doing nothing scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.

Amazon

AI management simulation models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Separated the Models

The headline finding is almost uncomfortable in its simplicity: all models spotted every crisis and refused every manipulation attempt. The differences showed up at the finish line. Only two of five signed the €55,000 deal their own analysis had earned — “Same diagnosis, same pitch — no signature.”

The deal turned on a buried fact. The decisive competitor weakness wasn’t in the customer event at all; it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 MRR in the company’s real money mechanics.

Then there was the social engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was strikingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI security and trust analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

K3’s Near-Perfect Run

So what earned the newcomer its 93? K3 found the buried security needle, won the €55k deal, saved the churning customer, and resisted all three baits — with just one deviation across the entire week, the cleanest discipline in the field. In a scoring system where a single breach of trust caps everything, that discipline is exactly what buys you a top score.

Amazon

AI experiment platforms for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Most Thorough Model Came Last

The most sobering profile belongs to Opus 4.8. It was the most thorough participant — +80 learned rules, the deepest analyses — yet finished last at 73. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Effort and thoroughness, it turns out, are not the same as finishing the job.

Fairness Footnote

One important caveat: K3 ran without an effort parameter (API default) while the other models ran at xhigh reasoning effort. In other words, the newcomer’s near-win came without the extra reasoning budget its rivals had. That makes the result more striking, not less — but it’s worth stating plainly.

It’s Not a Slide Deck — It’s Running Right Now

What makes this more than a benchmark row is that the company is live: 13 synthetic employees, real money mechanics — a burn of €105k/month against €2.3k MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, and full plain-language results are on the benchmarks page.

There’s even a game angle: 242 real, unedited management decisions power a “guess the model” quiz — closer to blind A/B listening tests than any marketing chart. And for enterprises, the same wargame can run against a read-only export of your own business; nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Why Creators Should Care

Anyone who’s compared audio tools by spec sheet and then been surprised in a blind test knows the drill: the polished demo doesn’t tell you how a tool behaves on your material, under pressure. AI models are the same. If an agent will ever touch your customer inbox, your payments or your release calendar, the question isn’t whether it writes well — it’s whether it finishes what it starts, reads your files first, and stays honest when someone tries to trick it.

The Crucible’s answer: even frontier models differ sharply on exactly those traits, and the newcomer near the top ran with less reasoning budget than its rivals. Test before you trust. The league table is one data point; your own worst week is the one that counts.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Equal Loudness Curves: Why Bass Disappears at Low Volume

Curious why bass disappears at low volume? Equal loudness curves reveal how your ears perceive sound differently at various levels, affecting your listening experience.

A Sound Engineer Who Never Touches the Fader Still Gets Paid: The AI Benchmark With a Floor at 26

One AI benchmark publishes its floor: a do-nothing manager scores 26, not 0. Here’s why partial progress counts and one breach of trust caps everything.

Music Subwoofer Basics for Stereo Setups

Discover the essential subwoofer setup tips to enhance your stereo bass, and learn how to optimize your system for powerful, balanced sound.

How to Compare Listening Setups Without Fooling Yourself

Want to accurately compare listening setups and truly understand their strengths? Keep reading to avoid biased judgments and improve your sound experience.