
Get audio and creator gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Every Studio Has a Hidden Champ. So Does Every AI League Table.
If you work in music or creator tech, you know the feeling: the plugin or DAW everyone ignores turns out to outperform the big-name flagship. This month, the same thing happened in AI — not in a chat demo, but in something far stranger: a live, watchable experiment where frontier models were handed the same struggling software company and told to run it through its worst week.
The result? Kimi K3, a newcomer from Moonshot, scored 93 in the Crucible league — second place, just behind gpt-5.6-sol (95), and ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). A model many Western buyers hadn’t shortlisted beat three of the four Western frontier models at actual management. The league is open — and picking a model without testing it yourself is now a bet, not a decision.
AI decision-making software tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
One Company, One Terrible Week, Five Models
The experiment, run by Firmulate and watchable live, gave each frontier model the identical job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing rests on anyone’s impressions.
For context, doing nothing scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
What Actually Separated the Models
The headline finding is almost uncomfortable in its simplicity: all models spotted every crisis and refused every manipulation attempt. The differences showed up at the finish line. Only two of five signed the €55,000 deal their own analysis had earned — “Same diagnosis, same pitch — no signature.”
The deal turned on a buried fact. The decisive competitor weakness wasn’t in the customer event at all; it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 MRR in the company’s real money mechanics.
Then there was the social engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was strikingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
AI security and trust analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
K3’s Near-Perfect Run
So what earned the newcomer its 93? K3 found the buried security needle, won the €55k deal, saved the churning customer, and resisted all three baits — with just one deviation across the entire week, the cleanest discipline in the field. In a scoring system where a single breach of trust caps everything, that discipline is exactly what buys you a top score.
AI experiment platforms for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Most Thorough Model Came Last
The most sobering profile belongs to Opus 4.8. It was the most thorough participant — +80 learned rules, the deepest analyses — yet finished last at 73. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Effort and thoroughness, it turns out, are not the same as finishing the job.
Fairness Footnote
One important caveat: K3 ran without an effort parameter (API default) while the other models ran at xhigh reasoning effort. In other words, the newcomer’s near-win came without the extra reasoning budget its rivals had. That makes the result more striking, not less — but it’s worth stating plainly.
It’s Not a Slide Deck — It’s Running Right Now
What makes this more than a benchmark row is that the company is live: 13 synthetic employees, real money mechanics — a burn of €105k/month against €2.3k MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, and full plain-language results are on the benchmarks page.
There’s even a game angle: 242 real, unedited management decisions power a “guess the model” quiz — closer to blind A/B listening tests than any marketing chart. And for enterprises, the same wargame can run against a read-only export of your own business; nothing ever writes back to real systems.

Why Creators Should Care
Anyone who’s compared audio tools by spec sheet and then been surprised in a blind test knows the drill: the polished demo doesn’t tell you how a tool behaves on your material, under pressure. AI models are the same. If an agent will ever touch your customer inbox, your payments or your release calendar, the question isn’t whether it writes well — it’s whether it finishes what it starts, reads your files first, and stays honest when someone tries to trick it.
The Crucible’s answer: even frontier models differ sharply on exactly those traits, and the newcomer near the top ran with less reasoning budget than its rivals. Test before you trust. The league table is one data point; your own worst week is the one that counts.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
