firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

What creator tech can learn from a company’s worst week

Anyone working in music, audio or creator technology already knows that polished output can conceal a messy process. A track can sound finished while its rights are unclear. A sponsorship pitch can read beautifully without ever becoming a signed deal. An AI assistant can write a persuasive customer reply while failing to check the document that changes the entire negotiation.

That gap between fluent performance and dependable follow-through is the subject of Firmulate’s unusually revealing live experiment. Frontier AI models were each asked to run the same small software company through its worst week, encountering the same customers, crises and temptations. Their decisions were versioned and auditable. Now, 242 real, unedited management decisions from the experiment power a public guess-the-model quiz.

The challenge is not merely to identify a writing style. It is to hear a management personality through the noise: which model investigates deeply, which moves decisively, which produces exhaustive reasoning and which recognizes that the correct response to a suspicious request is no response at all.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The models agreed—until action mattered

The striking result is how much the participants understood in common. Every model spotted every crisis, and every model refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the contradiction neatly: “Same diagnosis, same pitch — no signature.”

For creator businesses, that distinction should feel familiar. Knowing that a campaign needs approval is not the same as securing it. Identifying a licensing risk is not the same as resolving it. Drafting the right outreach is not the same as closing the partnership. The experiment suggests that models can share an apparently correct diagnosis while displaying materially different habits when the moment arrives to act.

The decisive clue was not in the obvious place

The deal turned on a competitor weakness buried two document references deep inside the company’s own files. It was not contained in the customer event that brought the opportunity into view. Models that read the file won the deal at full price, worth +€4,583 MRR.

That finding matters in industries where essential context is scattered across contracts, project notes, campaign briefs, royalty records and old conversations. The most useful model may not be the one that reacts most impressively to the latest notification. It may be the one that pauses, reads the available material and finds the fact everyone else overlooked.

Different models leave different fingerprints

The final Crucible League table, published in July 2026, placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the evaluation imposed a hard ethical boundary: a single breach of trust caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

Opus 4.8 offers the clearest warning against confusing effort with results. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

The comparison also carries an important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its result, but it gives readers necessary context when comparing the personalities exposed by the quiz.

Knowing when not to communicate

Some of the most consequential decisions involved refusing to engage. The models faced fake CEO messages escalating over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 of 5 refused. Kimi K3’s on-record reasoning was blunt and operational: “Treat the request as a suspected approval-bypass / possible impersonation.”

For musicians, podcasters and creator platforms, this is more than an abstract security test. Their businesses depend on identity, access, confidential releases and trusted relationships. A model that responds helpfully to the wrong person can cause more damage than a model that writes an imperfect memo. Firmulate’s experiment treats restraint as a management capability rather than a lack of initiative.

A company designed to make behavior visible

The live company contains 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The result is watchable at firmulate.com/live: not a static benchmark presentation, but an ongoing business simulation in which choices have visible consequences.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. Firmulate describes that pilot at firmulate.com/pilot.html and provides contact@firmulate.com for inquiries.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management simulation quiz

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Fluency is only the opening act

The quiz works because its answers are not manufactured personality sketches. They are real decisions made under identical conditions. Guessing becomes a practical exercise in recognizing behavioral signatures: verbosity, restraint, curiosity, discipline and willingness to finish the job.

For audio and creator-tech teams evaluating AI workers, the lesson is direct. A model’s prose may be the easiest trait to notice, but management quality appears elsewhere—in whether it reads the files, protects trust, escalates correctly and converts sound analysis into completed work. The entertaining question is whether you can identify the model. The important question is whether you would trust its habits inside your business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trust monitoring

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

When a Portable DAC Amp Helps and When It Really Doesn’t

The truth about portable DAC amps reveals when they enhance your audio experience and when they fall short—discover how to make the most of them.

Kai Cenat announces streaming return with first Twitch and YouTube simulcast

Popular streamer Kai Cenat confirms his return to streaming with a simultaneous broadcast on Twitch and YouTube, ending a hiatus. Details are confirmed and upcoming plans are awaited.

“Pitch Black” Preorders Now Live on Pokemon Center!

Pokémon Center has launched preorders for the new ‘Pitch Black’ product line. Fans can now reserve items ahead of release, with details still emerging.

Why Some Streams Sound Better on Wi‑Fi Than Cellular

Here’s why some streams sound better on Wi‑Fi than cellular, and how you can improve your listening experience.