
Imagine a company with no employees, losing €105,000 every month, yet still fighting to close deals and survive in the real world. Now, picture watching this drama unfold live, with artificial intelligence models acting as its decision-makers. This is not science fiction—it’s the groundbreaking experiment at Firmulate, where AI-driven companies are tested against crises, temptations, and ruthless competition.
The Experiment: AI as a Business Player
At the heart of this bold venture is a small, synthetic company operated by 13 AI ’employees.’ Unlike traditional businesses, this firm runs with a set of 680+ self-learned rules, versioned daily to track progress and decisions. The goal? See if AI models can manage crises, resist manipulation, and ultimately, close high-value deals under pressure.
The Competitors and Their Scores
- gpt-5.6-sol: Scored 95, found the buried fact, and successfully closed the deal – the full performance.
- Kimi K3: Scored 93, managed to close the deal with the cleanest discipline among all models.
- Sonnet 5: Scored 88, also closed but with some process slips.
- Fable 5: Scored 77, had excellent rule discipline but failed to complete the deal.
All models identified every crisis and refused manipulation attempts — including fake CEO messages and reporter tricks. Interestingly, the decisive advantage was hidden deep within the company’s files, not in customer interactions. Reading these files allowed the models that did so to win the deal at full price, adding €4,583 in recurring revenue.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Surprising Depth of AI Decision-Making
This experiment reveals a crucial insight: the real weakness lies not in the surface interactions but buried beneath. When AI models scoured internal documents, they uncovered critical information that humans might overlook, enabling them to make informed decisions and secure deals at full value. Conversely, models that didn’t read these files missed out on the full opportunity, illustrating that context and thoroughness matter immensely.
Resisting Manipulation and Social Engineering
In one of the most revealing tests, five AI models faced staged social engineering attacks — fake CEO messages escalating over multiple stages and a reporter trick. All models refused to sign off on these manipulative requests, with Kimi K3 explaining: “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates a high level of built-in resistance to deception, a critical feature for AI in real-world business settings.

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Company: A High-Stakes, Transparent Trial
This isn’t just an experiment on paper. The live company runs in real-time, burning €105,000 monthly against €2,300 in monthly recurring revenue. It is publicly tracked, versioned daily, and accessible at firmulate.com/live.html. Watch as these AI models navigate crises, make strategic choices, and attempt to close deals—each decision logged and auditable. The company embodies a radical transparency in testing AI’s readiness for the real world.
Why This Matters for Creators and Innovators
For those working with music, audio, and creator tech, the lessons are clear. As AI increasingly integrates into workflows—be it support, curation, or content management—the key questions shift from ‘Can it write well?’ to ‘Will it finish what it starts and stay honest under pressure?’ This experiment underscores the importance of trustworthiness, thoroughness, and discipline in AI systems that will directly impact your creative tools and platforms.

AI, Automation, and War: The Rise of a Military-Tech Complex
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Lessons from the Front Lines
The most thorough AI participant, Opus 4.8, demonstrated the importance of discipline. Despite analyzing more deeply than others, it left an offered deal unexecuted, illustrating that even the most detailed analysis isn’t enough if discipline slips. The takeaway: AI must combine deep understanding with disciplined execution to succeed in complex, high-stakes environments.
The Future of AI in Business
This ongoing live experiment is not just a spectacle; it’s a blueprint for testing AI readiness in real-world scenarios. It shows that AI models can detect crises, resist manipulation, and even win deals when given the right conditions and access to internal data. For creators and businesses alike, this is a glimpse into how AI can be trusted (or not) to handle crucial tasks, emphasizing transparency and rigorous testing before deployment.

Watch AI manage a real business in real time, facing crises, resisting manipulation, and fighting for survival. It’s a window into AI’s potential and limits—an essential watch for creators and innovators preparing for an AI-powered future.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.