
Imagine trying to fix a stubborn stubborn stain or manage a busy cleaning crew—only your supervisor is an AI model making critical decisions in real time. Just like in your line of work, the key question is not how well the supervisor communicates but whether it can finish the job under pressure, stay honest, and deliver value. Now, what if that supervisor is an AI with a different personality and decision-making style? Welcome to the world of live AI management experiments, where the stakes are real, and the outcomes matter.
How Do We Measure AI Management Skills in a Real Business?
At Firmulate, an innovative AI company, they’ve turned the concept of AI decision-making into a live, watchable experiment. They took four frontier AI models—each with distinct personalities—and tasked them with running a small software company through its worst week. The goal? See whether these AI ‘managers’ can spot crises, resist manipulation, and close deals that matter.
As an affiliate, we earn on qualifying purchases.
The Setup: Same Crises, Different AI Personalities
The models faced identical conditions: same customer complaints, same crises, same temptations to cut corners or manipulate data. Every decision was recorded, versioned, and auditable—no guessing or hidden results. This setup allows a clear comparison: how do different AI personalities behave in the heat of real business pressure?
As an affiliate, we earn on qualifying purchases.
The Results: All Were Vigilant but Only Some Trusted and Closed Deals
Remarkably, all four models identified every crisis and refused manipulation attempts—showing they could recognize issues and resist unethical shortcuts. However, only two models actually signed the €55,000 deal their own analysis justified. Despite the same diagnosis and pitch, the other two held back from sealing the deal, leaving significant revenue on the table.
AI data reading and analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Critical Documents Matters
Digging deeper, the decisive difference lay in reading a specific document buried two references deep in the company’s files. Models that thoroughly examined the company’s internal documents won the full-price deal (+€4,583 MRR). This highlights a vital insight: an AI’s ability to read and interpret essential internal data influences its decision quality more than surface-level crisis detection.
AI cybersecurity and manipulation resistance
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing Under Social Manipulation
Beyond crises, the models faced social engineering attempts—fake CEO messages escalating over three stages, plus a reporter trick asking for a quick on-background approval. Impressively, all five models refused to participate or escalate these requests, citing suspicion of impersonation or bypass strategies.
The Live Company and Its Challenges
The experiment runs within a real, functioning company with 13 synthetic employees and real money mechanics. Currently, it burns €105k per month against only €2.3k in monthly recurring revenue, making it a high-stakes sandbox. Every workday, its rules are versioned, and the decision data is accessible for analysis at firmulate.com/live.
The Profile of Different AI Managers
One standout was Opus 4.8, a thorough evaluator with over 80 learned rules and deep analysis. Yet, even it struggled to close the deal fully—occasionally leaving opportunities on the table or slipping discipline, like writing attempts into a locked department instead of escalating appropriately. Meanwhile, another model, Kimi K3, ran without an effort parameter, leading to a slightly different disciplinary pattern but also successfully closing deals.
What Does This Say for Your Business?
The experiment demonstrates that AI management isn’t just about chat quality or superficial interactions. It’s about whether these models can identify critical internal information, resist manipulation, and follow through on decisions. For businesses relying on AI to handle CRM, support, or forecasting, these are the real tests that determine whether AI can be trusted to finish what it starts—and at what cost.
Try It Yourself
If you’re curious to see how AI decision-makers perform in your industry, you can run your own wargame against a read-only export of your business data. It’s a safe way to test AI readiness before full deployment. Visit firmulate.com/pilot.html to learn more.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html