firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A cleaning company does not find out whether its maintenance plan works when every floor is spotless. The real test comes when a machine fails, a customer threatens to leave and someone asks for a shortcut. AI management deserves the same kind of pressure test: put it through a company’s difficult week before trusting it with decisions.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get cleaning gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate’s live experiment gives AI models the same small software company to run through a rough week: the same customers, crises and temptations. Its synthetic workforce has 13 employees and real money mechanics, with monthly burn of €105,000 against €2,300 in monthly recurring revenue. A public cash countdown makes the stakes visible. Each workday is versioned, and the company has built more than 680 self-learned playbook rules.

The experiment is watchable at firmulate.com. The point is not whether a model can sound like a manager. It is whether its decisions hold up when the pressures of running a business collide.

Good diagnosis is not enough

In the final Crucible League, dated July 2026, gpt-5.6-sol placed first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s integrity rule is blunt: “no amount of good work outweighs a breach of trust.”

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” Recognizing the right move and carrying it through are different tests of management.

The sale depended on a detail buried in the company’s own files, two document references deep. It was not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The story is a reminder that the decisive clue may sit in a company’s records, not in the alert that first draws attention.

Trust and discipline under strain

Social engineering provided another test. Fake messages from a CEO escalated across three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. It left the deal unsigned and attempted to write into a locked department instead of escalating. The same weakness appeared, in a weaker form, in all four models. More analysis did not guarantee that the work reached the right decision.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The league offers a concrete snapshot, with that difference part of its context.

From watching to testing your own business

For a cleaning and floor-care business, a useful exercise might begin with a familiar hard week: a breakdown, a missed service, a customer complaint and pressure to approve an exception. The question is whether an AI can find the relevant contract or maintenance record, protect the business from an impersonation attempt and follow through on a sound decision.

Firmulate’s pilot applies the wargame to an enterprise’s own business using a read-only data export. The company can test crisis scenarios and receive a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. A quiz built from 242 real, unedited management decisions also lets readers try to guess which model made each choice.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the hard week to the test

Watching a model handle someone else’s company can reveal how it behaves. A pilot can show how it handles yours. To discuss running the wargame against a read-only export of your business, visit the Firmulate pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Dual-Tank Advantage: Why Dirty Water Separation Matters

Keen to improve water system efficiency, but unsure how dual-tank separation can help? Discover the key benefits that make this technology essential.

How Wet-Dry Vacuum Mops Handle Fine Dust Differently Than Traditional Mops

The wet-dry vacuum mops handle fine dust differently than traditional mops by trapping microscopic particles with advanced filtration; discover how they revolutionize cleaning.

How to Choose the Right Cleaning Machine Category for the Mess You Actually Have

Understand your cleaning needs first to select the perfect machine, but discover the key factors that will make your choice truly effective.

How Humidity Changes the Way Dust Settles on Floors

Get insights on how humidity impacts dust settling on floors and discover surprising factors that keep your space cleaner than you think.