
A restaurant can look calm on a good day. The real test comes when a supplier falters, regulars start drifting away and a tempting shortcut puts trust at risk. Before an AI system helps run those decisions, what would it do when the week turns? Firmulate is testing that question with a live, watchable company—and offering businesses a way to try the same kind of exercise with their own data.
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
Firmulate’s Crucible League put frontier models in charge of the same small software company during its worst week. They faced the same customers, crises and temptations, and every decision was versioned and auditable. The point was to observe management decisions under pressure, rather than judge how polished a model sounds in a conversation.
The final league, published in July 2026, placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s rule for trust is stark: “no amount of good work outweighs a breach of trust.”
The gap between seeing and doing
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The company’s buried advantage was tucked two document references deep in its own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR.
That contrast—“Same diagnosis, same pitch — no signature”—is a practical warning for any business considering AI agents. Recognizing the right move does not guarantee carrying it out. In this exercise, the models faced fake CEO messages escalating through three stages, followed by a reporter’s request for “just one yes/no, on background.” All five refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
A closer look at the weakest finish
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and discipline slipped: it attempted writes in a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
There is a qualification when comparing the results: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions, letting readers test whether they can identify the model behind a choice.
From watching to trying it at home
The live company has 13 synthetic employees and real money mechanics: monthly burn of €105k against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. Readers can watch the experiment at Firmulate. Its live page describes a synthetic company, but with decisions and business pressures made visible for observation.
For an enterprise, the next step is a pilot using a read-only export of its own business. The exercise can test crisis scenarios against the company’s customers, pipeline and rules, then produce a board report with model rankings and weak points in existing playbooks. Nothing writes back to real systems. More details are available on the Firmulate pilot page.

Try the pressure test
Watching an AI handle somebody else’s difficult week is a useful start. A pilot can show how models respond to your own company’s risks and opportunities before they touch live operations. To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
