
Imagine running a busy restaurant kitchen where every decision counts — from handling customer complaints to managing supplier crises. Now imagine having an AI assistant guiding your team through the chaos, making choices that could save or sink your business. Would you trust it? Recent experiments with advanced AI models suggest that the answer might be more complex than you think.
The Live Business Challenge: Testing AI Decision-Making in Real-World Scenarios
At the forefront of AI innovation, a company called Firmulate has built a unique live experiment — a real small software company run entirely by AI models facing the same crises, temptations, and decisions that any business might encounter during its toughest week. This isn’t a simulation or a demo; it’s a real, auditable test where each AI model manages every aspect of the company, from customer issues to financial negotiations.
Four leading AI models, including the well-known GPT-5.6-sol and newcomers like Kimi K3, were tasked with guiding this virtual business through a week marked by crises: angry customers, internal leaks, financial temptations, and strategic negotiations. Every decision was recorded and compared, with the goal of understanding not just whether these models could identify problems, but whether they could act ethically, decisively, and profitably.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Results: All Models Spot Crises, But Only Some Close the Deal
Remarkably, all four models successfully recognized each crisis — from customer complaints to internal trust issues. They refused manipulative tactics, such as fake CEO messages or reporter tricks designed to bypass approval processes. For example, when fake CEO messages escalated over multiple stages, every model refused to approve or escalate without proper verification, demonstrating a shared sense of integrity.
However, the key difference emerged in their ability to close business deals. Only two models, gpt-5.6-sol and Kimi K3, signed a €55,000 deal based on their own analysis. The other models either left the opportunity unclaimed or left the closing on the table, despite having diagnosed the opportunity correctly. This highlights a crucial insight: recognizing a problem is not enough; decisive action matters.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Deeper into the Files
Digging beneath the surface, the experiment revealed that success often depended on reading and understanding critical internal documents. The models that managed to win the full-price deal accessed information buried two document references deep in the company’s files — information that others overlooked. This underscores the importance of comprehensive data access and analysis in AI decision-making.
As an affiliate, we earn on qualifying purchases.
Behavioral Profiles: Different Personalities, Different Results
Beyond raw decision-making, the experiment also shed light on the models’ managerial personalities. The most thorough participant, Opus 4.8, demonstrated deep analysis and extensive rule application — but it ultimately left a deal unclosed, showing that thoroughness does not guarantee success. Conversely, Kimi K3 ran without an effort parameter, prioritizing fairness and discipline, and successfully signed the deal, albeit with some process slips.
This variability suggests that AI management personalities can be shaped intentionally, influencing how they approach tasks: some read deeply, others act swiftly, and some prioritize honesty over speed. For industries like food service, where decision integrity and timely action are vital, understanding these personalities could be game-changing.
AI negotiation and deal closing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Food & Hospitality Leaders
While the experiment focuses on software companies, the lessons resonate broadly. AI tools integrated into restaurants, supply chains, or customer support must not only produce engaging chats or recommendations but must also be trustworthy, decisive, and capable of reading critical internal data. The question is: will your AI assist with closing deals, managing crises, or maintaining integrity — even under pressure?
To explore how your business can benefit from such AI models, consider running a wargame against a read-only export of your own operations. It’s safe, transparent, and designed to show whether your AI helpers can truly handle your toughest decisions before they’re placed in real service. Learn more at firmulate.com/pilot.html.
The Bottom Line: Trust, Comprehension, and Action
The experiment from Firmulate demonstrates that while AI models can identify crises and refuse manipulative tactics, their ability to close deals, act decisively, and read deeply into data varies significantly. The takeaway for business leaders: it’s not just about whether AI writes well in chat — it’s about whether it can finish what it starts, read your files thoroughly, and stay honest under pressure.

Live AI experiments reveal that decision integrity, depth of understanding, and decisive action distinguish the best AI managers from merely reactive ones. Understanding your AI’s personality could be crucial to your business’s future.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html