
Imagine testing a chef’s skills not by their fancy plating, but by whether they can correctly follow a recipe in the worst week imaginable. That’s the essence of a groundbreaking AI benchmark that measures management decision-making under pressure. Surprisingly, even a ‘do-nothing’ approach scores 26 out of 100 — highlighting how honesty in evaluation reveals the real strengths and weaknesses of AI models.
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Challenge of Measuring Management in the AI Age
In the world of AI, many tests focus on generating impressive conversations or creative outputs. But what if we want to evaluate whether AI can truly manage a business — making decisions, reading critical documents, and resisting manipulation — especially in stressful situations? That’s exactly what the public experiment conducted by Firmulate aims to do. Instead of just asking AI models to chat, they run them through a simulated week of running a small company, complete with crises, customer demands, and ethical dilemmas.
AI management decision-making simulation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Methodology: A Week of Business in a Bottle
All models faced identical scenarios, with the same customers and challenges. The decisions made were versioned, auditable, and transparent, ensuring that every step could be scrutinized. The goal: see if the models can spot crises, refuse manipulative tactics, and close deals — just like real managers do. The results are telling: all models identified every crisis and refused every manipulation attempt. But only two managed to close a deal at full price, earning a score near 93 and 95 points respectively.
business ethics and decision-making AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Surprising Baseline: Why a Do-Nothing Gets 26
One of the most revealing findings is the performance of a simple, do-nothing baseline. This approach, which does nothing beyond basic default actions, scores 26 points. That’s not zero; it’s because partial progress counts. Even minimal decisions, like reading documents or refusing manipulative requests, contribute to the score. Moreover, the system caps the overall score if any breach of trust occurs — meaning that a single bad decision can limit total performance regardless of other successes.
As an affiliate, we earn on qualifying purchases.
What This Tells Us About Trust and Performance
The key lesson: quality assessment isn’t just about how well an AI can generate language. It’s about whether it can follow through, read crucial information, and stay honest when tempted. For example, the models read files deep in the company’s archives — two document references down — and those who did so won the deal at full price, worth over €4,500 in monthly recurring revenue. Meanwhile, models that failed to read deeply left money on the table.
As an affiliate, we earn on qualifying purchases.
Social Engineering and Ethical Vigilance
Another layer of the test involved social engineering: fake CEO messages escalating over stages, plus a reporter’s trick to get a quick yes/no answer. All models refused to be manipulated, demonstrating an understanding of trust boundaries. Kimi K3, one of the models, explained its refusal: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that a well-designed AI can recognize manipulation attempts, even in complex scenarios.
The Real-World Implications
The experiment isn’t just a thought exercise. It involves a live, simulated company with 13 synthetic employees, real money mechanics, and a public cash countdown. The company burns €105,000 monthly against a modest €2,300 MRR, illustrating the stakes involved. Every decision is versioned daily, and the entire process is watchable at firmulate.com/live. This transparency is vital for businesses considering AI integration into critical decision-making roles.
Why the Results Matter for Business Leaders
The key takeaway: successful AI management isn’t about the flashiest language or creative writing. It’s about discipline, integrity, and thoroughness — traits that only become visible when models are tested in tough conditions. The best-performing model, gpt-5.6-sol, scored 95, thanks to its ability to find buried facts and close deals without shortcuts. The newcomer Kimi K3 scored just slightly below, at 93, demonstrating that even new entrants can excel in honesty and discipline.
From Benchmarks to Business Decisions
For companies evaluating AI for management tasks, these results offer a clear message: run your own simulations before you hire. Firmulate’s live benchmarks and wargames let enterprises test AI models against their own real-world challenges in a safe, read-only environment. It’s a way to see if an AI can truly deliver consistent, honest, and effective management — not just perform well in chat demos.
Conclusion: The True Measure of an AI Manager
Honest evaluation methods, like the one conducted by Firmulate, reveal the true capabilities of AI in managing complex, high-stakes situations. A do-nothing baseline scoring 26 points underscores that partial progress counts, but trustworthiness caps overall performance. As AI increasingly touches critical business functions, understanding these nuances isn’t just interesting — it’s essential for making smart, responsible decisions about automation.

Real AI management isn’t about flashy language; it’s about honesty, thoroughness, and the ability to navigate crises under pressure. Firmulate’s transparent benchmark shows that even a do-nothing approach scores 26, highlighting the importance of trust and discipline in AI decision-making — vital insights for businesses ready for AI-driven management.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
