AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine testing a chef’s skills not by their fancy plating, but by whether they can correctly follow a recipe in the worst week imaginable. That’s the essence of a groundbreaking AI benchmark that measures management decision-making under pressure. Surprisingly, even a ‘do-nothing’ approach scores 26 out of 100 — highlighting how honesty in evaluation reveals the real strengths and weaknesses of AI models.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Challenge of Measuring Management in the AI Age

In the world of AI, many tests focus on generating impressive conversations or creative outputs. But what if we want to evaluate whether AI can truly manage a business — making decisions, reading critical documents, and resisting manipulation — especially in stressful situations? That’s exactly what the public experiment conducted by Firmulate aims to do. Instead of just asking AI models to chat, they run them through a simulated week of running a small company, complete with crises, customer demands, and ethical dilemmas.

Amazon

AI management decision-making simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Methodology: A Week of Business in a Bottle

All models faced identical scenarios, with the same customers and challenges. The decisions made were versioned, auditable, and transparent, ensuring that every step could be scrutinized. The goal: see if the models can spot crises, refuse manipulative tactics, and close deals — just like real managers do. The results are telling: all models identified every crisis and refused every manipulation attempt. But only two managed to close a deal at full price, earning a score near 93 and 95 points respectively.

Amazon

business ethics and decision-making AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Baseline: Why a Do-Nothing Gets 26

One of the most revealing findings is the performance of a simple, do-nothing baseline. This approach, which does nothing beyond basic default actions, scores 26 points. That’s not zero; it’s because partial progress counts. Even minimal decisions, like reading documents or refusing manipulative requests, contribute to the score. Moreover, the system caps the overall score if any breach of trust occurs — meaning that a single bad decision can limit total performance regardless of other successes.

Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Tells Us About Trust and Performance

The key lesson: quality assessment isn’t just about how well an AI can generate language. It’s about whether it can follow through, read crucial information, and stay honest when tempted. For example, the models read files deep in the company’s archives — two document references down — and those who did so won the deal at full price, worth over €4,500 in monthly recurring revenue. Meanwhile, models that failed to read deeply left money on the table.

Amazon

management skills assessment AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering and Ethical Vigilance

Another layer of the test involved social engineering: fake CEO messages escalating over stages, plus a reporter’s trick to get a quick yes/no answer. All models refused to be manipulated, demonstrating an understanding of trust boundaries. Kimi K3, one of the models, explained its refusal: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that a well-designed AI can recognize manipulation attempts, even in complex scenarios.

The Real-World Implications

The experiment isn’t just a thought exercise. It involves a live, simulated company with 13 synthetic employees, real money mechanics, and a public cash countdown. The company burns €105,000 monthly against a modest €2,300 MRR, illustrating the stakes involved. Every decision is versioned daily, and the entire process is watchable at firmulate.com/live. This transparency is vital for businesses considering AI integration into critical decision-making roles.

Why the Results Matter for Business Leaders

The key takeaway: successful AI management isn’t about the flashiest language or creative writing. It’s about discipline, integrity, and thoroughness — traits that only become visible when models are tested in tough conditions. The best-performing model, gpt-5.6-sol, scored 95, thanks to its ability to find buried facts and close deals without shortcuts. The newcomer Kimi K3 scored just slightly below, at 93, demonstrating that even new entrants can excel in honesty and discipline.

From Benchmarks to Business Decisions

For companies evaluating AI for management tasks, these results offer a clear message: run your own simulations before you hire. Firmulate’s live benchmarks and wargames let enterprises test AI models against their own real-world challenges in a safe, read-only environment. It’s a way to see if an AI can truly deliver consistent, honest, and effective management — not just perform well in chat demos.

Conclusion: The True Measure of an AI Manager

Honest evaluation methods, like the one conducted by Firmulate, reveal the true capabilities of AI in managing complex, high-stakes situations. A do-nothing baseline scoring 26 points underscores that partial progress counts, but trustworthiness caps overall performance. As AI increasingly touches critical business functions, understanding these nuances isn’t just interesting — it’s essential for making smart, responsible decisions about automation.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Real AI management isn’t about flashy language; it’s about honesty, thoroughness, and the ability to navigate crises under pressure. Firmulate’s transparent benchmark shows that even a do-nothing approach scores 26, highlighting the importance of trust and discipline in AI decision-making — vital insights for businesses ready for AI-driven management.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Appliance Accessories Deserve More Attention

Great accessories can transform your appliances, but their true potential is often overlooked—discover why paying attention can make all the difference.

The Easiest Way to Keep Coffee Gear Cleaner

Here’s a simple tip to keep your coffee gear cleaner effortlessly and enjoy better-tasting brews—discover the secret next.

Beat the Summer Heat with a Perfect Iced Espresso Using Ninja Luxe™ Café Pro

Learn how to make a refreshing iced espresso at home with the Ninja Luxe™ Café Pro, co-designed with David Beckham, perfect for summer days.

Why Thermal Carafes Matter More Than Hot Plates

Thermal carafes matter more than hot plates because they preserve coffee’s flavor and aroma, ensuring a better taste experience you won’t want to miss.