
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
When Volume Doesn’t Equal Victory: The Limits of AI Diligence
Imagine preparing a complex recipe — meticulously gathering ingredients, following every step, yet still missing the secret spice that makes the dish unforgettable. In the world of AI, diligence and thoroughness alone aren’t enough to close the deal. Even the most disciplined models may falter when it matters most.
AI decision-making analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Firmulate Experiment: Putting AI Through Its Paces
Recently, a public experiment by Firmulate put four advanced AI models in the same challenging scenario: managing a small software company’s worst week. Every model faced the same crises—customer issues, internal temptations to manipulate data, and social engineering tricks—designed to test their integrity, focus, and decision-making.
Each AI was tasked with navigating these storms, with their decisions fully recorded and auditable. The models ranged from the well-known gpt-5.6-sol to newer contenders like Kimi K3, and the results reveal some surprising truths about AI performance and discipline.
The Scores and Outcomes
- gpt-5.6-sol led with a score of 95, successfully uncovering buried information in the company’s data files and sealing a €55,000 deal—a full performance achievement.
- Kimi K3 scored 93, demonstrating the cleanest discipline and also closing the deal.
- Sonnet 5 scored 88, with a few process slips but still managing to close the deal.
- Fable 5 scored 77, also closing but with more slips along the way.
- The baseline, a do-nothing approach, scored 26, showing minimal progress and poor decision-making.
Remarkably, all models identified every crisis and refused manipulation attempts, including social engineering tricks like fake CEO messages and reporter requests. The key difference lay in the depth of analysis and discipline: only two models signed the deal they had earned through their own analysis, while the other two left opportunities on the table due to lapses in focus and escalation discipline.
The Hidden Weakness: Reading the Files
Further analysis revealed that the decisive edge was the models’ ability to access and comprehend the company’s internal documentation. Those that read and understood deeper references in the company’s files managed to close the deal at full price, adding €4,583 MRR. It’s a stark reminder that volume of effort isn’t enough—prioritization and targeted focus matter more.
Social Engineering and Trust
In the social engineering trial, all models correctly refused to execute harmful requests, treating suspicious messages as potential impersonation or approval bypass attempts. Kimi K3 explicitly stated: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a shared understanding that discipline and skepticism are vital in real-world AI applications.
The Live Company Experiment
Firmulate’s live setup simulates a real company with 13 synthetic employees, real money mechanics—burning €105k/month against a modest €2.3k MRR—and a public cash countdown. Every day, decisions are versioned and analyzed, providing a transparent view of AI performance under pressure. This ongoing experiment underscores the importance of discipline, prioritization, and focused analysis over sheer effort volume.
AI focus and discipline training software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Lessons for AI and Business
What does this mean for enterprises deploying AI? The core takeaway is straightforward: diligence alone does not guarantee success. Models that focus on reading key documents, understanding context, and resisting manipulative tactics are more likely to close deals, maintain integrity, and contribute real value.
Furthermore, the experiment shows that even the most thorough participant (like Opus 4.8 with over 80 learned rules and deep analyses) can falter if discipline slips—such as failing to escalate instead of writing into a locked department. This highlights that discipline, prioritization, and strategic focus are critical elements alongside diligence.
Why This Matters for Business Leaders
As AI begins to touch sales, support, and decision-making workflows, it’s tempting to assume that volume of effort or complexity is enough. But the Firmulate experiment demonstrates that the real challenge lies in focus and integrity—reading the right information, resisting shortcuts, and sticking to disciplined decision protocols. In practice, AI models that excel in these areas can unlock full value, closing high-value deals and avoiding costly mistakes.
To explore how your business can simulate these scenarios and prepare your AI workforce, consider running a wargame based on your own operations. This approach offers a safe, risk-free way to identify weaknesses before deploying AI at scale.
Discover the live experiment and see the models in action at firmulate.com/live, where transparency and real-world testing continue to shape the future of AI in business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI internal documentation comprehension tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI social engineering detection systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.