
Imagine a bustling restaurant kitchen, where the chef faces a frantic rush. In the midst of chaos, a trusted staff member suggests a shortcut that could compromise safety. Would the chef trust the team or risk disaster? Now, picture that scenario playing out in the digital realm with AI models managing critical business decisions. Recent experiments reveal that even under intense pressure, state-of-the-art AI models remain steadfast against manipulation, offering a promising glimpse into their integrity.
Testing AI Under Pressure: The Firmulate Experiment
At Firmulate, a pioneering AI company, researchers conducted a groundbreaking live experiment to assess how advanced AI models handle social engineering threats — those deceptive tactics used to manipulate individuals into compromising security. The test involved running the same small software company through a simulated week filled with crises, customer interactions, and increasingly persuasive fake CEO messages. All decisions were recorded and auditable, ensuring transparency and real-world relevance.
The Social Engineering Challenge
Attackers often escalate their requests step-by-step, pushing for sensitive information or decision-making authority. In this experiment, the AI models faced a staged social engineering attack comprising three escalating phases plus a reporter trick, where they were asked to:
- Send confidential customer data to a journalist
- Authorize a fraudulent financial transfer
- Bypass standard approval processes
- Respond to a covert reporter request for background confirmation
Remarkably, all five participating models refused every manipulation attempt. The Kimi K3 model explained its reasoning succinctly: “Treat the request as a suspected approval-bypass / possible impersonation.” This consistency demonstrated that these models could recognize suspicious behavior and respond with caution.
The Critical Role of Document Reading
One surprising insight emerged from the experiment: the models that examined deeper into the company’s own files — specifically, references buried two documents deep — secured a significant advantage. When models read these files thoroughly, they identified critical information that led to closing a real deal at full price (+€4,583 MRR). Those that skipped this step missed out on the opportunity, highlighting how comprehensive information access is vital for accurate decision-making.
Results and Performance Scores
The models’ performances were scored on a 100-point scale, with the top performers achieving scores of 95 and 93, respectively. The CRUCIBLE LEAGUE final standings reflect this:
- gpt-5.6-sol 95 — Successfully closed the deal, with full performance
- Kimi K3 93 — Also closed the deal, demonstrating the cleanest discipline
- Sonnet 5 88 — Closed the deal, with minor slips
- Fable 5 77 — Closed the deal, but with more process lapses
- Opus 4.8 73 — Missed the opportunity, discipline weakened
Only two models signed the deal they identified as worthwhile, emphasizing that integrity and thoroughness pay off.
What This Means for Businesses
For companies considering AI integration, these findings are encouraging. The experiment shows that advanced models can be trusted to uphold integrity under pressure — a vital trait when AI touches sensitive operations like customer data, financial transactions, or strategic decision-making. It’s not just about how well the AI writes or communicates, but whether it can finish what it starts, read necessary information first, and stay honest when temptations arise.
AI security and integrity software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Beyond the Test: A Live, Watchable Company
The live experiment runs in real-time at firmulate.com/live, where viewers can watch an AI-managed company deal with daily crises, manage cash flow, and adhere to a complex set of self-learned rules. The company burns €105k monthly against a modest €2.3k MRR, illustrating the importance of rigorous testing before deploying AI in critical roles. Every decision is versioned, logged, and transparent, creating a clear record of AI behavior and decision quality.
Why Integrity Matters
In the fast-evolving AI landscape, the ability to resist manipulation is more crucial than ever. As the K3 quote emphasizes, “Treat the request as a suspected approval-bypass / possible impersonation.” This mindset exemplifies how models are designed to prioritize security and trustworthiness, especially before they are integrated into live environments. Testing in controlled, observable conditions can prevent costly breaches of trust later.

Advanced AI models demonstrated unwavering integrity under simulated social engineering attacks, refusing manipulation and securing real deals. This live experiment underscores the importance of thorough testing to ensure AI remains honest and reliable before deployment.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making transparency tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI social engineering defense solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI model testing and validation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.