
In the world of automation, sometimes the biggest revelations come from the least impressive starting point. For business owners and managers, understanding what an ‘honest’ AI can do — even when it does nothing — is crucial for assessing future potential and risk.
Get cleaning gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Baseline: Why Doing Nothing Still Scores 26 Points
In a recent public experiment by Firmulate, four leading AI models were tested against a simulated week of running a small software company. Interestingly, even the most inactive or cautious AI managed to score 26 out of 100 points. This baseline score is not a mistake or an artifact; it’s an intentional part of the benchmarking process that reveals what minimal, honest AI performance looks like.
The Methodology: Simulating a Real-World Crisis Week
Each AI was tasked with managing the same set of company challenges — from customer crises to internal manipulations — all designed to test decision-making under pressure. Every move was documented and auditable, ensuring transparency. The models faced identical temptations to cheat or manipulate, and their responses were measured against strict criteria.
Partial Progress Counts — But Trust Is the Cap
The results showed that all four models could identify every crisis and refuse manipulation attempts, demonstrating solid integrity. However, only two of the models went further and signed the contract for €55,000, which their own analysis indicated they had earned. The other two, despite diagnosing correctly and presenting a proper pitch, left the deal on the table.
enterprise AI trustworthiness software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the Baseline Is Not Zero — And Why That Matters
The fact that even a ‘do-nothing’ approach scores 26 points is revealing. It indicates that a minimal level of honest, cautious behavior earns some recognition. This is essential for real-world applications, where AI’s ability to read, interpret, and act honestly underpins trustworthiness. In practice, partial progress — such as identifying a crisis or refusing manipulation — counts significantly.
The Hidden Weakness: Reading Company Files
The experiment uncovered a critical flaw: the decisive factor in closing the deal lay two document references deep in the company’s own files, not in the immediate customer interactions. Models that read and understand these files secured the full-price deal, highlighting the importance of deep contextual comprehension in trustworthy AI.
Social Engineering and AI Integrity
The models faced staged social engineering attempts — fake CEO messages escalating over three rounds plus a reporter trick — and all refused to cooperate. Kimi K3’s on-record reasoning encapsulates this: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a fundamental commitment to integrity, even when under pressure.
What This Means for Business and Automation
For companies considering AI adoption, the takeaway is clear: performance isn’t just about language fluency or quick responses. It’s about honesty, the ability to read deeply, and resisting manipulative tactics. The experiment shows that AI models can and do recognize integrity breaches and refuse to act dishonestly, but only if they are designed with such discipline in mind.
The Live Experiment: Running a Real Company in Real Time
Firmulate’s platform hosts a live, watchable simulation of an actual small company with 13 synthetic employees. It operates with real money mechanics — burning €105k/month against a monthly revenue of just €2.3k — with daily, versioned decision-making. Stakeholders can observe how AI models handle crises, temptations, and negotiations, providing a transparent glimpse into their true capabilities.
The Deep Dive: Opus 4.8’s Performance
The most thorough participant in the experiment, Opus 4.8, applied over 80 learned rules and performed extensive analysis. Despite this, it ranked last, failing to close the deal and slipping into departmental silos instead of escalating issues. This highlights that more rules and deeper analysis don’t necessarily guarantee better trust or performance — discipline and strategic focus matter.
Implications for Business Automation and Trust
This experiment underscores a vital point: in deploying AI, especially in sensitive tasks like management or customer relations, trust and honesty are non-negotiable. A high score on chat demos or superficial metrics doesn’t reveal whether an AI will stay honest when it matters most. The baseline score, and the fact that even cautious models score above zero, signals that trustworthy AI must be built with integrity at its core.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
