AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In the world of automation, sometimes the biggest revelations come from the least impressive starting point. For business owners and managers, understanding what an ‘honest’ AI can do — even when it does nothing — is crucial for assessing future potential and risk.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get cleaning gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Baseline: Why Doing Nothing Still Scores 26 Points

In a recent public experiment by Firmulate, four leading AI models were tested against a simulated week of running a small software company. Interestingly, even the most inactive or cautious AI managed to score 26 out of 100 points. This baseline score is not a mistake or an artifact; it’s an intentional part of the benchmarking process that reveals what minimal, honest AI performance looks like.

The Methodology: Simulating a Real-World Crisis Week

Each AI was tasked with managing the same set of company challenges — from customer crises to internal manipulations — all designed to test decision-making under pressure. Every move was documented and auditable, ensuring transparency. The models faced identical temptations to cheat or manipulate, and their responses were measured against strict criteria.

Partial Progress Counts — But Trust Is the Cap

The results showed that all four models could identify every crisis and refuse manipulation attempts, demonstrating solid integrity. However, only two of the models went further and signed the contract for €55,000, which their own analysis indicated they had earned. The other two, despite diagnosing correctly and presenting a proper pitch, left the deal on the table.

Amazon

enterprise AI trustworthiness software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the Baseline Is Not Zero — And Why That Matters

The fact that even a ‘do-nothing’ approach scores 26 points is revealing. It indicates that a minimal level of honest, cautious behavior earns some recognition. This is essential for real-world applications, where AI’s ability to read, interpret, and act honestly underpins trustworthiness. In practice, partial progress — such as identifying a crisis or refusing manipulation — counts significantly.

The Hidden Weakness: Reading Company Files

The experiment uncovered a critical flaw: the decisive factor in closing the deal lay two document references deep in the company’s own files, not in the immediate customer interactions. Models that read and understand these files secured the full-price deal, highlighting the importance of deep contextual comprehension in trustworthy AI.

Social Engineering and AI Integrity

The models faced staged social engineering attempts — fake CEO messages escalating over three rounds plus a reporter trick — and all refused to cooperate. Kimi K3’s on-record reasoning encapsulates this: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a fundamental commitment to integrity, even when under pressure.

What This Means for Business and Automation

For companies considering AI adoption, the takeaway is clear: performance isn’t just about language fluency or quick responses. It’s about honesty, the ability to read deeply, and resisting manipulative tactics. The experiment shows that AI models can and do recognize integrity breaches and refuse to act dishonestly, but only if they are designed with such discipline in mind.

The Live Experiment: Running a Real Company in Real Time

Firmulate’s platform hosts a live, watchable simulation of an actual small company with 13 synthetic employees. It operates with real money mechanics — burning €105k/month against a monthly revenue of just €2.3k — with daily, versioned decision-making. Stakeholders can observe how AI models handle crises, temptations, and negotiations, providing a transparent glimpse into their true capabilities.

The Deep Dive: Opus 4.8’s Performance

The most thorough participant in the experiment, Opus 4.8, applied over 80 learned rules and performed extensive analysis. Despite this, it ranked last, failing to close the deal and slipping into departmental silos instead of escalating issues. This highlights that more rules and deeper analysis don’t necessarily guarantee better trust or performance — discipline and strategic focus matter.

Implications for Business Automation and Trust

This experiment underscores a vital point: in deploying AI, especially in sensitive tasks like management or customer relations, trust and honesty are non-negotiable. A high score on chat demos or superficial metrics doesn’t reveal whether an AI will stay honest when it matters most. The baseline score, and the fact that even cautious models score above zero, signals that trustworthy AI must be built with integrity at its core.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best Bissell Carpet Cleaner for Stairs (2026) — Guide 12

Discover the top Bissell cleaning devices in 2026. Find the best overall, value, and specialized options to match your cleaning needs today.

Rug Pads and Mold: How to Avoid Trapping Moisture

Discover how to prevent mold under rug pads and ensure your space stays safe and dry. Keep reading for essential tips and tricks.

Best Bissell Carpet Cleaner for Stairs: Top Picks for Every Need

Discover the best Bissell models for deep cleaning, portability, and pet stain removal in 2026. Find the perfect fit for your cleaning needs today.