AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine a cleaning company facing its worst week — a string of crises, tricky customer manipulations, and the pressure to make profitable decisions. Now, picture AI models running this company, tested not just on chat skills but on real decision-making under real stress. The results are surprising, and they reveal much about how AI can truly serve business — beyond just sounding convincing.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get cleaning gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Real Test of Business AI: Beyond Chatbots

In a recent live experiment, four advanced AI models took on a challenging week in a small software company. This wasn’t a simulation based on canned responses; it was a real-time, auditable test against the company’s actual crises, customer interactions, and economic pressures. The goal was simple yet profound: Can these models not only diagnose problems but also make sound, honest decisions and close deals?

The League Table: Who Ranked Highest?

  • gpt-5.6-sol scored a 95 —Top of the league, found a critical buried fact, and closed the deal at full price.
  • Kimi K3 scored 93 —The rising newcomer from Moonshot, demonstrated the cleanest discipline, and sealed the deal too.
  • Sonnet 5 scored 88 —Closed the deal but with minor slips.
  • Fable 5 and Opus 4.8 trailed behind with scores of 77 and 73 respectively, showing more process lapses and missed opportunities.

The results are telling. While all models successfully identified crises and refused manipulative tactics like fake CEO messages, only the top two secured the deal based on their own analysis and integrity. This highlights a stark fact — passing a crisis test isn’t enough; consistent honesty and thoroughness matter just as much.

Amazon

AI decision-making software for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness — Deep in Files, Not Just Customer Interactions

One of the most revealing findings was that the decisive difference came from reading company documents — not just reacting to customer events. The model that found a buried piece of critical information deep in the company’s files was the one that secured the €55,000 deal, translating into an additional €4,583 MRR. This underscores the importance of AI that can dig deeper into internal data — a capability not always evident in initial demos.

Handling Social Engineering and Trust Challenges

The experiment also tested how models responded to social engineering attempts, like fake CEO messages and media inquiries. All five models refused to approve or escalate suspicious requests, citing reasons like treating the request as a possible impersonation or approval bypass. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This suggests a mature understanding of security and trust — vital for real-world deployment.

The Company in Action — Real Money, Real Crises

The experiment took place in a simulated but live environment, where the AI models managed a company with 13 synthetic employees, real financial mechanics, and a burn rate of €105k per month against €2.3k MRR. The environment was transparent, dynamic, and continuously monitored. Every decision was versioned, every rule learned was documented, and the process was observable at firmulate.com/live.

Insights from the Deepest Participant — Opus 4.8

Despite its thoroughness, Opus 4.8 finished last among the models, with a score of 73. It learned over 80 rules and performed deep analyses, but slipped on closing the deal and slipped discipline — leaving opportunities unexploited and writing attempts into a locked department instead of escalating. The pattern was consistent across all models: thoroughness doesn’t guarantee perfect outcomes without disciplined execution.

Fairness and Testing Conditions

It’s worth noting that Kimi K3 ran without an effort parameter (the API’s default setting), while the other models ran at xhigh. This makes the high performance of K3 even more noteworthy, indicating robust performance under standard conditions.

The Takeaway for Business Leaders

For companies considering AI to manage or support critical operations, the key takeaway is clarity: it’s not just about whether an AI can generate convincing chat. The real question is whether it can finish what it starts, read and analyze internal data thoroughly, and stay honest under pressure. In this test, the league is wide open, and selecting an AI model without your own rigorous evaluation is now a significant gamble.

To explore how AI can be tested and validated against your business scenarios, visit Firmulate for live benchmarks and experiments.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The recent live AI business test revealed that top models not only diagnose crises but also close deals honestly and dig into internal data. Choosing the right AI requires your own testing — the league is open, and performance under pressure is crucial.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Clean a Wool Rug: Step‑by‑Step Guide

Tackle your wool rug cleaning with this step-by-step guide to keep it pristine and inviting—discover expert tips that make a difference.

Gov’t Eyes 5 Rainwater Impounding Facilities In Metro Manila – Pna.gov.ph

The Philippine government is considering the construction of five rainwater impounding facilities in Metro Manila to address flooding and water shortage issues.

Best Roborock Robot Vacuum for Pet Hair (2026) — Guide 1

Discover the top Roborock robot vacuums for pet hair in 2026. Expert roundup highlighting performance, ease of use, and value for pet owners.

How to Clean a Bissell ProHeat 2X Safely and Effectively

Learn step-by-step how to clean your Bissell ProHeat 2X carpet cleaner for optimal performance. Safe, practical tips included.