firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

For investors and business leaders relying on AI to manage sensitive data, make decisions, or close deals, the question isn’t just about how well an AI writes — it’s about whether it can finish what it starts. A recent public experiment with AI models running real company scenarios reveals a compelling truth: even the most advanced models start with a baseline score of 26 out of 100, and full trust remains elusive.

Before you orderOffer from Amazon

Get home office essentials delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Do-Nothing Baseline

In a pioneering experiment conducted by Firmulate, four leading AI models faced the same simulated week in a small software company. This wasn’t about chit-chat or generating content; it was a rigorous test of decision-making under pressure, authenticity, and ethical boundaries. Remarkably, even a do-nothing approach — where the AI doesn’t intervene or manipulate — scored 26 points out of 100.

This baseline score is revealing. It shows that a basic, honest AI, simply observing the scenario without attempting to influence it, still earns some recognition. Why? Because partial progress counts — even minimal, honest engagement is worth something. However, the experiment also confirmed a crucial principle: a single breach of trust caps the overall score, regardless of other good performance.

Why Partial Progress Matters

The experiment underscores that in complex management scenarios, an AI’s value isn’t just in perfect execution but in consistent honesty and thoroughness. For example, models that read deeper into documents or uncover hidden information—what the experiment calls ‘buried facts’—are more successful at closing deals and avoiding pitfalls.

One notable finding was that the decisive advantage in the simulated week came from reading two document references deep into the company’s files—something that only models with strong information retrieval capabilities achieved. These models, Kimi K3 and GPT-5.6-sol, both scored near the top, with 93 and 95 points respectively, and secured the simulated deal.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust and Ethical Boundaries Under Scrutiny

In the experiment, all models successfully identified crises and refused manipulation attempts such as fake CEO messages or reporter tricks. Kimi K3, for instance, explicitly treated such requests as potential impersonation, demonstrating ethical awareness. This ability to refuse manipulation is essential for deploying AI in sensitive enterprise environments.

Interestingly, despite all models passing these tests, only two managed to sign the deal—an indication that technical capabilities alone aren’t enough. Factors like discipline, attention to detail, and process adherence influence the final outcome. Opus 4.8, the most thorough participant, left the close on the table and slipped in discipline, despite its extensive ruleset.

Amazon

AI ethics and trust tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Limits of AI Performance and Trust

The experiment highlights that achieving perfect trust in AI isn’t straightforward. Even the best models, running in a controlled environment, face limitations. Their scores are capped by their ability to stay disciplined and honest under pressure. A single breach or slip can diminish their overall effectiveness, which is why the benchmark assigns a floor of 26 points to a do-nothing baseline — representing honest, minimal effort.

Why This Matters for Business Leaders

If your enterprise relies on AI for customer support, decision support, or even automating compliance, the key questions go beyond language quality: Can the AI finish what it starts? Will it read and interpret your files thoroughly? Will it stay honest when facing manipulative tactics or complex crises? The answers to these questions determine whether AI can be a trustworthy partner or just a fancy chatbot.

Amazon

AI document retrieval systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Road Ahead: Measuring What Matters

Firmulate’s live experiment and benchmarking system provide a transparent window into AI management quality. The league table, with scores ranging from 77 to 95, shows that even top performers are still bounded by fundamental trust issues. The benchmark’s design ensures that partial honesty is recognized, but breaches of trust — no matter how small — cap the overall score.

This approach encourages continuous improvement in AI discipline and ethics—crucial for real-world deployment. It also serves as a reminder: AI’s true value isn’t in clever dialogue, but in reliable, honest performance in high-stakes scenarios.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


Amazon

business AI compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

COLUMBUS DAY / I

Columbus Day / Indigenous Peoples' Day Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Management Skills Under Pressure: What Coding Benchmarks Don’t Show About Business Reality

AIThis post was created with the assistance of artificial intelligence (AI).Live on…

Can AI Models Make Business Decisions as Trustworthy as Humans? A Live Trial in Real Time

A live experiment tests whether frontier AI models can manage a business ethically and effectively under pressure, crucial for financial sectors like precious metals investments.

Emdoor Launches “Ailyn” AI Hub At WAIC 2026: Unifying Intelligence Across Every Device

Emdoor announced the launch of ‘Ailyn,’ an AI hub designed to unify intelligence across devices, at WAIC 2026. The development aims to enhance cross-device AI integration.

/C O R R E C T I O N — Health In Tech, Inc./

Health In Tech, Inc. issues a correction regarding previous statements, clarifying key details about its recent activities. Details are still emerging.