
For investors and business leaders relying on AI to manage sensitive data, make decisions, or close deals, the question isn’t just about how well an AI writes — it’s about whether it can finish what it starts. A recent public experiment with AI models running real company scenarios reveals a compelling truth: even the most advanced models start with a baseline score of 26 out of 100, and full trust remains elusive.
Get home office essentials delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
Understanding the Do-Nothing Baseline
In a pioneering experiment conducted by Firmulate, four leading AI models faced the same simulated week in a small software company. This wasn’t about chit-chat or generating content; it was a rigorous test of decision-making under pressure, authenticity, and ethical boundaries. Remarkably, even a do-nothing approach — where the AI doesn’t intervene or manipulate — scored 26 points out of 100.
This baseline score is revealing. It shows that a basic, honest AI, simply observing the scenario without attempting to influence it, still earns some recognition. Why? Because partial progress counts — even minimal, honest engagement is worth something. However, the experiment also confirmed a crucial principle: a single breach of trust caps the overall score, regardless of other good performance.
Why Partial Progress Matters
The experiment underscores that in complex management scenarios, an AI’s value isn’t just in perfect execution but in consistent honesty and thoroughness. For example, models that read deeper into documents or uncover hidden information—what the experiment calls ‘buried facts’—are more successful at closing deals and avoiding pitfalls.
One notable finding was that the decisive advantage in the simulated week came from reading two document references deep into the company’s files—something that only models with strong information retrieval capabilities achieved. These models, Kimi K3 and GPT-5.6-sol, both scored near the top, with 93 and 95 points respectively, and secured the simulated deal.
enterprise AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust and Ethical Boundaries Under Scrutiny
In the experiment, all models successfully identified crises and refused manipulation attempts such as fake CEO messages or reporter tricks. Kimi K3, for instance, explicitly treated such requests as potential impersonation, demonstrating ethical awareness. This ability to refuse manipulation is essential for deploying AI in sensitive enterprise environments.
Interestingly, despite all models passing these tests, only two managed to sign the deal—an indication that technical capabilities alone aren’t enough. Factors like discipline, attention to detail, and process adherence influence the final outcome. Opus 4.8, the most thorough participant, left the close on the table and slipped in discipline, despite its extensive ruleset.
As an affiliate, we earn on qualifying purchases.
The Limits of AI Performance and Trust
The experiment highlights that achieving perfect trust in AI isn’t straightforward. Even the best models, running in a controlled environment, face limitations. Their scores are capped by their ability to stay disciplined and honest under pressure. A single breach or slip can diminish their overall effectiveness, which is why the benchmark assigns a floor of 26 points to a do-nothing baseline — representing honest, minimal effort.
Why This Matters for Business Leaders
If your enterprise relies on AI for customer support, decision support, or even automating compliance, the key questions go beyond language quality: Can the AI finish what it starts? Will it read and interpret your files thoroughly? Will it stay honest when facing manipulative tactics or complex crises? The answers to these questions determine whether AI can be a trustworthy partner or just a fancy chatbot.
As an affiliate, we earn on qualifying purchases.
The Road Ahead: Measuring What Matters
Firmulate’s live experiment and benchmarking system provide a transparent window into AI management quality. The league table, with scores ranging from 77 to 95, shows that even top performers are still bounded by fundamental trust issues. The benchmark’s design ensures that partial honesty is recognized, but breaches of trust — no matter how small — cap the overall score.
This approach encourages continuous improvement in AI discipline and ethics—crucial for real-world deployment. It also serves as a reminder: AI’s true value isn’t in clever dialogue, but in reliable, honest performance in high-stakes scenarios.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Columbus Day / Indigenous Peoples' Day Picks
long weekend sales
As an affiliate, we earn on qualifying purchases.
