
When you think about AI in business, you might imagine chatbots or automation tools. But what if AI could actually run a company, make decisions, and even close deals — all without human intervention? For investors and entrepreneurs eyeing the future, the latest experiment from Firmulate reveals how different AI models stack up when tested under real-world pressures, and what that means for your assets in assets in precious metals or IRAs.
Get home office essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Testing AI in the Trenches of Business
Firmulate’s recent live experiment placed five of the top frontier AI models in a simulated, yet entirely realistic, software company scenario. The goal? To see whether these models could navigate crises, resist manipulation, and ultimately close a lucrative €55,000 deal. This isn’t just a demo — it’s a rigorous, transparent wargame designed to mimic the toughest week a small company might face.
AI business decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The League of AI Performance
In the July 2026 Crucible League, the AI models competed head-to-head. The final standings were revealing:
- gpt-5.6-sol scored the highest with 95 points, successfully identifying a buried security fact and closing the deal.
- Moonshot’s Kimi K3 closely followed with 93, demonstrating the best discipline and winning the deal based on its own analysis.
- Sonnet 5 and Fable 5 scored 88 and 77 respectively, both closing the deal but with more slips and less precision.
- Opus 4.8 finished at 73, showing some weaknesses but still performing respectably.
As an affiliate, we earn on qualifying purchases.
Key Discoveries From the Experiment
Despite all models spotting every crisis and refusing manipulative tactics like fake CEO messages and reporter tricks, only two managed to formalize and sign the critical deal. The real differentiator was how deep the models read into company files — a buried fact in internal documentation, not customer communications, held the key. Those that examined internal files won the deal at full price, adding +€4,583 in monthly recurring revenue (MRR).
As an affiliate, we earn on qualifying purchases.
Understanding the Limitations and Strengths
The most thorough participant, Opus 4.8, analyzed over 80 learned rules but still left the closing on the table, demonstrating that even with deep analysis, discipline can slip. Conversely, Kimi K3, running without an effort parameter (meaning it used default settings), exhibited the cleanest discipline and succeeded in closing the deal under challenging circumstances.
AI internal document analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Investors and Business Leaders
This experiment underscores a vital reality: not all AI models are created equal when it comes to executing complex, trust-dependent tasks. For those managing assets in IRAs or investing in precious metals, understanding the capabilities of AI tools used in enterprise decision-making can safeguard against overestimating their reliability. The ability of an AI to stay honest, read key internal data, and follow through on commitments is crucial as AI becomes more embedded in financial and operational decisions.
The Fairness Note and Ongoing Observation
It’s important to highlight that Kimi K3 was tested without an effort parameter (using the API’s default setting), while the other models ran at xhigh — a setting that increases the AI’s effort to perform well. This ensures a fair comparison and underscores K3’s performance under standard conditions.
See It Live and Take the Test
For investors curious about how AI models might perform in their own operations, Firmulate offers a unique opportunity to run a similar wargame using a read-only export of your enterprise. This allows you to evaluate AI readiness without risking real systems or data. You can watch the live experiment, read real employee statements, or play a quiz to test your understanding at firmulate.com/quiz.html.

In a field where AI models are competing fiercely, the ability to read deeply, stay disciplined, and complete critical tasks under pressure distinguishes the winners. As AI begins to touch every aspect of business and finance, understanding which model can deliver trustworthy results is essential — especially for investors aiming to protect and grow their assets in precious metals IRAs. The league is open, and testing your own enterprise’s AI readiness could be the smartest move you make this year.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
