
Get everyday essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
How Do You Measure an AI’s Business Sense? The Surprising Score of a Do-Nothing Baseline
In the world of AI, especially when it comes to managing real businesses, trust and decision-making are everything. Yet, the simplest baseline—doing almost nothing—scores a surprising 26 points out of 100 in a recent unfiltered benchmark, revealing a lot about how we evaluate AI performance today.
As an affiliate, we earn on qualifying purchases.
Understanding the Benchmark: More Than Just Words
Firmulate’s live experiment pits four frontier AI models against the challenge of managing a small software company through its worst week. Each model is tasked with handling crises, customer interactions, and manipulations, all under the same conditions. The goal? See whether these models can complete real management work, not just generate convincing chat responses.
Remarkably, all four models identified every crisis and refused every manipulation attempt. This demonstrates a fundamental level of honesty and discipline that’s critical when AI systems are involved in sensitive business processes.
The Curious Case of the Do-Nothing Baseline
What’s striking isn’t just their honesty—it’s the baseline score. A do-nothing approach, which essentially makes no effort and simply passes on decisions without analysis, scores 26 points. This score reflects that even minimal engagement—such as reading some documents—can generate partial progress. Importantly, a single breach of trust, like signing a fraudulent deal, caps the total score at this level, regardless of other successes. This design ensures the evaluation emphasizes honesty and trustworthiness over superficial performance.
AI business decision management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Deep Roots of Success: Reading the Right Files
The real differentiator among models wasn’t just surface-level decision-making. Instead, success came from reading deeply into company documents. The strongest model, gpt-5.6-sol, found critical information buried two document references deep and closed a lucrative deal worth over €4,583 monthly recurring revenue. Models that skimmed or missed these references couldn’t close the deal at full price, illustrating that thorough information processing matters more than flashy responses.
Trust in Action: Refusing Manipulation
All models demonstrated integrity by refusing social engineering tricks—a staged escalation involving fake CEO messages and reporter inquiries. Each refused to sign off on fraudulent requests, with Kimi K3 explicitly treating such requests as potential impersonation. This disciplined response is vital for AI systems in real-world applications, where trust breaches can be costly.
AI cybersecurity and fraud prevention software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Limitations and Lessons from the Experiment
The most comprehensive participant, Opus 4.8, showed something telling—despite its depth of analysis, it slipped on discipline during closing negotiations, leaving money on the table. This highlights an important point: even advanced AI models can falter in consistency and discipline, especially under pressure.
Another interesting note is that the models’ performance varied based on their configuration. K3, which ran without the default effort parameter, performed at a high level, closing deals without extra effort, whereas others with higher effort settings faced more challenges. The key takeaway? Effort levels influence performance, but honesty and thoroughness are fundamental regardless.
AI enterprise trust and integrity solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Implication: Trust Over Fluency
This experiment underscores a critical insight for businesses considering AI: it’s not just about how well an AI writes or chats, but whether it can follow through, access relevant information, and maintain honesty under pressure. These qualities are essential for AI systems integrated into customer relationship management, support, or forecasting roles.
For enterprise leaders, the message is clear: a high score doesn’t mean an AI is trustworthy or effective in real business contexts. The true measure lies in its ability to complete real work, read deeply, and stay honest—features that current benchmarks only begin to capture.

Key Takeaways
- The baseline score for doing almost nothing is 26, emphasizing that partial progress counts in AI evaluation.
- A single breach of trust caps the total score, encouraging honesty over superficial success.
- Deep reading and thorough information processing are more important than flashy responses for closing deals.
- Trustworthiness and discipline under pressure are critical qualities for AI in business, not just language fluency.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
