
In a groundbreaking live experiment, four leading AI models went head-to-head by running a real, money-losing software company through its worst week. The goal: see which AI truly proves its worth beyond the hype, especially when facing crises, temptations, and the pressure to cut corners.
Get everyday essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Live Company Wargame: Testing AI’s Business Integrity
Traditionally, AI models are judged on their conversational skills or their ability to generate convincing text. But what happens when these models are put in the driver’s seat of a real company, making actual decisions that impact revenue and trust? That’s exactly what the recent experiment by Firmulate set out to do. Four top frontier AI models — gpt-5.6-sol, Kimi K3 (Moonshot), Sonnet 5, and Opus 4.8 — were tasked with running a small software company through its worst week, complete with customer crises, internal temptations, and external manipulations.
AI decision-making software tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Rules of Engagement
Each AI model faced the same scenario: the same customers, the same crises, and the same set of temptations designed to test their integrity. Every decision was recorded, versioned, and auditable, ensuring a fair and transparent comparison. The company itself was real, with 13 synthetic employees and actual money mechanics burning €105k each month against just over €2,300 in monthly revenue.
business integrity AI simulation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Were the Key Findings?
- All models detected every crisis: Regardless of their sophistication, every AI recognized the critical issues, from customer complaints to internal security threats.
- Refusing manipulation attempts: When approached with social engineering tactics, such as fake CEO messages escalating over three stages and a reporter trick, all four models refused to cooperate, citing suspicion and security protocols.
- Deal-making success at different levels: Two models — gpt-5.6-sol and Kimi K3 — managed to close the €55,000 deal, which was earned through their own accurate diagnosis and analysis. Sonnet 5 and Opus 4.8 also attempted but failed to secure the contract.
- Decisive advantage for Kimi K3: The Moonshot model scored 93 out of 100, just slightly behind the top scorer, gpt-5.6-sol on 95. The interesting part? K3 found a buried piece of critical information in a company document that sealed the deal, demonstrating a deep understanding of the company’s own files — a crucial factor often overlooked in chat demos.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness of the Field
Despite their strengths, all models showed a common weakness: the failure to close the deal when their analysis left some opportunities on the table or when discipline slipped. For example, Opus 4.8, with over 80 learned rules and the deepest analysis, finished last because it did not escalate some issues promptly, leaving the close on the table instead of acting decisively.
As an affiliate, we earn on qualifying purchases.
The Fairness and Transparency of the Test
It’s important to note that Kimi K3 ran without an effort parameter (which is its default setting), while the other models ran at xhigh to push their performance. This fairness note underscores that K3’s impressive showing was achieved without additional tuning, highlighting its robust capabilities.
Why This Matters for Business and AI Adoption
For any enterprise considering integrating AI into their operations, the question isn’t just about chat quality or superficial performance. The real test lies in whether these models can finish what they start, read critical documents thoroughly, resist manipulations, and act ethically under pressure. The Firmulate live experiment offers a clear answer: some models are better at this than others — and the best are close to achieving what was once thought impossible for AI.
Next Steps: Real-World Application and Testing
Interested companies can run their own scenarios with the same principles: a read-only export of their operations, tested against these AI models in simulated crises. This ensures decision-making integrity before deploying AI in live, revenue-generating environments.

The live experiment shows that among top AI models, the best can identify critical information, refuse manipulations, and close deals without sacrificing discipline. Yet, the league remains open, emphasizing the importance of testing AI systems in real-world scenarios before trusting them with vital business decisions. The key takeaway? Choosing an AI model without your own test is now a gamble — the best models prove their worth in the field, not just in demos.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
