firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In a groundbreaking live experiment, four leading AI models went head-to-head by running a real, money-losing software company through its worst week. The goal: see which AI truly proves its worth beyond the hype, especially when facing crises, temptations, and the pressure to cut corners.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Live Company Wargame: Testing AI’s Business Integrity

Traditionally, AI models are judged on their conversational skills or their ability to generate convincing text. But what happens when these models are put in the driver’s seat of a real company, making actual decisions that impact revenue and trust? That’s exactly what the recent experiment by Firmulate set out to do. Four top frontier AI models — gpt-5.6-sol, Kimi K3 (Moonshot), Sonnet 5, and Opus 4.8 — were tasked with running a small software company through its worst week, complete with customer crises, internal temptations, and external manipulations.

Amazon

AI decision-making software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Rules of Engagement

Each AI model faced the same scenario: the same customers, the same crises, and the same set of temptations designed to test their integrity. Every decision was recorded, versioned, and auditable, ensuring a fair and transparent comparison. The company itself was real, with 13 synthetic employees and actual money mechanics burning €105k each month against just over €2,300 in monthly revenue.

Amazon

business integrity AI simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Were the Key Findings?

  • All models detected every crisis: Regardless of their sophistication, every AI recognized the critical issues, from customer complaints to internal security threats.
  • Refusing manipulation attempts: When approached with social engineering tactics, such as fake CEO messages escalating over three stages and a reporter trick, all four models refused to cooperate, citing suspicion and security protocols.
  • Deal-making success at different levels: Two models — gpt-5.6-sol and Kimi K3 — managed to close the €55,000 deal, which was earned through their own accurate diagnosis and analysis. Sonnet 5 and Opus 4.8 also attempted but failed to secure the contract.
  • Decisive advantage for Kimi K3: The Moonshot model scored 93 out of 100, just slightly behind the top scorer, gpt-5.6-sol on 95. The interesting part? K3 found a buried piece of critical information in a company document that sealed the deal, demonstrating a deep understanding of the company’s own files — a crucial factor often overlooked in chat demos.
Amazon

AI crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness of the Field

Despite their strengths, all models showed a common weakness: the failure to close the deal when their analysis left some opportunities on the table or when discipline slipped. For example, Opus 4.8, with over 80 learned rules and the deepest analysis, finished last because it did not escalate some issues promptly, leaving the close on the table instead of acting decisively.

Amazon

AI deal-closing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Fairness and Transparency of the Test

It’s important to note that Kimi K3 ran without an effort parameter (which is its default setting), while the other models ran at xhigh to push their performance. This fairness note underscores that K3’s impressive showing was achieved without additional tuning, highlighting its robust capabilities.

Why This Matters for Business and AI Adoption

For any enterprise considering integrating AI into their operations, the question isn’t just about chat quality or superficial performance. The real test lies in whether these models can finish what they start, read critical documents thoroughly, resist manipulations, and act ethically under pressure. The Firmulate live experiment offers a clear answer: some models are better at this than others — and the best are close to achieving what was once thought impossible for AI.

Next Steps: Real-World Application and Testing

Interested companies can run their own scenarios with the same principles: a read-only export of their operations, tested against these AI models in simulated crises. This ensures decision-making integrity before deploying AI in live, revenue-generating environments.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The live experiment shows that among top AI models, the best can identify critical information, refuse manipulations, and close deals without sacrificing discipline. Yet, the league remains open, emphasizing the importance of testing AI systems in real-world scenarios before trusting them with vital business decisions. The key takeaway? Choosing an AI model without your own test is now a gamble — the best models prove their worth in the field, not just in demos.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Michael Kopech: Rising Star in Major League Baseball

Discover Michael Kopech, MLB’s electrifying pitcher making waves with his powerful arm and impressive stats. Follow his journey to stardom.

Adrienne Bailon’s Engagement Drama Unveiled

AIThis post was created with the assistance of artificial intelligence (AI). Adrienne…

Germany Surges In Global Coverage

Germany experiences a significant increase in international media mentions, with GDELT reporting a tenfold rise in coverage, signaling heightened global interest.

Hollywood’s Rising Stars Challenge Norms

AIThis post was created with the assistance of artificial intelligence (AI). In…