📊 Full opportunity report: Using Management Tests To Decode AI’s Work Behavior on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
A live experiment evaluates AI models on managing a simulated company crisis, uncovering how different models perform in decision-making, trust, and action. Results show significant variation in operational discipline and trustworthiness, impacting AI management potential.
Researchers have launched a live, public experiment testing five AI management models on a simulated company crisis, revealing how these models handle real-world business decisions, trust, and execution. This approach is discussed in the original analysis. This testing approach provides concrete insights into the operational capabilities of AI in management roles, beyond theoretical or benchmark assessments.
The experiment, hosted on firmulate.com, involved five AI models running a small software company through its worst week, with identical crises, customer issues, and internal dilemmas. This testing methodology aligns with the insights from the original analysis. Each model’s decisions were recorded, auditable, and scored based on their effectiveness, diligence, and ability to complete critical actions. The models included gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8. The results, published in July 2026, ranked gpt-5.6-sol first with 95 points and Opus 4.8 last with 73, illustrating notable differences in operational discipline and trustworthiness. For more on AI management testing, see AI’s Management Deficit Revealed Only After Correct Outcomes.
Despite all models recognizing crises and refusing manipulation attempts, only two successfully closed a crucial €55,000 deal, highlighting that analysis alone does not guarantee action. The experiment underscores that effective management requires both understanding and execution, with some models excelling in analysis but failing to follow through, while others demonstrated disciplined action. The models’ performance in security and trust-related tests was consistent, with all models correctly refusing manipulative requests, indicating strong risk recognition.
Implications for AI Management and Business Automation
This experiment demonstrates that AI models differ significantly in their ability to translate analysis into action, a vital consideration for enterprises deploying AI in operational roles. The findings suggest that evaluating AI solely on analytical accuracy is insufficient; operational discipline, trustworthiness, and follow-through are equally critical. For businesses, these results highlight the importance of testing AI models in realistic, high-pressure scenarios before granting them decision-making authority, reducing the risk of failures that could undermine trust or cause financial loss.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on AI Management Testing and Industry Relevance
Traditional AI assessments focus on benchmarks measuring language understanding or problem-solving skills, often in isolated tasks. However, applying AI to management and operational roles requires assessing how models perform in dynamic, multi-faceted situations involving decision-making, trust, and follow-up actions. Recent developments in AI management experiments, such as the Firmulate live tests, aim to fill this gap by simulating real-world crises and measuring models’ operational behaviors. These efforts respond to industry concerns about AI’s readiness for autonomous decision-making in complex environments, emphasizing the need for rigorous, scenario-based testing.
“This testing approach reveals that AI’s ability to analyze is not enough; effective management depends on follow-through and trustworthiness.”
— Live experiment organizer
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Management Performance
It remains unclear how these models will perform in longer-term or more complex real-world scenarios outside the controlled experiment. The impact of different operational settings, integration with human teams, and evolving AI capabilities are still being studied. Additionally, questions about how to best train or fine-tune models for operational discipline are ongoing, and whether these results generalize across industries or more diverse business contexts remains to be seen.
As an affiliate, we earn on qualifying purchases.
Future Directions for AI Management Testing and Deployment
Researchers plan to expand these live management tests to include more models, longer scenarios, and varied business environments. Companies are encouraged to adopt similar scenario-based evaluations before deploying AI in critical decision-making roles. Further developments may include integrating these tests into AI development pipelines, refining metrics for operational discipline, and establishing industry standards for AI management readiness. The goal is to create a robust framework for certifying AI models’ suitability for autonomous management tasks.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does this experiment reveal about current AI models?
The experiment shows that while AI models can recognize crises and avoid manipulation, their ability to follow through with decisive actions varies significantly. Some models excel in analysis but fail to complete critical operational tasks, highlighting the importance of operational discipline.
Why is following through important for AI in management?
Effective management requires not only understanding problems but also executing necessary actions. An AI that analyzes well but does not act risks undermining trust and failing to deliver business results.
Can these testing methods be applied to other AI systems?
Yes, scenario-based, real-world testing can be adapted for different AI applications, especially those involving operational decision-making, customer interactions, or security-critical tasks.
What are the limitations of this experiment?
The current tests are limited to simulated crises within controlled environments. The models’ performance in ongoing, real-world operational settings remains to be validated.
What should companies do before deploying AI for management tasks?
Companies should conduct scenario-based tests similar to this experiment to evaluate AI models’ ability to follow through with decisions, maintain trust, and handle pressure before full deployment.
Source: ThorstenMeyerAI.com