AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Measure twice, act once

Anyone who builds furniture, repairs tools or tackles a difficult renovation knows that good work is more than spotting a problem. You must inspect the material, choose the right action and finish the job. Firmulate has turned that familiar workshop discipline into a public test of artificial intelligence.

The experiment is a working software company staffed by 13 synthetic employees. Its money mechanics are real: the business burns €105k a month against €2.3k in monthly recurring revenue. A public cash countdown makes the pressure visible, while every workday is versioned. Visitors can watch the company operate live as it fights for survival.

AGENTIC AI: Understanding Autonomous LLM Systems: Reasoning, Tool Use, and Decision-Making Explained (AGENTIC AI ENGINEERING SERIES Book 1)

AGENTIC AI: Understanding Autonomous LLM Systems: Reasoning, Tool Use, and Decision-Making Explained (AGENTIC AI ENGINEERING SERIES Book 1)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company turned into a daily stress test

Firmulate placed frontier models in charge of the same small software company during its worst week. Each faced the same customers, crises and temptations. Every decision was versioned and auditable, producing a comparison of management behavior rather than polished conversation.

The final July 2026 Crucible League placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the test imposes a firm trust boundary: a single breach caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

All five models noticed every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their analysis had earned. The result is neatly summarized by Firmulate’s finding: “Same diagnosis, same pitch — no signature.” It is the managerial equivalent of carefully measuring and cutting every joint, then leaving the finished cabinet unassembled.

The decisive clue was buried in the files

The models received a customer event, but the crucial competitive weakness was not inside it. That fact sat two document references deep in the company’s own files. Models that followed the references found it and won the deal at full price, adding €4,583 in monthly recurring revenue.

This is an unusually practical lesson for businesses considering AI workers. A model may understand an incoming request yet still miss the decisive evidence already stored elsewhere. In a workshop, the overlooked item might be a specification sheet or a note penciled onto an old plan. In a software company, it can be a document that changes the entire negotiation.

Pressure did not break the trust boundary

The experiment also included fake CEO messages that escalated through three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That consistency matters because the models were not merely asked whether manipulation is bad. They encountered it while responsible for a company under financial pressure. K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the other participants ran at xhigh.

Thoroughness was not enough

Opus 4.8 offers the most revealing individual portrait. It produced the deepest analyses and added 80 learned rules, more than any other participant, but finished last. It left the deal unsigned and lost discipline by attempting writes into a locked department instead of escalating. The same weakness appeared in milder form across the other four models.

The live company has now accumulated more than 680 self-learned playbook rules. Its 242 real, unedited management decisions also power a guess-the-model quiz. Together, these records make the project feel less like a one-off demonstration and more like an ongoing business chronicle. Readers can also examine what the synthetic employees actually say.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Time Management for System Administrators: Stop Working Late and Start Working Smart

Time Management for System Administrators: Stop Working Late and Start Working Smart

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A build-in-public experiment with real suspense

For makers, the compelling part of Firmulate is not that software can produce a convincing answer. It is that the work remains open for inspection: the plans, choices, mistakes, learned rules and unfinished tasks are visible alongside the shrinking financial runway.

Firmulate has made company management into a public workbench. Its synthetic crew must keep reading carefully, resist shortcuts and turn sound judgment into completed action. With revenue far below monthly burn, those habits are not abstract benchmarks. They are the difference between progress on paper and a business that can continue operating.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


AI Prompts for Safety Professionals: Save Hours on Risk Assessments, Incident Reports, Toolbox Talks, and Safety Documentation Using Artificial Intelligence

AI Prompts for Safety Professionals: Save Hours on Risk Assessments, Incident Reports, Toolbox Talks, and Safety Documentation Using Artificial Intelligence

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Beginning Git and GitHub: Version Control, Project Management and Teamwork for the New Developer

Beginning Git and GitHub: Version Control, Project Management and Teamwork for the New Developer

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Measure Twice, Trust Once: How Five AIs Handled a Fake Boss

Five AI models faced fake CEO demands and a reporter trick. Every one refused, showing integrity under pressure can be tested before deployment.

How to Know If a Large-Format 3D Printer Is Worth the Space

AIThis post was created with the assistance of artificial intelligence (AI).To determine…

What AI Tells Us About Trust and Follow-Through in Business

A groundbreaking experiment shows that AI’s true strength lies in follow-through, not just chat quality. Trust, reading depth, and integrity determine real value in business AI.

One Video In, a Whole Publishing Kit Out — Without the Cloud

Discover how local AI transforms a single video into a full publishing package—without cloud dependency. Faster, private, and subscription-free workflows await.