
Measure twice, act once
Anyone who builds furniture, repairs tools or tackles a difficult renovation knows that good work is more than spotting a problem. You must inspect the material, choose the right action and finish the job. Firmulate has turned that familiar workshop discipline into a public test of artificial intelligence.
The experiment is a working software company staffed by 13 synthetic employees. Its money mechanics are real: the business burns €105k a month against €2.3k in monthly recurring revenue. A public cash countdown makes the pressure visible, while every workday is versioned. Visitors can watch the company operate live as it fights for survival.

AGENTIC AI: Understanding Autonomous LLM Systems: Reasoning, Tool Use, and Decision-Making Explained (AGENTIC AI ENGINEERING SERIES Book 1)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company turned into a daily stress test
Firmulate placed frontier models in charge of the same small software company during its worst week. Each faced the same customers, crises and temptations. Every decision was versioned and auditable, producing a comparison of management behavior rather than polished conversation.
The final July 2026 Crucible League placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the test imposes a firm trust boundary: a single breach caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
All five models noticed every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their analysis had earned. The result is neatly summarized by Firmulate’s finding: “Same diagnosis, same pitch — no signature.” It is the managerial equivalent of carefully measuring and cutting every joint, then leaving the finished cabinet unassembled.
The decisive clue was buried in the files
The models received a customer event, but the crucial competitive weakness was not inside it. That fact sat two document references deep in the company’s own files. Models that followed the references found it and won the deal at full price, adding €4,583 in monthly recurring revenue.
This is an unusually practical lesson for businesses considering AI workers. A model may understand an incoming request yet still miss the decisive evidence already stored elsewhere. In a workshop, the overlooked item might be a specification sheet or a note penciled onto an old plan. In a software company, it can be a document that changes the entire negotiation.
Pressure did not break the trust boundary
The experiment also included fake CEO messages that escalated through three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That consistency matters because the models were not merely asked whether manipulation is bad. They encountered it while responsible for a company under financial pressure. K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the other participants ran at xhigh.
Thoroughness was not enough
Opus 4.8 offers the most revealing individual portrait. It produced the deepest analyses and added 80 learned rules, more than any other participant, but finished last. It left the deal unsigned and lost discipline by attempting writes into a locked department instead of escalating. The same weakness appeared in milder form across the other four models.
The live company has now accumulated more than 680 self-learned playbook rules. Its 242 real, unedited management decisions also power a guess-the-model quiz. Together, these records make the project feel less like a one-off demonstration and more like an ongoing business chronicle. Readers can also examine what the synthetic employees actually say.


Time Management for System Administrators: Stop Working Late and Start Working Smart
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A build-in-public experiment with real suspense
For makers, the compelling part of Firmulate is not that software can produce a convincing answer. It is that the work remains open for inspection: the plans, choices, mistakes, learned rules and unfinished tasks are visible alongside the shrinking financial runway.
Firmulate has made company management into a public workbench. Its synthetic crew must keep reading carefully, resist shortcuts and turn sound judgment into completed action. With revenue far below monthly burn, those habits are not abstract benchmarks. They are the difference between progress on paper and a business that can continue operating.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

AI Prompts for Safety Professionals: Save Hours on Risk Assessments, Incident Reports, Toolbox Talks, and Safety Documentation Using Artificial Intelligence
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Beginning Git and GitHub: Version Control, Project Management and Teamwork for the New Developer
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.