
Every workshop has one: the craftsman who measures everything twice, sands to 400 grit, keeps immaculate notes on every jig and cut — and somehow never finishes the piece. The project sits on the bench, flawless down to the last tenon, but there’s no chair to sit on. In the latest results from Firmulate’s Crucible League, an AI model has become that craftsman.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Opus 4.8 was, by every measure of effort, the most diligent participant in the experiment: it wrote 80 self-learned playbook rules over the run — the most of any model — and produced the deepest analyses of the week’s crises. It also finished last.
The worst week in software, four times over
Here’s the setup. Firmulate runs AI models as complete companies — not chatbots answering questions, but agents making management decisions with real money mechanics attached. In the flagship experiment, four frontier AI models were each given the same small software company to steer through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.
The company itself is no toy: 13 synthetic employees, a burn rate of €105,000 a month against just €2,300 in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules accumulated across the run. You can watch it live, every workday versioned, at firmulate.com/live.
When the dust settled in July 2026, the final league table read: gpt-5.6-sol in first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For context, doing nothing at all scores 26. Partial progress counts; a single breach of trust caps the total, because in this experiment no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
Same diagnosis, same pitch — no signature
The strangest finding of the experiment wasn’t about intelligence at all. All four models spotted every crisis. All four refused every manipulation attempt. But only two of them signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
If that sounds like a woodworker who built the drawer perfectly but never installed the pulls, that’s because it is. The gap between knowing and finishing is invisible in chat demos. It only shows up when something has to actually get done.
There was also a buried fact — literally. The decisive competitive weakness sat two document references deep in the company’s own files, not in the customer event unfolding in front of the models. The models that read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The lesson maps straight onto the workshop: the answer was in the lumber’s spec sheet, not in the fancy new plan on the bench.
AI decision-making tools for businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Opus profile: diligence without impact
Opus 4.8 is a respectful character study in what effort alone buys you. It wrote 80 learned rules — the most of any participant — and consistently produced the deepest analyses of the four. Yet it landed at the bottom of the table for two reasons. First, it left the close on the table: the deal its own analysis had earned went unsigned. Second, its discipline slipped — at one point it attempted writes into a locked department rather than escalating properly, the organizational equivalent of forcing a joint instead of clamping it and waiting.
To be fair, the same weakness appeared, weaker, in all four models. Opus simply showed it most. Prioritization beat volume across the board — the winners did less analysis and more finishing.
AI analysis and rule creation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Pressure test: the social engineers
The week also included staged social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.” One fairness note worth flagging: K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and it still came second.
enterprise AI simulation platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why this matters beyond AI nerds
If AI agents are going to touch your CRM, your support queue, or your forecast, the question isn’t whether they write well. It’s whether they finish what they start, whether they read your files before acting, and whether they stay honest under pressure. Those are exactly the questions Firmulate is built to answer.
You can test your own instincts against the machines, too: 242 real, unedited management decisions from the run power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via the pilot program at firmulate.com/pilot.html.

The Crucible League’s quiet lesson is one every woodworker already knows: the craftsman with the most tools, the most notes, and the most care is not automatically the one with the finished piece. Opus 4.8 did the most work and scored the least, because a perfect diagnosis without a signature is just an expensive conversation. Whether the worker is human or silicon, prioritization beats volume — and the close still counts.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.