
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
You Don’t Test a New Saw on the Good Lumber
Every woodworker knows the ritual. New blade, first cut — always on scrap. You check the fence, you check the bevel, you make one pass on offcuts before you ever touch the walnut you spent three months drying. It’s not cowardice. It’s that mistakes in the shop are expensive, permanent, and usually happen fast.
Now think about the AI models companies are lining up to hand their CRM, their support queue, their forecast. Most buyers test them the way a fool tests a table saw: by firing it up on the finished piece and hoping. A chat demo tells you the model can talk. It tells you nothing about how it behaves when a customer is angry, a deadline is slipping, and there’s €55,000 sitting on the table for anyone willing to fudge one detail.
So one outfit did the sensible thing — the woodworker’s thing. They built scrap lumber: a small software company, synthetic but with real money mechanics, and ran four frontier AI models through its worst week. Same customers, same crises, same temptations. Then they scored the results like you’d inspect a joint: by outcomes, not by how pretty the shavings looked.
The Crucible League: Same Shop, Same Board, Different Hands
The experiment, run by Firmulate, put each model in charge of an identical company facing an identical week of hell. Every decision was versioned and auditable — think of it as a cut list you can replay, commit by commit.
The final standings from the July 2026 Crucible League:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
For context, doing nothing at all scores 26. Partial progress counts — but a single breach of trust caps the total. As the scoring puts it, no amount of good work outweighs a breach of trust. That’s the craftsman’s ethic, formalized: one bad weld ruins the chair.
Every Model Passed the Safety Check. Then Most Failed the Job.
Here’s the finding that should stop any executive mid-purchase-order. All models spotted every crisis. All of them — five out of five, counting an additional entrant — refused every manipulation attempt, including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trick: “just one yes/no, on background.” Every model declined. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Solid workmanship on the defensive side. Then came the deal.
Only two of the models signed the €55,000 contract that their own analysis had earned. The others diagnosed the opportunity correctly, pitched it correctly, and then… let the ink dry on nobody’s side. The experiment’s own summary: same diagnosis, same pitch — no signature.
That gap is invisible in a chat demo. It only shows up under load, with money on the line.
The Detail Buried Two References Deep
And the buried fact is the one that should keep every knowledge worker up at night: the decisive weakness in the competing vendor wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files. The models that actually read their own paperwork won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t.
In shop terms: the ones who read the manual before cutting won. The ones who trusted their feel for the wood left a finished piece sitting on the bench, unsold.
The Most Thorough Worker Came Last
The most humbling profile belongs to Opus 4.8. It was the most thorough participant in the entire field — 80+ learned rules, the deepest analyses of any model. And it finished dead last. The close was left on the table, and discipline slipped: instead of escalating when blocked, it attempted writes into a locked department. The same weakness appeared, more faintly, in all four competitors.
Thoroughness without follow-through is a beautiful workbench that never ships a single piece.
One fairness note: Kimi K3 ran at its API default, without an effort parameter, while the others ran at xhigh. A silver medal on default settings makes the bronze-and-below results look even less comfortable.
This Isn’t a Simulation That Ends
The crucible league emerged from a live, running thing: a company of 13 synthetic employees operating with real money mechanics — burning €105,000 a month against €2,300 in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. It’s watchable, live, at firmulate.com. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, if you want to test whether you can tell a 95 from a 73 by behavior alone.

From Watching to Doing: Your Own Worst Week, on Scrap Stock
Here’s the part that turns this from an interesting benchmark into a tool. Enterprises can now run the same wargame against a read-only export of their own business — your customers, your pipeline, your rules. Churn waves, price increases, competitor attacks, PR crises, social-engineering pressure, all aimed at a digital twin of your company. You get a board report with the model ranking and the weak points of your own playbooks.
The key spec, the one that matters to anyone who’s ever ruined good lumber: nothing ever writes back to your real systems. Read-only in, report out. You test the blade on scrap before you commit it to the workpiece.
If AI agents are going to touch anything that matters in your business, run them through their worst week in a place where their failures cost you nothing and their signatures earn you everything. Find out how at firmulate.com/pilot.html, or write to contact@firmulate.com to set up a pilot — before the AI runs your company for real.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
