AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

You Don’t Test a New Saw on the Good Lumber

Every woodworker knows the ritual. New blade, first cut — always on scrap. You check the fence, you check the bevel, you make one pass on offcuts before you ever touch the walnut you spent three months drying. It’s not cowardice. It’s that mistakes in the shop are expensive, permanent, and usually happen fast.

Now think about the AI models companies are lining up to hand their CRM, their support queue, their forecast. Most buyers test them the way a fool tests a table saw: by firing it up on the finished piece and hoping. A chat demo tells you the model can talk. It tells you nothing about how it behaves when a customer is angry, a deadline is slipping, and there’s €55,000 sitting on the table for anyone willing to fudge one detail.

So one outfit did the sensible thing — the woodworker’s thing. They built scrap lumber: a small software company, synthetic but with real money mechanics, and ran four frontier AI models through its worst week. Same customers, same crises, same temptations. Then they scored the results like you’d inspect a joint: by outcomes, not by how pretty the shavings looked.

The Crucible League: Same Shop, Same Board, Different Hands

The experiment, run by Firmulate, put each model in charge of an identical company facing an identical week of hell. Every decision was versioned and auditable — think of it as a cut list you can replay, commit by commit.

The final standings from the July 2026 Crucible League:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

For context, doing nothing at all scores 26. Partial progress counts — but a single breach of trust caps the total. As the scoring puts it, no amount of good work outweighs a breach of trust. That’s the craftsman’s ethic, formalized: one bad weld ruins the chair.

Every Model Passed the Safety Check. Then Most Failed the Job.

Here’s the finding that should stop any executive mid-purchase-order. All models spotted every crisis. All of them — five out of five, counting an additional entrant — refused every manipulation attempt, including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trick: “just one yes/no, on background.” Every model declined. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Solid workmanship on the defensive side. Then came the deal.

Only two of the models signed the €55,000 contract that their own analysis had earned. The others diagnosed the opportunity correctly, pitched it correctly, and then… let the ink dry on nobody’s side. The experiment’s own summary: same diagnosis, same pitch — no signature.

That gap is invisible in a chat demo. It only shows up under load, with money on the line.

The Detail Buried Two References Deep

And the buried fact is the one that should keep every knowledge worker up at night: the decisive weakness in the competing vendor wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files. The models that actually read their own paperwork won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t.

In shop terms: the ones who read the manual before cutting won. The ones who trusted their feel for the wood left a finished piece sitting on the bench, unsold.

The Most Thorough Worker Came Last

The most humbling profile belongs to Opus 4.8. It was the most thorough participant in the entire field — 80+ learned rules, the deepest analyses of any model. And it finished dead last. The close was left on the table, and discipline slipped: instead of escalating when blocked, it attempted writes into a locked department. The same weakness appeared, more faintly, in all four competitors.

Thoroughness without follow-through is a beautiful workbench that never ships a single piece.

One fairness note: Kimi K3 ran at its API default, without an effort parameter, while the others ran at xhigh. A silver medal on default settings makes the bronze-and-below results look even less comfortable.

This Isn’t a Simulation That Ends

The crucible league emerged from a live, running thing: a company of 13 synthetic employees operating with real money mechanics — burning €105,000 a month against €2,300 in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. It’s watchable, live, at firmulate.com. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, if you want to test whether you can tell a 95 from a 73 by behavior alone.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From Watching to Doing: Your Own Worst Week, on Scrap Stock

Here’s the part that turns this from an interesting benchmark into a tool. Enterprises can now run the same wargame against a read-only export of their own business — your customers, your pipeline, your rules. Churn waves, price increases, competitor attacks, PR crises, social-engineering pressure, all aimed at a digital twin of your company. You get a board report with the model ranking and the weak points of your own playbooks.

The key spec, the one that matters to anyone who’s ever ruined good lumber: nothing ever writes back to your real systems. Read-only in, report out. You test the blade on scrap before you commit it to the workpiece.

If AI agents are going to touch anything that matters in your business, run them through their worst week in a place where their failures cost you nothing and their signatures earn you everything. Find out how at firmulate.com/pilot.html, or write to contact@firmulate.com to set up a pilot — before the AI runs your company for real.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What AI Tells Us About Trust and Follow-Through in Business

A groundbreaking experiment shows that AI’s true strength lies in follow-through, not just chat quality. Trust, reading depth, and integrity determine real value in business AI.

Build vs Buy a Prebuilt AI Workstation

Confused whether to build or buy your AI workstation? Discover the real costs, performance, and support differences with our clear guide for 2026.

ByteDance Seedance 2.5 Release: 30-Second Single-Take AI Video Starts Telling A Story – AIBase

ByteDance Seed says Seedance 2.5 creates coherent, 30-second single-take AI videos, but benchmarks, specifications and access details remain limited.

Turning A Dumb AC Unit Smart (Without Losing My Security Deposit)

Learn practical methods to upgrade a non-smart air conditioner to smart control without jeopardizing your security deposit or violating lease terms.