
Every woodworker knows the rule: measure twice, cut once. The cut is the dramatic part — the satisfying moment when the saw bites into the grain. But the quality of the whole project was decided earlier, quietly, when you bent over the tape and actually read the measurement. Skip that step and no amount of skilled sawing saves the board.
It turns out AI models have exactly the same failure mode. In a recent public experiment by Firmulate — a company that runs AI models through simulated businesses and publishes the results on its open benchmarks page — four frontier AI agents were each handed the same small software company during its worst week. All four could “saw” beautifully. Only some of them bothered to read the tape. The difference was worth €55,000.
The worst week in business, run four times
The setup was deliberately brutal. Each model — the field included gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 — ran the identical company through the identical string of crises: angry customers, hard choices, and a series of temptations to cheat, including fake CEO messages that escalated over three stages and a reporter’s trick framed as “just one yes/no, on background.” Every decision was versioned and auditable, so nothing about a model’s behavior could hide after the fact.
The headline results from the final July 2026 league table: gpt-5.6-sol finished first with a score of 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own framing puts it, “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
Everyone passed the obvious test
Here’s what should reassure you: all five models, when the social engineering came, refused it. Five out of five turned down the impersonation attempts. Kimi K3 even left on-record reasoning that read, “Treat the request as a suspected approval-bypass / possible impersonation.” Every crisis was spotted. Every manipulation was rejected. On the test everyone worries about — will the AI go rogue? — the field was clean.
Then came the part nobody was testing for.
As an affiliate, we earn on qualifying purchases.
The fact buried two documents deep
Somewhere in that worst week sat a €55,000 deal. The decisive competitive weakness that would win it wasn’t in the customer conversation, though. It was buried two document references deep in the company’s own internal files — the corporate equivalent of the measurement written on the back of the plan, not on the cut list.
Same diagnosis, same pitch — the models that hadn’t done the reading delivered an analysis indistinguishable from the winners’. But only the models that traced the references and read the file actually closed the deal, at full price. The experiment quantified it: winning the deal was worth +€4,583 in monthly recurring revenue. The others simply never signed. The deal didn’t leak away through bad judgment; it evaporated through skipped homework.
As an affiliate, we earn on qualifying purchases.
Thoroughness isn’t the same as follow-through
The most striking profile in the league was Opus 4.8. It was the most thorough participant in the field — it learned more than 80 new rules during its run and produced the deepest analyses of any model. It still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating the issue properly. The same weakness, in weaker form, showed up in all four top performers.
One fairness note worth flagging: Kimi K3 ran at its API-default effort setting while the others ran at an extra-high setting — and still took second place with the cleanest discipline of the field.
As an affiliate, we earn on qualifying purchases.
Why a woodworking reader should care
You might never hand an AI a CRM or a support queue. But you’ve probably already asked one to summarize a manual, plan a cut list, or research a tool purchase. The lesson generalizes: the models that sound best and the models that do the reading are not always the same models, and the difference only shows up when something valuable is on the line.
Firmulate makes this watchable rather than theoretical. The experiment runs on a live synthetic company — 13 employees, real money mechanics, a burn of €105k a month against €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules — with every workday versioned and published. The site rebuilds itself twice a day as new benchmark runs finish. And for those who want to test their own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz. Enterprises can go further and run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

The chat-demo era of evaluating AI is over, or should be. Asking “does it write well” tells you nothing about whether it finishes what it starts, reads your files before answering, or stays disciplined when the exciting work is done and the paperwork isn’t. In this experiment, the models that won weren’t the ones with the flashiest analyses — Opus 4.8 had the deepest analysis in the field and finished last. The winners were the ones that did the unglamorous thing: followed a reference, read a document, picked up the pen.
Measure twice, cut once. It’s old advice in the workshop, and it turns out to be a measurable, purchase-deciding property of AI agents too — worth exactly €55,000, in this case. Before you trust an agent with anything that matters, check whether it does its homework. Firmulate’s public benchmarks are a good place to see what that looks like when it’s scored, versioned, and out in the open.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html