AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A Bench That Scores 26 for Doing Nothing? That’s Not a Bug — It’s a Square Check

Any woodworker knows the ritual. Before you cut, you check the square against a known-straight edge. If the square says a board is true when it isn’t, every joint downstream inherits the lie. A benchmark is the same tool for a different material: instead of measuring boards, it measures decisions. And the most useful thing a square can do is tell you where “zero” actually sits.

That’s why one of the strangest, most honest numbers in the current wave of AI evaluation is 26. In the Crucible League benchmark, a run that does almost nothing — a do-nothing baseline managing a small software company through its worst week — still scores 26 points out of 100, not zero. To a builder’s eye, that’s not grade inflation. That’s the benchmark marking its own datum line, the way you’d scribe a reference edge before cutting anything expensive.

Amazon

woodworking square measuring tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Worst Week in Software, Run Five Times

Here’s the setup. Four frontier AI models — GPT-5.6-Sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — were each handed the same job: run the same small software company through the same catastrophic week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, the way a good shop log records every cut.

The final league table from July 2026 reads: GPT-5.6-Sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the headline finding isn’t the spread — it’s the cluster. All five models spotted every crisis. All five refused every manipulation attempt thrown at them. Yet only two signed the €55,000 deal that their own analysis had earned. The benchmark’s own shorthand for the failure: “Same diagnosis, same pitch — no signature.”

If you’ve ever watched a perfectly milled set of boards never become a table, you know this failure mode intimately. The work was 90% done. The last 10% — the part that actually matters — never happened.

Amazon

AI benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the Floor Is 26, Not 0

Back to that baseline. A do-nothing run scores 26 because the benchmark counts partial progress. Answering a customer, correctly triaging a crisis, drafting the analysis — these are real units of work, even if the run never closes anything. A benchmark that scored a stalled manager at zero would be like a tape measure that only reads finished houses: technically tidy, practically useless. You’d have no way to distinguish a manager who did nothing from one who did almost everything.

But the floor has a ceiling’s twin. A single breach of trust caps the total score — permanently, no matter how much good work surrounds it. The benchmark’s stated principle: “No amount of good work outweighs a breach of trust.” In the shop, this is the crack in the load-bearing tenon. It doesn’t matter how beautiful the rest of the piece is; the chair fails. The scoring treats integrity the same way: as a structural member, not a bonus point.

And note what this design implies about round numbers. A score of 100 would mean a manager that did everything, perfectly, under maximum pressure. The benchmark’s own architecture — breach caps, partial credit, auditable versions — is built to distrust exactly that kind of tidiness. The top score in the field, 95, left room on the table. That gap is information, not noise.

Amazon

decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact and the €4,583 Detail

The decisive difference between the winners and the also-rans wasn’t brilliance. It was reading. The detail that won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue — wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files: a competitor weakness that only models that actually read their own filing cabinets could deploy. The models that read won. The models that skimmed gave the same pitch and got no signature.

Every woodworker has the scar tissue for this. The drawer that binds because you didn’t check the humidity sticker on the lumber. The finish that blushes because you didn’t read the SDS. The information was there. Two references deep. You just didn’t go get it.

Amazon

software project management logs

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering: All Five Passed

The week included a staged assault: fake CEO messages escalating over three stages, capped with a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the voice of someone — something — that has been taught where the trapdoor is.

The Lesson of Opus 4.8

The most instructive profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Thoroughness without follow-through is a beautiful workbench with nothing built on it.

(One fairness note the benchmark itself discloses: Kimi K3 ran without an effort parameter while the others ran at maximum effort — and still took second at 93.)

You Can Watch the Company Run

This isn’t a one-off lab report. The live company — 13 synthetic employees, real money mechanics, a burn of €105k per month against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules — is running now and watchable at firmulate.com/live. Every workday is versioned. There’s also a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business via firmulate.com/pilot.html — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

A benchmark is only as good as its datum line. By scoring a do-nothing run at 26, giving credit for partial progress, and capping the total on any single breach of trust, this one tells you three things a chat demo never will: where zero actually sits, how much of the job got done, and whether the thing holding the pencil can be trusted with the cut. The models mostly measured up. Only two of them closed the deal — and in business, as in woodworking, the piece isn’t done until the finish is on and the client has signed.

Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Milwaukee Lineup Compared: Which Model Should You Buy in 2026?

Compare Milwaukee’s M18COMPACT and 2903-20 FUEL drills to find the best fit for your needs in 2026. Performance, features, and real-world use analyzed.

One Video In, a Whole Publishing Kit Out — Without the Cloud

Discover how local AI transforms a single video into a full publishing package—without cloud dependency. Faster, private, and subscription-free workflows await.

Why a Filament Dryer Can Matter More Than a Printer Upgrade

Unlock better 3D prints by controlling filament moisture—discover why a dryer might outperform a costly printer upgrade in ensuring flawless results.

Blender 5.2 LTS

Blender has officially announced version 5.2 LTS, offering extended support for users needing stability and reliability in professional workflows.