AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

The Workshop Test That Tool Makers Understand

Any woodworker knows the drill — literally. A router that feels perfect in the store can chatter through hard maple. A chisel that gleams in its packaging may not hold an edge past the first dovetail. You don’t judge a tool by its shelf appeal; you judge it by the work it finishes, under load, on a bad day, when the clamp slips and the grain fights back.

Right now, businesses are buying AI agents the way amateurs buy tools: off the spec sheet, dazzled by the demo. Coding leaderboards and chat arenas measure whether a model answers well. But the jobs companies actually want done — running a support queue, closing a deal, keeping a forecast honest — are workshop jobs. They’re measured by what gets finished, not how polished the shavings look.

That’s the gap a live public experiment at Firmulate set out to expose, and its first finished season produced a result every maker would recognize instantly.

Amazon

AI model testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Shop, Same Week from Hell

The setup is elegantly controlled, like a proper jig. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — were each handed the same small software company and told to run it through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing rests on anecdotes.

Think of it as the AI equivalent of handing five apprentices the same tangled board and the same broken plane, then watching who actually delivers the finished piece.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone Passed the Eye Test. Only Two Delivered the Furniture.

The headline finding from the final league table is a paradox. All the models spotted every crisis. All of them refused every manipulation attempt thrown at them. And yet only two signed the €55,000 deal that their own analysis had earned. As the project’s own summary puts it: “Same diagnosis, same pitch — no signature.”

The final standings: gpt-5.6-sol finished first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For calibration, doing nothing at all still scored 26, because partial progress counts — but a single breach of trust caps the total. In this workshop, “no amount of good work outweighs a breach of trust.”

Amazon

AI support automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact: Models That Read the Files Won

Here’s the detail that should haunt anyone buying an AI agent. The decisive weakness in the competing vendor — the fact that won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue — wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files.

The models that actually read their own shop’s paperwork found it and closed. The ones that didn’t, didn’t. It’s the oldest lesson in the trades: measure twice, and actually look at what’s in front of you before you cut.

Under Pressure, the Social Engineering Failed

The week included fake CEO messages that escalated over three stages, plus a reporter trying the classic “just one yes/no, on background” trick. All five models refused, every time. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” Honesty under pressure, it turns out, may be the easier skill to teach than follow-through.

The Most Thorough Player Came Last

Opus 4.8 is the cautionary tale of the season. It was the most thorough participant — over 80 learned rules added, the deepest analyses of the field — and still finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Diligence without delivery is a beautiful workbench with nothing built on it.

One fairness note worth flagging: K3 ran without an effort parameter (API default) while its rivals ran at xhigh — and still took second at 93.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Not a Slide Deck — a Company You Can Watch Lose Money

What makes this more than a one-off paper is that the company is live. It has 13 synthetic employees, real money mechanics — burning €105,000 a month against €2,300 in MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. You can watch it run at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Management Quality, Not Chat Quality

The category Firmulate is proposing — and demonstrating in public — is management quality, not chat quality. Not: can the model write a lovely email? But: does it finish what it starts, does it read your files first, does it stay honest when a fake CEO comes knocking, and does it hold its discipline when the week goes sideways?

For the buyer, the lesson is the one every craftsman learns early: the demo tells you how a tool talks. Only the job, run end to end under real pressure, tells you how it works. Before you hand an agent your CRM, your support queue, or your forecast — put it in the shop, hand it the worst board in the pile, and see what actually comes off the bench.

The first season’s answer was blunt: spotting every crisis and refusing every temptation wasn’t enough. Two of five finished the job. Would the tool on your shelf?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Blender 5.2 LTS

Blender has officially announced version 5.2 LTS, offering extended support for users needing stability and reliability in professional workflows.

Why a Filament Dryer Can Matter More Than a Printer Upgrade

Unlock better 3D prints by controlling filament moisture—discover why a dryer might outperform a costly printer upgrade in ensuring flawless results.

Measure Twice, Trust Once: How Five AIs Handled a Fake Boss

Five AI models faced fake CEO demands and a reporter trick. Every one refused, showing integrity under pressure can be tested before deployment.

Collection Of Digital Clock Designs

A new collection highlights diverse digital clock designs, emphasizing innovation in digital time displays for enthusiasts and designers.