The AI Corner

The AI Corner

The Harness Is the Product: 6 Decisions That Turn an AI Agent Into a Worker You Can Leave Alone

DoorDash runs 130,000 agent tasks a month and OpenAI shipped 1,500 pull requests with three engineers, while Anthropic paid 22x more to get an app that worked.

Ruben Dominguez's avatar
Ruben Dominguez
Oct 04, 2026
∙ Paid

Your agent says the job is done, and the tests never ran.

It’s sharp for twenty minutes, then forgets a rule you gave it at the start. You retype the same three instructions every session. And you can’t walk away, because there’s always another approval waiting for a click.

Swap in a better model and every one of those failures stays put, because they live in the software wrapped around the model: what it gets told, what it keeps, what it’s allowed to touch, and who checks the result. That wrapper has a name now, the harness, and you already have one. Whether anyone designed it is a separate question.

The term took off in February 2026, when Mitchell Hashimoto used it on his blog and OpenAI published its harness engineering post a few days later. Since then, the companies getting the most out of agents have converged on the same idea from different directions: the model is the engine, and the harness decides whether you can leave the room.

TL;DR

  1. A harness is everything between the model and the work: the loop, the tools, the memory, the saved state, the permissions, and the definition of done. Half came from your vendor. The other half is yours, and it exists whether you designed it or not.

  1. DoorDash, OpenAI, and Anthropic built three completely different harnesses (a platform, a repository, a role split) and got production-grade agents out of all three. Every piece was built by the team running the agent.

  2. Anthropic’s harnessed run cost $200 and six hours against $9 and twenty minutes unharnessed. A harness pays off on work you couldn’t hand off at all, and loses money on work where you wanted to save twenty minutes.

  3. Premium covers the six decisions with copy-paste templates for each, a 12-question harness scorecard, the pass-cubed eval method, and the five-piece weekend build.

What a Harness Actually Is

Six things sit between the model and the work:

▫️ The loop that keeps it going, and the rule for when it stops

▫️ The tools it can see, and what they say when they fail

▫️ What stays in its context window, and what gets summarized away

▫️ What survives when the session dies

▫️ What it’s allowed to touch

▫️ Who decides the job is done

Your vendor owns the inner loop, the built-in tools, and the memory handling, and you can change little of it. Everything else is yours: the instructions file, the tests, the permissions, and the definition of done. Every rule you keep retyping into chat is part of your harness. So is every permission you clicked through because clicking was faster than reading.

Three Harnesses That Work

1. DoorDash built a platform

DoorDash moved its engineering agents off developers’ laptops and onto Flux, an internal cloud platform. In a single month in 2026, Flux automated 130,000 engineering tasks, and it now runs more than 25,000 automated code reviews a week across 300-plus reusable playbooks.

The design has four parts. Every agent gets its own sandbox. Every call to an internal system passes through one MCP gateway that grants only the permissions a job declared and logs everything. The work itself is written as YAML playbooks, which DoorDash describes as the Docker container for agent skills, mixing agent steps with ordinary deterministic code. The line from their engineering blog worth taping to a wall: getting an agent to write code is mostly solved, and the hard part is the environment around it.

2. OpenAI built a repository

OpenAI’s harness team built an internal product with zero hand-written code over five months. Three engineers, later seven, merged roughly 1,500 pull requests, about 3.5 per engineer per day, and throughput went up as the team grew.

No platform. The repo is the harness. Their first attempt, one giant AGENTS.md file, failed in the way every giant instruction file fails: stale rules piling up until nobody, human or model, read them. The fix was a roughly 100-line AGENTS.md that works as a table of contents into a structured docs folder, plus architectural rules enforced by custom linters instead of prose. Their own summary of the lesson: “give Codex a map, not a 1,000-page instruction manual.”

3. Anthropic built a role split

Anthropic’s Prithvi Rajasekaran borrowed the structure of a GAN. One agent, the planner, turns a one-sentence brief into a full product spec, and a second agent builds it. A third drives the finished app in a live browser through Playwright and grades it, catching bugs like wrong route ordering that would pass normal CI. The three talk only by writing files to each other.

Anthropic published the receipts, which almost nobody does. A solo agent built a retro game maker in about 20 minutes for $9, and the central feature was broken. The full three-agent harness took 6 hours and $200, and the game worked. Later, on a newer model, a simplified version built a browser music app in 3 hours 50 minutes for $124.70. The planner cost 46 cents.

So the harnessed version cost over twenty times more and produced the only version that worked, which is the trade every vendor demo leaves out.

Three completely different shapes, and the same lesson under all of them: the team using the agent built the part that made it work.


That’s the idea and the proof. What follows is how to build yours.

Premium subscribers get the full harness manual:

▫️ The six decisions, each with what to do, a copy-paste template, what it’s worth, and when to skip it

▫️ The four-file state system: SPEC, PLAN, PROGRESS, and DECISIONS templates ready to drop into any repo

▫️ The structured-error schema that turns “invalid request” into something an agent can act on

▫️ The pass-cubed eval method, and why one green run tells you almost nothing

▫️ The 12-question harness scorecard, plus the five-piece weekend build and the rule for when a harness costs more than it earns

Get the harness manual


Keep reading with a 7-day free trial

Subscribe to The AI Corner to keep reading this post and get 7 days of free access to the full post archives.

Already a paid subscriber? Sign in
© 2026 The AI Corner · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture