SyndaiSign inStart free
← Back to blog

The model or the harness: which one is doing the work?

·Syndai Team·6 minutes
ai-agents
coding-agents
harness
context-engineering

Ask three coding agents to build the same feature and the results rhyme. The models underneath Claude Code, Codex, and Cursor sit close enough now that the model is rarely the reason one run ships and another stalls. The reason is the harness: the code around the model that decides what it sees, what it can touch, and whether its output is allowed to land.

This is the shift worth naming. For two years the model was the product. You picked the smartest one and the rest followed. That era is closing. When the models converge, the harness is where the outcome is decided.

What a harness actually is

A harness is everything that is not the model weights. Give an agent a task and the harness answers four questions before the model ever helps:

  • What goes in the context window. Which files, which prior turns, which docs, which errors. Assemble it well and the model looks brilliant. Assemble it badly and the same model flails.
  • What tools it can call, and how. Shell, file edits, a browser, a database. Whether those calls run as JSON descriptions or as executable code. Whether a call can reach the network or is sealed off.
  • How work is checked. Whether output is run against tests, types, and a review pass, or trusted because it looked plausible.
  • Whether the result is allowed to land. The gate between "the agent produced a diff" and "the diff is on main."

None of that is the model. All of it decides whether the run succeeds.

Why the model stopped being the differentiator

Two things happened at once. Frontier models got good enough that the gap between the top few narrowed to something most tasks cannot feel. And open-weight models closed most of the remaining distance on agentic coding, at a fraction of the cost. So "use the best model" is no longer a moat anyone can hold for long.

So the variance moved. Every agent now has access to a capable model. The one that wins feeds the model the right context, routes tools cleanly, and refuses to ship code it has not checked. Those are harness properties. You can run the exact same model inside two harnesses and get a clean merged PR from one and a thrash loop from the other.

Context assembly is the first job

Most agent failures are context failures. The model was never given the file that mattered, or was handed five hundred files and lost the one that did. A harness earns its keep here first. Pull the right slice of the repo. Carry the errors forward. Drop the turns that no longer matter. Keep the window dense with signal instead of noise.

This is why "context engineering" became a discipline in its own right. It is the machinery that builds the prompt, every turn, from a codebase far larger than any window. A prompt is one string you type; context assembly is the system that writes that string for you, continuously.

Tool routing decides speed and safety

How an agent calls tools is not a detail. An agent can call tools as executable code instead of round-tripping a JSON description per call. That changes how fast it moves, and how much it can chain in one step. It also changes the blast radius. A tool call that can reach the network is a tool call that can exfiltrate a secret if the agent is steered by untrusted input. The harness decides whether that call is possible at all.

The recent run of coding-agent vulnerabilities landed exactly here: untrusted input reaching an agent that had more reach than it needed. That is a harness design failure, not a model failure. The model did what a compromised instruction told it to. The harness let the reach exist.

Verification is what separates a demo from a tool

A model will hand you code that looks right. Looking right is not the bar. The harness is where "looks right" gets tested against "is right". Run the tests. Check the types. Run a review pass that reads the diff as a stranger would. Loop until the findings stop. A run that skips this ships plausible bugs. A run that enforces it ships or fails honestly.

This is also where a report from the agent is not evidence. "I fixed it" is a claim. The green test is the evidence. A harness that treats claims as results is a harness that ships broken code with confidence.

The gate is the point

The last job is the smallest to describe and the hardest to get right: the boundary between produced and landed. A diff that passed review on a stale base is not a diff that passes on the current one. A green badge that registered before all the checks attached is not a green run. The gate is where an agent's work becomes real, and a gate that is easy to fool is worse than no gate, because it launders bad work as trusted.

So which one is doing the work?

The model writes the code. The harness decides whether that code was fed the right context, called tools safely, got checked, and was allowed to land. When models were far apart, picking the model was most of the battle. Now that they are close, the harness is the battle. If you are choosing or building a coding agent in 2026, judge the harness, because that is the part that still varies. One fixed model can score differently as its harness changes. See Prime Agent and the self-improving harness for a worked example.

Common questions

Does the model still matter? Yes, as a floor. A weak model caps what any harness can do. Above that floor, on most real coding tasks, the harness decides the outcome.

Is this just prompt engineering? No. A prompt is one string. A harness is the system that assembles that string every turn, routes the tools, verifies the output, and gates the result.

Can I put a cheaper or open model in a good harness and get frontier results? For many tasks, closer than you would expect. Verification and context assembly matter more as the model gets cheaper, not less, because a good harness catches what a cheaper model gets wrong.

What should I evaluate when choosing a coding agent? Context assembly, tool safety and routing, how it verifies its own output, and how it decides what is allowed to merge. Benchmarks measure the model. Those four measure the harness. For the current field, see agent frameworks and harnesses.

Where does MCP fit into the harness? MCP is the transport and tool layer around the model, and its 2026-07-28 revision went stateless, which is a harness-level change that leaves the model untouched.

Do agent skills live in the model or the harness? In the harness. A skill is instructions the harness loads and follows when a task matches, not anything baked into the weights, so it travels with the agent. We cover writing ones that port across Claude Code, Codex, and Cursor in portable agent skills.