SyndaiSign inStart free
← Back to blog

Prime Agent and the self-improving harness: what RLM actually means

·Syndai Team·5 minutes
ai-agents
coding-agents
harness
prime-agent

On August 5, 2026, PrimeIntellect open-sourced Prime Agent. It is an MIT-licensed coding and research agent. It is built around two ideas: the Recursive Language Model and the Continual Harness. The launch is worth understanding on its own terms. It is also a clean case study in a shift we have written about before. When the model is fixed, the harness around it decides the result.

The reason to look closely is the headline number. PrimeIntellect reports a striking result. Run inside Prime Agent, Opus 5 scores above the human-expert baseline on a reasoning benchmark. The same model, run bare, does not. Whatever you make of the number, the claim is about the scaffolding rather than the weights.

What Prime Agent is

In PrimeIntellect's own words, Prime Agent is "an open-source coding and research agent for general and long-running work." It is released under the MIT License. The design rests on two abstractions.

The Recursive Language Model (RLM)

The RLM is the interesting part, and it is two moves.

Context is a variable. Most agents reread an ever-growing transcript each turn. Prime Agent instead keeps context as a variable inside a persistent IPython (Python REPL) kernel. The model can act on its own history through code. It can store, slice, and reference what it has seen as data. It is never handed the whole thing again at each step.

Subagents are function calls. Delegating a sub-task needs no separate orchestration layer. It is a call: await rlm("sub-task") spawns a full child session. The child gets its own model, its own kernel, and its own history. It hands back the result as a value. Recursion is the point. A task can spawn sub-tasks that spawn sub-tasks. Each one is a real agent, composed like a function.

The second abstraction is the Continual Harness. It is what PrimeIntellect means by "self-improving." The harness is built to carry and refine its own scaffolding across work. It does not reset each run.

The result, and how to read it

PrimeIntellect reports 95.5% (RHAE Best@1) on ARC-AGI-3, using Opus 5 inside Prime Agent. It states that this is above ARC's reported human-expert baseline of 95.4%.

Read that carefully. It is a self-reported result from the tool's authors, not an independently checked score. The honest way to cite it is with that attribution attached. What makes it interesting goes beyond the decimal. PrimeIntellect credits the gain to the harness. The model is a frontier model anyone can call. The delta they claim comes from how the harness assembles context, delegates, and composes calls around it.

That is the "harness over model" thesis stated as a product. Suppose a fixed model produces one result bare, and a much stronger one inside a particular harness. Then the harness is the variable that moved. Judging the agent means judging that harness.

Where self-improving harnesses get hard

The same property that makes this exciting is where the caution lives. A harness that refines its own scaffolding changes its behavior over time. That runs into two things a coding agent has to answer for.

Reproducibility. Say the scaffolding differs on run two from run one. Then "it worked last time" carries less weight. A result you cannot reproduce is a result you cannot fully trust. And a harness that edits itself makes exact reruns harder by design.

Review. Recursive subagent spawning is powerful. It is also a larger surface to supervise. Each rlm(...) call is a real agent with its own reach. The more of them run without a human reading the diff, the more review has to live in the system itself. A person at the end cannot cover it all.

Neither is a reason to avoid building this way. They are the questions that decide what a self-improving harness ships. Reliable work, or just impressive demos. The answers are the boring parts. Deterministic gates. Checks that do not depend on the agent's own report. A boundary that decides what is allowed to land.

Common questions

What is Prime Agent? An open-source (MIT) coding and research agent from PrimeIntellect. It is built around a Recursive Language Model plus a Continual Harness. The RLM treats context as a variable and subagent calls as functions, inside a persistent Python REPL. The Continual Harness refines its own scaffolding.

What is a Recursive Language Model (RLM)? PrimeIntellect's abstraction for context and delegation. The agent keeps its context as a variable it can act on through code. It spawns subagents by calling rlm(...). Delegation is a function call that hands back a result, with no separate orchestration layer.

Did Prime Agent beat the human baseline on ARC-AGI-3? PrimeIntellect reports a score above ARC's stated human-expert baseline, using Opus 5 inside Prime Agent. That is the authors' own reported result, credited to the harness rather than the model. Read it as a self-reported claim rather than an independent benchmark.

Why does this matter if the model is the same? Because it is evidence that the harness moved the result. The same model can score far higher inside one harness than another. So the harness is the thing to judge.