Skip to main content

Healer

The Healer fixes what the Evaluator finds. It turns weak scores into specific context changes, back-tests each one against your run history, and proposes only the changes that clear a blast-radius threshold, each for an engineer to review.

What the Healer changes​

A weak score localizes a problem, and the Healer proposes the narrowest change that addresses it. Typical changes are:

ChangeWhen it applies
Updated skillA skill the agent followed is stale or missing a step
AGENTS.md correctionRepository instructions that coding agents read are wrong or out of date
New memoryA run surfaced knowledge that the next run should recall; see Memories
Narrower tool scopeThe agent wasted steps or reached for tools it didn't need
Different modelThe task is handled as well by a cheaper model, or needs a stronger one

Each change arrives as its own pull request with the scores that motivated it, so you can accept one fix and decline another.

Back-testing and the blast-radius threshold​

Before a change reaches review, the Healer back-tests it against your run history.

1
Select

The Healer picks past runs from your own environment that the change would affect, including the ones the agent got right and the ones it got wrong. The failures show whether the change fixes something; the successes show whether it breaks something.

2
Replay

Those runs are re-run with the proposed change and scored with the same rubrics the Evaluator uses on live runs.

3
Gate

The change is promoted to a pull request only if it clears the blast-radius threshold, meaning it improves the target runs without regressing too many others.

Back-testing only covers situations that have already happened, so a change can pass it and still behave differently on a kind of work your history doesn't contain. That is one reason every change still needs a human approval.

Engineer approval​

Every change the Healer proposes needs an engineer's approval to go live. Skill, AGENTS.md and configuration changes are version-controlled, and you review each one like a pull request: read the diff, check the back-test results, and accept or close it. A proposed memory goes to the review queue on the Memories page instead, where the same auto-accept rules apply as for any other memory. The pull request goes to the repository and reviewers that own the context it changes. See Sovereignty for how approval works for agent tool calls.

Because the change lands in shared context, an approved fix helps every agent that reads that context, including coding agents working in the same repositories.