Self-Improvement
Deploying an agent is only step one. As codebases and production architecture shift, APIs evolve and new tools arrive, agents drift, and a prompt or skill that worked last quarter quietly starts producing worse results. Keeping agents accurate and cost-efficient takes continuous evaluation and adaptation, and in Autoheal two agents do that work: the Evaluator and the Healer.

How the loop works
The loop runs through four stages, and each one feeds the next.
- Engineering Context Graph. The ECG holds what your agents know: skills, memories and the catalog. Agents read it through tool calls.
- Coding agents. Your existing coding tools, such as Claude Code, Codex or Cursor, call the ECG for context while they write code, and that code ships. See the CLI for how they connect.
- Factory worker agents. Once the code is live, Autoheal's worker agents handle what happens to it: incidents, vulnerabilities, and coding cost and context. Every run they make produces learnings about how the shipped code, and the context behind it, held up.
- Self-improvement agents. The Evaluator scores each run against what happened downstream, the Healer turns weak scores into context changes, and an engineer approves each change before it goes back into the ECG.
The loop closes when the next coding agent or worker agent reads the improved context.
Why downstream outcomes are the signal
In the SDLC one agent's outcome is another agent's input, so a run can be judged by what happened after it. A code change can be scored by its review comments and CI re-runs, and an incident triage can be scored by whether alerts kept firing and what the incident's actual root cause turned out to be. The Evaluator builds its rubrics from those signals, so scores reflect real consequences in your environment rather than a generic benchmark. See Evaluator.
Engineers stay in control
Every change the Healer proposes needs an engineer's approval to go live: skill, AGENTS.md and configuration changes are version-controlled and reviewed like a pull request, and proposed memories go through the Memories review queue. Before a change reaches review it has been back-tested against your run history and cleared a blast-radius threshold, so the diff arrives with evidence attached. Agent tool calls that write to your systems are governed separately, by approval policies.
Improvements compound
Because the context is shared, an improvement made for one agent benefits every agent that reads the same context, and the next coding agent to read it starts from better context, which the Evaluator measures in its next scores. The loop targets higher accuracy, faster execution and lower cost per successful task.
The loop only improves what the Evaluator can measure. A failure mode that leaves no downstream signal, or one that shows up only in a single rare incident, will take longer to surface than one that repeats. As agents prove reliable on the signals you do have, engineers can expand their autonomy, for example by letting an agent act on a class of incidents it has handled correctly many times.
Related
- Evaluator: how every run is scored
- Healer: how weak scores become reviewed context changes
- Engineering Context Graph: the context the loop improves