Evaluator
The Evaluator scores every agent run. It grades each run's trajectory and outcome against rubrics generated from what happened downstream, and it runs continuously on your own data, in your isolated tenant (SaaS) or your own environment (BYOC). Its scores are what the Healer acts on.
Scoring by downstream outcomes
In the SDLC, one agent's outcome is another agent's input, which means most runs have a verdict waiting for them later in the pipeline. The Evaluator uses those later events as ground truth:
| Signal | What it says about the run |
|---|---|
| Review comments | Whether reviewers had to correct or push back on code the agent helped produce |
| CI re-runs | Whether the change passed cleanly or needed retries and follow-up fixes |
| Fired alerts | Whether the problem the agent claimed to fix or triage kept alerting |
| The incident's actual root cause | Whether a triage or investigation pointed at the cause the team eventually confirmed |
Rubrics are generated from these signals for your environment, so a run is judged against what your reviewers, pipelines and on-call engineers did rather than a generic benchmark.
What gets scored
Every run is graded on two axes, and they fail independently:
| Axis | Question it answers | What it catches |
|---|---|---|
| Trajectory | How did the agent work? | Wrong data sources, wasted or redundant steps, an unsound diagnostic path, tools queried in a senseless order |
| Outcome | What did the agent produce, and did it hold up? | A root cause that isn't evidence-backed, an incomplete RCA, a fix that didn't address the defect, a change that drew review corrections |
An agent can reach the right conclusion through a bad trajectory, and that keeps working until the day it doesn't. It can also follow a sound path and still write up a conclusion the evidence doesn't support. Scoring both makes the signal worth acting on.
Scores are tracked per agent over time, so the Healer works from trends rather than a single bad run.
Evaluation is private to your tenant. It runs in your environment against your runs, and your run data is never used to train models.
Limits
The Evaluator can only score what leaves a signal. Downstream outcomes arrive late, since an incident's confirmed root cause may land days after the triage run, so some scores settle only once that evidence exists. A failure mode that produces no review comment, re-run, alert or confirmed root cause stays invisible until someone notices it in a run.
Related
- Healer: what happens after a weak score
- Self-Improvement overview: how the Evaluator fits in the factory loop
- Sovereignty: where evaluation runs and what data leaves your environment