AI Coding Cost Efficiency
The AI Coding Cost Efficiency agent lowers the cost per task of the coding agents your developers already use, such as Claude Code, Codex or Cursor, without changing which agents they use. It reads their session traces, works out where the tokens went, tests cheaper configurations against real tasks from your own sessions, and proposes the changes that hold up as pull requests: a skill for a development team, or a settings change for platform engineering.
Collecting session traces from coding agents is in development (coming soon). The Autoheal CLI installs the Autoheal skill into Claude Code, Codex and GitHub Copilot today, but doesn't yet upload their session traces.
The agent is typically owned by a DevEx or platform engineering team. It doesn't run your coding agents or sit in the request path; it works from their traces after the fact, and every change it proposes goes through review like any other PR.
How a run works
The agent reads session traces from each connected coding agent: the prompts, the tool calls and their results, the model and reasoning effort used for each step, and the token count at each turn. Traces from different coding agents are normalized into one format, so one analysis covers all of them.
Each session's tokens are attributed to what consumed them: tool schemas loaded into context before the first prompt, tool calls made one at a time that could have been batched, repeated exploration of code a previous session already worked out, reasoning effort above what the task needed, and steps that ran on a larger model than required. The result is a breakdown of cost per task by cause, per team and per repository.
The agent selects completed tasks from your own sessions, with their known outcomes, and turns them into a benchmark set. Tasks are drawn from real work rather than a public suite, so a configuration is judged on the kind of changes your engineers actually make.
Candidate configurations are run against the benchmark side by side with the current one: a different model, a lower reasoning effort, a smaller helper model for routine subtasks, lazily loaded tool schemas, or a new skill that replaces repeated exploration. Nothing reaches developers during this stage.
Each candidate is scored on three measures: tasks completed correctly, cost per task, and time per task. A candidate that is cheaper but completes fewer tasks correctly, or slows engineers down, is rejected.
Candidates that hold accuracy while lowering cost are proposed as pull requests to the repository that owns the configuration, with the benchmark results attached. An engineer approves before anything changes for developers.
What the agent changes
Proposals land in one of two places, depending on who owns the fix.
| Change | Owner | Examples |
|---|---|---|
| Skill or AGENTS.md update | The development team that owns the repository | A skill that documents how the build works, so sessions stop rediscovering it; a correction to AGENTS.md that stops a recurring wrong turn |
| Settings change | Platform engineering | Default model or reasoning effort for a class of task; a smaller helper model for routine subtasks; deferred loading of tool schemas; tool-call batching; MCP server scoping |
Each PR states the cause it addresses, the benchmark tasks it was tested on, and the before-and-after scores on accuracy, cost and time, so a reviewer can judge the tradeoff rather than take a percentage on trust.
Where the spend usually goes
These are the causes the attribution step looks for, and the change each one usually leads to:
| Cause | What it looks like in a trace | Typical fix |
|---|---|---|
| Tool schemas loaded up front | Many MCP tool definitions in context before the developer's first prompt | Load schemas on demand, or scope which MCP servers a repository enables |
| Unbatched tool calls | Long chains of single reads or searches that could run in one call | Settings or instructions that encourage parallel or batched calls |
| Repeated exploration | Different sessions reading the same files to answer the same question | A skill or AGENTS.md section that records the answer |
| Effort above need | High reasoning effort on routine edits | A lower default effort for that class of task |
| Model above need | The largest model used for every step | A smaller helper model for routine subtasks, with the larger model kept for planning and hard reasoning |
Reports
The agent keeps a running view of coding-agent cost:
- Cost per task over time, by team, repository and coding agent.
- Spend by cause, from the attribution step.
- Change history: each accepted PR with the scores it was accepted on, and the measured cost per task after it went live.
That history is what answers the question of what a change bought: each configuration change is tied to a benchmark result and a measured outcome, not only to a higher count of pull requests.
How it improves
What the agent learns about your codebase and your sessions is kept as memories, so later runs start from known causes rather than re-attributing the same waste. Because skills and AGENTS.md files live in shared context, a skill proposed for one team's sessions is available to every coding agent that reads that repository.
Models and pricing change, so the agent re-runs the benchmark when a new model becomes available on your approved list, and proposes a switch only when the new model holds accuracy at a lower cost. Each of its own runs is scored by the Evaluator, and the Healer proposes changes to the agent's own context.
Access and safety
- Read-only on traces. The agent reads session traces and repository content; it doesn't run inside developers' sessions or change their tools directly.
- Changes only as PRs. Every change reaches developers through a reviewed pull request. A governance policy on write operations pauses the agent's writes, such as pushing a branch, for a person to approve.
- Where traces are processed. Traces can contain source code and prompts, so they are processed in your isolated tenant (SaaS) or your own environment (BYOC), never shared across tenants. See Sovereignty.
- Your models. Benchmark candidates are drawn only from models on your approved list.
Limits
- The benchmark reflects past work. Tasks come from completed sessions, so a configuration can pass the benchmark and still do worse on a kind of work your history doesn't contain. Watch the post-change cost and accuracy in the change history.
- Correctness needs an outcome signal. A task counts as done correctly only when its outcome is known, such as a merged PR that passed CI without follow-up fixes. Sessions with no clear outcome are left out of the benchmark.
- Traces must be available. The agent can only analyze coding agents whose session traces are connected. Sessions run where traces aren't collected aren't counted.
Get started
- Connect your coding agents' session traces once trace collection is available (coming soon).
- Connect the repositories where skills, AGENTS.md files and coding-agent settings live, with a credential that can push branches and open pull requests. The GitHub integration's own tools are read-only; the agent pushes branches and opens PRs with
gitandghin its sandbox, so the connected credential needs Contents and Pull requests write permission. On GitLab, the integration includes branch and merge-request tools. - Choose the models and reasoning efforts on your approved list that the agent may test.
- Run the agent once to produce the first spend-by-cause report, then review its first proposals before putting it on a schedule.