Skip to main content

Build Your Own Agents

A custom agent automates a repeatable post-coding task that no pre-built agent covers. You give Autoheal a trigger, a goal, a budget and the tools an agent may reach, and the agent works that task, ships something an engineer reviews, and is scored on every run.

You own every part of the agent: its behavior, tools, models and planning. Evals calibrate its performance over time, so you can see whether a change, from a model swap to splitting work into subagents, actually helped. Recurring patterns become new skills, and every change requires your approval.

Custom agents are created and managed in the Agents area of Autoheal: the Library tab lists every agent and workflow in your tenant, and the Sessions tab shows their runs.

Define an agent​

An agent is defined by four parts. Together they fix when it runs, what it is trying to achieve, what it may spend, and what it may touch.

PartWhat you set
TriggerA webhook, a schedule, a PR event or a chat command
GoalThe outcome you want, written as a prompt
BudgetDefault model, reasoning ceiling and spend per run
ToolsThe systems it may reach and the calls it may make

Trigger​

A trigger decides when the agent runs, and an agent runs whenever any of its triggers fire. You can trigger on:

  • A schedule, for recurring work such as a Monday cleanup or a nightly scan.
  • A PR event from GitHub, GitLab, Azure DevOps or Bitbucket, such as a pull request merged on a specific repository.
  • A webhook, on an inbound POST from any system that can call one, including your CI pipeline.
  • A chat command in Slack or Microsoft Teams.

You can also start a run manually, from the CLI, or over MCP.

Goal​

The goal is the outcome you want, written as a prompt. In the agent builder it is the agent's Instructions. A regression watcher's goal might read: after each merge to main, find any production regression the change caused by comparing error rate, p99 latency and saturation for the affected services against the deploy timestamp, checking Sentry for new error signatures, and reporting the likely cause with a recommended rollback or fix.

The goal is stable within a run, but it isn't frozen. The Evaluator scores every run against downstream outcomes, and when scores drop the Healer proposes a context change, back-tested against your own past runs. No change goes live until an engineer approves it, so the agent improves over time without changing under you unannounced.

Budget​

The budget sets what a run may cost:

  • Default model. Each agent picks its model from your approved list. Subagents can use a separate default, so you keep the frontier model for the reasoning that needs it and route routine steps to a cheaper one.
  • Reasoning ceiling. How much thinking the model may do per step.
  • Spend per run. A cap on what a single run may spend, alongside a budget per agent.

Tools​

Tools are the systems the agent may reach and the calls it may make. Access is deny-by-default: nothing outside the agent's grants is reachable, and the platform enforces that boundary rather than the prompt.

Grants are per integration, such as GitHub, Datadog, Grafana or LaunchDarkly, and each grant is one of:

GrantWhat the agent can do
ReadOnly read operations on that integration
Specific toolsOnly the named tools you select
AllEvery tool the integration exposes, including writes

The built-in Engineering Context Graph (your catalog, memories and skills) is always available. Write access, such as opening a pull request or posting a comment, is granted deliberately for the agents that need it, and governance policies can require a person's approval before any matching call runs, such as every write. Many integrations ship only read tools, so for those the Read and All grants give the same access.

Example jobs​

These are the kind of jobs a custom agent takes off a Platform Engineering, DevEx, DevOps or SRE team.

JobWhat the agent does
Stale feature flag cleanupEvery Monday, finds flags with no evaluations in 60 days and opens one PR per flag with the dead branch removed and tests updated
Deprecated internal APIWhen an API is sunset but its callers sit in forty repos nobody has time to touch, finds every caller and opens one migration PR per repo with tests run
CODEOWNERS and on-call driftStale CODEOWNERS stall reviews and page departed teammates, so on offboarding the agent finds every entry, suggests successors from commit history, and opens a PR
Flaky test quarantineWhen CI passes on retry, quarantines the flaky test with logs attached, and opens a fix PR once the cause reproduces

If a task can be written as a goal and grounded in your engineering context, an agent can run it.

What it does once running​

What happensWhere it's covered
OutputEach run produces something a person can review, reject or merge: a pull request with tests run, or a finding with its evidence and confidence scoreRuns and results
ScoringEach run is scored against your evals, and weak scores lead to proposed context changes that an engineer approvesEvaluator, Healer
Shared contextMemories, skills and learned integration paths are available to other agents, so a second agent for a similar job starts with that contextMemories, Skills
ControlsThe same controls as every other agent: tool grants, approval policies, and a budget per run and per agentGovernance Policies

Evaluation​

Evals give each change to an agent a measurable result, so a model swap or a restructuring into subagents is judged on scores rather than impressions.

FeatureWhat it does
Built-in evalsScore every run's trajectory and outcome without any setup
Custom judgesRubric-based judges you write and attach to an agent, scored as a number, a pass/fail or a category
Score historyDaily trends per eval, with comparison across agent versions
Proposed skillsRecurring patterns, such as a retry storm, can become skills, which go live only after an engineer approves them

Evals only measure what your rubrics and downstream signals capture, so a failure mode nobody wrote a rubric for stays invisible until someone notices it in a run.

Where it runs​

A custom agent runs on the same platform as the Incident Response, Vulnerability Remediation and AI Coding Cost Efficiency agents, in your isolated tenant (SaaS) or your own environment (BYOC), on models you have approved. See Sovereignty for SaaS and BYOC, connected or airgapped.

  • Built-in harness. Agents run on Autoheal's built-in harness, with each run's work isolated in a sandbox. Support for running agents on the coding agents your teams already use is in development (coming soon).
  • Bring your own models. Route each task to the right model from your approved list, with per-agent budgets, a separate default for subagents, and an allowlist the platform enforces.
  • Invoke it from where the team works. On a schedule, from a webhook or CI, the CLI, over MCP, or from the Slack or Teams channel where the team already tracks the work.

Runs and results​

Every run is recorded under Sessions. A run shows its status (completed, waiting for input, failed or stopped), the trigger that started it, the tools it used, its duration, and a result summary such as "Opened 12 PRs for stale flags" or "3 new flaky tests in the checkout suite". Opening a run shows the full trace: each step the agent took, the context it loaded, the tools it called, and the conclusion it reached.

When a run needs a decision, it pauses as waiting for input instead of acting unattended. Agents post their outcomes where the work already happens, such as a pull request, a PR comment, or a message in a Slack or Microsoft Teams channel.

Get started​

  1. Open Agents → Library and create an agent.
  2. Connect the integrations the agent needs.
  3. Add a trigger: a schedule, PR event, webhook or chat command.
  4. Write the goal.
  5. Set the budget: default model, reasoning ceiling and spend per run.
  6. Grant the tools the agent may use.
  7. Attach evals, then test the agent and activate it.