DevOps8 min read

Telemetry for AI Coding Agents: The 7 Signals We Page On

An AI coding agent is a production system. The seven telemetry signals that, when they regress, mean something is materially wrong, with the dashboards and alert thresholds.

EnsureFix DevOps Team · Platform & DevOps Engineers, EnsureFix
Telemetry for AI Coding Agents: The 7 Signals We Page On, EnsureFix

A Production System Without A Page Order

Most teams treat their AI coding agent like a CI plugin. They look at the dashboard occasionally. They notice when it is obviously broken. They do not page on it.

For mature deployments, this is the wrong posture. An agent producing dozens of PRs per day is a production system. It has a SLO. It has failure modes that compound silently. It deserves the same telemetry rigor as any service touching the codebase.

This post is the seven signals we have found worth paging on, with the thresholds that work across deployments.

Signal 1: Acceptance Rate

Definition: percentage of AI-generated PRs merged within their target window.

Alert: drops more than 8 percentage points week-over-week.

Why it matters: acceptance rate degradation is the canonical leading indicator of something wrong upstream, model version drift, prompt change, retrieval regression. The change that broke it is recent.

What to check first: model version, prompt version, recent infrastructure changes. The cause is almost always one of these three.

Signal 2: Median Confidence

Definition: median of the calibrated confidence score across all AI-generated PRs in the last 24 hours.

Alert: drops more than 0.05 day-over-day.

Why it matters: the agent's own self-assessment is regressing. Either the ticket pool is harder, or the agent is less capable than it was. Either is information.

What to check first: ticket pool composition. Did intake change? If not, the agent regressed.

Signal 3: Escape Rate

Definition: percentage of merged AI-generated PRs that produced an incident or hotfix within 7 days.

Alert: above 1.5% in any 7-day window.

Why it matters: this is the canary for harm. Most teams have an escape rate target of well under 1%. Crossing 1.5% is a strong signal to pause auto-apply categories and tighten human review.

What to check first: the escape category. Is it concentrated in one repo, one bug class, one model version? Pause that bucket, investigate.

Signal 4: Validator Block Rate

Definition: percentage of AI-generated changes blocked by the pre-PR validator (not just flagged).

Alert: drops more than 30% week-over-week, or rises more than 50%.

Why it matters: a sudden drop suggests the validator stopped firing, broken integration, misconfiguration, or rules silently turned off. A sudden rise suggests the agent is producing worse code or a new rule is over-firing.

What to check first: validator health. Is it actually running? Did anyone change rules?

Signal 5: Per-Repo Cost Drift

Definition: 7-day moving average of cost per ticket, per repo.

Alert: 30% above the trailing 60-day average for any individual repo.

Why it matters: a single repo's cost can spike from a single bad pattern, usually retrieval is pulling in too many files, or the agent is repeatedly retrying on a category it cannot handle. The drift surfaces before the monthly bill arrives.

What to check first: token usage per ticket in the affected repo. Compare against neighbors. The deviation is usually obvious.

Signal 6: Time-To-First-Plan

Definition: 95th percentile of seconds from ticket pickup to plan output.

Alert: above 90 seconds for two consecutive hours.

Why it matters: this is the user-facing latency. When it stretches, reviewers context-switch away from the AI's workstream, and the human approval loop stalls.

What to check first: model serving capacity, retrieval latency, queue depth.

Signal 7: Reviewer-Override-To-Approval Ratio

Definition: number of reviewer overrides on AI rejections divided by number of clean approvals.

Alert: above 0.15.

Why it matters: a high override rate suggests the AI is rejecting things the team disagrees with, either over-cautious, or mis-calibrated on a specific bug class. Either way, trust is being eroded.

What to check first: which categories drive the overrides. Tune the validator or the confidence threshold for those categories.

What We Stopped Paging On

Three signals we initially paged on and turned off:

Per-ticket cost. A single expensive ticket is not an incident. Aggregate drift matters; individual outliers do not.

Plan word count. We thought a degradation in plan verbosity would signal model regression. It did not, plan length varies too much with ticket type. Median confidence is a better signal.

Model API latency. The model vendor's latency is real but not under our control. We page on time-to-first-plan, which is the integrated latency including retrieval and orchestration.

Dashboard Layout

A single dashboard, three rows:

  • Row 1: acceptance rate, escape rate, median confidence. The "are we still working" row.
  • Row 2: validator block rate, cost per ticket, time-to-first-plan. The "is the pipeline healthy" row.
  • Row 3: per-repo and per-category breakdowns of the row 1 metrics. The "where is the problem" row.

Anyone on call can read this dashboard in 30 seconds and form an opinion about whether the agent is healthy.

The Alert Routing

Three tiers:

  • Page. Escape rate above 1.5%, acceptance rate drop above 8 points, time-to-first-plan above 90s for over 2 hours. Wake someone up.
  • Slack. Median confidence drop, cost drift, validator block rate change. Same-day review.
  • Email. Weekly summary of all signals, week-over-week. Trend review, not incident.

The pages are calibrated to actual incidents, not noise. In our deployments, page rate is about one per quarter. Each page corresponds to a real issue.

What This Adds Up To

An AI coding agent is not "set and forget." It is a production system with telemetry, SLO, and a page order. Teams that treat it as such catch problems early. Teams that do not catch them when the quarterly review notices that acceptance is half what it used to be.

For the metrics framework underlying this, see 12 metrics for AI coding agent success. For how the signals tie back to the learning loop, see self-improving AI from code reviews.

Frequently asked questions

What metrics should I monitor for an AI coding agent?

Track seven signals worth paging on: acceptance rate, median calibrated confidence, escape rate, validator block rate, per-repo cost drift, time-to-first-plan, and the reviewer-override-to-approval ratio. Arrange them on one dashboard in three rows ('are we still working,' 'is the pipeline healthy,' and 'where is the problem'), so anyone on call can form an opinion in 30 seconds. For the broader framework these signals sit on, see 12 metrics for AI coding agent success.

What is a good escape rate for AI-generated code?

Most teams target an escape rate (merged AI PRs that cause an incident or hotfix within 7 days), of well under 1%. Crossing 1.5% in any 7-day window is a strong signal to pause auto-apply categories and tighten review. When it spikes, check whether the escapes concentrate in one repo, one bug class, or one model version before doing anything else. For the data behind escape-rate reduction, see how AI code review cuts bug escape rate.

How do I know if my AI coding agent is regressing?

Acceptance rate is the earliest tell: an 8-percentage-point week-over-week drop is the canonical leading indicator that something upstream broke recently. Median calibrated confidence dropping more than 0.05 day-over-day is a second signal, either the ticket pool got harder or the agent got weaker. Both point you at model version, prompt version, or retrieval as the likely cause.

Should you set up on-call alerts for an AI coding agent?

Yes, a mature deployment producing dozens of PRs a day is a production system with silent, compounding failure modes. Use three tiers: page for escape rate above 1.5%, acceptance drops over 8 points, or time-to-first-plan above 90s for two hours; Slack for confidence, cost, and validator changes; and a weekly email for trend review. Calibrated well, the page rate lands around one genuine incident per quarter.

What is time-to-first-plan and why does it matter?

Time-to-first-plan is the 95th-percentile seconds from ticket pickup to plan output, the user-facing latency of the agent. When it stretches past 90 seconds, reviewers context-switch away from the AI's workstream and the human approval loop stalls. It is a better page signal than raw model API latency because it captures the integrated cost of retrieval and orchestration, not just the vendor's response time.

EnsureFix DevOps Team

Platform & DevOps Engineers, EnsureFix

The EnsureFix DevOps team runs the distributed worker pool and CI integrations behind the pipeline, and writes about cycle time, MTTR, and delivery metrics.

observabilitytelemetrymonitoringAI operationsSRE

From reading to running

Ready to automate your tickets?

Watch EnsureFix take a real item from your backlog all the way to a pull request.