Why Raw Confidence Is Not Useful
Ask a frontier model "how confident are you?" and you will get a number. The number is usually around 0.85 regardless of how confident the model should actually be. Models are systematically overconfident at low difficulty and underconfident at high difficulty, and the bias varies by task type. Raw self-reported confidence is not actionable.
Calibration is the work of turning the raw number into a number that tracks the actual probability of correctness. Without it, your decision engine (the part that routes high-risk changes to humans), is making decisions on noise.
This is the method we use.
The Calibration Set
You need a labeled outcome set. For AI coding, the labels we use:
- True positive: Change merged, no escape into production.
- False positive: Change merged, caused an incident or hotfix within 7 days.
- True rejection: Change blocked or rejected, would have caused issues.
- False rejection: Change blocked or rejected, was actually correct.
We collect these per ticket category. Bug fixes are calibrated separately from dependency bumps, which are calibrated separately from refactors. The calibration curve is meaningfully different across categories, the same raw confidence score means different things in different contexts.
Minimum sample size: 50 samples per category before calibration is meaningful. Most teams need 3-6 weeks of operation to build the set.
The Calibration Curve
For each category, we bin raw confidence into deciles and compute the actual accuracy in each bin. A perfectly calibrated model has accuracy = confidence in every bin. In practice the curves look like:
- Raw confidence 0.85-0.95: actual accuracy 0.62.
- Raw confidence 0.95-1.0: actual accuracy 0.79.
- Raw confidence 0.5-0.6: actual accuracy 0.48 (close to calibrated).
- Raw confidence below 0.5: model rarely reports this, but when it does, accuracy is around 0.25.
The model is systematically overconfident in the high range. The calibration function pulls those scores down.
Applying The Curve
Two ways to apply:
Isotonic regression. Fit a monotone function from raw to calibrated. Robust to small sample sizes, no parametric assumptions. This is our default.
Platt scaling. Fit a sigmoid from raw to calibrated. Smoother extrapolation but more assumptions. We use this when sample sizes are very small.
The calibrated score is what feeds the decision engine. Raw scores are kept for debugging but never used for routing.
What The Decision Engine Does With It
The calibrated score routes the change:
- Above 0.85: auto-apply candidate (subject to category gates).
- 0.65-0.85: human review, standard cadence.
- 0.45-0.65: human review, elevated cadence, flag specific concerns for the reviewer.
- Below 0.45: refuse or escalate.
The thresholds are tunable per team. A team early in rollout uses higher thresholds (more human review). A team late in rollout, with strong learning loop data, can lower them.
What Breaks Calibration
Three patterns that degrade calibration:
Model version changes. A new model version reshapes the confidence distribution. Recalibrate within 2 weeks of any model swap. We block model updates in pipelines until the calibration set has been refreshed.
Distribution drift. The mix of tickets shifts, suddenly more refactors, fewer bug fixes. The category-level curves still hold, but the aggregate looks worse. Always calibrate per category.
Reviewer leniency drift. If reviewers start accepting marginal changes they would have rejected a quarter ago, the "merged with no escape" label becomes noisier. Periodically sample merged PRs and ask a senior reviewer to re-grade. If their grade differs from the merged-without-issue label, the labels are noisy.
A Specific Number That Matters
The ratio of "calibrated confidence above 0.85 changes that escape" to "calibrated confidence above 0.85 changes that ship clean" is the strongest indicator of decision engine health. We target 1:50 or better. Above 1:30, we tighten the threshold. Below 1:100, we can consider loosening.
Raw confidence cannot give you this kind of signal. Calibrated confidence can.
Where Calibration Does Not Help
Three failure modes calibration cannot fix:
The model is wrong in a way it cannot detect. If the model is missing context (a related file was not retrieved), its confidence will be high and wrong, and no calibration curve will catch it. The fix is upstream, in retrieval, not in calibration.
The label distribution is bimodal. If a category has both "trivial typo fix" and "complex multi-file refactor" tickets mixed in, the curve averages them and is mis-calibrated for both. Subcategorize.
Adversarial inputs. If someone is deliberately crafting tickets to manipulate the agent's confidence, calibration fitted on benign inputs will fail. Not a current concern for most teams, but worth flagging.
What To Build First
- Start logging raw confidence per ticket. If you are not already.
- Tag merge-time outcomes per PR.
- After 50+ samples per category, fit the first calibration curve. Plot it. Notice how off the model is.
- Wire calibrated score into your decision engine. Keep raw score for debugging.
- Refit monthly. Recalibrate immediately after model swaps.
Calibration is the difference between a decision engine that feels arbitrary and one that earns trust. It is also one of the cheaper engineering investments to make.
For the decision engine architecture, see anatomy of an autonomous pull request. For how this plays into ticket refusal, see why your agent should reject tickets.
Frequently asked questions
Can you trust an AI model's self-reported confidence score?
Not raw. Models cluster their self-reported confidence around 0.85 almost regardless of actual correctness, and they're systematically overconfident on easy tasks and underconfident on hard ones. Feeding raw confidence into a decision engine means routing high-risk changes on noise, you need to calibrate it against real outcomes first.
What is confidence calibration for AI coding agents?
It's the work of turning a model's raw confidence number into one that tracks the actual probability of correctness. You collect labeled outcomes per ticket category (merged with no escape, merged then caused an incident, correctly blocked, wrongly blocked), bin the raw scores, and fit a function that maps raw to calibrated. The calibrated score is what should drive routing decisions.
How much data do you need to calibrate AI confidence scores?
At least 50 samples per ticket category before the calibration is meaningful, which most teams reach after 3 to 6 weeks of operation. Calibrate categories separately (bug fixes apart from dependency bumps apart from refactors), because their curves differ, and refit monthly, with an immediate recalibration after any model swap.
How do calibrated confidence scores route code changes?
The calibrated score maps to review tiers: above 0.85 is an auto-apply candidate subject to category gates, 0.65 to 0.85 goes to standard human review, 0.45 to 0.65 goes to elevated review with specific concerns flagged for the reviewer, and below 0.45 is refused or escalated. The thresholds are tunable, teams early in rollout run them higher. This connects directly to intake, covered in why your AI agent should reject most tickets.
What can calibration not fix in an AI coding agent?
Three things: errors the model can't detect because it's missing context (a related file wasn't retrieved), so its confidence is high and wrong, fix retrieval, not calibration; bimodal label distributions where trivial and complex tickets are mixed in one category, which averages the curve and requires subcategorizing; and adversarial inputs crafted to manipulate confidence. For the decision-engine and pull-request context this feeds, see the anatomy of an autonomous pull request.
EnsureFix Research Team
The EnsureFix research team studies model behavior, confidence calibration, and evaluation methodology for autonomous coding agents, translating findings into the production pipeline.