The Class Of Ticket Single-Agent Models Cannot Handle
Some tickets are simple: one file, one repo, one change. Single-agent LLMs handle them well. Other tickets touch three services across two repos, require coordinated schema changes, and need to be deployed in a specific order. Single-agent LLMs fail at these in characteristic, repeatable ways.
This post is about why they fail and what architecture handles them coherently.
The Failure Modes
Three patterns we see when a single-agent LLM is given a cross-service ticket:
Context overflow. The agent loads all three repos into its context. Even with a 200k token window, the relevant code from three services plus their tests plus their schemas easily blows past the productive range. Accuracy regresses sharply past 90k tokens (see context window engineering).
Plan-execution mismatch. The agent plans changes in three services in parallel, then executes them sequentially, but loses track of which assumptions it made in the plan. Service B is implemented based on what Service A was supposed to do, but Service A's actual implementation diverged from the plan.
Deployment-order blindness. The agent produces three correct diffs that, if deployed in the wrong order, cause an outage. The schema migration needs to ship before the code that references the new column, but the agent does not know that.
These failures are deterministic. Give the same single-agent system the same cross-service ticket, and it will fail the same way.
What A Multi-Agent Architecture Provides
A pipeline approach decomposes the problem:
- Cross-repo planner. Reads the ticket and identifies which repos and services are involved. Produces a per-service plan with explicit interfaces and dependencies.
- Per-service planners. Each gets the per-service plan plus the interface contracts. They plan within their service.
- Per-service coders. Each writes its own diff against the per-service plan. They do not see each other's repos.
- Cross-service validator. Receives all three diffs plus the interface contracts. Validates that the interfaces match what each service expects.
- Deployment-order analyzer. Reads the diffs and produces a deployment order with constraints: "deploy service A's schema migration first, then service B, then service C."
The architecture is more complicated than a single agent. It produces correct outcomes on the cases where the single agent fails.
The Interface Contract
The key data structure is the interface contract. When the cross-repo planner says "service A will expose a new endpoint /users/bulk with this request and response shape," that contract is what each per-service plan operates against.
Without explicit contracts, each per-service plan is implicit about its assumptions. With explicit contracts, mismatches are detectable before any code is written.
In our data, 30% of cross-service tickets fail at the contract validation step before any code is written. Catching the mismatch in the planning phase saves three coding attempts and three CI runs.
What The Deployment-Order Analyzer Does
The analyzer reads the diffs and identifies:
- Schema changes: must deploy first.
- Producer/consumer changes: producers deploy before consumers for new fields, consumers deploy before producers for removed fields.
- Feature-flagged changes: can deploy in any order if the flag controls activation.
- Breaking changes: must coordinate with the consuming team if there is no compat layer.
The output is an ordered list of merge-and-deploy steps with optional manual checkpoints. The human reviewer sees the plan as part of the PR description.
The Specific Cases Where It Helps Most
Three ticket shapes where multi-agent decisively beats single-agent:
API version bumps that propagate. Service A bumps from v1 to v2. Services B and C call it. The change requires coordinated updates and a backward-compat window. Single-agent: produces correct v2 in A, breaks B and C silently. Multi-agent: produces v2 in A with a compat layer, updates B and C, plans deprecation of compat layer for next quarter.
Schema migrations with code changes. New required column. Migration must add the column, code must populate it, and existing rows need backfill. Single-agent: writes the code expecting the column to exist, ships before migration runs, breaks production. Multi-agent: plans the migration sequence with explicit phases.
Distributed transaction changes. A workflow that spans three services and needs to remain atomic. The retry semantics in one service interact with the timeout semantics in another. Single-agent: reasons about each service in isolation. Multi-agent: maintains the global view through the planner.
What It Does Not Help With
Two patterns that multi-agent does not fix:
Truly novel architecture. If the cross-service work involves designing a new system pattern that the team has not used before, the agent does not have learned context for it. A human architect leads, and the agent assists.
Cross-team coordination. Service A's team and service B's team disagree about who owns the interface. The agent cannot mediate. A tech lead does, and the agent works downstream of the resolution.
The Cost
Multi-agent on cross-service tickets is more expensive than single-agent on single-service tickets. The per-stage routing (see token economy) keeps it reasonable. A typical cross-service ticket costs $4-8 versus $1-2 for a single-service ticket.
The cost-per-correct-PR ratio is favorable because single-agent on cross-service tickets has a 20-30% acceptance rate (we measured this), and multi-agent on the same tickets is 70-80%. Three multi-agent attempts produce a correct PR. Single-agent might need ten attempts and still fail.
Where To Use This
Three signals you are in cross-service territory:
- The ticket text mentions more than one service by name.
- The estimated change touches more than one repository.
- The plan involves any kind of schema or contract change.
If any apply, route to the multi-agent pipeline. If none apply, the single-service pipeline is faster and cheaper.
For the underlying architecture rationale, see multi-agent AI architecture and why single-agent LLMs fail in enterprise code.
Frequently asked questions
Why do single-agent LLMs fail on cross-service code changes?
They fail in three deterministic ways. Context overflow: loading three repos plus their tests and schemas blows past the productive range even in a 200k window, and accuracy regresses sharply past about 90k tokens. Plan-execution mismatch: the agent plans in parallel but executes sequentially and loses track of assumptions, so Service B is built against what Service A was supposed to do rather than what it became. And deployment-order blindness: it produces three correct diffs that cause an outage if deployed in the wrong order. See why single-agent LLMs fail at enterprise code.
How does AI handle a code change that spans multiple repositories?
A multi-agent pipeline decomposes the work. A cross-repo planner identifies the services involved and produces a per-service plan with explicit interfaces and dependencies. Per-service planners and coders work within their own repo against those contracts, a cross-service validator checks that all the diffs' interfaces match, and a deployment-order analyzer produces an ordered merge-and-deploy sequence. The architecture is more complex than a single agent but produces correct outcomes on the cases where the single agent breaks. See multi-agent AI architecture.
What is an interface contract in multi-agent code generation?
The interface contract is the explicit shape each service exposes, for example, 'service A will expose /users/bulk with this request and response.' It is what every per-service plan operates against. Without explicit contracts, each plan makes implicit assumptions that only collide at runtime; with them, mismatches are detectable before any code is written. In practice 30% of cross-service tickets fail at contract validation in the planning phase, which saves three coding attempts and three CI runs.
How does multi-agent AI determine the correct deployment order?
The deployment-order analyzer reads the diffs and applies rules: schema changes deploy first; for producer/consumer changes, producers go before consumers for new fields and consumers before producers for removed fields; feature-flagged changes can deploy in any order when the flag controls activation; and breaking changes without a compat layer must be coordinated with the consuming team. The output is an ordered list of merge-and-deploy steps with optional manual checkpoints, shown to the reviewer in the PR description.
How much does a cross-service AI ticket cost compared to a single-service one?
A typical cross-service ticket costs $4-8 versus $1-2 for a single-service ticket, with per-stage model routing keeping the multi-agent cost reasonable. The cost-per-correct-PR ratio still favors multi-agent: single-agent acceptance on cross-service tickets is 20-30% and might need ten attempts while still failing, whereas multi-agent lands 70-80% and produces a correct PR in about three attempts. See the token economy of per-stage model routing.
When should you route a ticket to a multi-agent pipeline instead of a single agent?
Use three signals: the ticket text names more than one service, the estimated change touches more than one repository, or the plan involves any schema or contract change. If any apply, route to the multi-agent pipeline; if none do, the single-service pipeline is faster and cheaper. Multi-agent will not fix everything, though, truly novel architecture and cross-team ownership disputes still need a human architect or tech lead, with the agent working downstream of the resolution.
EnsureFix Research Team
The EnsureFix research team studies model behavior, confidence calibration, and evaluation methodology for autonomous coding agents, translating findings into the production pipeline.