How the cascade actually works
The cascade starts with a blocked event, not a blocked agent. In a five-service saga orchestration, each service represents a local transaction, and the saga kicks off in Service 1 by emitting an event to Kafka (readmedium.com). Services 2 through 5 are not "waiting on Service 1" in any abstract sense — they are waiting on that specific event. If a tool call inside Service 1 fails and the retry budget is spent, the event is never emitted, and every downstream consumer sits idle. The failure is not that Service 1 gave up; it is that Services 2–5 never got a turn.
The reason a shared retry counter makes this worse is the same reason a shared integer counter loses updates under concurrency. The C# Interlocked tutorial at zetcode.com documents the failure mode precisely: eight tasks each incrementing a shared counter 500,000 times should produce 4,000,000, but the run printed 692,864. The statement counter++ is not one operation — it is a read, an add, and a write, and when two threads interleave those three steps, both read the same value and the second write overwrites the first. The lost updates are never reapplied. A global retry counter behaves identically: two agents that both read "2 retries used" and both write "3 retries used" have collectively consumed one retry slot, not two.
That lost-update pattern is the mechanism by which one bad tool call propagates. Agent 1 fails a call, retries, and writes its increment back to the shared counter. Agent 2, running concurrently, reads a stale value, retries its own call, and writes back a value that erases Agent 1's increment. The counter now understates consumption, so the pipeline keeps retrying past the intended ceiling — or, in the opposite interleaving, Agent 1's final increment lands after Agent 2 has already read the budget as exhausted, and Agent 2 fails fast on a budget it never actually spent. Either way, the counter no longer describes any single agent's behavior.
The check is straightforward: instrument the retry counter with an atomic increment and log the value each agent reads before and after its attempt. If the read-before and read-after values do not advance by exactly one per attempt, the counter is shared and interleaved, and the budget is fiction. The rule that follows is to give each agent its own counter, cap it at three tries, and trip a circuit breaker when any agent hits the cap — failing the whole pipeline immediately rather than letting a downstream agent inherit an exhausted budget it did not spend.
| Signal | What it means | Action |
|---|---|---|
| Counter read-after minus read-before ≠ 1 | Interleaved updates are being lost | Move to per-agent counters |
| One agent's attempts consume another's budget | Budget is shared, not owned | Trip circuit breaker on cap |
| Downstream agent fails with zero attempts | Upstream exhaustion propagated | Fail pipeline fast, do not retry downstream |

The evidence: 3 tries is the cliff
The strongest argument for a per-agent retry cap is not a design preference — it is arithmetic. The zetcode.com Interlocked tutorial runs a race-condition demonstration that should make any pipeline author pause: eight concurrent tasks, each incrementing a shared counter 500,000 times, produced an expected total of 4,000,000 but an actual value of 692,864. That is a loss of 3,307,136 increments, or roughly 82.7% of the expected writes. Recompute it yourself: 4,000,000 minus 692,864 equals 3,307,136, and 3,307,136 divided by 4,000,000 is 0.8268. The lost updates were never applied and never retried, because nothing in the code knew they had been dropped.
That result matters here because of how the counter is built. The same source documents that Interlocked is lock-free: each call executes as one indivisible atomic operation, the calling thread never blocks, and it cannot deadlock. A retry counter built on Interlocked therefore fails in the quiet direction. It will not hang your pipeline; it will under-count. A tool call that should have consumed one unit of retry budget may consume none, and the agent proceeds as if it still had attempts in reserve. The check for a reader is blunt: if your retry counter is lock-free and you have never reconciled its final value against the number of attempts you actually issued, assume it under-counts.
Now widen the frame. The readmedium.com saga write-up describes orchestrating a process across five services, each representing a local transaction, with the saga kicked off in Service 1 by emitting an event to Kafka. Five is the fan-out width that turns one stuck retry into five stalled agents. When the shared budget is exhausted by a single bad tool call, every downstream agent inherits an empty account — not because each of them failed, but because the counter they all read from was already at its limit. The failure is positional, not behavioral.
Three tries is the cliff because it is the smallest number that still looks generous while being trivially exhaustible. One bad call plus its two retries consumes the entire allowance, and the agents behind it get nothing. The rule that follows is mechanical: cap retries at three per agent, counted locally, and trip a circuit breaker when any agent hits that cap so the whole pipeline fails fast instead of dribbling attempts into a dead chain.
Two checks make this operational. First, verify that each agent owns its own counter — if two agents can read the same variable, you have a global budget wearing a per-agent label. Second, verify that the breaker fires on the cap, not after it, so no downstream agent ever starts work against an exhausted upstream budget.

Options compared: global vs per-agent cap
A global retry budget looks tidy on a whiteboard: one counter, one ceiling, every agent draws from the same pool. It fails in practice for a mechanical reason. The zetcode.com Interlocked tutorial is explicit that the class "works on one location at a time" — each call is a single atomic operation on a single variable. A shared retry counter plus the agent's own state (which agent is running, whether it has already burned a try, whether it is mid-tool-call) is two locations, and Interlocked cannot update both as one indivisible unit. The tutorial's race-condition demo makes the cost concrete: eight tasks incrementing one shared counter half a million times each should produce 4,000,000, and the run printed 692,864. That is the shape of the bug you inherit when five agents share one retry counter — lost updates, and a budget that drains faster or slower than any single agent can observe.
Compare the three options directly:
| Option | Counter scope | Failure mode | Verdict |
|---|---|---|---|
| Global retry budget | One shared counter across all 5 agents | Interlocked cannot atomically update counter plus agent state (zetcode.com); one bad tool call in Agent 1 drains the pool before Agent 4 runs | Reject |
| Per-agent cap of 3 | One counter per agent | None at the budget layer — a bad tool call in Agent 1 cannot touch Agent 4's budget | Winner |
| Per-agent cap plus circuit breaker | One counter per agent, plus a trip state | Agent stops retrying after 3 failures and the saga emits a compensating event instead | Winner when the pipeline is a saga (readmedium.com) |
The per-agent cap of 3 is the default because it is the only option where the counter and the agent's state live in the same scope. Each agent owns its own integer, so the atomicity problem disappears: there is nothing to keep in sync across agents, and no downstream agent can inherit an upstream agent's exhausted budget. The check is simple — before any tool call, read that agent's counter; if it is already at 3, do not call. If it is below 3, call, and increment on failure only.
Add a circuit breaker when the pipeline is a saga. The readmedium.com orchestration writeup describes five services, each a local transaction, with the saga kicked off in Service 1 by emitting an event to Kafka. In that topology, retrying a failed step is not neutral — it holds the saga open while downstream services wait on an event that may never arrive. After the third failure, the agent should trip: stop retrying, mark itself open, and let the orchestrator emit a compensating event. The rule to enforce is that the breaker trips on the same count as the cap, so there is exactly one threshold to reason about, not two.
Pick per-agent cap plus breaker for saga pipelines, per-agent cap alone for straight-line chains. Either way, the global counter is the option to strike first — it is the one that lets a single bad tool call in Agent 1 spend Agent 4's budget.

Costs and numbers that matter
The arithmetic of retry cost is where most pipeline budgets quietly break. Five agents, each allowed three tries, is fifteen attempts before the pipeline can declare failure. If every attempt costs one tool call, that is fifteen tool calls spent to reach a verdict — and that number is the ceiling you should be budgeting against, not the average you hope to see.
Now compare that ceiling to what a shared counter actually reports. The zetcode.com Interlocked tutorial demonstrates the lost-update problem with eight concurrent tasks incrementing a shared counter half a million times each; the expected total is 4,000,000, but the run prints 692,864. That is a loss rate of 82.7 percent — roughly 3.3 million increments that were read, added, and then overwritten before anyone observed them. Apply that same rate to a shared retry counter across a five-agent chain and the reporting gap becomes concrete: 82.7 percent of 15 attempts is about 12.4 attempts that never register, leaving the counter showing roughly 2.6 of 15. The pipeline has burned fifteen calls while the dashboard insists it has barely started.
That gap is not a rounding error. It is the difference between a budget that looks healthy and a budget that is already exhausted. A per-agent cap of three tries keeps the accounting honest because each agent owns its own counter, so no interleaved write from a sibling agent can erase a retry that was actually spent. The check is simple: after any failed run, sum the per-agent attempt counts and confirm the total equals the number of tool calls your logs show. If the sum is lower, you are reading a shared counter and it is lying to you.
A circuit breaker changes the blast radius, not just the accounting. Trip it after three failures and the chain stops at three calls per agent instead of grinding through fifteen across the pipeline. The downstream agents never start, so they never inherit an exhausted budget from an upstream peer — which is the failure mode a global counter invites.
| Scope | Attempts before failure | Counter reliability |
|---|---|---|
| Global retry counter, 5 agents | 15 | Under-reports by ~12.4 attempts at the 82.7% loss rate |
| Per-agent cap, 3 tries each | 15 | Accurate per agent; no cross-agent overwrite |
| Per-agent cap plus circuit breaker | 3 per agent, chain halts on trip | Accurate and bounded |
The rule to carry forward: cap retries at three per agent, verify the sum of per-agent attempts against your tool-call log, and trip the breaker on the third failure so the pipeline fails fast rather than spending its way to the same conclusion.

What the evidence does NOT establish
The available sources does not give a measured cascade rate for five-agent pipelines. The 82.7% figure comes from the zetcode.com Interlocked tutorial's race-condition demonstration, where eight concurrent tasks each increment a shared counter 500,000 times and the expected total of 4,000,000 lands at 692,864 instead — a lost-update rate of roughly 82.7% on that single counter test. That is a concurrency artifact in an 8-task benchmark, not a measurement of how often one bad tool call propagates through a chain of five agents. Treat the number as a warning about shared mutable state, not as your pipeline's expected failure rate.
The available sources also does not establish that 3 is the optimal retry count. Three is the method threshold this guide sets, chosen because it is the smallest budget that still absorbs a transient fault while keeping worst-case cost bounded. Your own pipeline's failure distribution is the only authority on whether 3 fits. Before adopting the cap, log every retry attempt per agent for a representative window, then plot how many failures resolve on attempt 2 versus attempt 3 versus never. If almost nothing resolves on the third try, lower the cap; if a meaningful share does, keep it. The rule is a starting point you verify, not a constant you inherit.
The rule breaks in a specific, identifiable case: when agents share state beyond the retry counter. Interlocked works on one memory location at a time and takes no lock, so it cannot protect several variables that must change together (zetcode.com). If your agents coordinate a retry counter alongside a shared status flag, a queue position, or a budget field, atomic increments on the counter alone leave the other fields exposed to the same interleaving that produced the 692,864 result. In that configuration, replace Interlocked with a lock around the whole state transition, or move the shared state behind a single owner agent that serializes updates.
Two further edge cases deserve a check before you trust the cap. First, if a downstream agent can independently trigger the same failing tool call, a per-agent cap does not stop the pipeline from re-entering the same fault through a different door — verify by tracing whether any tool is reachable from more than one agent. Second, if your circuit breaker trips on a per-agent counter but the orchestrator restarts the saga, the breaker must persist across restarts or it resets to zero and the cascade repeats. Confirm where breaker state lives before you rely on it.
Where the rule still wins: any pipeline where each agent owns a disjoint slice of state, retries are idempotent, and the orchestrator can fail the whole run fast. In that shape, a per-agent cap plus a breaker contains the blast radius to the agent that actually failed, and no downstream agent inherits an exhausted budget it never spent.

5-agent saga, 3-try cap
Walk the saga forward and watch the counters. Five agents, A1 through A5, each holding its own Interlocked retry counter capped at 3. The saga kicks off in A1, which emits an event to Kafka to start the chain (readmedium.com). At the moment of dispatch, every counter sits at zero and the pipeline is healthy.
Checkpoint 1: A1's tool call fails. A1 increments its own counter to 1. A2, A3, A4, and A5 remain at 0. The pipeline is still alive — one retry consumed, four agents untouched. This is the state a per-agent cap is designed to preserve: a single agent's bad call does not spend anyone else's budget.
Checkpoint 2: A1 fails again. A1's counter reads 2. Now run the same sequence against a shared counter and the number stops being trustworthy. The zetcode.com Interlocked tutorial demonstrates the failure mode directly: eight concurrent tasks each incrementing a shared counter 500,000 times should produce 4,000,000, but the observed run printed 692,864 — lost updates, because read, add, and write are three separate steps that interleave (zetcode.com). Scale that race down to a saga and the reported value at checkpoint 2 could read 0 or 1 instead of 2. A downstream agent that reads that stale value believes the budget is fresh and retries into a wall.
Checkpoint 3: A1 fails a third time. A1's counter hits 3 — the cap. The circuit breaker trips, and the whole pipeline fails fast rather than letting A2 through A5 inherit an exhausted budget. The rule for your own pipeline: cap retries at 3 per agent, trip the breaker on the third failure, and fail the saga immediately. Never let a downstream agent read an upstream agent's counter.
| Checkpoint | A1 counter | A2–A5 counters | Pipeline state |
|---|---|---|---|
| 1 — first failure | 1 | 0 | Alive |
| 2 — second failure | 2 | 0 | Alive; shared counter may misreport as 0 or 1 |
| 3 — third failure | 3 (cap) | 0 | Breaker trips; saga fails fast |
The check to run before your next deploy: instrument each agent's counter independently and assert that no agent's counter is ever read by another. If your saga shares one counter across all five agents, the zetcode.com race result is your early warning — the number you log will not be the number you think you have.
Decision rules
Rule one: if more than one agent shares a single retry counter, then split that counter into per-agent Interlocked counters before the next run. The zetcode.com Interlocked tutorial demonstrates that unsynchronized concurrent increments to a shared variable lose updates because the read-add-write sequence interleaves across threads — the actual result falls short of the expected total on every run. The same pattern applies to a shared retry counter in an agent pipeline, where one agent's retries silently consume another agent's budget. System.Threading.Interlocked.Increment performs the read-modify-write as one indivisible, lock-free unit, so no thread can interleave its update with another's and every agent's count remains accurate.
Rule two: if an agent reaches 3 retries, then trip its circuit breaker and emit a compensating event rather than retrying. The readmedium.com saga pattern describes a five-service orchestration where each service represents a local transaction; when a local transaction cannot complete after its budget is exhausted, the saga coordinator must reverse the effects of previously completed transactions by sending compensating events — not by retrying the failing agent a fourth time. The circuit breaker stays open until an explicit reset, preventing a flapping tool from re-entering the pipeline on every oscillation.
Rule three: if the pipeline is a saga across 5 services, then treat each service as a local transaction with its own budget. The readmedium.com source describes the saga kicking off in Service 1 by emitting an event to Kafka, with Services 2 through 5 each acting as an independent local transaction. Each local transaction carries its own retry cap, so a failure in Service 3 does not draw down the retry budget assigned to Service 5, and the coordinator can compensate only the services that actually committed.
Rule four: if a downstream agent observes an upstream failure, then fail fast without retrying — never inherit an exhausted retry budget. A downstream agent that retries after its upstream has already exhausted its budget is attempting work on data that may be stale or missing, and each retry wastes budget that belongs to a different agent. The correct behavior is to propagate the failure immediately so the saga coordinator can begin compensation.
Rule five: if a compensating event fires, then the saga coordinator rolls back completed local transactions in reverse order rather than restarting the pipeline from the beginning. This preserves the work that succeeded before the failure and avoids re-triggering side effects that were already committed to external systems.
What to do next
| Step | Action | Why it matters |
|---|---|---|
| 1 | Define your specific needs and budget | Narrows options to what actually fits |
| 2 | Compare top 3 options side by side | Reveals the best value for your situation |
| 3 | Check current pricing and availability | Prices change frequently — verify before committing |
| 4 | Book directly with the provider | Often gets better terms than third parties |
| 5 | Set a reminder to review in 6 months | Policies and pricing shift — stay current |
Also worth reading: Agent tool failure recovery: 95% success with retry-first vs replan 2026: Agent tool failure recovery: 95% · From simple chains to interlocked workflows: a practical migration guide: From simple chains to interlocked · Agent pipeline failure recovery 2026: 8-second interlock retry vs restart: Agent pipeline failure recovery 2026: