A downstream service is degraded. Retrying every call keeps the upstream pinned and prevents recovery. Worse, the agent burns tokens on retries that would have failed.
When a tool depends on a single external service that has well-understood failure modes and a viable fallback (a secondary provider, a graceful 'service unavailable' message).
When there is no fallback and the service is critical to the agent's output. Better to surface the failure than to fake a result.
Describe this pattern to the builder assistant and it wires it up for you, brains, tools and triggers. It's vibe coding for agents, no config to write.
Walking through the pattern one piece at a time so the design is clear, not memorised.
The runtime exposes session.consecutive_failures.{tool_name}, tracked natively - no extra module needed to count them.
The condition checks that count before the call is even attempted. Once it reaches three, the gate action blocks the next attempt outright.
The gate's reason string tells the model exactly why the call was blocked, so it can explain the situation to the user instead of retrying blindly.
A successful web.fetch resets the consecutive-failure count to zero, so the circuit closes itself as soon as the service recovers - no manual reset needed.
The pattern above is not the only answer. Here is when something else is the right call.
Classic CB pattern: after the cooldown, allow one request through. If it succeeds, close the circuit; if it fails, reopen for another cycle. Slightly more logic, much smoother under intermittent failures.
Simpler. Costs more in degraded scenarios because every call still tries.
Engineering notes from the Digitorn team. No marketing, no launch announcements, no "10 prompts that will change your life". Just the things we write that we'd want to read.