An agent in a loop can hammer an expensive API and spike your bill. You want a hard ceiling on calls per session, not a soft promise in the prompt.
Any tool with per-call cost or upstream rate limits. LLM providers, paid search APIs, third-party integrations via connectors.
Free internal services where rate limiting just adds latency without preventing real harm.
Describe this pattern to the builder assistant and it wires it up for you, brains, tools and triggers. It's vibe coding for agents, no config to write.
Walking through the pattern one piece at a time so the design is clear, not memorised.
Telling the model 'don't make too many calls' fails a meaningful fraction of the time. The runtime hook is deterministic: at the threshold, the next call is gated.
An inject_message hook tells the agent it has used half its budget. Models react to this signal and consolidate their calls.
When the cap fires, the gate action returns a clear reason instead of a raw failure. The agent works with what it already has instead of crashing.
Worst-case tool calls per session are known upfront. You can compute upper bounds for monthly spend without staring at a dashboard.
The pattern above is not the only answer. Here is when something else is the right call.
security.behavior.profile: coding already ships a max_sequential_same_tool rule. Coarser, no per-tool granularity, but zero extra YAML.
Put a proxy in front of the external API that enforces a global rate. Works for many agents at once, doesn't help with single-session cost.
Engineering notes from the Digitorn team. No marketing, no launch announcements, no "10 prompts that will change your life". Just the things we write that we'd want to read.