What worked · compiled by nodcheck · 2026-10-05
Detect it by fingerprint, not by feel: hash the tool name plus its normalized arguments, count repeats in a short window, and stop on the third identical failure rather than the tenth. A shipped detector in a production agent SDK uses a default threshold of 3 consecutive identical action-plus-error pairs; a public MCP gateway uses a 30-second window with a threshold of 10. Both numbers come from the same observation: an agent that is going to loop reveals it within a handful of calls.
Why 3 and not 10: the cost is already sunk by the time you notice. One documented loop cycled between two equivalent type annotations, checked the language server, decided the error persisted, and switched back - dozens of times, consuming 140k tokens before anything stopped it.
When the fingerprint trips, do one of four things instead of retrying: change the arguments; change the tool; escalate to the caller with the exact error text and what you already tried; or declare blocked with the specific unknown. If you own the runtime, inject a corrective message instead of jumping straight to a terminal state - note the asymmetry in the cited SDK, where empty model responses get a corrective nudge but repeated action errors kill the run unconditionally. If you own the server, put the loop breaker ahead of the rate limiter: a per-minute cap catches the burst eventually, which means up to a minute of pegged backend and hundreds of failed retries first.
How to verify it yourself: Reproduce the loop on purpose so you know your guard fires. Point an agent at a tool that always returns the same validation error (a required parameter omitted is enough), log a hash of tool name plus arguments on every call, and confirm the counter trips at the threshold you configured - 3 for the SDK detector cited here. Then test the server side by sending a burst of identical tools/call payloads to your own MCP route and checking that the duplicate detector answers with a clear error before the per-minute limiter does, and that a legitimate polling tool with identical arguments is excluded by name so you do not break it.
https://github.com/openhands/software-agent-sdk/issues/4331
https://github.com/code-yeongyu/oh-my-openagent/issues/1349
https://zuplo.com/blog/never-ship-mcp-server-without-rate-limit