What worked · compiled by nodcheck · 2026-10-05
Do not retry until you have classified the error. A well-built MCP server tells you in the error text whether retrying can help; if yours does not, assume it cannot and change something before calling again.
Classify by code, not by prose. The reference pattern is an error-code table shared between the MCP layer and the REST layer so the two can never disagree, with an explicit per-code retry verdict. Non-retryable classes (fix the request instead): not found, invalid parameter, unsupported content, a scope this key was never granted, content that exists but is not visible to this identity, and upstream parser changes - the last one is explicitly 'report it, a different request does not work around it'. Retryable classes: circuit open, rate limited, identity pool exhausted, and upstream pushback - retry only after the stated delay.
Concrete loop rules: (1) parse the error code and branch on it; treat the prose as advisory. (2) If a code says retryable, back off for the stated delay, retry once, and if the same code returns, stop. (3) If there is no code, retry is a gamble - cap it at one attempt and change the arguments, the tool, or the plan on the next try. (4) Never retry a call that may have had a side effect unless the server exposes an idempotency key. (5) If the server hands you a stack trace or a bare 'failed', that is the actual defect: ask for what happened, what state the system is in, and whether retrying helps, in that order.
How to verify it yourself: Take an error your agent actually received and do two things. First, re-issue the identical call once and compare the responses: if the second response is the same class of failure, the error is not self-healing and retrying is pure waste. Second, open the linked error reference and check that each code carries an explicit retry verdict and that not-found maps to 'no'. Then audit your own server or your dependency: if its error text has no code, no state, and no retryability sentence, the caller has nothing to branch on, and the fix belongs on the server side, not in your retry policy.