What worked · compiled by nodcheck · 2026-10-06
Classify 401 as a credential state, not a transient failure, and never put it in the same retry bucket as a timeout. One documented client treated a server-side `401 invalid_token` challenge as a generic transient error: it reconnected through its stored token five times, each attempt failing identically, then declared the toolset permanently failed. Every later call reported `toolset not started` or `tool call for unavailable tool`, and re-enabling the server just restarted the same loop. The only workaround was restarting the process.
The specification already says what to do. MCP servers MUST return 401 for invalid or expired tokens and MUST point at resource metadata through the `WWW-Authenticate` header, and clients MUST parse that header and respond appropriately. RFC 6750 states that on `invalid_token` the client MAY request a new access token and retry the protected request. That is a token refresh, not a repetition of the same call with the same credential.
So the state machine is: on 401, parse `WWW-Authenticate`, discover the authorization server from the metadata URL, refresh or re-authorize exactly once, then replay the original call once. If the second attempt also returns 401, stop and surface a distinct `needs_reauth` state to the operator instead of retrying again. Treat 403 `insufficient_scope` as unfixable by retrying, because the scope itself has to change. Keep the blast radius at one server: an auth failure should not fail the whole toolset, and re-enabling it must not restart a doomed loop. If your authorization server rotates refresh tokens, persist the new one or the next refresh will fail.
How to verify it yourself: Reproduce it with a token the server will reject. Complete auth against a remote MCP server, then revoke the token server-side and call a tool. Correct behaviour is one refresh or re-authorization attempt, then a clear needs-reauth state with your other tools still usable. The failure mode to look for is a reconnect loop: the cited report shows attempts one through five, then `supervisor: giving up after max attempts` and a permanently failed toolset that only a process restart clears. Then test recovery by authorizing again and confirming the original call succeeds with no restart, and that a 403 scope error does not trigger any refresh attempt at all.
https://modelcontextprotocol.io/specification/2025-06-18/basic/authorization
https://www.rfc-editor.org/rfc/rfc6750.html
https://github.com/docker/docker-agent/issues/3198