Agentic AI Foundation Logo
Stateless MCP changes your security model: how to harden your MCP servers

Stateless MCP changes your security model: how to harden your MCP servers

Chaitanya RahalkarOctober 7, 2026

The latest MCP specification removed protocol-level sessions. This post covers where the security work they did has moved, and what to check before you migrate a server.

TL;DR: In the stateless MCP release, every request carries everything needed to handle it, which makes servers easier to scale. It also moves several security decisions out of the transport and into your code: who owns a state handle, whether a requestState value was tampered with, whether a header matches the body it describes, and who may share a cached tool list. This post collects the spec's rules for each in one place, with examples.

MCP sessions were never meant to be a security feature. The previous spec said outright that servers must not use them for authentication. But they still held context that security decisions depended on. A session tied a series of requests to one client and gave the server a place to remember what had already happened.

The latest MCP specification (version 2026-07-28) removes the initialize handshake and the Mcp-Session-Id header. Every request now carries its own protocol metadata, and requests to protected HTTP servers carry authorization. As AAIF reported in September, that helped providers handle this summer's surge in tool calls, because any request can now go to any server instance behind an ordinary load balancer.

The security work didn't go away with the session. It moved into tool arguments, into values the server sends out and the client returns unchanged, into HTTP headers, and into cache hints. The spec has rules for each of these, but they're spread across several pages. Here's what to review if you're migrating a server.

Every request has to prove who it's from

The spec now says plainly that "no state should be inferred from previous requests, even those on the same connection or stream." An open connection doesn't count as a session, not even a stdio process. For HTTP servers, that means validating the bearer token and its audience on every request. The spec already required this. The difference is that there's no session left to rely on between checks.

Also know what not to trust. Clients should now send io.modelcontextprotocol/clientInfo with every request, which makes it tempting to allowlist clients by name. The spec warns that this information is "self-reported by the sender and [is] not verified by the protocol," and that implementations "SHOULD NOT rely on [it] for security decisions." Use it for logs, not access control.

State handles are names, not keys

Servers that need state across calls, like a shopping cart or a database transaction, now return an explicit handle from a creation tool. The model then passes it back as an ordinary argument:

// create_basket returns
{ "structuredContent": { "basket_id": "bsk_a1b2c3" } }

// a later call
{ "name": "add_item", "arguments": { "basket_id": "bsk_a1b2c3", "sku": "..." } }

This changes who can see the identifier. A session ID lived in an HTTP header. A handle lives in the model's context, so it can also end up in transcripts, traces, retry queues and prompts passed to subagents. Assume anything that can read those places can reuse the handle.

The security best practices now cover this risk in a section called State Handle Hijacking. Its core rule: servers "MUST NOT treat possession of a state handle as authentication." Generate handles randomly and let them expire. Bind each one on the server to the user identified by the verified token, and check ownership on every call. A server without authentication has no verified user to bind the handle to. Long, random, expiring handles reduce guessing and replay risk, but they don't provide the ownership guarantees of authenticated binding.

One detail matters here. The spec gives <user_id>:<handle> as an example storage key. If your user IDs can contain a colon, as multi-tenant IDs often do, and the handle comes straight from a tool argument, simple concatenation can produce the same key for two different users. A community write-up demonstrated this. The fix is to store the pair as a composite key, or to encode it so the two parts can't run together:

python
import json

# Collides: ("tenant-a:svc", "H") and ("tenant-a", "svc:H")
# both become "tenant-a:svc:H"
key = f"{principal}:{handle}"

# Doesn't collide: the boundary between the two values is preserved
key = json.dumps([principal, handle])

Also return the same "not found" error for a handle that doesn't exist and one that belongs to someone else. Otherwise the ownership check tells an attacker which handles are real.

Treat requestState as attacker-controlled input

A server sometimes needs something in the middle of a call, like a confirmation or a missing parameter. Before this release, it sent its own request down a stream it held open. The new Multi Round-Trip Requests (MRTR) pattern replaces that. The server returns a result with resultType: "input_required", its questions in inputRequests, and optionally a requestState value that only the server can interpret. The client collects the answers and retries the original call, sending back inputResponses and the same requestState.

This means a stateless server can ask "Are you sure you want to delete this project?" before doing something irreversible. But the state that makes the retry work has passed through the client. The MRTR section of the spec is direct: servers "MUST treat requestState as an attacker-controlled input." If the value affects authorization, resource access or business logic, the server has to protect it from tampering, for example with a keyed hash (HMAC) or authenticated encryption (AEAD), and reject anything that fails the check. To limit replay, the spec recommends including the authenticated user, a short expiry time, and an identifier for the original request in the protected payload. That identifier can be the method name plus a hash of the key arguments.

That last item is easy to skip, so here's why it matters. Suppose delete_project signs its confirmation state but doesn't tie it to the arguments. A confirmation given for deleting a scratch project could then be reused to delete production. A minimal version that prevents this:

python
import base64, hashlib, hmac, json, time
 
def _b64(b):
    return base64.urlsafe_b64encode(b).decode()
 
def _digest(args):
    return hashlib.sha256(json.dumps(args, sort_keys=True).encode()).hexdigest()
 
def _sign(key, body):
    return _b64(hmac.new(key, body.encode(), "sha256").digest())
 
def mint_state(key, principal, method, args, ttl=300):
    payload = {
        "sub": principal,                      # from the verified access token
        "req": f"{method}:{_digest(args)}",    # ties state to this exact request
        "exp": int(time.time()) + ttl,         # short expiry
    }
    body = _b64(json.dumps(payload).encode())
    return body + "." + _sign(key, body)
 
def verify_state(key, state, principal, method, args):
    body, _, tag = state.rpartition(".")
    if not hmac.compare_digest(tag, _sign(key, body)):
        raise PermissionError("requestState failed verification")
    p = json.loads(base64.urlsafe_b64decode(body))
    if (p["sub"] != principal or p["exp"] < time.time()
            or p["req"] != f"{method}:{_digest(args)}"):
        raise PermissionError("requestState does not match this caller or request")
    return p

Keep two limits in mind. These checks shorten the window for replay, but they don't make the state single-use. If an action must happen at most once, enforce that on the server. Also, an elicitation result is the client's report of what the user chose. It guards against a model acting on its own, but it doesn't prove a person clicked anything.

For secrets and payment credentials, the elicitation spec forbids form-mode elicitation and requires URL mode, which sends the user to a secure page outside the MCP client. That page has to confirm that the person who opens it is the same user who triggered the request. Otherwise, an attacker could forward the link to someone else as a phishing lure.

Headers are for routing; the body is the source of truth

Requests over the Streamable HTTP transport must now include an Mcp-Method header. Tool calls, resource reads and prompt fetches must also include Mcp-Name. Servers can mark tool parameters with x-mcp-header, and clients then copy those values into Mcp-Param-* headers. Gateways and web application firewalls can then route and authorize each tool call without parsing the JSON body.

That creates a familiar gap: two components making decisions from different copies of the same fact. If a gateway allows Mcp-Name: read_file but the server executes delete_file from the body, the gateway's policy accomplished nothing. The transport spec requires any server that processes the body to reject such a mismatch with HTTP 400 and the HeaderMismatch error (-32020). It also says gateways that enforce policy based on these headers should reject requests from protocol versions that predate this check.

So make sure the comparison happens in your server, not only in your gateway. And don't mark sensitive parameters with x-mcp-header. The spec specifically lists passwords, API keys, tokens and personal data, because every proxy and log along the path can read headers.

Cache scope is an access-control decision

List results such as tools/list, along with resources/read, now carry two caching fields: ttlMs and cacheScope. Any client, gateway or proxy may store a "public" result and serve it to any user, even if it came from an authenticated endpoint. A "private" result belongs to one authorization context and must not be shared across authorization contexts.

tools/list is allowed to vary with the caller's permissions. So if you filter tools per user and label the result "public", a shared cache can show one user's tools to another. Label filtered lists and user-specific resources "private". As the caching spec notes, the label doesn't enforce access by itself, so check permissions on every call.

Retries are new requests

The release also removes the ability to resume a broken stream. If a response stream breaks, the in-flight request is lost, and the client must send it again as a new request. So a tool that charged a card or sent a message might receive the same call twice. Design tools with side effects to handle repeats safely, for example by recording completed work against a state handle and skipping it on a repeat call.

Smaller authorization fixes that close real gaps

Most of the authorization changes covered here apply to clients and authorization servers. They address known OAuth attack paths and client-registration risks:

  • Issuer check (RFC 9207). Before exchanging an authorization code for a token, clients must now check the iss value in the response against the authorization server they expected. If that server advertises iss support but the value is missing, the client must reject the response. This blocks "mix-up" attacks, where a malicious authorization server tricks a client into handing over a code issued by an honest one.
  • Credentials tied to their issuer. Clients must not reuse stored credentials with a different authorization server.
  • New client registration method. Dynamic Client Registration is deprecated in favor of Client ID Metadata Documents (CIMD), where a client identifies itself with a URL pointing to a JSON description of itself. If you run an authorization server, fetching those URLs creates a server-side request forgery (SSRF) risk, so block requests to internal addresses. A metadata document also can't prove which local program owns a localhost redirect, so show users the redirect hostname.

A migration checklist

  • Validate the token and its audience on every request. Never base access decisions on clientInfo.
  • Make handles random and expiring, bind them to the verified user with a key that can't collide, and check ownership on every call.
  • Sign requestState and tie it to the user, an expiry time and the original request. Enforce single use for actions that must happen only once.
  • Collect secrets and payment credentials only through URL-mode elicitation, and confirm the user who opens the link.
  • Reject header/body mismatches in the server itself, and keep sensitive values out of x-mcp-header.
  • Mark per-user results cacheScope: "private", and check permissions at call time.
  • Make tools with side effects safe to retry.

What's still open

Some servers used session IDs to track resources created by anonymous users, without showing that identifier to the model. In the public discussion of SEP-2567, participants acknowledged this release doesn't support that case directly, and suggested it may need support in the authorization layer in a future release. The Transports Working Group is also discussing whether MCP needs a Conversation ID or Thread ID to group a logical sequence of requests.

If you run stateless MCP in production, your experience will help settle these questions. Join the Transports Working Group discussions on GitHub, or bring your lessons to AGNTCon + MCPCon North America in San Jose on October 22-23. Without sessions, MCP is easier to scale. With these checks in place, the security decisions around state, routing and caching are explicit in code you can inspect.

Share

Author

subscription section bg
Subscribe

Subscribe to the AAIF Briefing

Weekly signal on standards, governance, and the people building the future. No fluff. Just what matters.

About AAIF