Stateless MCP eliminates fixed sessions on AWS servers
AWS detailed how the latest version of the Model Context Protocol removes sessions at the protocol level, letting any instance serve any request and simplifying the horizontal scaling of agent servers.
AWS published this week, on the AWS Architecture Blog, an analysis of how the latest changes to the Model Context Protocol (MCP) specification affect the deployment of remote servers. InfoQ reported the story on September 25, 2026: the core of the change is the removal of protocol-level sessions, which eliminates the requirement for sticky sessions and for a session store shared across instances.
For anyone who has already run an MCP server in production, this directly addresses one of the most annoying points of the architecture: maintaining session affinity (routing based on Mcp-Session-Id) requires extra routing layers, and synchronizing session state across replicas tends to become a headache in scenarios with aggressive autoscaling or blue-green deploys.
What the specification removes (and what it adds)
The new MCP version eliminates the initialize/initialized handshake and the Mcp-Session-Id header, which until then identified a persistent session between client and server. Without this handshake, a request can be routed to any available instance behind a conventional load balancer, with no special affinity logic.
In place of the mandatory handshake, the specification introduces an optional operation, server/discover, for clients that need to know the server's capabilities before firing tool calls. According to the authors of the AWS post, Anand Komandooru, Steven DeVries, and Haleh Najafzadeh, this replaces the persistent session model with a request-response model closer to what's already used in traditional REST APIs.
Another central piece of the change is what the specification calls MRTR: instead of the server keeping an open stream to be able to initiate requests to the client during a multi-step interaction, the flow now uses input_required type responses followed by new requests from the client. In practice, this trades a long-lived connection for a sequence of short, independent calls, which is exactly the pattern that serverless infrastructure and conventional load balancers handle well.
Why this reduces infrastructure cost and complexity
The practical point raised by the AWS authors is that teams currently running an MCP server in production can eliminate components built solely to sustain protocol-level sessions: routing with session affinity becomes conventional routing, and the storage used exclusively to hold MCP session state is no longer necessary. This doesn't mean the application becomes stateless altogether, only that the protocol no longer imposes that obligation. As Michael Madsen summarized, commenting on the specification on LinkedIn and quoted by InfoQ:
"The protocol is stateless. Your application doesn't have to be."
In other words: if your application needs to remember context between calls (conversation history, the result of a previous tool), that state remains your responsibility, except now it lives explicitly in the application layer, and is no longer imposed by MCP's transport. This gives you freedom to choose where to store that state (database, cache, or not storing it at all, if the use case allows), instead of being forced to keep connections alive and sessions pinned to a specific instance.
A direct effect of this change is that the authors now identify AWS Lambda as a deployment option that fits this request-response model, since the protocol no longer requires persistent session connections. For those who currently avoid Lambda for MCP servers precisely because of the stateful nature of sessions, this reopens the door to a pay-per-invocation cost model instead of keeping instances always active to preserve session.
Routing and observability get their own headers
The specification also adds the Mcp-Method and Mcp-Name headers, designed to let gateways perform routing and throttling based on the method and the name of the tool being called, without needing to inspect the request body. This is relevant for anyone operating multiple MCP servers behind a single gateway who needs to apply rate limiting per tool or route specific calls to specialized instances.
For distributed tracing, the specification adopts W3C Trace Context, which makes it easier to integrate MCP calls with existing observability stacks based on OpenTelemetry. And for caching, the ttlMs and cacheScope controls give the server a standardized way to indicate how long and in what scope a response can be cached, which helps reduce cost in scenarios with repeated calls to tools whose results don't change often.
What gets harder: stream resumability and idempotency
Not everything is gained without a trade-off. The specification removed stream resumability, a feature that allowed a client to resume an interrupted operation from where it stopped. Without it, an interrupted operation needs to be restarted from scratch by the client, which InfoQ points out as a factor that increases the importance of ensuring idempotency in tool calls that produce side effects (writing data, triggering an external action, and so on). Anyone building MCP tools that alter external state now needs to treat retries as a normal scenario, not an exception, and design tools so that a second execution of the same call doesn't duplicate the effect.
Migration: coexistence between old and new versions
The transition isn't instantaneous. The Apify project, cited in the AWS post, is implementing stateless support in parallel with its existing sessionful MCP server, with conformance tests covering both versions of the protocol simultaneously. This gives a hint of what migration should look like in practice for anyone running MCP servers in production today: it's not a clean break, it's running both versions in parallel until traffic from legacy clients disappears.
AWS itself recommends tracking the protocol version at the gateway and keeping session infrastructure in place until legacy traffic has been eliminated, rather than turning everything off at once. The MCP project, in turn, established a feature lifecycle policy that defines a set migration period for deprecated capabilities, giving predictability to those who need to plan when they can actually retire session infrastructure.
AWS has mapped these changes to its Well-Architected guidance for agentic AI, covering monitoring, tracing, security, and tool integration, which signals that the protocol change is already being treated as a reference architecture standard, not as an isolated implementation detail.
What remains open
The stateless specification solves the session affinity problem at the protocol level, but it pushes back onto whoever builds the application the decision of how and where to store context between calls when the use case requires conversation or execution memory. Teams that today rely heavily on persistent sessions for business logic will need to design this state layer explicitly, likely with their own cache or database, instead of depending on MCP's transport for it. And as long as legacy MCP clients remain in production, the promise of infrastructure simplification is only partially realized: session infrastructure keeps existing until legacy traffic reaches zero.
Translated from the Brazilian Portuguese original · Read the original
Supabase has 16,000 databases with exposed personal data, research finds
A survey by UpGuard found names, addresses, passwords and tokens publicly accessible in thousands of projects hosted on the platform, almost always due to a configuration mistake by the developer themselves.