TL;DR: The clearest message from Amsterdam was how much of agent engineering now depends on what sits around the model: protocols, identity, state, evaluation and workflows that keep human review manageable.
Across two days in Amsterdam, the same engineering questions kept surfacing: how MCP handles longer-running and distributed work, what identity an agent should carry, where state and context should live, how to test agents against real tasks, and what happens when coding agents can produce work faster than humans can review it.
MCP is growing far beyond tool calls
Connectivity becomes more important as agents take on longer and broader tasks, MCP co-creator David Soria Parra argued in his keynote, “MCP and the Era of Connectivity.” A coding agent can get useful feedback from a compiler or test suite, but agents working across research, finance or other knowledge-heavy tasks need access to external systems, data and services.That puts different pressure on the protocol.
David highlighted the stateless core introduced in the July MCP specification as one important change. Removing the assumption that session state stays attached to a particular server instance makes it easier to distribute and scale MCP services across infrastructure.
The operational effects of that shift were the focus of Shaun Smith’s keynote, “Getting to Stateless MCP: In Production.” Drawing on Hugging Face's experience running MCP remotely, he described problems with session management, timeouts and repeated protocol traffic. The July release addressed some of those issues by removing lifecycle handshakes, adding protocol-level caching to reduce repeated calls, and changing elicitations to require less client-server coordination.
Other sessions looked at the next set of problems around the protocol, including longer-running Tasks, conformance, server discovery, MCP Apps and how MCP sits alongside interfaces such as WebMCP and CLI tools.
MCP increasingly has to support distributed, asynchronous and long-running interactions, not just individual tool calls.
Agent identity needs more than a user or workload
Christian Posta's identity talk, “What IS an Agent's Identity?”, looked at whether we can reuse the identity mechanisms we already have.
Sometimes, yes. If an agent is working directly for a user in a tight loop, acting with that user's credentials can be enough. But that model starts to break down when an agent has only part of a user's authority, works autonomously, collaborates with other agents or encounters an action that needs separate human approval.
Take the case of workload identity. Mapping an agent to the Kubernetes workload running it can work initially, but becomes less useful if several agents share the same workload, or if an agent is suspended and later resumed somewhere else.
He argued for an identity that can remain attached to the agent or agent instance independently of the infrastructure currently running it. That identity can then become the anchor for authorization, attribution, policy and revocation.
"Who is running this process?" and "which agent is taking this action, with whose authority?" are not always the same questions.
As agents move between systems and take actions on behalf of people or organizations, identity and authorization have to travel with the work rather than being inferred solely from the process executing it.
Not all context belongs in the model window
Does a long-running agent need to stay running when there’s nothing to do? Clare Liguori explored that question in her keynote, “Reactive Agents: Your Agent Doesn't Need to Be Always On.” An agent might wait minutes for infrastructure to be provisioned, days for a person to respond, or weeks for an external event. Keeping a process running throughout that time ties the agent's state to compute that may barely be doing any work.
She laid out a reactive-agent model that replaces that waiting with hibernation. When there is no work to do, the process disappears and the agent becomes data – conversation history and state that can be loaded again when something happens. In an experimental Amazon project, she said the team was able to hold 75,000 hibernated agents in the memory of a Raspberry Pi because there was no running process for each one.
Bloomberg's talk, “The Unix Philosophy for AI Agents: File Systems as the Context Primitive,” asked what happens when the amount of information available to an agent is much larger than the context it needs right now?
The team explored using the file system as a common workspace instead of continually adding information to the model context. Information can leave the active context without disappearing, agents can retrieve it when needed, and outputs can become named artifacts that survive between steps or move between agents.
This approach gives different sources a common interface. Instead of adding another collection of tool descriptions, schemas and instructions to the model context every time a new source appears, an agent can navigate files and bring only the relevant material into its working context.
Agent hibernation and file-system offloads both show that persistent agent state and active model context do not have to be the same thing, and separating them can improve performance and speed while reducing cost.
Evals need to reflect the work agents actually do
The eval sessions in Amsterdam moved quickly away from generic benchmarks and toward the tasks agents are expected to perform.
GitHub's evals session, “Testing Agents and Their Tools: Offline Evaluation, Synthetic Tasks, and A/B Experiments,” described three layers of evaluation around MCP and agentic features: testing the MCP server itself, testing agents against controlled synthetic tasks, and evaluating behavior against real traffic.A protocol or tool-selection change can work correctly in isolation without telling you whether an agent completes the wider task successfully, while offline tests cannot reproduce everything that happens when real users interact with a system.
Datadog's session, “From Vibes To Data: Evaluating Agents on Your Real Work,” took the idea further by building evals around work its engineers actually perform. One lesson from the project was that more evals were not necessarily better. The team started getting useful signals from a relatively small set, then grew to a few hundred signals, before realizing that volume was difficult to manage.
Their next step was to use evidence from real engineering work to identify gaps in the eval set, then use the results to improve the context available to agents without changing the task simply to produce a better score.
The result of Datadog’s work was a more practical eval loop. First, makes the eval loop much more practical: observe the work, then find where agents struggle, and finally turn useful cases into tests with continuous learning and feedback to reshape the process.
Coding agents shift the bottleneck to human attention
Several coding agent talks highlighted that generating more output is useful only if the rest of the engineering workflow can absorb it.
In “I Was the Bottleneck, Not the Agent,” Datadog engineer Vincent Ysmal described running eight coding agents in parallel. The agents could work concurrently, but every completed task still came back to human engineers with code to read, deploy and test. The slowest component in the system became the person supervising them.
In his talk “AI Is Not Your Peer,” Dylan Ratcliffe described how his team moved much of its human review earlier in the process, reviewing the plan before an agent starts implementing it rather than trying to reconstruct every decision afterwards from a large generated pull request.
His reasoning was that a short plan is easier for another person to understand, challenge and refine than thousands of lines of generated code. Humans still make the architectural choices and tradeoffs, while implementation can be checked against the agreed plan.
Marlene Mhangami's keynote “A New Way To Build Software: How GitHub Is Evolving For a Human, Agent Future,” looked at how GitHub is adapting software workflows for human-agent collaboration. Agents can attach screenshots to issues as evidence of their work, stacked pull requests can make larger changes easier to review, and maintainers have more control over how automated contributions reach their projects.
Across all three talks, faster agent output put more pressure on review, coordination and maintainer controls.
From Amsterdam to San Jose
We'll be looking more closely at several of these talks in follow-up posts as more session recordings are released.
And many talks at AGNTCon + MCPCon North America in San Jose on October 22-23 will cover the same questions, with sessions across MCP, agent infrastructure, security, coding agents and the wider agentic stack. Use code COMMUNITY25 for 25% off registration
Share
Author




