Consider the banking, insurance, payments and travel industries. Now list the agentic use cases with the clearest returns: incident triage, capacity planning, fraud investigation, claims adjudication, settlement break resolution or irregular-operations rebooking. Every one of these use cases lives in those industries and needs to reason over the system of record to be worthy of any value. And most of these systems run on mainframes.
According to IBM, 44 of the world's top 50 banks use mainframes, which process 87% of all transactions worldwide. [1]
These are also the industries under the heaviest regulatory scrutiny of how data is handled: PCI DSS for cardholder data, GLBA and state privacy laws like CCPA/CPRA, and NYDFS cybersecurity regulations. These rules don't generally dictate where data must physically sit, but they impose demanding requirements for security controls, third-party oversight, and auditability.In response, many organizations have adopted internal policies and, sometimes, contractual commitments to clients under which production data does not leave their own environment, full stop.
So the constraint isn't an edge case to be handled later. It sits on top of the highest-value agentic workloads in the economy. If your agent architecture assumes the data can travel, you have excluded many of the world's largest enterprises from it.
Move the reasoning, not the records
The dominant agentic pattern assumes context flows toward the model and most enterprises desire to use frontier models in the cloud. Extract the data, chunk it, embed it, index it, retrieve it, put it in the prompt. Every stage of that pipeline is a copy, and each copy is a new place the data lives, a new thing to secure, and a new artifact someone can subpoena.
When copying is not an option, the architecture inverts. The model never receives the record set. It receives a task, a set of tools, a schema of the data. A local execution layer inside the security perimeter (aka the harness) does the actual reading, analysis and acting. The raw data stays put. Only results scoped to what the requester was already entitled to see travel back to the model.
It is worth being precise about why this is harder than it sounds, because agent frameworks quietly assume something the mainframe does not provide.In many coding-oriented agent environments, the agent can be given a shell. That shell can act as a general fallback: if no tool exists for what the model wants, it can list a directory, grep a log, curl an endpoint, or write a script on the spot.
There is no direct equivalent for accessing many mainframe-native resources. The data that matters lives in VSAM data sets, DB2 tables, IMS databases, CICS regions or SMF records, reached through subsystem interfaces, not through a filesystem you can walk. A UNIX shell does exist on the platform, but it does not reach much of that. An agent can hardly improvise its way to a mainframe record, and more capable model reasoning does not close that gap.
That sounds like a limitation. Treat it as a forcing function. Shell access can significantly broaden an agent's execution surface, although the host identity, permissions, sandboxing, and other controls still determine what it can do. On the mainframe, you cannot hand one over even by accident. Every capability the agent has must be deliberately built, named, scoped, and authorized, because many mainframe-native resources must be exposed explicitly.
This shifts where your engineering effort goes. Here that burden moves to the harness that sits on-premises between the model's reasoning and the capabilities that touch the system of record. It has to be considerably smarter than a retrieval front-end. It translates a plan expressed in general language into the specific calls a particular system will accept, scopes every one of those calls to what the requester is actually entitled to, and decides how much of what comes back ever crosses the boundary. The model supplies the reasoning. The harness supplies the judgment about what may be asked, of what, by whom, and what may be said back.
Let the host enforce the boundary
There is a lot of good work happening right now on agent authorization: token exchange, scoped credentials, policy engines, identity brokers, gateways that sit in front of tool calls and decide what an agent may invoke. If you are building agents on cloud infrastructure, you have a growing menu of these to choose from. However, they do not automatically understand mainframe-native controls.
That is not a criticism, it is a statement of scope. These solutions tackle OAuth scopes, cloud IAM roles, API endpoints, and Kubernetes service accounts. Without integration with mainframe security controls, they may not understand a RACF resource class, a data set profile, a DB2 secondary authorization ID, or a CICS transaction the user is or isn't permitted to run. Point one of them at a mainframe and it may be able to tell you which tool the agent may call but it cannot tell you whether the person behind that call is allowed to see the account the tool is about to read.
In environments that require end-user accountability, one pattern is for the call to execute under the requester's mainframe identity rather than inherit a shared component identity. That makes the component performing identity assertion the most security-sensitive thing in your architecture. It can potentially cause work to run as any user. Treat it accordingly: it should be small, single-purpose, and boring. A fixed, enumerable set of operations. No general command passthrough, no arbitrary data set paths, no code submitted from outside.
Read live when freshness matters
Because mainframe data is different and hard to access, traditional systems often operate on a copy of the data extracted periodically. Duplicate data can introduce two distinct risks: stale information and a larger security surface.
For operational decisions, stale data can be worse than no answer. An agent that confidently reports a position from a snapshot taken an hour ago has not saved anyone time; it has introduced a plausible wrong answer into an incident. Reading live can reduce those risks: there may be no separate vector store or re-indexing pipeline, less divergence between the agent's view and current system state, and no additional persistent copy of the source data.
It costs you something. Live reads mean a latency budget, careful rate limiting, and real thought about what happens when the agent's curiosity meets a production workload at peak. Those are engineering tradeoffs to manage, but they may be preferable to reconciling a stale copy of a core banking system.
Design for the audit before you design for the agent
Regulated industries may be asked to explain their agents after the fact, and that explanation is only as good as the evidence captured while the agent was working. The exact audit and control requirements vary by organization and regulatory framework, but teams may need to reconstruct what happened and under whose authority.
Mainframe shops had a head start here that went unappreciated for years. The mainframe’s instrumentation service (SMF) has been emitting structured records of system activity for decades. But it is only half the record. Depending on configuration, SMF can tell you that a data set was accessed at 03:14 under a given identity. It does not tell you that a human asked why a settlement file failed, that the model decided reading that dataset would answer the question, or that six other tools were called first. The model-side trace has the opposite gap: it can capture request context, outputs, and the tool-call sequence, but does not by itself show what happened on the mainframe.
In this architecture, the harness is the component that sees both, which makes auditing a first-class responsibility of the mediation layer rather than a by-product of it. It is the one place where the human's request, the model's plan, each tool call issued, the identity asserted for it, the system's response, and the final answer can be written down as a single causally-linked record. It also knows something nothing else in the architecture knows: what actually crossed the boundary.
Bring the hard environments into the standards
Everything above generalizes. Healthcare, defense, public sector, and organizations subject to data residency or transfer restrictions can face the same shape of problem: the highest-value reasoning targets the data you are least free to move.
At Geniez AI, we've been building under these constraints in production mainframe environments. One lesson has been how much of the architecture generalizes beyond mainframes: explicit capabilities, host-native authorization, live access where freshness matters, and traceability across the boundary.
Share
Author

Gil Peleg
Co-founder, CEO at Geniez AI



