I already think per-action pricing is a bad fit for durable execution. For AI agents, I think it’s much worse.
The reason is simple:
With agents, the distributed systems machinery is often hidden inside the agent harness.
The developer doesn’t explicitly write every state transition, retry, callback, timer, tool invocation, checkpoint, resume, or recovery path. The framework does.
So if the platform charges per “action,” the person paying the bill may have little way to predict how many actions a run will trigger. That’s a disastrous pricing model for something as dynamic as an agent.
Where per-action pricing actually works
Per-action pricing works well for basic cloud infrastructure. It makes sense when an “action” is something a developer explicitly wrote code to trigger, and can therefore count, or at least estimate, in advance.
That’s true for serverless functions, API gateways, queues, and database operations, which are often priced this way, and it makes sense. The developer’s code is the source of truth for how many actions a run will generate, so reading the code tells you almost exactly what you’re going to pay. The unit is small and uniform, and the platform’s own costs scale with it too, so the incentives line up on both sides. You pay for what you triggered, and the platform charges for what it had to run.
Workflow orchestration / durable execution is where this model becomes harder to reason about, because an action the code triggered can result in multiple additional actions that are difficult to predict in advance.
An agent is not a predictable workflow
Traditional workflows are usually fairly easy to reason about. You define the steps: call service A, then service B, wait for approval, continue to service C. Even if the runtime makes the workflow durable underneath, the developer still has a decent mental model of the execution path.
Agents are different.
An agent may decide to call one tool on one run and fifteen tools on the next. It may inspect a result and decide to try again. It may hand work to another agent. It may ask for clarification. It may pause for a human. It may discover new information halfway through the task and completely change its plan.
That non-determinism is part of the value. But it also means the developer often doesn’t know the execution path in advance.
Now put per-action pricing underneath that. You’re effectively asking customers to predict the internal execution behavior of a system whose behavior is intentionally dynamic.
What exactly am I paying for?
This gets confusing very quickly. Suppose I ask an agent:
Investigate this customer issue and resolve it if you can.
The agent might:
- reason about the request,
- query a database,
- call an API,
- wait on another service,
- retry a failed tool,
- invoke another agent,
- checkpoint state,
- pause for human approval,
- resume later,
- call three more tools,
- and eventually complete.
How many billable actions was that? Five? Twenty? Fifty?
And more importantly:
How would I know before I ran it?
With agent frameworks, the application developer may not even be the one creating some of those execution primitives. The harness may automatically checkpoint. The runtime may retry. A tool wrapper may generate additional transitions. An orchestration layer may break one logical operation into several durable steps. Recovery after failure may introduce even more work. All of that may be completely invisible from the agent code.
So now the pricing model depends on implementation details several layers below the thing the customer actually built. That’s a bad abstraction.
The harness hides the thing you’re billing for
This is the part that makes agent pricing fundamentally different. In a normal distributed application, I can at least inspect the code and reason about how many operations I’m performing.
With an agent, much of the execution behavior is delegated to the framework. The harness decides how to coordinate tools. The durable runtime decides how to checkpoint and recover. The orchestration system decides when to suspend and resume. The model itself may decide what path to take.
Those layers are supposed to remove complexity from the developer. But if your bill is tied to actions generated inside those layers, then the customer is financially exposed to complexity they may not be able to predict in advance.
That’s the opposite of abstraction. A good abstraction hides implementation details. Per-action pricing turns hidden implementation details into financial variables.
Agent behavior is already hard enough to predict
AI infrastructure already has plenty of variable costs. Token usage varies. Model choice matters. Context windows grow. Agents may loop. Tool calls are dynamic. External APIs have their own pricing.
Teams are already trying to figure out how much an agent task actually costs. Durable execution shouldn’t add another opaque meter underneath it. Imagine trying to answer a CFO asking:
What does it cost us when this agent handles a support case?
And the answer is:
It depends how many internal durable actions the agent framework happens to generate.
That’s not useful. The business didn’t ask to execute 37 durable state transitions. It asked the agent to handle a support case. The pricing unit should get as close as possible to that level of abstraction.
Failures make the model even worse
Durability exists because failures happen. A tool times out. A downstream API returns 500. A worker crashes. A process restarts. An agent needs to resume after waiting six hours for a human.
Those should be runtime concerns. The user’s intent hasn’t changed. The task is still:
Resolve this support case.
But under per-action pricing, recovery can generate additional billable activity. So the system becomes less predictable precisely when the environment becomes less reliable. That’s a terrible property for infrastructure pricing.
A customer shouldn’t have to wonder:
Did that outage just increase my bill?
Reliability mechanisms should protect the customer from failure. They shouldn’t make the economics of the workload harder to understand.
Price the agent execution instead
The cleaner model is to charge for the thing the customer actually initiated:
the agent execution.
An agent starts a task.
It reasons.
It invokes tools.
It waits.
It retries.
It recovers.
It gets human input.
It resumes.
It eventually finishes.
That is one execution.
The internal machinery required to make that execution durable should be the platform’s concern. This creates a unit people can actually reason about.
“We ran 500,000 agent executions this month.” That’s understandable. “We generated 84 million billable actions across 500,000 agent executions.” That’s infrastructure trivia.
Of course, not every agent execution has the same cost. One might finish in two seconds. Another might run for a day and invoke hundreds of tools.
That’s fine. A long or tool-heavy execution can still be priced higher, through duration-based tiers, a metered rate on compute or wall-clock time past some threshold, or a cap that kicks in for extreme cases. What makes that different from per-action pricing is what’s being measured. Duration and compute are things a customer already has a feel for. They know whether they’re kicking off a quick lookup or a long-running investigation, and they can watch the meter move while it runs. An internal action count can be harder to watch and estimate against, because some of that activity is generated by the harness rather than directly by the application code. The primary pricing abstraction should still match the thing the customer believes they’re running. It just doesn’t have to be a flat rate to do that.
Pricing should move up the abstraction stack
This is really the bigger principle. As infrastructure moves up the abstraction stack, pricing should move with it.
Cloud providers exposed compute, so we paid for compute. Serverless exposed functions, so we paid for invocations and duration. Durable execution exposes workflows. Agent platforms expose agent runs and tasks.
If a user is thinking in terms of agents, but the invoice is thinking in terms of internal state-machine operations, there’s a mismatch. And the more autonomous agents become, the worse that mismatch gets.
The whole point of an agent harness is that I don’t need to micromanage every transition. The whole point of a durable runtime is that I don’t need to micromanage recovery. So don’t make me understand those things just to understand the bill.
Hidden complexity shouldn’t define the bill
I think this is going to matter a lot as agents move from experimentation into production. Enterprises will want predictable economics.
They’ll want to understand the cost per task. Cost per customer interaction. Cost per investigation. Cost per completed business process.
Those are useful numbers. “Cost per internal durable action” usually isn’t.
If the user has no realistic way of predicting how many billable actions a run will trigger, then “pay for what you use” stops being a benefit. Because the user doesn’t actually know in advance what they’re using.
For agents, the pricing unit should be the execution.
One agent task starts.
The runtime does whatever it needs to do.
The task completes.
Everything underneath that should stay where it belongs: outside the primary billing abstraction.
Share
Author

Yaron Schneider
Yaron is the Co-Founder of Diagrid and Chair of the Workflow Working Group at the Agentic AI Foundation
View All Posts



