Why cost per business outcome may matter more than cost per token.
For the last few years, enterprise AI cost conversations have revolved around a familiar number:
price per million tokens.
Teams compare model providers. They negotiate discounts. They shorten prompts. They add caching. They route simpler requests to cheaper models.
All useful.
But AI agents are changing the economics.
An agent doesn't necessarily make one model call and return an answer.
It can reason, retrieve data, invoke tools, call APIs, retry failed operations, ask another model for verification, launch sub-agents and repeat parts of the workflow before producing one useful result.
The result is a new cost stack:
Token cost → Agent-run cost → Workflow cost → Business-outcome cost
And I think enterprises are approaching a point where cost per token becomes one of the least interesting numbers in that chain.
The more useful question is:
How much did it cost the AI system to produce one successful business outcome—and was that outcome worth the money?
That is the emerging problem of AI Agent Economics.
The warning signs are already visible
Some unusually useful numbers appeared this week.
OpenAI disclosed that by mid-August, its median researcher was consuming more than $600 per day of inference at API-equivalent pricing through coding agents. Its heaviest research users exceeded $7,000 per day. The important part, however, is that OpenAI does not automatically regard this spending as waste: the company associates heavier agent use with more experiments and greater research throughput. (Business Insider)
That distinction matters.
Expensive AI isn't necessarily bad AI.
An agent costing $100 that creates $5,000 of value could be excellent economics.
An agent costing $2 that creates $0.50 of value is expensive.
We are therefore looking at the wrong denominator when we ask only:
“How many tokens did we consume?”
Agentic AI introduces a different cost structure
Imagine an enterprise claims-processing agent.
The employee sees something simple:
“Review this claim and recommend the next action.”
Behind the scenes, the workflow might look more like:
Claim received
↓
Document extraction
↓
Classification
↓
Customer-history lookup
↓
Policy retrieval
↓
Coverage validation
↓
Fraud-risk check
↓
Reasoning
↓
External API
↓
Exception detected
↓
Additional reasoning
↓
Tool retry
↓
Recommendation
↓
Human approvalEach layer can create cost.
Not only inference.
There may be:
model calls + context + reasoning + retrieval + vector/database queries + APIs + tools + retries + observability + infrastructure + human review.
And one of the most interesting characteristics of agent architecture is that the execution path can vary from one task to another.
Traditional software might execute a relatively predictable sequence.
An agent may decide that one claim needs five steps and another needs twenty.
That means architecture increasingly determines economics.
We can already see this in commercial agent pricing
Consider AWS DevOps Agent pricing.
AWS charges $0.0083 per active agent-second for investigations, evaluations and on-demand SRE tasks.
AWS's own examples illustrate how quickly the unit changes with usage:
A small team running ten monthly investigations costs about $39.84/month.
An active team running investigations and on-demand tasks reaches about $343.62/month.
Its enterprise example—500 incidents, 40 evaluations and 30 custom SRE-agent tasks—reaches $2,365.50/month before considering some connected-service costs. (Amazon Web Services, Inc.)
None of those numbers is inherently expensive or cheap.
If a $4 agent investigation prevents an engineer from spending 45 minutes diagnosing an infrastructure incident, it could be exceptionally valuable.
If hundreds of unnecessary investigations run continuously, the economics change.
Again:
cost without outcome tells us very little.
What companies are likely to get wrong
1. Optimizing model price instead of workflow economics
Imagine two models.
Model A
$0.20/run 80% workflow success rate
Model B
$0.30/run 98% workflow success rate
Model A looks cheaper.
But if failed executions trigger retries or human intervention, Model B could easily produce a lower cost per successful outcome.
The cheapest model is therefore not necessarily the cheapest system.
2. Ignoring agent amplification
An ordinary LLM request might involve:
User → Model → Answer
An agent workflow might involve:
User → Agent → Model → Tool → Model → API → Model → Tool → Retry → Model → Answer
Now add multiple agents:
Coordinator Agent
│
┌──────────────┼──────────────┐
▼ ▼ ▼
Research Agent Data Agent Validation Agent
│ │ │
Models Tools Models
Search APIs Rules
Data Data APIsA seemingly inexpensive unit price can multiply across the execution graph.
This week provided an extreme illustration: a large-scale OpenAI mathematical research experiment reportedly used about 10,000 concurrent agents, 2.7 million inter-agent messages and roughly 88 hours of work before verification. Whatever one thinks of the specific experiment, it demonstrates how radically agentic workloads can amplify computation relative to a conventional prompt-response interaction. (Business Insider)
Enterprise workflows will usually be vastly smaller.
The architectural principle remains.
3. Treating every successful run as equally valuable
Suppose two agents each cost $3 per execution.
One:
$3 → creates an internal meeting summary
The other:
$3 → resolves a customer support ticket
Same technical cost.
Potentially very different economic value.
This is why organizations eventually need to connect AI telemetry to business-process telemetry.
4. Cutting AI usage simply because spending increases
This could become another mistake.
If AI spending doubles while business throughput quadruples, cutting the budget may destroy value.
The better question is:
Did marginal AI spending produce greater marginal business value?
OpenAI's own internal usage provides a useful example precisely because extremely high agent consumption is being tolerated where the company believes research productivity justifies it. (Business Insider)
AI FinOps therefore cannot simply become:
“How do we reduce tokens?”
It needs to become:
“How do we maximize economic value per AI dollar?”
The metric needs to move up the stack
I think enterprise AI economics will evolve through four levels.
Level 1 — Cost per token
Useful for model procurement.
Model
↓
Tokens
↓
$Necessary.
But increasingly insufficient.
Level 2 — Cost per agent run
Now measure:
Model calls
+ reasoning
+ retrieval
+ tool calls
+ APIs
+ retries
+ infrastructure
────────────────
Agent-run costBetter.
Level 3 — Cost per successful workflow
Now include reliability:
Total workflow operating cost
─────────────────────────────
Successful workflowsThis exposes an important fact:
failed AI isn't free AI.
Retries, escalations and human intervention all have economic consequences.
Level 4 — Cost per business outcome
This is where the metric becomes useful to executives.
Examples might include:
Financial services: cost per claim processed or fraud investigation completed.
Manufacturing: cost per quality exception or maintenance diagnosis resolved.
SaaS: AI cost per support resolution, customer or AI-powered feature interaction.
Consulting: cost per research task or deliverable produced.
Engineering: cost per accepted pull request, incident resolved or development task completed.
Now AI expenditure can finally be compared with business value.
A practical framework: DABOG
Organizations don't need another giant AI transformation program to begin doing this.
They need better instrumentation and operating discipline.
I would use five steps.
D — Discover
First understand the AI estate.
Map:
Agents → workflows → models → tools → APIs → data services
Many organizations will eventually discover that the model invoice represents only part of the real operating cost.
A — Attribute
Move from provider-level spending:
OpenAI = $200K
toward:
Business Unit
↓
Use Case
↓
Workflow
↓
Agent
↓
Model / Tools / APIs
↓
CostWithout attribution, optimization becomes guesswork.
B — Benchmark
Do not benchmark cost alone.
Measure:
Cost + Quality + Latency + Success Rate
A cheaper workflow that fails more frequently may be economically worse.
This is where AI engineering and FinOps begin converging.
O — Optimize
Only after understanding the workflow should teams optimize.
Typical opportunities include:
- model routing
- smaller models for simple steps
- prompt/context reduction
- caching
- retrieval optimization
- unnecessary reasoning removal
- retry limits
- duplicate tool-call elimination
- sub-agent fan-out control
- batch processing
- workflow redesign
Notice something important:
Most of these are architecture decisions, not procurement decisions.
G — Govern
Finally, establish economic guardrails.
For example:
Agent budget
Cost/run threshold
Retry ceiling
Context limit
Tool-call limit
Cost anomaly detection
Business-unit showback
Outcome KPIThe objective is not to prevent agents from spending money.
It is to prevent them from spending money without producing proportional value.
Why this problem deserves enterprise attention
From an XTechStacks problem-selection perspective, AI Agent Economics scores unusually well:
DimensionAssessmentPain severityHighFrequencyExtremely HighCost impactHigh at scaleBusiness impactVery HighRiskMedium–HighUrgencyHigh and increasingWillingness to payLikely High once AI spend becomes materialCross-industry applicabilityVery HighRepeatabilityVery HighProductization potentialVery High
The frequency dimension is especially important.
This isn't a governance exercise performed once per year.
The economic event occurs every time an agent runs.
And successful AI adoption makes the problem larger.
The broader FinOps market is already showing the shift: recent reporting on the FinOps Foundation's 2026 survey says 98% of surveyed practitioners now manage AI costs, up substantially from two years earlier, while respondents characterize AI pricing as more variable and less transparent than conventional cloud spending. (MarketWatch)
Meanwhile, enterprise AI platforms are increasingly exposing token consumption, GPU utilization, model benchmarking and cost-control capabilities directly in their operating layers. (Express Computer)
That suggests AI economics is moving from an experimental concern toward an enterprise platform capability.
The architecture I expect to emerge
Eventually I think organizations will need something resembling an Agent Economics Observatory:
BUSINESS OUTCOME
▲
│
VALUE
▲
│
Business Unit → Workflow → Agent
│
┌─────────────┼─────────────┐
▼ ▼ ▼
Model Tools Data
│ │ │
Tokens APIs Retrieval
└─────────────┼─────────────┘
▼
COST
│
▼
COST PER SUCCESSFUL OUTCOMENot another dashboard showing:
“You consumed 8.4 billion tokens.”
But something capable of saying:
“This workflow processed 18,400 claims, successfully completed 16,900, required human escalation on 1,500, cost $31,000 to operate and avoided approximately X hours of manual work.”
That is a conversation a CIO and CFO can actually have.
What technology leaders can do now
Instrument workflows, not just models. Capture model, tool, retrieval, retry and infrastructure costs at the agent/workflow level.
Measure success explicitly. Every production agent should have a definition of successful completion.
Connect AI telemetry to business telemetry. A technically successful agent call is not necessarily a successful business outcome.
Set economic guardrails. Control runaway retries, excessive context, unnecessary tool calls and poorly designed agent loops.
Optimize value, not simply cost. A more expensive agent can be the better system if it produces substantially better outcomes.
The bigger lesson
We spent the first phase of enterprise GenAI asking:
Which model should we use?
Then:
How do we build agents?
The next question may increasingly be:
Are these agents economically worth running at scale?
And answering that question requires more than a token dashboard.
It requires connecting:
architecture → execution → cost → outcome → business value.
The organizations that build that capability early will have an important advantage.
They won't simply know how much AI they're consuming.
They'll know where AI is actually creating economic value.
And that may ultimately be the metric that matters most.
A question for technology leaders
If your AI spending doubled next quarter, could you tell your CFO which business outcomes improved because of it?
If the answer isn't clear, I'd be interested in comparing notes on where the hardest gap is today—cost attribution, workflow visibility, outcome measurement or ROI.
XTechStacks Newsletter team
Originally published in XTechStacks Weekly on LinkedIn on September 13, 2026. Read the original edition.

