Why cheaper AI can still create more expensive enterprise systems.
Last week, I wrote about a problem enterprises will increasingly face as they begin operating hundreds of AI agents:
Who owns the agent? What can it access? What is it allowed to do?
But once an organization gives an agent permission to operate, another question appears:
Is this agent economically worth running?
That question is becoming more important because something counterintuitive is happening.
AI inference is getting cheaper. Yet agentic AI systems can become more expensive to operate.
Gartner predicts that AI inference cost per agentic workflow will increase more than fivefold through 2028. Its explanation is what it calls the Inference Paradox: better unit economics make more sophisticated AI applications possible, but those applications consume more reasoning, more tokens and more inference overall. (Gartner)
Gartner separately estimates that agentic models can require 5–30 times more tokens per task than a standard GenAI chatbot. (Gartner)
So perhaps we are optimizing the wrong number.
The enterprise does not ultimately buy tokens.
It buys outcomes.
The real cost isn't the model call
A conventional AI interaction may look simple:
Question
↓
Model
↓
AnswerBut an agentic workflow may look more like this:
Business Request
↓
Plan
↓
Reason
↓
Retrieve Context
↓
Call Tool
↓
Call API
↓
Reason Again
↓
Validate
↓
Failure?
↙ ↘
Retry Continue
↓
Business OutcomeEach step can consume resources.
So the real economic boundary may be:
Inference
+ Context
+ Retrieval
+ Databases
+ APIs
+ Tools
+ Orchestration
+ Retries
+ Observability
+ Infrastructure
+ Human Escalation
────────────────────
True Workflow CostThis is where the enterprise measurement problem begins.
Most organizations can eventually answer:
“How much did we spend with our AI providers?”
Far fewer can answer:
“How much did it cost us to successfully resolve one customer issue, process one claim, investigate one incident or complete one engineering task?”
That distinction becomes increasingly important as AI moves from assistants into systems that actually execute work.
We are already seeing this at enterprise scale
McKinsey provides an interesting real-world example.
As of May 2026, the firm reported processing roughly five trillion AI tokens per month. But consumption was highly concentrated: approximately 10% of users accounted for around 65% of total token usage. McKinsey says it is now adding guidelines and guardrails as it prepares for autonomous agents consuming intelligence on behalf of users, and describes demand management as a permanent capability rather than a one-time cost exercise. (McKinsey & Company)
This is not only a consulting-company issue.
The FinOps Foundation's 2026 survey—based on 1,192 respondents representing more than $83 billion in annual cloud spending—found that 98% of respondents now manage AI spend, up from 31% two years earlier. AI cost management is also the number-one skillset teams say they need to develop. (FinOps Data)
The problems practitioners report are revealing:
visibility into AI costs, allocation of those costs to business units, and determining AI value or ROI. (FinOps Data)
The industry is therefore moving beyond:
“How much AI did we consume?”
toward:
“What value did that consumption create?”
Expensive AI isn't necessarily bad AI
Consider two hypothetical agents.
Agent A
Cost per execution: $0.20
Successful outcomes: 80%Agent B
Cost per execution: $0.30
Successful outcomes: 98%If we measure only inference cost, Agent A appears better.
But suppose failures trigger retries or ten minutes of employee intervention.
Agent B may actually have the lower cost per successful outcome.
Now consider:
Agent X
$1 AI cost
↓
Produces an internal summary
that nobody usesversus:
Agent Y
$8 AI cost
↓
Diagnoses an infrastructure incident
↓
Avoids 45 minutes of engineering workAgent Y consumed eight times more AI.
It may also have dramatically better economics.
This is why AI FinOps cannot simply become:
“Use fewer tokens.”
The more useful objective is:
Produce more useful business value for every AI dollar consumed.
Where companies can get this wrong
1. Optimizing model price instead of workflow economics
The cheapest model is not necessarily the cheapest system.
Model choice affects quality, retries, latency, human intervention and successful completion.
Gartner argues that routine, high-frequency work should increasingly be routed toward efficient smaller or domain-specific models, while expensive frontier intelligence should be reserved for tasks where the economics justify it. (Gartner)
The optimization problem therefore becomes:
Cost + Quality + Latency + Success Rate
—not cost alone.
2. Treating technical success as business success
An API returning:
200 OK
does not mean:
Customer issue resolved
An agent can execute successfully while producing something a human rejects.
A coding agent can generate code that never gets merged.
A claims agent can complete a workflow that still requires manual investigation.
Technical execution and business outcome are different things.
3. Ignoring the surrounding cost stack
Inference is easy to notice because providers give us a bill.
But production agents also consume retrieval, storage, databases, APIs, monitoring, orchestration, infrastructure and human review.
The economic boundary therefore needs to move from:
model
to:
workflow.
4. Cutting AI usage simply because spending rises
Consider:
AI expenditure +100%
Successful outcomes +250%
Manual effort -35%
Cycle time -50%Would reducing AI consumption automatically be the right decision?
Probably not.
The more useful question is:
What did the incremental AI expenditure produce?
Move the metric up the stack
I think AI economics will gradually mature through four levels.
Level 1 — Cost per token
Useful for procurement and model comparison.
Necessary, but increasingly insufficient.
Level 2 — Cost per agent run
Now include:
Models
+ Tools
+ APIs
+ Retrieval
+ Retries
+ InfrastructureBetter.
Level 3 — Cost per successful workflow
Now reliability matters:
Total workflow operating cost
─────────────────────────────
Successful workflowsBecause failed AI isn't free AI.
Level 4 — Cost per business outcome
Now technical telemetry connects to the process the organization actually cares about.
For example:
Financial services
Cost per claim successfully processed
Manufacturing
Cost per maintenance diagnosis completed
SaaS
Cost per customer issue resolved
Consulting
Cost per research or delivery task completed
Engineering
Cost per accepted PR or incident resolved
At this point, the CTO, CIO, CFO and business leader can finally discuss the AI system in the same economic language.
A practical framework: TRACE
Organizations don't need perfect ROI accounting before they start.
I would begin with five steps.
T — Track
Understand the complete execution path:
Agent
↓
Workflow
↓
Models
↓
Tools / APIs
↓
Data
↓
InfrastructureIf you cannot see what the workflow consumes, you cannot understand its economics.
R — Relate
Connect consumption to:
Business Unit
↓
Use Case
↓
Workflow
↓
Agent
↓
CostInstead of:
“Our AI providers cost $300,000 this month.”
move toward:
“The claims workflow consumed $42,000.”
And eventually:
“A successfully processed claim costs $X.”
A — Assess
Evaluate:
Cost + Quality + Latency + Success Rate
together.
That prevents cost optimization from quietly damaging the usefulness of the system.
C — Control
Now optimize the architecture.
Possible levers include:
- routing simple work to smaller models
- reducing unnecessary context
- caching repeated information
- improving retrieval
- limiting retries
- removing redundant tool calls
- controlling agent loops
- controlling sub-agent fan-out
- batching suitable workloads
Notice that many of these aren't finance decisions.
They are architecture decisions.
E — Evaluate value
Finally ask:
What did the workflow cost?
↓
Did it succeed?
↓
What outcome occurred?
↓
What was that outcome worth?Perfect attribution may be impossible.
But useful proxies usually exist:
hours avoided, cycle-time reduction, successful resolutions, additional throughput, defects prevented, revenue influenced, incidents resolved or manual steps removed.
That is already much more valuable than a token counter.
This is becoming an architecture problem
One of the more important shifts here is organizational.
AI economics cannot belong exclusively to Finance.
Nor should it belong entirely to the AI engineering team.
The FinOps Foundation has already broadened its mission from managing the value of cloud to managing the value of technology, while AI spend has become nearly universal among its surveyed practitioners. (FinOps)
Agent economics sits directly at the intersection of:
Architecture
+
Platform Engineering
+
AI Engineering
+
FinOps
+
Business OperationsThat means organizations may eventually need an economic telemetry layer resembling:
Agent Execution
↓
Workflow Trace
↓
Resource Consumption
↓
Cost Attribution
↓
Successful Outcome
↓
Business ValueNot another dashboard saying:
“You consumed 8 billion tokens.”
But something capable of saying:
“This workflow processed 16,900 cases successfully, cost $1.85 per successful outcome, required human intervention in 8% of cases, and these three execution steps account for 41% of its operating cost.”
That is a very different conversation.
And it is one a CFO can participate in.
What technology leaders can do now
You don't need to wait for perfect tooling.
Instrument workflows, not just models. Track retrieval, tools, APIs, retries and infrastructure alongside inference.
Define success before deployment. Every production agent should have an observable definition of successful completion.
Allocate AI consumption to use cases. Even imperfect attribution is better than one undifferentiated AI bill.
Measure cost and quality together. Otherwise optimization may make the system cheaper—and worse.
Introduce economic guardrails. Retry ceilings, model-routing policies, anomaly alerts and budget thresholds should increasingly become normal agent-platform capabilities.
Connect technical telemetry to business events. That is ultimately what makes cost-per-outcome possible.
The bigger shift
Enterprise AI started with:
Which model should we use?
Then:
How do we build agents?
Then:
How do we secure and govern them?
The next question is becoming:
Are these agents economically worth operating at scale?
The answer will not come from token pricing alone.
It requires connecting:
architecture → execution → cost → outcome → business value
And this may be the most important distinction:
The organizations that manage AI economics well will not necessarily be the ones that spend the least.
They may be the ones that know where spending more creates disproportionate value—and where it doesn't.
One question for technology leaders
If your AI spending doubled next quarter, could you tell your CFO:
Which business outcomes improved because of it?
If that answer is difficult today, I'd be interested in comparing notes on where the hardest gap is proving to be:
cost attribution, workflow visibility, outcome measurement, or connecting AI activity to business value?
— XTechStacks
Originally published in XTechStacks Weekly on LinkedIn on September 20, 2026. Read the original edition.

