Skip to newsletter
XTechStacksTECHNOLOGY · AI · STRATEGY · IMPACT
Menu
Subscribe to newsletter
XTECHSTACKS WEEKLY

September 20, 2026

Stop Measuring AI Agents by Tokens. Measure Cost per Business Outcome

Why cheaper AI can still create more expensive enterprise systems.

By Atul Yadav · 8 min read

Stop Measuring AI Agents by Tokens. Measure Cost per Business Outcome

Why cheaper AI can still create more expensive enterprise systems.

Last week, I wrote about a problem enterprises will increasingly face as they begin operating hundreds of AI agents:

Who owns the agent? What can it access? What is it allowed to do?

But once an organization gives an agent permission to operate, another question appears:

Is this agent economically worth running?

That question is becoming more important because something counterintuitive is happening.

AI inference is getting cheaper. Yet agentic AI systems can become more expensive to operate.

Gartner predicts that AI inference cost per agentic workflow will increase more than fivefold through 2028. Its explanation is what it calls the Inference Paradox: better unit economics make more sophisticated AI applications possible, but those applications consume more reasoning, more tokens and more inference overall. (Gartner)

Gartner separately estimates that agentic models can require 5–30 times more tokens per task than a standard GenAI chatbot. (Gartner)

So perhaps we are optimizing the wrong number.

The enterprise does not ultimately buy tokens.

It buys outcomes.


The real cost isn't the model call

A conventional AI interaction may look simple:

Question
   ↓
Model
   ↓
Answer

But an agentic workflow may look more like this:

Business Request
      ↓
     Plan
      ↓
    Reason
      ↓
Retrieve Context
      ↓
  Call Tool
      ↓
   Call API
      ↓
Reason Again
      ↓
  Validate
      ↓
 Failure?
  ↙    ↘
Retry  Continue
      ↓
Business Outcome

Each step can consume resources.

So the real economic boundary may be:

Inference
+ Context
+ Retrieval
+ Databases
+ APIs
+ Tools
+ Orchestration
+ Retries
+ Observability
+ Infrastructure
+ Human Escalation
────────────────────
True Workflow Cost

This is where the enterprise measurement problem begins.

Most organizations can eventually answer:

“How much did we spend with our AI providers?”

Far fewer can answer:

“How much did it cost us to successfully resolve one customer issue, process one claim, investigate one incident or complete one engineering task?”

That distinction becomes increasingly important as AI moves from assistants into systems that actually execute work.


We are already seeing this at enterprise scale

McKinsey provides an interesting real-world example.

As of May 2026, the firm reported processing roughly five trillion AI tokens per month. But consumption was highly concentrated: approximately 10% of users accounted for around 65% of total token usage. McKinsey says it is now adding guidelines and guardrails as it prepares for autonomous agents consuming intelligence on behalf of users, and describes demand management as a permanent capability rather than a one-time cost exercise. (McKinsey & Company)

This is not only a consulting-company issue.

The FinOps Foundation's 2026 survey—based on 1,192 respondents representing more than $83 billion in annual cloud spending—found that 98% of respondents now manage AI spend, up from 31% two years earlier. AI cost management is also the number-one skillset teams say they need to develop. (FinOps Data)

The problems practitioners report are revealing:

visibility into AI costs, allocation of those costs to business units, and determining AI value or ROI. (FinOps Data)

The industry is therefore moving beyond:

“How much AI did we consume?”

toward:

“What value did that consumption create?”


Expensive AI isn't necessarily bad AI

Consider two hypothetical agents.

Agent A

Cost per execution: $0.20
Successful outcomes: 80%

Agent B

Cost per execution: $0.30
Successful outcomes: 98%

If we measure only inference cost, Agent A appears better.

But suppose failures trigger retries or ten minutes of employee intervention.

Agent B may actually have the lower cost per successful outcome.

Now consider:

Agent X

$1 AI cost
↓
Produces an internal summary
that nobody uses

versus:

Agent Y

$8 AI cost
↓
Diagnoses an infrastructure incident
↓
Avoids 45 minutes of engineering work

Agent Y consumed eight times more AI.

It may also have dramatically better economics.

This is why AI FinOps cannot simply become:

“Use fewer tokens.”

The more useful objective is:

Produce more useful business value for every AI dollar consumed.


Where companies can get this wrong

1. Optimizing model price instead of workflow economics

The cheapest model is not necessarily the cheapest system.

Model choice affects quality, retries, latency, human intervention and successful completion.

Gartner argues that routine, high-frequency work should increasingly be routed toward efficient smaller or domain-specific models, while expensive frontier intelligence should be reserved for tasks where the economics justify it. (Gartner)

The optimization problem therefore becomes:

Cost + Quality + Latency + Success Rate

—not cost alone.

2. Treating technical success as business success

An API returning:

200 OK

does not mean:

Customer issue resolved

An agent can execute successfully while producing something a human rejects.

A coding agent can generate code that never gets merged.

A claims agent can complete a workflow that still requires manual investigation.

Technical execution and business outcome are different things.

3. Ignoring the surrounding cost stack

Inference is easy to notice because providers give us a bill.

But production agents also consume retrieval, storage, databases, APIs, monitoring, orchestration, infrastructure and human review.

The economic boundary therefore needs to move from:

model

to:

workflow.

4. Cutting AI usage simply because spending rises

Consider:

AI expenditure       +100%
Successful outcomes  +250%
Manual effort          -35%
Cycle time             -50%

Would reducing AI consumption automatically be the right decision?

Probably not.

The more useful question is:

What did the incremental AI expenditure produce?


Move the metric up the stack

I think AI economics will gradually mature through four levels.

Level 1 — Cost per token

Useful for procurement and model comparison.

Necessary, but increasingly insufficient.

Level 2 — Cost per agent run

Now include:

Models
+ Tools
+ APIs
+ Retrieval
+ Retries
+ Infrastructure

Better.

Level 3 — Cost per successful workflow

Now reliability matters:

Total workflow operating cost
─────────────────────────────
Successful workflows

Because failed AI isn't free AI.

Level 4 — Cost per business outcome

Now technical telemetry connects to the process the organization actually cares about.

For example:

Financial services

Cost per claim successfully processed

Manufacturing

Cost per maintenance diagnosis completed

SaaS

Cost per customer issue resolved

Consulting

Cost per research or delivery task completed

Engineering

Cost per accepted PR or incident resolved

At this point, the CTO, CIO, CFO and business leader can finally discuss the AI system in the same economic language.


A practical framework: TRACE

Organizations don't need perfect ROI accounting before they start.

I would begin with five steps.

T — Track

Understand the complete execution path:

Agent
 ↓
Workflow
 ↓
Models
 ↓
Tools / APIs
 ↓
Data
 ↓
Infrastructure

If you cannot see what the workflow consumes, you cannot understand its economics.

R — Relate

Connect consumption to:

Business Unit
      ↓
Use Case
      ↓
Workflow
      ↓
Agent
      ↓
Cost

Instead of:

“Our AI providers cost $300,000 this month.”

move toward:

“The claims workflow consumed $42,000.”

And eventually:

“A successfully processed claim costs $X.”

A — Assess

Evaluate:

Cost + Quality + Latency + Success Rate

together.

That prevents cost optimization from quietly damaging the usefulness of the system.

C — Control

Now optimize the architecture.

Possible levers include:

Notice that many of these aren't finance decisions.

They are architecture decisions.

E — Evaluate value

Finally ask:

What did the workflow cost?
            ↓
Did it succeed?
            ↓
What outcome occurred?
            ↓
What was that outcome worth?

Perfect attribution may be impossible.

But useful proxies usually exist:

hours avoided, cycle-time reduction, successful resolutions, additional throughput, defects prevented, revenue influenced, incidents resolved or manual steps removed.

That is already much more valuable than a token counter.


This is becoming an architecture problem

One of the more important shifts here is organizational.

AI economics cannot belong exclusively to Finance.

Nor should it belong entirely to the AI engineering team.

The FinOps Foundation has already broadened its mission from managing the value of cloud to managing the value of technology, while AI spend has become nearly universal among its surveyed practitioners. (FinOps)

Agent economics sits directly at the intersection of:

Architecture
     +
Platform Engineering
     +
AI Engineering
     +
FinOps
     +
Business Operations

That means organizations may eventually need an economic telemetry layer resembling:

Agent Execution
      ↓
Workflow Trace
      ↓
Resource Consumption
      ↓
Cost Attribution
      ↓
Successful Outcome
      ↓
Business Value

Not another dashboard saying:

“You consumed 8 billion tokens.”

But something capable of saying:

“This workflow processed 16,900 cases successfully, cost $1.85 per successful outcome, required human intervention in 8% of cases, and these three execution steps account for 41% of its operating cost.”

That is a very different conversation.

And it is one a CFO can participate in.


What technology leaders can do now

You don't need to wait for perfect tooling.

Instrument workflows, not just models. Track retrieval, tools, APIs, retries and infrastructure alongside inference.

Define success before deployment. Every production agent should have an observable definition of successful completion.

Allocate AI consumption to use cases. Even imperfect attribution is better than one undifferentiated AI bill.

Measure cost and quality together. Otherwise optimization may make the system cheaper—and worse.

Introduce economic guardrails. Retry ceilings, model-routing policies, anomaly alerts and budget thresholds should increasingly become normal agent-platform capabilities.

Connect technical telemetry to business events. That is ultimately what makes cost-per-outcome possible.


The bigger shift

Enterprise AI started with:

Which model should we use?

Then:

How do we build agents?

Then:

How do we secure and govern them?

The next question is becoming:

Are these agents economically worth operating at scale?

The answer will not come from token pricing alone.

It requires connecting:

architecture → execution → cost → outcome → business value

And this may be the most important distinction:

The organizations that manage AI economics well will not necessarily be the ones that spend the least.

They may be the ones that know where spending more creates disproportionate value—and where it doesn't.


One question for technology leaders

If your AI spending doubled next quarter, could you tell your CFO:

Which business outcomes improved because of it?

If that answer is difficult today, I'd be interested in comparing notes on where the hardest gap is proving to be:

cost attribution, workflow visibility, outcome measurement, or connecting AI activity to business value?

— XTechStacks

Originally published in XTechStacks Weekly on LinkedIn on September 20, 2026. Read the original edition.

Share on LinkedIn ↗

Keep a clearer view of what’s next.

Subscribe to XTechStacks Weekly for practical technology insight.