Key Takeaways:
- Context growth, redundant reasoning, agent fan-out, refinement loops, and overuse of frontier models can drive unnecessary AI token costs.
- Spend can be controlled through deterministic workflows, execution limits, model routing, caching, context management, and usage monitoring.
- Measuring cost per completed outcome helps teams connect AI spend to business value and identify where further optimization is needed.
The push to use AI is coming from every direction. Organizations are encouraging engineers to experiment, automate, and put AI into production with a “go, go, go” mentality. As a result, AI usage is growing fast, and so are the bills. According to Gartner, worldwide AI-optimized infrastructure as a service (IaaS) spending is projected to grow 96% through 2026, reaching $42 billion.
Frontier models are expensive to run, yet today’s prices don’t fully reflect their true cost. Much of that cost is still subsidized, including through enterprise contracts. As those subsidies begin to go away, relying on frontier models for every task will become more expensive. Analysts say that, by 2028, the frontier model cost per completed task will more than quadruple as AI vendors move toward profitability.
At the same time, 95% of enterprise AI still runs on the priciest frontier models even for simple tasks. This is where AI inference costs come in. They are the costs an organization incurs each time a model processes input and generates output in production. And while inference can become one of the largest ongoing expenses of an AI program, a surprising amount of that spend may be unnecessary.
Netflix senior engineer Tejas Chopra estimates that as much as 90% of the tokens sent to large language models are redundant. In January 2026, he created Project Headroom to reduce that waste by pruning agent instructions before they reach the model. By May, the project was estimated to have saved its users around $700,000, freeing up roughly 200 billion tokens for other work.
If a small division at Netflix can save such a huge amount of tokens and spending, imagine the volume of context tokens being wasted across the enterprise when that kind of pruning isn’t in place.
McKinsey says that organizations that take a more deliberate approach to AI consumption can save 20-30% on their AI costs. The key is understanding where the waste comes from and putting the right controls in place before it becomes part of the architecture.
Here, you’ll find five ways AI spending can add up quickly and five strategies to keep those costs in check.
Five Ways AI Inference Costs Get Out of Control
1. Context grows with every LLM exchange.
Every LLM chat or LLM-based agent action assembles a context window from enterprise documents, tool schemas, system instructions, and prior conversation. Each subsequent exchange reingests that window along with whatever the last step added.
At its core, the exchange acts as a growing loop. In this design, the same request made at the fourth request uses significantly more tokens than step one for identical work. Accuracy degrades as the window grows, so the organization pays more and gets less.
2. Reasoning tokens are billed and hardly seen.
Chain-of-thought models generate internal reasoning that is billed at output rates and doesn’t appear in the response. Organizations must either actively monitor token usage at the API level, or this cost stays invisible until the invoice arrives. On complex tasks, it can exceed the visible output several times over.
3. Agent chains fan out and recurse without a ceiling.
One agent calls another, which spins up three more, each opening its own session. The chain has no way to track how deep it has gone or how many other agents are running alongside it. A fan-out storm doesn’t produce any errors and shows normal response times, but shows up prominently in a bigger bill.
For example, a customer sends an email asking why their invoice went up. The support agent picks it up but can’t answer the question, so it hands the request to a billing agent. The billing agent needs data it doesn’t have, so it calls a contract agent, a usage agent, and an identity agent. That’s what fan-out looks like: one handoff turns into several sessions, and each one loads its own copy of the customer file using tokens along the way.
Then the contract agent decides the customer record looks outdated and calls the identity agent again. The identity agent has now run twice for the exact same request.
The customer eventually gets an answer and is happy with it. But over the source of the exchange, their email has engaged with five agents and triggered a dozen model calls. This was one customer, a single request. Many organizations today are being billed running the same pattern across every email in the support queue every day, throughout the month.
4. Refinement loops run past the point of value.
Agents working through open-ended objectives keep improving work that was already solid. “Keep refining until it’s right.” “Check this against everything we’ve covered.” Each pass re-reads the accumulated context, re-pays for every token in it, and has no objective threshold at which the work is declared good enough.
Retries, re-prompts, and repeated tool calls re-solve problems the system already paid to solve. Gartner places most of the agentic bill in redundancy rather than useful work, and identifies context accumulation as the primary driver.
5. Routine tasks default to frontier models.
Classification, extraction, formatting, and lookups get sent to the largest model because that is what the integration was pointed at initially. The default sends them to the provider’s newest flagship, priced at the top of its range with the strongest capabilities. But this is overkill for a lot of tasks. The work could be sufficiently completed by an older model for much cheaper.
The price gap between tiers is significant, and it can widen as vendors release new frontier models. For example, Anthropic’s model pricing (at the point of publishing), shows a fivefold difference between Claude Sonnet 5 (an earlier model), which costs $2 per million input tokens and $10 per million output tokens, and Claude Fable 5.1 (the newest model), which costs $10 and $50, respectively.
The AI Token Trap: Why the Real Cost of AI Isn’t What You Think
Learn MoreFive Controls That Bring the AI Spend Back
Each cause is addressed at a different layer, which is why a single lever rarely moves the bill far.
Table 1. AI inference cost drivers and the controls that stop them
| Cause | Control | Where it acts |
| Routine work on frontier models | Keep deterministic work out of the model and right-size the rest | Architecture |
| Unbounded workflow spend | An execution envelope set at instantiation | Before the first call |
| Chain fan-out and recursion | Limits on request rate, token count, tool calls, and context | At the model call |
| Context accumulation and reasoning overhead | Routing, caching, compression and output limits | At the model call |
| Rework and unclear value | Cost per completed outcome | After execution |
1. Keep deterministic work out of the model.
If a task can be done with a simple rule, lookup, or piece of code, there’s no reason to spend inference on it. Use deterministic programming for that work and save model calls for tasks that require reasoning.
Apply the same principle to context and knowledge. An agent shouldn’t have to keep re-reading a growing context window to access information it already needs. Keep reusable knowledge and agent memory available outside the model’s context window, then bring in only what the agent needs for the task at hand. That reduces repeated tokens while giving agents a more consistent view of the work they are doing.
The OneReach.ai GSX platform’s Communication Fabric provides a shared event model across channels and sessions, creating a “super-session” that gives agents continuity across interactions, allowing any type of session to be recorded reliably as memory. Its Agent Registry also provides a shared layer for managing agents and their access, so agents can collaborate across channels, sessions, and time without each interaction having to rebuild the full context from scratch or to flood the context windows with unnecessary waste. This acts in a similar function to the Netflix example cited earlier.
2. Set the spend ceiling before the workflow starts.
Every workflow should have a clear limit on how much it can spend before it starts. The right ceiling depends on the task: routine, predictable workflows can have tighter limits, while complex workflows may need more room. Start with the expected cost of a successful run, add some headroom for retries and unexpected steps, and adjust the limit as you collect real usage data.
These ceilings can sit at the session and workflow level and act similar to circuit breakers in an electrical run. If a workflow reaches its limit, it automatically stops before the spend keeps growing.
As Robb Wilson, CEO and co-founder of OneReach.ai, says about agentic cost bloat, “The controls have to exist from the first call, never after the first invoice.”
3. Limit request rate, token count, tool calls, and context
The model provider exposes several controls that can help keep inference spend in check. For example, OpenAI provides rate limits, output-token limits, tool-call limits, and context truncation controls.
Table 2. OpenAI controls that can limit inference spend
| Control | What it does | What it helps to prevent |
| Rate limits | Sets limits on requests and tokens over a given period | Sudden spikes in model usage |
| max_output_tokens | Sets the maximum number of tokens a response can generate, including reasoning tokens | Overly long and expensive responses |
| max_tool_calls | Sets the maximum number of built-in tool calls that can be processed in a response | Excessive tool use within a single run |
| Context truncation | Limits how much conversation history is retained when a context grows too large | Paying to carry an unnecessarily large context from turn to turn |
OpenAI also exposes usage data for input, output, cached input, and reasoning tokens, giving teams visibility into what is actually driving consumption. Similarly, Google’s usage data reports input, output, thinking, cached, tool-use, and total tokens, while Anthropic’s models report input, output, cache write, cache hit, and refresh tokens.
To address input tokens’ domination of the cost structure due to repeated context transmission (a “communication tax”), Gartner recommends minimizing repeated context loading through AI usage gateways, caching mechanisms, or redesigned interaction protocols.
4. Route every call to the cheapest applicable model.
Two enterprise cases published in August 2026 provide good examples of how this works.
AT&T cut the cost of coding and some other advanced AI tasks by as much as 56% by routing employee queries to cheaper models where appropriate, with performance quality declining by only 2%.
Databricks reports that its AI Gateway Smart Router consistently reduces average task cost by more than 30% while roughly matching the quality of the most expensive model in the working set. Tuning harness and caching settings produced an almost 50% reduction in the number of generated tokens and associated costs, with no observed quality degradation for developers.
Caching is the fastest win in this layer. McKinsey reports that prompt caching reduces repeated input-token costs by up to about 90%, especially for retrieval-augmented generation (RAG) and agents with large, stable prefixes.
OneReach.ai’s GSX platform brings this approach to runtime through its Cognitive Orchestration capability, which decides where a given call goes while the work is running. It allows teams to optimize for both cost savings and model efficacy at the solution level, or even for a specific session or run within each solution.
5. Measure cost per completed outcome.
A monthly AI provider bill tells you what you spent, but it doesn’t tell you which agent spent it, on what task, or whether the work was worth the tokens. Tag agents at build time into your registry, so every step carries its team, region, and use case before it runs. Meter at execution, so token counts, model calls, and tool invocations are captured as they happen. Then report those costs against the result the agent was built to deliver.
Knowing the value per outcome can be difficult, but even a rough estimate of the value per successful execution can help determine whether a production agent is economically viable. Over time, as you collect more data, you can develop more detailed and accurate values.
McKinsey says as CIOs build up their FinOps muscles for AI, they need to prioritize establishing a centralized AI control plane that provides a single source of truth for spend, usage, performance, and business outcomes (e.g., cost per resolved claim, cost per completed onboarding, cost per closed ticket), not just token consumption.
Know What You’re Paying For
The economics of AI are getting harder. IDC projects that agent use across the G2000 will increase tenfold by 2027, while token and API call loads rise a thousandfold. By 2029, IDC expects more than a billion active agents worldwide, executing 217 billion actions a day and consuming 3.7 trillion tokens daily.
This means cheaper inference doesn’t automatically mean lower AI spend. At enterprise scale, how agents are designed, routed, constrained, and measured will matter more than the price of any individual token.
That makes agentic FinOps less about cutting the bill and more about understanding what the bill is buying: what ran, what it cost, what it delivered, and whether the next run is worth paying for. Like human resources or infrastructure, organizations need to weigh their investments and determine if they are within targeted margins or even net positive. Once teams can answer these questions, they can grow their agent estate without being upside down on the economics.
Track every token. Own every dollar spent.
GSX AI Agent FinOpsFAQs
- What are the biggest drivers of AI inference costs?
The biggest drivers include unnecessary context, reasoning tokens, unbounded agent chains, refinement loops, and using expensive frontier models for routine work. These costs can quickly compound as agents make more calls, carry more context from one step to the next, and repeat work that has already been done.
- How can enterprises reduce AI inference costs without hurting quality?
Start by keeping deterministic work out of the model, then route each task to the cheapest model that meets the required quality level. Set limits on request rates, token counts, tool calls, and context for agent workflows, and use caching and context compression to reduce repeated work. Evaluate quality before and after each change to make sure cost savings don’t come at the expense of results.
- What is the best way to measure the cost of AI agents?
Don’t measure AI spend only by tokens or monthly provider bills. Track each agent’s model calls, token usage, tool invocations, team, region, and use case, then connect that spend to a completed business outcome. Metrics such as cost per resolved claim, cost per completed onboarding, or cost per closed ticket show whether an agent is actually delivering value for its cost.