AI Is Getting Cheaper. So Why Might Enterprise AI Costs Keep Growing?

AI token costs have fallen sharply, but that does not necessarily mean enterprise AI costs will fall with them. This article explores how model efficiency, compute constraints, provider economics and rising consumption could shape the future cost of enterprise AI.

AI Is Getting Cheaper. So Why Might Enterprise AI Costs Keep Growing?
Modern glass office building with reflective panels, representing enterprise AI infrastructure, compute capacity and the future cost of AI.

I was asked recently whether the cost of AI is likely to go up or down over the next few years. More specifically, the question was whether AI token costs will continue to fall as models and infrastructure become more efficient, or whether increasingly capable models, enormous infrastructure investment and growing demand for compute could push costs in the other direction.

It sounds like a simple question, but the answer is less straightforward.

There is strong historical evidence that the cost of accessing a given level of AI capability has fallen dramatically. At the same time, AI adoption is expanding, usage patterns are changing, infrastructure requirements are growing and enterprise AI pricing is becoming more complex.

There is also a separate question about whether the prices users pay today fully reflect the long-term economics of providing AI at scale.

That creates an important distinction between the unit cost of AI, the cost of supplying AI, and the total amount organisations ultimately spend on it.

Those measures do not necessarily move in the same direction.

The Cost of Equivalent AI Capability Has Fallen Quickly

One of the clearest examples comes from the Stanford Institute for Human-Centered Artificial Intelligence’s 2025 AI Index.

Stanford reported that the cost of querying a model achieving approximately GPT-3.5-level performance on the MMLU benchmark fell from around US$20 per million tokens in November 2022 to US$0.07 per million tokens by October 2024. That represents a reduction of more than 280 times in less than two years.

Stanford also reported significant improvements in the underlying economics of AI hardware. Machine-learning hardware performance increased by around 43% annually, costs at a given level of performance fell by around 30% per year, and energy efficiency improved by approximately 40% annually over the period examined.

The significance of this change is substantial. A level of AI capability that was relatively expensive in 2022 could be accessed for a fraction of the cost less than two years later.

There is, however, an important qualification.

The Stanford comparison looks at the cost of achieving a comparable level of capability. It does not tell us what the newest or most capable AI model will cost three or five years from now.

The Capability Frontier Keeps Moving

AI pricing is not simply the same product becoming cheaper every year. The product itself keeps changing.

As smaller and lower-cost models improve, new models appear with stronger reasoning, larger context windows, multimodal capabilities, tool use and increasingly sophisticated forms of automation.

A task that once required one of the most capable models available may later be handled by a smaller and cheaper model. At the same time, a new category of higher-capability models can emerge above it.

There are therefore two different questions hidden inside any discussion about the future of AI costs.

The first is: what will today’s level of AI capability cost in the future?

The second is: what will the most capable AI available at that future point cost?

Historically, the answer to the first question has often been “substantially less”. The second is much harder to predict.

Current AI pricing also illustrates why there is no longer one meaningful “price of AI”. Anthropic’s current API pricing, for example, differentiates between models, input and output tokens, prompt caching, batch processing and some forms of tool use. Batch processing is currently discounted relative to standard processing, while cached input is priced differently again. Some server-side tools can also introduce separate usage-based charges.

I have covered the mechanics of these charging structures in more detail in Enterprise AI API Pricing: Token Cost Modelling and Budget Considerations.

Enterprise AI pricing is therefore becoming more segmented rather than converging around a single token rate.

But Are Today’s AI Prices Sustainable?

There is another side to the argument.

Falling inference costs do not necessarily mean that every AI product being sold today is already operating at mature, sustainable economics.

There have been documented examples where the price charged to users did not cover the level of usage being generated. In January 2025, OpenAI CEO Sam Altman said the company was losing money on its US$200-a-month ChatGPT Pro subscription because customers were using the service more heavily than expected. TechCrunch reported Altman’s comments at the time.

That example does not establish that AI services generally are being sold below cost. It does show that customer pricing and the underlying cost of providing AI can diverge, particularly while products, usage patterns and business models are still developing.

Frontier AI development is also extremely capital intensive. In September 2026, Reuters, citing the Financial Times, reported that OpenAI expects cumulative cash burn of approximately US$278 billion between 2026 and 2030 as it expands computing capacity and infrastructure.

That is a company forecast reported by external media rather than an indication of what AI prices themselves will do. It does, however, illustrate the scale of capital associated with frontier AI development.

This creates a credible counterargument to the assumption that AI prices can only move down. If some current pricing reflects a period of aggressive investment, competition or customer acquisition, future commercial structures do not necessarily have to follow the historical inference-cost curve.

Equally, large investment does not prove that prices will rise. Providers may achieve substantial efficiency gains, competition may constrain pricing, and new technology may continue reducing the amount of compute required to deliver a given level of capability.

The eventual outcome remains uncertain.

Compute Is Not an Unlimited Resource

A related question is whether the physical infrastructure required for AI could itself become a constraint.

AI systems ultimately depend on specialised chips, servers, data centres, electricity, cooling, networking and suitable land. Those resources can expand, but they cannot necessarily expand instantly.

Microsoft provides a useful real-world example. In its 2025 Annual Report, the company said its data centres depend on the availability of permitted and buildable land, predictable energy, networking equipment, servers and GPUs. Microsoft also reported that continued investment in cloud and AI infrastructure would increase operating costs and could reduce operating margins.

The capacity constraint has not been purely theoretical.

During Microsoft’s fiscal 2026 third-quarter earnings call, the company said it expected to invest roughly US$190 billion in capital expenditure during calendar 2026. Despite efforts to bring additional GPU, CPU and storage capacity online, Microsoft said it expected capacity to remain constrained at least through 2026.

At the same time, the company reported that hardware and software optimisation had improved inference throughput for its most-used Copilot models.

This illustrates the competing forces particularly well.

More infrastructure is being built. The infrastructure is becoming more efficient. Yet demand can still exceed available capacity.

Electricity adds another dimension. The International Energy Agency’s Energy and AI analysis projects global data-centre electricity consumption to grow by around 15% per year between 2024 and 2030 in its base case. Electricity consumption from accelerated servers, which the IEA says is mainly driven by AI adoption, is projected to grow by around 30% annually over the same period.

None of this means that electricity, GPUs or data-centre capacity will simply “run out”. It does mean that the physical supply of AI compute has its own economics, investment cycles and potential constraints.

Those forces could place upward pressure on the cost of some AI services even while the technology itself becomes more efficient.

AI Consumption May Matter More Than Token Price

There is then the demand side of the equation.

Early enterprise generative AI usage was relatively easy to conceptualise. A person entered a prompt, a model processed it and a response came back. There might be follow-up questions, but the relationship between a human interaction and the underlying model activity remained relatively visible.

Agentic AI can change the cost structure substantially.

A user might initiate a single task while the system performs multiple operations in the background. It could retrieve information, analyse a result, call a tool, evaluate the response, retry a failed step and make further model calls before producing an output.

To the user, that can still appear to be one task. At the infrastructure level, it can involve many separate consumption events.

This is where declining AI token costs can become misleading when viewed in isolation.

If the price per token falls by 50% but token consumption increases fourfold, overall token expenditure rises. If AI becomes embedded into more applications, workflows and automated processes, the number of model interactions could increase even while the price of each interaction continues to decline.

Other outcomes are also possible. Smaller models, improved caching, more efficient architectures or changing usage patterns could reduce the cost of particular workloads.

The central point is that AI token pricing does not determine total enterprise AI costs on its own.

A Token Is No Longer the Whole Meter

The cost picture becomes more complicated because input and output tokens are increasingly only part of the commercial model.

Contemporary AI pricing can include cached tokens, batch processing, tool calls, searches, different processing modes and different model tiers. Two applications processing a similar number of visible user requests can therefore have very different underlying economics.

One workload might repeatedly process a large amount of context. Another might reuse cached information. One might generate lengthy outputs. Another might rely heavily on tools. Some systems may use smaller models for routine activity and invoke more capable models only for particular tasks.

This creates an important distinction between cost per token and cost per completed task.

Consider two hypothetical models. Model A has a lower token price but requires several attempts, more context and more model calls to complete a task. Model B has a higher token price but completes the same task using fewer calls and fewer tokens overall.

The cheaper token does not necessarily produce the cheaper completed task.

Neither measure is inherently wrong. They measure different aspects of AI cost.

Two AI workflows show how lower-priced tokens with more calls and retries can cost more per completed task than higher-priced tokens requiring fewer calls.

Several Forces Are Moving at the Same Time

The future cost of AI can therefore be viewed as the result of several competing forces.

Efficiency can push costs down. Models, hardware and infrastructure have historically become more efficient at delivering equivalent capability.

Capability moves independently. New frontier models can require substantially more computation while performing tasks previous generations could not.

Infrastructure economics introduce another variable. Chips, data centres, electricity and networking require very large capital investment and can periodically face capacity constraints.

Provider economics also matter. Customer prices do not necessarily have to equal the underlying cost of supplying a product at every stage of its development.

And then there is consumption. Even if the cost of each individual unit falls, substantially greater usage can increase total expenditure.

These forces can move independently and, at times, in opposite directions.

That is what makes longer-term AI cost modelling difficult. The market’s broader shift from per-seat pricing towards consumption-based AI cost models adds another layer to that uncertainty.

A straight-line extrapolation of today’s token price says relatively little about the future total cost of an enterprise AI environment.

It also explains why enterprise AI total cost of ownership can evolve differently from the published price of an individual model.

The Economics of Abundance

There is a broader economic idea behind all of this.

When a technology becomes cheaper, people do not necessarily spend less on it. In some markets, lower unit costs make entirely new forms of consumption economically viable.

AI may follow a similar pattern.

If machine-generated reasoning, analysis, content and automation continue to become cheaper, AI could appear in more applications and operate more frequently. Rather than being something an employee deliberately opens and uses, AI could increasingly operate within software, workflows and automated processes.

Under that scenario, each individual unit of AI capability becomes cheaper while the total quantity being consumed increases substantially.

The future of AI costs could therefore become less a story about expensive technology becoming cheap and more a story about previously scarce computational capability becoming abundant enough to be used routinely.

From an enterprise AI procurement perspective, this distinction changes the nature of the cost question. The published price of a model represents only one part of the economics of an AI-enabled environment.

So, Is AI Going to Get Cheaper?

For equivalent levels of capability, the historical evidence says it already has dramatically.

Whether that pace of decline continues is uncertain.

There are credible forces pushing costs lower: better hardware, more efficient models, competition, caching, model routing and improvements in inference infrastructure.

There are also forces that could create upward pressure: enormous capital requirements, constrained compute capacity, growing electricity demand, more compute-intensive frontier models and commercial pricing that may evolve as the economics of the industry mature.

Then there is consumption.

As AI moves from standalone interfaces into software, workflows and increasingly autonomous systems, the relationship between one human action and the amount of underlying AI activity becomes less direct.

That means the original question may actually contain two separate questions.

Will it become cheaper to produce a given amount of AI capability?

Historical evidence suggests that it can.

But the second question is:

What will organisations ultimately pay when capability, infrastructure, provider economics and consumption are all changing at the same time?

That answer is much less certain.

And it may matter far more than the future price of one million tokens.

This article provides general information and commentary only. It does not constitute legal, financial, accounting, tax, procurement, investment, technical or other professional advice. Future AI pricing, technology developments and usage patterns are uncertain, and vendor pricing and product structures may change over time.

Last reviewed: 20 September 2026.