NVDA 223.96 ▲2.27%GOOGL 354.30 ▼0.96%MSFT 499.99 ▲0.03%AMD 483.36 ▼1.21%INTC 101.65 ▲1.84%TSMC 420.04 ▲0.44%AMZN 274.48 ▲0.82%META 592.10 ▲0.37%AAPL 313.33 ▲0.29%PLTR 172.01 ▲10.32%
Markets at last close

OpenAI · Infrastructure

LLM inference costs reshape enterprise AI economics

·1 min read

When GPT-4 launched in March 2023, running a single long conversation cost roughly $0.06 per thousand output tokens, making enterprise use cases such as contract review, support triage, and document summarization expensive at scale. By mid-2026, comparable compute costs have fallen by a factor of somewhere between 80 and 150, depending on provider and workload, and the decline is still accelerating.

The drop is being driven by more efficient model architectures, broader hardware availability, and open-weight competition. Mixture-of-Experts models reduce compute demands by activating only part of a model for each input, while techniques such as speculative decoding, flash attention, and quantization improve efficiency further. Hardware supply has also expanded through TSMC capacity, AMD’s MI300X and MI325X, Google’s TPU v5, AWS Trainium2, and Microsoft’s Maia.

Lower inference costs are weakening per-query pricing models that dominated AI startups in 2023 and 2024. Flat-rate, seat-based software pricing with AI included is becoming more attractive as inference begins to resemble a routine infrastructure cost rather than a scarce resource. At $0.0004 per thousand tokens, AI can run continuously across forms, reports, and user interactions, making embedded intelligence a default design assumption rather than a premium add-on.

Originally reported by digitalfrontier.newsRead the source →
Related coverage
All OpenAI news →