AI inference token prices hit annual low
Average prices for AI inference have dropped to their lowest level of the year, according to research by investment bank Jeffries citing Silicon Data. Inference prices per million tokens ranged from $1.16 (£0.86)to $1.18 from 6 to 8 August, down from the $2.04 average on 31 May and the $1.45 average of late July.
The decline followed the release of V4-Flash-0731 from Hangzhou-based DeepSeek, priced at $0.03 per task and described by Jeffries as materially cheaper than domestic and international rivals. The model led global token consumption rankings on OpenRouter, accounting for 27 percent of total processing volume on Monday, ahead of Google’s 25 percent.
DeepSeek’s model was also the most used model last week, processing 8.22 trillion tokens, followed by Tencent’s Hy3 at 7.13 trillion and an April version of V4-Flash at 6.05 trillion. Jeffries analysts said the pricing trend reflects growing emphasis on cost efficiency in both the US and China, even as enterprise AI costs rise because agentic tools can consume large and unpredictable numbers of tokens.