NVDA 225.16 ▼0.06%GOOGL 345.90 ▼0.13%MSFT 495.40 ▼0.30%AMD 514.39 ▲6.50%INTC 102.50 ▼1.97%TSMC 426.35 ▼0.96%AMZN 262.65 ▼0.94%META 589.85 ▼0.86%AAPL 305.93 ▲0.22%PLTR 174.04 ▼2.78%
Markets at last close

Google · Models

Price becomes the key LLM benchmark

·1 min read

Price is becoming the most important benchmark for large language models as enterprise buyers focus on token budgets and practical returns. Google’s Gemini 3.7 Flash, positioned as a flagship workhorse model for coding, knowledge work and agents, is priced at 75 cents per 1 million input tokens and $3.75 per 1 million output tokens, with promotional pricing running to January 1, 2027.

The pressure is spreading across the market. Writer said Writer Agent operates at a 52% lower cost with its Palmyra X6 model and agent harness. Meta is using open weight models as part of a pricing-led comeback, Nvidia is targeting enterprise customization with Nemotron, SpaceXAI is pushing Grok 4.6 with aggressive pricing, and DeepSeek’s V4-Pro is priced at 43 cents per 1 million input and 87 cents per 1 million output tokens.

Enterprise finance and operations leaders are treating token consumption as a managed cost rather than an experimental expense. Executives at Palantir, Moody’s, Synchrony Financial and Block described stricter monitoring, workload routing, model selection and process-level cost analysis. Falling prices may benefit customers, but they also raise questions for Anthropic and OpenAI as the economics of foundation models face closer scrutiny.

Originally reported by constellationr.comRead the source →
Related coverage
All Google news →