Why Cheap AI Inference Could Hurt Some Stocks and Help Others

Learn which chip, cloud, and app stocks gain or lose as AI query costs fall and usage surges.

Cheap AI inference hurts stocks tied to scarcity and helps stocks tied to volume. Inference is the paid work of running a trained model to answer each user query, and cheaper queries shift profit from high-priced compute to owners of scale and software. By 2025 inference passed training as the dominant AI compute cost, at over 60% of total spend, according to Ainvest market-structure analysis. That shift rewards suppliers and platforms that turn low prices into massive token volume.

Table of Contents

How fast did inference prices fall?

Stanford HAI AI Index 2025, reported by PYMNTS, puts GPT-3.5-level queries at $20 per million tokens in November 2022 and $0.07 by October 2024. That is a 280-fold drop in under two years, detailed in the PYMNTS price history.

Declines stayed uneven by task during 2025. Epoch AI tracking, reported by Brookings, found yearly declines from 9x to 900x, with a median near 50x.

Why do suppliers sell off first?

Efficient models can scare holders of chip and memory stocks. Startup Fortune, citing Bloomberg, ties DeepSeek's R1 release in January 2025, claimed at about $6 million in training cost, to about $589 billion taken from Nvidia in one session.

A similar pattern hit memory. Motley Fool reported on April 4, 2026 that Micron and Sandisk shares fell after a Google memory-saving inference advance, then recovered on Jevons-paradox demand arguments, explained in the Fool's Jevons paradox guide. Cheaper use per query can expand total queries enough to lift total hardware demand.

How can the leading chip supplier still gain?

Lower prices do not always mean lower chip revenue. FourWeekMBA reports Nvidia held about 74% inference share with $41B in inference revenue in Q1 2026, evidence that volume can offset price cuts, shown in the FourWeekMBA Nvidia breakdown.

Inference scale favors installed systems, software, and networking. Buyers reorder more tokens, developers ship more features, and the incumbent collects payment on each call.

Where does the saved margin go?

Hyperscalers keep more margin by moving work to captive chips. Ainvest puts AWS's chip business at a $25B-plus annualized run rate using Trainium-style parts for inference.

Application and model vendors also gain from heavier use. Menlo Ventures' 2025 mid-year update, via FinancialContent, puts enterprise LLM API spend at $8.4B by mid-2025, with Anthropic at about 40% and OpenAI at 27%, reported in the Menlo Ventures spend update. Investors can separate price fear from volume gain: Track those volume and share figures before assuming lower query prices mean lower profits.

  • Compare token-volume growth with per-token price declines
  • Check inference revenue share, not chip price alone
  • Watch captive-silicon run rates and enterprise API-spend splits

You Might Also Like