As AI model markets mature, the per-token price for AI inference will converge to roughly the underlying compute cost plus only about a 10–20% margin, leaving little additional economic rent for model providers on raw token usage.
“what we should expect is that the price of a token in AI land, you know, basically will be whatever the price of running the computers are. And maybe with like a, you know, plus ten, 20% margin.”
AI inference token pricing has fallen dramatically toward compute cost through 2025-2026 amid intense competition, broadly consistent with the prediction, though provider margins vary and haven't uniformly converged to just 10-20%.