$0.09 and $290.12 are both the price of 1M output tokens
An analysis of AI inference costs reveals a massive price disparity between different providers, driven largely by hardware choices rather than the models themselves. The author argues that the accelerator market is inefficiently priced and that self-hosting on specific hardware like AMD's MI355X can be significantly cheaper than using hosted APIs.
Why it matters
This insight challenges the assumption that AI costs are fixed and highlights the potential for significant savings through strategic hardware selection.
Nine cents. Two hundred and ninety dollars and twelve cents.
Those are both the price of one million output tokens. Same unit, same dataset, read off vendor pages in the same week. The first is gpt-oss-120b on a single MI355X. The second is Qwen3.5 on eight H100s.
Neither is a typo and neither is a trick. And almost none of the gap between them is the model.
So I went looking for what it actually is. Here is what I found, measured across 24 providers and 378 GPU rental rates, ranked smallest to largest:
Five levers. The two everyone talks about are the two smallest. This post walks up the list.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in