Article may be outdated

This article is 62 days old. Some details may have changed since publication.

Hacker News·5 min read·hard

$0.09 and $290.12 are both the price of 1M output tokens

O
OsamaJaber
$0.09 and $290.12 are both the price of 1M output tokens
✦AI Summary

An analysis of AI inference costs reveals a massive price disparity between different providers, driven largely by hardware choices rather than the models themselves. The author argues that the accelerator market is inefficiently priced and that self-hosting on specific hardware like AMD's MI355X can be significantly cheaper than using hosted APIs.

Why it matters

This insight challenges the assumption that AI costs are fixed and highlights the potential for significant savings through strategic hardware selection.

✦Dive DeeperCreate a free account to unlock

Nine cents. Two hundred and ninety dollars and twelve cents.

Those are both the price of one million output tokens. Same unit, same dataset, read off vendor pages in the same week. The first is gpt-oss-120b on a single MI355X. The second is Qwen3.5 on eight H100s.

Neither is a typo and neither is a trick. And almost none of the gap between them is the model.

So I went looking for what it actually is. Here is what I found, measured across 24 providers and 378 GPU rental rates, ranked smallest to largest:

Five levers. The two everyone talks about are the two smallest. This post walks up the list.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologybusiness
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in