Launch HN: Tokenless (YC S26) – Automatic model switching to save money
A new startup called Tokenless has launched a tool designed to reduce AI inference costs by automatically routing requests to the most efficient model. The service monitors model performance in real-time and cancels redundant requests to save users money.
Why it matters
As AI adoption grows, cost-optimization tools are becoming essential for businesses looking to scale their operations without excessive expenditure.
Token less The router that cuts your inference bill in half . A drop-in replacement for your API calls — always routed to the right model.
Most calls don't need a frontier model. Tokenless fans out your request to a group of models and watches them think. Once a model is clearly on track, we select it and cancel the other models, and you only pay for what you need.
We expose an OpenAI and Anthropic compatible endpoint. Point your models at us and get started today!
time → confidence, per model this request $0.0000 $0.0000 sent to Fable 5 $0.0000 $0.0000 − 52 % · $0.0101 never billed Measured , not marketed. Cost versus quality on public coding benchmarks — the same quality as Opus 4.8, at a fraction of the cost per task.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in