Show HN: Distill and serve small models with frontier quality for half the cost
A new developer tool called 'wmo' allows users to optimize and serve smaller AI models by using traces from larger frontier models. It provides a CLI for routing, distillation, and testing within simulated environments.
Why it matters
Offers a practical solution for companies looking to reduce AI inference costs while maintaining high performance through model distillation and routing.
wmo turns agent traces you already collect into continuous improvement. Start with a model endpoint at frontier quality with 40%+ lower cost. Keep improving it with world model simulations, meta-harness optimization, and model distillation.
pip install world-model-optimizer wmo providers set 2. Tune a router on your OTel traces.
wmo build --file traces.jsonl --name my-endpoint # Score every registered model on held-out tasks from your traces wmo optimize route sweep my-endpoint --traces traces.otel.jsonl # Turn those measurements into a routing policy wmo optimize route fit matrix.json --kind knn \ --out .wmo/models/my-endpoint/policy.json 3. Serve it.
wmo serve --name my-endpoint See what it bought you against the model you were using before:
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in