Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Magnitude is a new open-source inference engine designed to optimize large language models for specific hardware configurations. It features hand-optimized kernels that allow models to run significantly faster than standard implementations like llama.cpp while maintaining privacy.
Why it matters
This tool lowers the barrier for running high-performance AI models locally on consumer hardware, reducing reliance on cloud-based APIs.
Run open models as fast as your hardware allows
Magnitude is an open source inference engine for agents that optimizes itself for your exact hardware. It compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. One click connects the agent you already use (Pi, OpenCode, Hermes, Codex, and more). Works on Apple Silicon, NVIDIA, AMD, or nothing but a CPU.
Download Magnitude for macOS, Windows, or Linux
⭐ Help us reach more developers and grow the Magnitude community. Star this repo!
demo-9-29.mp4 Get started Download Magnitude , install it, and open the app. Choose a recommended model in Discover and download it. Connect your agent in Connections and start using it. The desktop app includes the magnitude CLI. No separate installation is needed.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in