Show HN: Lumabri – What if LLMs worked like Napster?
Lumabri is a new peer-to-peer engine that allows users to run large mixture-of-experts AI models across a swarm of machines. It is designed to operate efficiently on CPUs and SSDs, removing the strict requirement for high-end GPUs.
Why it matters
This technology democratizes access to large-scale AI inference by leveraging distributed computing power, potentially reducing reliance on centralized, expensive hardware.
Run huge mixture-of-experts models from a swarm of peers, with the colibri engine. Pure C, no dependencies.
One machine shares a model. Any other machine chats with it: nothing is downloaded up front, the bytes an inference actually touches arrive from the peer on first use and stay in a local mirror. The second question is served from local disk at full speed. The engine binary is unmodified.
The founding principle: any machine may join, GPU or not. The engine was built for CPU and SSD first; a GPU only makes it faster, never different, and the output is byte-identical either way. A swarm with no GPU at all is a working swarm. Networks that pool GPUs recruit from the few; lumabri recruits from everyone.
make On the machine that has a model (any colibri model directory):
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in