Article may be outdated

This article is 51 days old. Some details may have changed since publication.

Hacker News·3 min read·hard

Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots

H
HenryNdubuaku
Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
✦AI Summary

Needle2 is a new, highly compressed 45M-parameter AI model designed to run efficiently on low-power edge devices like wearables, robots, and budget smartphones. By focusing on tool calling and structured extraction, it provides functional AI capabilities without needing massive computational resources.

Why it matters

It demonstrates a shift toward 'small language models' that can bring sophisticated AI to the vast ecosystem of low-cost IoT hardware, bypassing the need for cloud-based processing.

✦Dive DeeperCreate a free account to unlock

Today we release Needle 2: an open 45M-parameter model for tool calling, device use and structured extraction. The whole model is a single 14MB binary that runs a full session in 28MB of RAM. It is built on our Simple Attention Network findings, compressed to CQ2-bit with Cactus Quants , and baked into its own engine.

On the tool call and mobile device use benchmarks, Needle 2 trades wins with other small models like FunctionGemma 270M, LFM2.5 230M and Apple FM, at 5× to 70× smaller, and 2 bits against their f16. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, between 400–1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges 300–700 on sub-$200 phones such as the Samsung A-Series. With a peak session RAM around 28MB, Needle runs on newer microcontrollers like ESP32-S3.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologystartups
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in