Train and run transformers directly on Apple's Neural Engine
A new software tool called Espresso allows developers to run transformer models directly on Apple's Neural Engine. It significantly outperforms standard CoreML and GPU-based inference methods.
Why it matters
Optimizing AI inference on consumer hardware is critical for running large language models locally and efficiently on personal devices.
Direct Neural Engine inference for transformers on Apple Silicon — 4.76x faster than CoreML.
Technical reporting on software performance benchmarks without political or social bias.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in