Benchmarking Pocket-Scale Inference

Artificial Analysis has launched a benchmarking initiative to evaluate the performance and intelligence of small AI models running on mobile hardware. The project focuses on models that fit within 8GB of memory, testing them on devices like the iPhone 17 Pro and Galaxy S26 Ultra to determine real-world utility.
Why it matters
As AI moves from the cloud to edge devices, understanding the efficiency and capability of 'pocket-scale' models is critical for the future of mobile computing.
Artificial Analysis K Artificial Analysis Models Coding Agents Speech, Image, Video Inference Leaderboards About AI Trends Arenas K 2–20W Mobile phones → 15W–2kW Laptops & workstations Late 2026 700W–600kW Datacenter systems → Benchmarking pocket-scale inference We benchmark small models on mobile phones. Artificial Analysis' testing covers model intelligence on a set of benchmarks chosen to represent real-world mobile device usage, and we partner with Liquid AI to gather real inference data measured on the devices themselves. Note: we have independently validated Liquid AI's inference measurement process.
“Small” models are all models that fit inside 8 GB of memory after quantization, including KV cache at 8K context. View all rules and our process in the methodology page .
Performance is benchmarked using llama.cpp on builds quantized to 4-bit or smaller. More models, quantizations, devices, inference frameworks and other variants coming soon.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in