Same model, same Q4_K_M label: 5.02, 5.07 and 5.27 bits per weight
A new Python tool called 'picchio' has been released to help developers accurately measure the performance of local Large Language Models. It provides detailed diagnostics on GPU usage, prefill/decode speeds, and CPU fallback to help users understand the true performance of their hardware.
Why it matters
Standard performance metrics for LLMs are often misleading; this tool provides transparency for developers running models on local hardware.
One Python file that measures local LLMs: effective bits per weight, the three tok/s lanes, and silent CPU fallback.
Install · Commands · Quant · Lanes · Measured · Examples
Most GPU speed claims are one tok/s number. That number can be correct and still tell you the wrong story. Three failure modes, each one command:
picchio splits prefill, decode and wallclock, reads the engine's log against the OS's GPU meter, and prints a verdict that says whether the GPU did the work, and why.
curl -fsSLO https://raw.githubusercontent.com/logxio/picchio/main/picchio.py python3 picchio.py With no arguments it finds your models (ollama tags, the current folder, the HF and LM Studio caches) and runs the one you pick. A .gguf path gets the full llama.cpp diagnosis; an ollama tag gets measurement mode.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in