Article may be outdated

This article is 84 days old. Some details may have changed since publication.

Hacker News·2 min read·hard

Same model, same Q4_K_M label: 5.02, 5.07 and 5.27 bits per weight

L
logickkk1
Same model, same Q4_K_M label: 5.02, 5.07 and 5.27 bits per weight
✦AI Summary

A new Python tool called 'picchio' has been released to help developers accurately measure the performance of local Large Language Models. It provides detailed diagnostics on GPU usage, prefill/decode speeds, and CPU fallback to help users understand the true performance of their hardware.

Why it matters

Standard performance metrics for LLMs are often misleading; this tool provides transparency for developers running models on local hardware.

✦Dive DeeperCreate a free account to unlock

One Python file that measures local LLMs: effective bits per weight, the three tok/s lanes, and silent CPU fallback.

Install · Commands · Quant · Lanes · Measured · Examples

Most GPU speed claims are one tok/s number. That number can be correct and still tell you the wrong story. Three failure modes, each one command:

picchio splits prefill, decode and wallclock, reads the engine's log against the OS's GPU meter, and prints a verdict that says whether the GPU did the work, and why.

curl -fsSLO https://raw.githubusercontent.com/logxio/picchio/main/picchio.py python3 picchio.py With no arguments it finds your models (ollama tags, the current folder, the HF and LM Studio caches) and runs the one you pick. A .gguf path gets the full llama.cpp diagnosis; an ollama tag gets measurement mode.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technology
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in