Article may be outdated

This article is 72 days old. Some details may have changed since publication.

Hacker News·4 min read·hard

Lossless model compression experiment: GLM-5.2 in 25% less memory

H
hambandit
Lossless model compression experiment: GLM-5.2 in 25% less memory
✦AI Summary

Researchers have successfully compressed the GLM-5.2 large language model by approximately 25-30% using a technique called K15 charged-format accounting. This method replaces standard 9-bit sign-and-exponent symbols with 4-bit codes, maintaining bit-for-bit accuracy while significantly reducing memory requirements.

Why it matters

Efficient model compression is critical for deploying massive AI models on hardware with limited memory, reducing costs and energy consumption.

✦Dive DeeperCreate a free account to unlock

A full GLM-5.2 scan found 30.168% K15 charged-format accounting. A separate byte-split representation was decoded bit-for-bit across all 59,509 BF16 tensors at 24.967% reduction. Those are distinct evidence classes.

GLM-5.2 753B · BF16 K15 ACCOUNTING 1,403 GiB −423 GiB 30.17% SMALLER [ FIG. 00 ] GLM-5.2 753B K15 CHARGED-FORMAT ACCOUNTING: 1,403.19 TO 979.87 GiB, 30.168%. SEPARATE BYTE-SPLIT EXACT INVERSE: 24.967%. How it works [ FIG. 01 ] 01 / Where the K15 estimate comes from Most weights share a few sign-exponent symbols Every BF16 weight is 16 bits: a sign, 8 exponent bits, 7 mantissa bits. In a trained model the exponents are wildly repetitive; a handful of values cover nearly every weight.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyscience
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in