Lossless model compression experiment: GLM-5.2 in 25% less memory

Researchers have successfully compressed the GLM-5.2 large language model by approximately 25-30% using a technique called K15 charged-format accounting. This method replaces standard 9-bit sign-and-exponent symbols with 4-bit codes, maintaining bit-for-bit accuracy while significantly reducing memory requirements.
Why it matters
Efficient model compression is critical for deploying massive AI models on hardware with limited memory, reducing costs and energy consumption.
A full GLM-5.2 scan found 30.168% K15 charged-format accounting. A separate byte-split representation was decoded bit-for-bit across all 59,509 BF16 tensors at 24.967% reduction. Those are distinct evidence classes.
GLM-5.2 753B · BF16 K15 ACCOUNTING 1,403 GiB −423 GiB 30.17% SMALLER [ FIG. 00 ] GLM-5.2 753B K15 CHARGED-FORMAT ACCOUNTING: 1,403.19 TO 979.87 GiB, 30.168%. SEPARATE BYTE-SPLIT EXACT INVERSE: 24.967%. How it works [ FIG. 01 ] 01 / Where the K15 estimate comes from Most weights share a few sign-exponent symbols Every BF16 weight is 16 bits: a sign, 8 exponent bits, 7 mantissa bits. In a trained model the exponents are wildly repetitive; a handful of values cover nearly every weight.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in