Hacker News·5 min read·hard

Sub-1-Bit LLM Compression via Latent Factorization

B
brainless
Sub-1-Bit LLM Compression via Latent Factorization
✦AI Summary

Researchers have introduced LittleBit-2, an advancement in sub-1-bit LLM compression that uses latent factorization and joint iterative quantization. This method allows for extreme model compression without requiring changes to the model architecture during inference.

Why it matters

This technology enables the deployment of large language models on resource-constrained hardware by drastically reducing memory requirements.

✦Dive DeeperCreate a free account to unlock

Sub-1-Bit LLM Compression via Latent Factorization Official implementation of LittleBit (NeurIPS 2025) and LittleBit-2 (ICML 2026).

LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment (ICML 2026) Banseok Lee, Youngmin Kim

LittleBit: Ultra Low-Bit Quantization via Latent Factorization (NeurIPS 2025) Banseok Lee*, Dongkyu Kim*, Youngcheon You, Youngmin Kim

LittleBit compresses large language models into the sub-1-bit regime by factorizing each dense weight matrix into low-rank latent factors, binarizing those factors, and restoring magnitude information through lightweight learned scales. This enables extreme compression, including the 0.1 bits-per-weight setting, while preserving the original model architecture at inference time.

LittleBit-2 improves this recipe by addressing latent geometry misalignment in the initialization stage. It applies Internal Latent Rotation with Joint Iterative Quantization (Joint-ITQ), aligning the SVD-derived latent factors with the binary hypercube before QAT. LittleBit-2 initialization is available as an opt-in ( --use_itq ) and produces no additional inference overhead.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyscience
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in