Hacker News·4 min read·hard

LensVLM-9B by Apple

N
nthypes
LensVLM-9B by Apple
AI Summary

Apple researchers have introduced LensVLM, a framework that improves the efficiency of Vision-Language Models by selectively expanding compressed visual text. This method allows models to maintain high accuracy while significantly reducing the computational cost of processing rendered text.

Why it matters

Efficient visual processing is critical for scaling AI capabilities in document and code analysis without requiring massive token sequences.

Dive DeeperCreate a free account to unlock

Papers arxiv:2605.07019 Copy markdown LensVLM: Selective Context Expansion for Compressed Visual Representation of Text Published on May 7 Upvote 3 Authors: Roy Xie , Dan Friedman , Donghan Yu , Bowen Pan , Christopher Fifty , Jang-Hyun Kim , Xianzhi Du , Zhe Gan , Vivek Rathod , Bhuwan Dhingra Abstract Vision-Language Models can process text as rendered images, but accuracy degrades with compression; LensVLM addresses this by scanning compressed images and selectively expanding relevant parts through learned tools, maintaining high accuracy even at high compression ratios.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyscience

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in