Frame selection is the whole game: notes on making LLMs watch video
This article explores the technical challenges of video processing for Large Language Models, specifically focusing on the importance of frame selection over uniform sampling. The author argues that intelligent frame extraction is essential for models to accurately interpret visual data rather than relying on human-curated summaries.
Why it matters
As multimodal AI becomes more prevalent, optimizing how models 'watch' video is critical for improving accuracy and reducing computational costs.
A vision LLM can realistically afford about 150 images per video. Which 150 you pick decides whether the model watched the video or just a slideshow about it. These are my notes from getting this wrong over and over while building claude-real-video , an MIT-licensed local pipeline.
Why feed a model a video at all, when a writeup of the same thing costs way fewer tokens?
Because an article is someone's compression of an event. A human watched it, decided what mattered, threw the rest away. The framing, the timing, the stuff on screen they didn't think was relevant. When an LLM reads the article, it learns inside that author's choices. It can't recover what got cut, and it can't disagree with a selection it never saw.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in