Article may be outdated

This article is 58 days old. Some details may have changed since publication.

Hacker News·6 min read·hard

Frame selection is the whole game: notes on making LLMs watch video

C
cortexosmain
✦AI Summary

This article explores the technical challenges of video processing for Large Language Models, specifically focusing on the importance of frame selection over uniform sampling. The author argues that intelligent frame extraction is essential for models to accurately interpret visual data rather than relying on human-curated summaries.

Why it matters

As multimodal AI becomes more prevalent, optimizing how models 'watch' video is critical for improving accuracy and reducing computational costs.

✦Dive DeeperCreate a free account to unlock

A vision LLM can realistically afford about 150 images per video. Which 150 you pick decides whether the model watched the video or just a slideshow about it. These are my notes from getting this wrong over and over while building claude-real-video , an MIT-licensed local pipeline.

Why feed a model a video at all, when a writeup of the same thing costs way fewer tokens?

Because an article is someone's compression of an event. A human watched it, decided what mattered, threw the rest away. The framing, the timing, the stuff on screen they didn't think was relevant. When an LLM reads the article, it learns inside that author's choices. It can't recover what got cut, and it can't disagree with a selection it never saw.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyai
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in