Article may be outdated

This article is 63 days old. Some details may have changed since publication.

Hacker News·3 min read·hard

Paging Through a Parquet File in DuckDB: File_row_number or Offset?

R
rustyconover
Paging Through a Parquet File in DuckDB: File_row_number or Offset?
✦AI Summary

This technical blog post compares the performance of using LIMIT/OFFSET versus file_row_number when paging through large Parquet files in DuckDB. It concludes that row-range filtering is significantly faster and more reliable for large datasets.

Why it matters

Optimizing data retrieval methods is essential for developers building scalable APIs that handle massive datasets in cloud environments.

✦Dive DeeperCreate a free account to unlock

Blog Paging Through a Parquet File in DuckDB: file_row_number or OFFSET? DuckDB July 30, 2026

You have a large Parquet file and an API that has to return its contents, but responses have a size ceiling, so the caller pages through it. LIMIT/OFFSET is the obvious way to write that and it is the wrong one. I measured file_row_number against it: 2.53x faster on a file with 163 row groups. Speed is the boring half of the answer. OFFSET will also hand back the wrong rows without telling you.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologybusiness
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in