Paging Through a Parquet File in DuckDB: File_row_number or Offset?

This technical blog post compares the performance of using LIMIT/OFFSET versus file_row_number when paging through large Parquet files in DuckDB. It concludes that row-range filtering is significantly faster and more reliable for large datasets.
Why it matters
Optimizing data retrieval methods is essential for developers building scalable APIs that handle massive datasets in cloud environments.
Blog Paging Through a Parquet File in DuckDB: file_row_number or OFFSET? DuckDB July 30, 2026
You have a large Parquet file and an API that has to return its contents, but responses have a size ceiling, so the caller pages through it. LIMIT/OFFSET is the obvious way to write that and it is the wrong one. I measured file_row_number against it: 2.53x faster on a file with 163 row groups. Speed is the boring half of the answer. OFFSET will also hand back the wrong rows without telling you.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in