DuckDB – Data power tools for your laptop, now in Clojure (2023)
This article discusses the integration of DuckDB into the Clojure-based tech.ml.dataset platform to improve data processing efficiency. It highlights how the transition from row-based to batched column-major processing allows for better handling of large datasets without needing complex cluster infrastructure.
Why it matters
It provides a practical solution for data scientists to perform high-performance analytics on local hardware, reducing reliance on expensive and complex distributed computing clusters.
Our in-memory column-major data processing platform, tech.ml.dataset (TMD), drives the future of functional data science. When data gets large enough not to fit in memory, one can continue with TMD by operating on samples of data, or otherwise filtering to relevant subsets to fit in bounds imposed by the working environment. Moreover, one can accomplish persistence, for data small and large, with nippy, arrow, or parquet .
When data becomes large enough, for example sets of .csv files on the order of ~100GB with relational aspects to them, the tools in their current state can become unwieldy. One is tempted to get involved in nonfunctional sparky cluster snafus. Of course, maintaining some level of transactional interaction and a simple disk IO model is still super-desirable. Local disks are big enough, and local chips are fast enough, no need to do anything rash.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in