Article may be outdated

This article is 57 days old. Some details may have changed since publication.

Hacker News·5 min read·hard

DuckDB – Data power tools for your laptop, now in Clojure (2023)

S
sourdecor
✦AI Summary

This article discusses the integration of DuckDB into the Clojure-based tech.ml.dataset platform to improve data processing efficiency. It highlights how the transition from row-based to batched column-major processing allows for better handling of large datasets without needing complex cluster infrastructure.

Why it matters

It provides a practical solution for data scientists to perform high-performance analytics on local hardware, reducing reliance on expensive and complex distributed computing clusters.

✦Dive DeeperCreate a free account to unlock

Our in-memory column-major data processing platform, tech.ml.dataset (TMD), drives the future of functional data science. When data gets large enough not to fit in memory, one can continue with TMD by operating on samples of data, or otherwise filtering to relevant subsets to fit in bounds imposed by the working environment. Moreover, one can accomplish persistence, for data small and large, with nippy, arrow, or parquet .

When data becomes large enough, for example sets of .csv files on the order of ~100GB with relational aspects to them, the tools in their current state can become unwieldy. One is tempted to get involved in nonfunctional sparky cluster snafus. Of course, maintaining some level of transactional interaction and a simple disk IO model is still super-desirable. Local disks are big enough, and local chips are fast enough, no need to do anything rash.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologybusiness
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in