Article may be outdated

This article is 86 days old. Some details may have changed since publication.

Hacker News·5 min read·hard

I trained a 113M-parameter earthquake LLM from absolute scratch

J
jzsfg
I trained a 113M-parameter earthquake LLM from absolute scratch
✦AI Summary

A developer documents the end-to-end process of training a 113M-parameter earthquake-focused language model from scratch. The project serves as an educational resource for understanding the full lifecycle of LLM development, including data cleaning, tokenization, and multi-GPU training.

Why it matters

It provides a transparent, reproducible blueprint for small-scale AI training, demonstrating that specialized models can be developed on consumer-grade hardware.

✦Dive DeeperCreate a free account to unlock

Train a small GPT for earthquake science — the entire LLM lifecycle, from a blank folder to a talking model, explained block by block.

Crawl → Clean → Tokenize → Model → Train → Infer, on 2× NVIDIA A30 (48 GB).

Six free data sources → crawl → clean/dedup → 16k BPE → 113M GQA+RoPE decoder → 2-GPU DDP training → streaming inference. Each stage is a section below.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyscienceai
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in