Hacker News·14 min read

Training a 3.8B LLM to 0.384 CORE for $998 – Hugo Vergnes

A
Anon84
Dive DeeperCreate a free account to unlock

Somewhere between “nanoGPT toy” and “you need a research lab” there’s a large, under-described region where one person with a few thousand dollars can train a meaningful model.

I wanted to see language and understanding emerge from random weights for myself, and to learn the parts you can only learn by starting from scratch. This project was written in the evenings, debugged on a 5090 and finished on rented B200s. It was heavily inspired by Andrej Karpathy’s nanochat .

The result is a 3.8B-parameter model scoring 0.384 on CORE , trained on 65B tokens in 43 hours for $998 .

What follows is what worked, what didn’t, and what I still don’t know.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in