Exploring Claude/GPT Knowledge Cutoffs and Pre-Training Timelines

This article explores methods for estimating the training timelines and technical specifications of frontier AI models like GPT-5 and Claude. It details how researchers use knowledge probes and data mixture inference to reverse-engineer the development stages of large language models.
Why it matters
Understanding the development cycles and capabilities of frontier AI models is critical for researchers and competitors to gauge the pace of technological advancement.
We can learn hidden facts about how frontier models were trained by “probing” them with carefully curated requests.
By scoring them on niche facts we can approximate how many parameters models like GPT-5 and Opus have, using “Incompressible Knowledge Probes”
By measuring how the models break down tokens we can reveal facts about the datasets mixtures they used to train the model (or at least the tokenizer) using “Data Mixture Inference”
By scoring them on date or self-identification related questions you can also estimate training timelines ( this post )
Everything here is an estimate. It’s possible that some speculation in this post is totally incorrect given there’s not a ton of publicly available ground truth to verify against.
The 3 main stages of model training. As a brief primer (see Alex Wa’s blog for more), how we train massive large language models has converaged into 3 stages:
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in