What happens when an LLM never sees material beyond fifth grade?
Researchers tested whether large language models can exceed the knowledge boundaries of their training data by training models exclusively on K-5 curriculum. The study found that interventions like scaling or fine-tuning do not improve performance on tasks outside the scope of the initial training data.
Why it matters
This research provides critical insights into the limitations of LLMs and suggests that pretraining data quality sets a hard ceiling on model capabilities.
The hosted 5B model, live in your browser. Open in a new tab ↗ if the chat doesn’t load below.
Modern LMs are trained on everything at once, so it is hard to tell whether a new skill was learned or merely elicited . We constrain the training distribution itself: an 88B-token corpus filtered to the U.S. elementary-school curriculum, with models trained from scratch on it and matched unfiltered controls.
An 88B-token corpus distilled from FineWeb-Edu through a five-stage filtering pipeline aligned with Common Core standards (K–5). Concepts, facts, and vocabulary taught above Grade 5 are explicitly excluded.
Three scales (0.6B / 1.3B / 5B) trained from scratch on LittleCurriculum: chattable models with an interpretable knowledge boundary. Each ships with a matched Unfiltered control for clean comparison.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in