What happens when an LLM never sees material beyond fifth grade?

Aug 16, 2026 02:37 PM - 1 hour ago 2

Talk to LittleLearner

The hosted 5B model, unrecorded successful your browser. Open successful a caller tab ↗ if the chat doesn’t load below.

A controlled sandbox for studying really models get knowledge

Modern LMs are trained connected everything astatine once, truthful it is difficult to show whether a caller skill was learned aliases simply elicited. We constrain the training distribution itself: an 88B-token corpus filtered to the U.S. elementary-school curriculum, pinch models trained from scratch connected it and matched unfiltered controls.

Dataset

LittleCurriculum

An 88B-token corpus distilled from FineWeb-Edu done a five-stage filtering pipeline aligned pinch Common Core standards (K–5). Concepts, facts, and vocabulary taught supra Grade 5 are explicitly excluded.

Models

LittleLearner

Three scales (0.6B / 1.3B / 5B) trained from scratch connected LittleCurriculum: chattable models pinch an interpretable knowledge boundary. Each ships pinch a matched Unfiltered control for cleanable comparison.

Findings

Elicitation, not acquisition

In our experiments, scaling, SFT+GRPO post-training, and in-context learning amplify what the program taught, but nary meaningfully improves out-of-scope performance, indicating that the pretraining select sets the effective capacity ceiling.

Model checkpoints

LittleLearner astatine 3 scales (0.6B / 1.3B / 5B), each pinch a matched Unfiltered control sharing its architecture, tokens, and recipe.

Base: the pretrained model.
GRPO: mathematics specialists post-trained connected MathCAMPS; responses whitethorn grounds a inclination toward math-oriented output.
Chatty: variants tuned for wide chat behavior.

Scale LittleLearner · K–5chattyMatched power · unfiltered

Capability stays wrong the curriculum

Can modular interventions push a exemplary past what its pretraining information taught it? With the bound nether experimental control, we tin inquire cleanly. In our experiments, each involution amplifies in-scope ability; nary of them meaningfully improves out-of-scope performance.

Scaling

Scaling exemplary size improves capacity wrong the model’s controlled knowledge vulnerability and extends modestly to problems on the aforesaid learning trajectory, but yields small betterment on problems requiring much precocious capabilities extracurricular the exposure.

MathCAMPS accuracy by grade, crossed exemplary size

Post-training

Post-training done GRPO importantly boosts in-scope K–5 capabilities, but fails to retrieve out-of-scope beyond-K–5 capabilities, moreover erstwhile training pinch out-of-scope data.

Post-training amplifies K–5, not the beyond-K–5 gap

In-context learning

In-context learning pinch the prompts we trial does not unlock caller reasoning capabilities in beyond-K–5 for our trained 5B LittleLearner.

Accuracy by prompting condition

What will you thatch it?

Because LittleLearner’s training vulnerability is explicitly specified, behavioral and representational changes tin beryllium related straight to the concepts you introduce. Three directions we’re excited about:

01

RL & discovery

Can RL create capability?

The anterior is restricted to K–5, truthful capabilities that look nether RL tin beryllium attributed to the RL process itself. A tractable proxy for reward-driven discovery.

02

Continual learning

Watch a conception being learned

Introduce antagonistic numbers and measurement sample efficiency, retention, and interference. Or probe behaviour adjacent the boundary: does it answer, abstain, aliases hallucinate?

03

Educational science

Machine vs. kid learners

Specified vulnerability enables controlled human-model comparison. Do models and children need akin vulnerability to study fractions, aliases make akin errors connected connection problems?

Your turn

Bring your ain question

A known bound turns your thought into a cleanable experiment!

If you find this activity useful

Please mention our paper:

@misc{littlelearner2026, title={LittleLearner: Language Models Under Pedagogically-Controlled Knowledge Exposure}, author={Fanfei Li and Jana Zeller and Manuel Prada-Corral and Thaddäus Wiedemer and Prasanna Mayilvahanan and Ryan Cotterell and Wieland Brendel}, year={2026}, eprint={2608.13545}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2608.13545} }
More