Talk to LittleLearner
The hosted 5B model, unrecorded successful your browser. Open successful a caller tab ↗ if the chat doesn’t load below.
A controlled sandbox for studying really models get knowledge
Modern LMs are trained connected everything astatine once, truthful it is difficult to show whether a caller skill was learned aliases simply elicited. We constrain the training distribution itself: an 88B-token corpus filtered to the U.S. elementary-school curriculum, pinch models trained from scratch connected it and matched unfiltered controls.
Dataset
LittleCurriculum
An 88B-token corpus distilled from FineWeb-Edu done a five-stage filtering pipeline aligned pinch Common Core standards (K–5). Concepts, facts, and vocabulary taught supra Grade 5 are explicitly excluded.
Models
LittleLearner
Three scales (0.6B / 1.3B / 5B) trained from scratch connected LittleCurriculum: chattable models pinch an interpretable knowledge boundary. Each ships pinch a matched Unfiltered control for cleanable comparison.
Findings
Elicitation, not acquisition
In our experiments, scaling, SFT+GRPO post-training, and in-context learning amplify what the program taught, but nary meaningfully improves out-of-scope performance, indicating that the pretraining select sets the effective capacity ceiling.
Model checkpoints
LittleLearner astatine 3 scales (0.6B / 1.3B / 5B), each pinch a matched Unfiltered control sharing its architecture, tokens, and recipe.
Base: the pretrained model.
GRPO: mathematics specialists post-trained connected MathCAMPS; responses whitethorn grounds a inclination toward math-oriented output.
Chatty: variants tuned for wide chat behavior.
Capability stays wrong the curriculum
Can modular interventions push a exemplary past what its pretraining information taught it? With the bound nether experimental control, we tin inquire cleanly. In our experiments, each involution amplifies in-scope ability; nary of them meaningfully improves out-of-scope performance.
Scaling
Scaling exemplary size improves capacity wrong the model’s controlled knowledge vulnerability and extends modestly to problems on the aforesaid learning trajectory, but yields small betterment on problems requiring much precocious capabilities extracurricular the exposure.
MathCAMPS accuracy by grade, crossed exemplary size
Post-training
Post-training done GRPO importantly boosts in-scope K–5 capabilities, but fails to retrieve out-of-scope beyond-K–5 capabilities, moreover erstwhile training pinch out-of-scope data.
Post-training amplifies K–5, not the beyond-K–5 gap
In-context learning
In-context learning pinch the prompts we trial does not unlock caller reasoning capabilities in beyond-K–5 for our trained 5B LittleLearner.
Accuracy by prompting condition
What will you thatch it?
Because LittleLearner’s training vulnerability is explicitly specified, behavioral and representational changes tin beryllium related straight to the concepts you introduce. Three directions we’re excited about:
01
RL & discovery
Can RL create capability?
The anterior is restricted to K–5, truthful capabilities that look nether RL tin beryllium attributed to the RL process itself. A tractable proxy for reward-driven discovery.
02
Continual learning
Watch a conception being learned
Introduce antagonistic numbers and measurement sample efficiency, retention, and interference. Or probe behaviour adjacent the boundary: does it answer, abstain, aliases hallucinate?
03
Educational science
Machine vs. kid learners
Specified vulnerability enables controlled human-model comparison. Do models and children need akin vulnerability to study fractions, aliases make akin errors connected connection problems?
+
Your turn
Bring your ain question
A known bound turns your thought into a cleanable experiment!
If you find this activity useful
Please mention our paper:
English (US) ·
Indonesian (ID) ·