LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LITTLECURRICULUM yields LITTLELE
Record details
Published: 13 August 2026
Source: arXiv cs.AI
Category: Research
Topics: Children & education
Retrieved: 14 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension
arXiv · 13 August 2026
Jointly Predicting Courses and Grades Using a Transformer-Based Model
arXiv cs.AI · 13 August 2026
Exploring Perceptions of Leveraging AI to Improve Outcomes in Maternal, Sexual, and Reproductive Health in Sub-Saharan Africa: Exploratory Qualitative Study
JMIR (Journal of Medical Internet Research) · 13 August 2026
Visibility Asymmetry: How Vendor Attention Shapes Which EdTech Breakdowns Become Product-Visible
arXiv cs.CY · 14 August 2026
Impact of introducing "Informatics I" to the common university entrance examination in Japan: a longitudinal study on students' perceptions of their information-related knowledge and skills from 2006 to 2026
arXiv cs.CY · 14 August 2026
Strategy-Oriented Feedback for Fostering Systematic Problem-Solving in Machine Learning Education
arXiv cs.CY · 14 August 2026
How to cite this record
ethics.ai (13 August 2026), “LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure,” evidence record 19434, https://ethics.ai/record/19434 (originally published by arXiv cs.AI).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.