Reusable Magnetic Sensor Uses SERS and AI to Detect Trace Uranium
SOURCE: BIOENGINEER.ORG
AUG 15, 2026
Forgetting may be the secret to better AI language learning
SOURCE: TECHXPLORE.COM
JUN 23, 2026
edited by Gaby Clark, reviewed by Andrew Zinin
Credit: Unsplash/CC0 Public Domain
Giving AI a human-like memory limitation may actually help it learn language better. In their new proof-of-principle study, Abishek Thamma (University of Amsterdam) and Micha Heilbron (Max Planck Institute for Psycholinguistics) show that small language models equipped with a transient memory learn grammar more efficiently when trained on child-scale amounts of language input. The findings demonstrate how insights from psycholinguistics can inspire new approaches to AI learning. The findings are published in the journal Transactions of the Association for Computational Linguistics.
The study builds on a longstanding idea in cognitive science: that limitations of human memory may actually support language learning. As people process language, the exact forms of words and sentences are quickly forgotten. Rather than being a disadvantage, this constraint may help learners focus on recurring patterns and acquire abstract grammatical knowledge.
To test whether this principle could also benefit artificial intelligence, the researchers introduced a human-like memory limitation into modern neural language models. While today's AI systems typically have access to much more detailed linguistic information than humans do, the results suggest that adding a transient memory can improve learning efficiency and grammatical generalization when training data are limited.
To address this, Thamma and Heilbron introduced a simple form of memory decay into Transformer language models, creating what they term fleeting memory transformers. Heilbron said: "The models were trained on the BabyLM benchmark, a data set designed to approximate the amount of linguistic input available to human learners during development. This enabled a controlled comparison between models with and without memory limitations under realistic data conditions."
The results provide consistent evidence that fleeting memory benefits language learning. Across training runs and model initializations, models equipped with memory decay achieved better language-modeling performance and stronger results on targeted evaluations of syntactic knowledge than standard Transformer models.
Heilbron continued: "Importantly, these benefits emerged only when memory decay was paired with a short 'echoic memory' buffer that preserved the most recent three to seven words. Together, these mechanisms appear to support learning by combining immediate access to local information with a gradual loss of more distant word forms."
The findings lend support to a longstanding proposal in cognitive science, dating to influential connectionist work by Elman (1993), that memory limitations can facilitate language acquisition rather than merely constrain it. They also suggest that the success of contemporary Transformer architectures does not imply that unrestricted memory is optimal for language learning.
At the same time, the study uncovered an unexpected dissociation, Thamma said: "Although fleeting memory improved language learning, it reduced the models' ability to predict human reading times using surprisal-based measures. This result runs counter to a common pattern in which improvements in language-modeling performance are associated with better prediction of human language processing behavior.
"Further analyses indicated that this discrepancy could not be explained by existing accounts of why stronger language models sometimes provide poorer fits to human reading-time data. The findings therefore suggest that the factors that support successful language learning may differ from those that support accurate prediction of online language processing."
Taken together, the study provides evidence that memory limitations can enhance language learning in modern neural networks, while also highlighting an important distinction between learning language effectively and modeling human behavior.
This study revisits a longstanding question in cognitive science through the lens of modern language models. The findings suggest that memory constraints continue to support language learning, even in contemporary neural networks, while also prompting new questions about how linguistic knowledge relates to the way humans process language.
More information
Abishek Thamma et al, Human-like Fleeting Memory Improves Language Learning but Impairs Reading Time Prediction in Transformer Language Models, Transactions of the Association for Computational Linguistics (2026). DOI: 10.1162/tacl.a.688
Key concepts
Large language modelsMachine learning methodologies
Provided by Max Planck Society
Who's behind this story?
MA in English, copy editor since 2021 with experience in higher education and health content. Dedicated to trustworthy science news. Full profile ?
Master's in physics with research experience. Long-time science news enthusiast. Plays key role in Science X's editorial success. Full profile ?
LATEST NEWS
Machine Learning
Machine Learning Identifies Predictors of Weight Loss in Adolescents With Obesity
AUG 15, 2026
WHAT'S TRENDING
Data Science
5 Imaginative Data Science Projects That Can Make Your Portfolio Stand Out
OCT 05, 2022
SOURCE: BIOENGINEER.ORG
AUG 15, 2026
SOURCE: BIOENGINEER.ORG
AUG 15, 2026
SOURCE: SCITECHDAILY.COM
AUG 01, 2026
SOURCE: QUANTUMZEITGEIST.COM
JUL 25, 2026
SOURCE: QUANTUMZEITGEIST.COM
JUL 25, 2026