Abstract
Overlearning is a form of learning: the continued practice of material or a skill after the point at which it can first be performed without error. Its amount is expressed as a percentage of the trials that reaching mastery originally took, so that a further block of practice equal to the original is 100 percent overlearning. A century of study, from Krueger's list-learning experiments to meta-analyses of skill retention, shows that overlearning does improve later retention, but with sharply diminishing returns against a steeply rising cost in time. Recent neuroscience finds a mechanism, a rapid shift toward inhibitory neurochemistry that hyperstabilizes the freshly learned skill. This article surveys what overlearning is, how it is measured, the size and limits of its effect, and why spacing is often the better use of the same practice.
Keywords: overlearning, retention, degree of learning, distributed practice, skill acquisition
Overlearning refers to practice that continues beyond the trial on which a learner first reaches a criterion of errorless performance (Krueger, 1929). The phenomenon sits at the centre of a practical question every student and trainer faces: once a skill can be performed correctly, is there value in continuing to drill it? The cognitive answer is a qualified yes. Extra practice does buy more durable memory, but each additional increment buys less than the last, and the same practice invested differently, spread across time rather than piled up in one session, usually buys more (Rohrer & Taylor, 2006).
- Overlearning is practice continued past the first errorless performance, quantified as a percentage of the trials mastery first required.
- It reliably improves later retention, but the benefit shows steep diminishing returns as the degree of overlearning rises.
- Its advantage decays over the retention interval, so overlearning helps most when the test is soon and matters less over long delays.
- Distributing the same practice across spaced sessions generally yields more durable memory than massing it as overlearning.
- At the neural level, overlearning rapidly makes processing inhibitory-dominant, hyperstabilizing the new skill against interference.
What Overlearning Is
Overlearning is defined by reference to a criterion of mastery. A learner practises until some standard of correct performance is first met, one perfect recitation of a word list, one flawless run of a procedure, and every trial after that point is overlearning. Because the baseline is the effort mastery took, overlearning is naturally measured as a proportion of that effort: if reaching the criterion took ten trials and the learner then performs ten more, the material has been overlearned by 100 percent; twenty further trials would be 200 percent (Postman, 1962).
The construct matters because the first errorless performance is a deceptively weak state of knowledge. Being just able to recite a list once does not mean the memory is robust; it means the memory has crossed a threshold. Overlearning addresses the gap between barely knowing something and knowing it so well that it survives a delay, resists interference, and can be executed automatically while attention is elsewhere. This is why overlearning is bound up with the durability of retention and with the automaticity that distinguishes a practised motor skill from a merely acquired one (Driskell, Willis, & Copper, 1992).
Measuring Overlearning
The classic method fixes a criterion, records the trials needed to reach it, and then administers a controlled amount of further practice defined relative to that number. Krueger's original design had participants learn lists of nouns to one perfect recitation and then gave separate groups 0, 50, 100, or 200 percent additional practice, with retention tested after intervals of days to weeks (Krueger, 1929). Because the overlearning is pegged to each learner's own trials-to-criterion, the manipulation equates degree of overlearning across fast and slow learners rather than giving everyone the same absolute amount of practice.
The degree of overlearning
Overlearning is practice continued past the first perfect performance. Its amount is expressed as a percentage of the trials that mastery first took. Here mastery arrives on trial 8; drag to add practice beyond it and watch the percentage and the study-time cost climb.
The design makes the central trade-off measurable. On one axis is the degree of overlearning; on the other, the cost, which is simply the extra time the additional trials consume. Reaching 100 percent overlearning doubles the study time spent on the item, and 200 percent triples it, so any retention benefit has to be weighed against a linear escalation in effort (Rohrer & Taylor, 2006). The method thus does not merely ask whether overlearning helps, but whether it helps enough to justify what it costs, a question that later work would answer largely in the negative for durable learning.
Figure 1
Retention as a Diminishing Function of the Degree of Overlearning
The Overlearning Effect and Its Limits
The basic effect is robust: overlearned material is retained better than material practised only to criterion. The logic of that comparison traces to Ebbinghaus, who first measured retention as the savings in relearning a list one had previously mastered and charted the forgetting curve along which those savings decay with time (Ebbinghaus, 1913). Overlearning acts on that curve by raising its starting height: Krueger found that 100 percent overlearning produced substantially better recall than no overlearning, and Postman replicated the ordering while sharpening its most important qualification, that the increment from 100 to 200 percent was small relative to the increment from 0 to 100 (Postman, 1962). The retention function is concave: it rises quickly and then flattens, so that the first units of extra practice are far more valuable than the last.
Diminishing returns of overlearning
More overlearning lifts later retention, but not in proportion to the effort. The first block of overlearning buys a large gain; a second, equal block buys much less. Set the degree of overlearning and read the modelled retention after 28 days.
The most careful quantitative summary comes from a meta-analysis of overlearning and retention, which confirmed a positive average effect while exposing its two decisive limits. First, the benefit is moderated by the retention interval: overlearning helps most at short delays and its advantage shrinks as time passes, so the extra practice buys a head start that erodes rather than a permanent gain. Second, the effect is smaller for the kinds of cognitive tasks most common in education than for simple physical or procedural tasks (Driskell, Willis, & Copper, 1992). Because the cost of overlearning is fixed and paid up front while its benefit decays, the return on the investment falls as the horizon lengthens.
| Study | Design | Central lesson |
|---|---|---|
| Krueger (1929) | Word lists learned to 0, 100, or 200% overlearning. | Overlearning improves retention, but 200% adds little over 100%. |
| Postman (1962) | Systematic variation of degree of overlearning. | The retention function is concave; diminishing returns are the rule. |
| Driskell et al. (1992) | Meta-analysis across tasks and retention intervals. | The benefit decays over time and is larger for physical than cognitive tasks. |
From Overlearning to Spacing
The limits of overlearning are best understood against a rival use of the same practice. Overlearning by definition masses the extra trials into a single session, immediately after criterion is reached. But identical trials distributed across separate sessions produce markedly more durable memory, a robust result known as the spacing or distributed-practice effect. In a direct comparison of the two strategies for retaining mathematics knowledge, spacing outperformed overlearning at a delayed test, leading to the conclusion that overlearning is an inefficient use of study time when durability is the goal (Rohrer & Taylor, 2006).
Same practice, massed or spaced
Overlearning packs the extra practice into one sitting. Spacing the very same trials across separate sessions is a desirable difficulty: it feels harder and slower, yet it leaves more behind at a delayed test. Choose how many extra trials to spend and compare the durable payoff.
This fits a broader principle. Conditions that make practice feel harder and slower in the moment, spacing, interleaving, and retrieval rather than rereading, often produce better long-term learning than conditions that make performance feel fluent and easy; Bjork calls these desirable difficulties, and the fluency of overlearned material is precisely the kind of easy, confidence-inflating performance that misleads learners about how well they will remember (Bjork, Dunlosky, & Kornell, 2013). Reviews of learning techniques accordingly rank distributed practice and practice testing among the most effective strategies, while treating overlearning as a weak technique whose main value is diagnostic rather than durable (Dunlosky, Rawson, Marsh, Nathan, & Willingham, 2013). The practical successor to overlearning is not more massed drilling but scheduling: relearning the material to criterion on spaced occasions, an approach that captures the durability overlearning cannot (Rawson & Dunlosky, 2011).
Worked Example
Suppose a student practises a vocabulary list and first recites it without error on trial 8. That eighth trial marks criterion; everything after is overlearning. If the student then completes 8 more trials, the degree of overlearning is 8 divided by 8, or 100 percent, for a total of 16 trials. Completing 16 further trials instead would be 200 percent overlearning, 24 trials in all.
Now attach the cost. At roughly two minutes per trial, reaching criterion took about 16 minutes. Achieving 100 percent overlearning brings the total to 16 trials, about 32 minutes, exactly double the baseline; 200 percent brings it to 24 trials, about 48 minutes, or triple. The retention side of the ledger runs the other way. Following the concave function that Krueger and Postman established, the jump from 0 to 100 percent overlearning yields a large gain in later recall, while the equally expensive jump from 100 to 200 percent yields only a small one (Postman, 1962). The lesson of the arithmetic is that the second block of overlearning costs as much as the first but returns far less, and that a student with 48 minutes to spend would very likely retain more by reaching criterion once and then relearning the list to criterion again on two later days than by piling all the practice into a single overlearned session (Rohrer & Taylor, 2006).
Current Directions
The behavioural picture of overlearning is a century old, but its neural basis is recent. Recording visual-cortex neurochemistry during a perceptual-learning task, one influential study found that overlearning triggers a rapid switch from excitatory to inhibitory dominance, chiefly a swift rise in GABA, that hyperstabilizes the just-acquired skill and shields it from interference by subsequent learning. Training only to criterion left the memory labile and vulnerable to being overwritten; a short bout of overlearning locked it in (Shibata et al., 2017). This gives overlearning a concrete function, converting a fragile new trace into a protected one, and reframes it as a mechanism of accelerated memory consolidation rather than mere repetition.
Follow-up work has connected this stabilization to the wider machinery of consolidation and reconsolidation, showing that the two processes share behavioural and neurochemical signatures and that the excitatory-inhibitory balance is a common lever on memory stability (Bang et al., 2018). On the applied side, a recent meta-analytic review of procedural skill retention and decay has updated the quantitative estimates that guide training in medicine, aviation, and the military, where the cost of a skill lapsing is high and overlearning has long been used deliberately to buy a margin of safety against decay (Tatel & Ackerman, 2025). The open questions now concern how the neurochemical account of hyperstabilization maps onto the behavioural diminishing-returns curve, and how best to combine a stabilizing dose of overlearning with the durability of spaced relearning.
Key Researchers
Hermann Ebbinghaus (1850-1909). University of Berlin; he founded the experimental study of memory and devised the savings method of relearning on which the very idea of overlearning depends, showing that practice beyond first mastery leaves measurable traces. Wikipedia
Robert A. Bjork (b. 1939). University of California, Los Angeles; he drew the distinction between learning and performance and formulated the theory of desirable difficulties, which explains why the fluency of overlearned material overstates how well it will be retained. ORCID - Wikipedia
Doug Rohrer. University of South Florida; he ran the mathematics-retention experiments that pitted overlearning against distributed practice and showed that spacing the same trials yields more durable memory. ORCID
Takeo Watanabe. Brown University; he led the perceptual-learning research showing that overlearning hyperstabilizes a newly acquired skill by rapidly shifting cortical processing toward inhibitory dominance. ORCID
Kazuhisa Shibata. RIKEN Center for Brain Science; as lead author of the hyperstabilization study he identified the rapid rise in the inhibitory-to-excitatory neurochemical ratio that fixes an overlearned skill in place. ORCID
Discussion
Overlearning is a rare case in which a full century of research has converged on a nuanced but stable verdict. The effect is real: practice past the point of first mastery does make memory more durable and performance more automatic, which is why it remains standard in domains where a lapse is costly, from emergency procedures to the fundamentals of a motor skill. Yet the same literature shows that overlearning is an inefficient way to spend practice when the goal is long-term retention, because its returns diminish steeply and its advantage decays with time, while the identical practice distributed across sessions does better on both counts (Rohrer & Taylor, 2006).
The tension between the behavioural and the neural accounts is where the topic is most alive. Behaviourally, overlearning looks like a weak, dispensable technique; neurally, a brief bout of it performs a specific and valuable job, stabilizing a fragile new trace against interference (Shibata et al., 2017). These are not contradictory. They suggest that a small dose of overlearning may be worth its cost precisely for protecting freshly learned material in the minutes after acquisition, while durability over weeks is better served by spaced relearning. The practical recommendation that follows is neither to drill endlessly nor to stop at the first correct performance, but to secure the new memory briefly and then return to it on a schedule (Dunlosky et al., 2013).
Glossary
- Automaticity.
- The capacity to perform a skill quickly and with little attention, a state that extended practice such as overlearning helps to reach.
- Criterion of mastery.
- The standard of correct performance, such as one errorless recitation, that marks the boundary between original learning and overlearning.
- Degree of overlearning.
- The amount of practice given past criterion, expressed as a percentage of the trials mastery originally required.
- Desirable difficulty.
- A condition that slows or hinders practice in the moment yet improves long-term learning, such as spacing or retrieval practice.
- Diminishing returns.
- The pattern in which each additional increment of overlearning adds less to retention than the increment before it.
- Distributed practice.
- Practice spread across separate sessions rather than massed together; generally more durable than overlearning for the same number of trials.
- Forgetting curve.
- The decline of retention as a function of the time since learning, along which overlearning shifts the starting point upward.
- Hyperstabilization.
- The rapid neural fixing of a just-learned skill by a shift toward inhibitory processing, making it resistant to interference.
- Massed practice.
- Practice concentrated in a single session, the form overlearning takes when extra trials follow immediately after criterion.
- Memory consolidation.
- The process by which a new memory becomes stable over time; overlearning appears to accelerate an early, neurochemical form of it.
- Overlearning.
- Practice of material or a skill continued past the point at which it can first be performed without error.
- Retention interval.
- The delay between the end of learning and the test of memory, over which the advantage of overlearning tends to shrink.
- Savings method.
- Ebbinghaus's technique of gauging retention by how much practice is saved on relearning, the measurement logic underlying overlearning.
- Successive relearning.
- Relearning material to criterion on repeated, spaced occasions; the durable successor to massed overlearning.
- Trials to criterion.
- The number of practice trials a learner needs to first reach the mastery standard, the denominator against which overlearning is scaled.
Frequently Asked Questions
What is overlearning?
Overlearning is practice that continues after material or a skill can first be performed without error; the extra practice is measured as a percentage of the trials mastery first required (Krueger, 1929).
How is the degree of overlearning calculated?
It is the number of trials given past criterion divided by the number of trials it took to reach criterion, expressed as a percentage; practising as long again as it took to master something is 100 percent overlearning (Postman, 1962).
Does overlearning actually improve retention?
Yes, on average overlearned material is retained better than material practised only to criterion, but the gains show steep diminishing returns and the advantage decays over the retention interval (Driskell, Willis, & Copper, 1992).
Why does overlearning show diminishing returns?
The retention function is concave: the first block of extra practice moves memory from a fragile threshold to a robust state, while later blocks refine an already-strong memory and so add much less (Postman, 1962).
Is overlearning better than spacing out practice?
Usually not for durable memory. Distributing the same trials across separate sessions typically produces better long-term retention than massing them as overlearning (Rohrer & Taylor, 2006).
Why does overlearned material feel so well known yet fade?
Overlearning inflates immediate fluency, which learners mistake for durable knowledge; this is why easy, confident performance is a poor guide to later retention (Bjork, Dunlosky, & Kornell, 2013).
What happens in the brain during overlearning?
Overlearning rapidly shifts cortical processing toward inhibitory dominance, a rise in GABA that hyperstabilizes the new skill and protects it from being overwritten by later learning (Shibata et al., 2017).
When is overlearning worth the cost?
It is most justified when a skill must not fail and the test may come soon, as in emergency or safety-critical training, or as a brief dose to stabilize a memory that will later be maintained by spaced relearning (Tatel & Ackerman, 2025).
References
Bang, J. W., Shibata, K., Frank, S. M., Walsh, E. G., Greenlee, M. W., Watanabe, T., & Sasaki, Y. (2018). Consolidation and reconsolidation share behavioural and neurochemical mechanisms. Nature Human Behaviour, 2(7), 507-513. https://doi.org/10.1038/s41562-018-0366-8
Bjork, R. A., Dunlosky, J., & Kornell, N. (2013). Self-regulated learning: Beliefs, techniques, and illusions. Annual Review of Psychology, 64(1), 417-444. https://doi.org/10.1146/annurev-psych-113011-143823
Driskell, J. E., Willis, R. P., & Copper, C. (1992). Effect of overlearning on retention. Journal of Applied Psychology, 77(5), 615-622. https://doi.org/10.1037/0021-9010.77.5.615
Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). Improving students' learning with effective learning techniques: Promising directions from cognitive and educational psychology. Psychological Science in the Public Interest, 14(1), 4-58. https://doi.org/10.1177/1529100612453266
Ebbinghaus, H. (1913). Memory: A contribution to experimental psychology (H. A. Ruger & C. E. Bussenius, Trans.). Teachers College, Columbia University. (Original work published 1885)
Krueger, W. C. F. (1929). The effect of overlearning on retention. Journal of Experimental Psychology, 12(1), 71-78. https://doi.org/10.1037/h0072036
Postman, L. (1962). Retention as a function of degree of overlearning. Science, 135(3504), 666-667. https://doi.org/10.1126/science.135.3504.666
Rawson, K. A., & Dunlosky, J. (2011). Optimizing schedules of retrieval practice for durable and efficient learning: How much is enough? Journal of Experimental Psychology: General, 140(3), 283-302. https://doi.org/10.1037/a0023956
Rohrer, D., & Taylor, K. (2006). The effects of overlearning and distributed practise on the retention of mathematics knowledge. Applied Cognitive Psychology, 20(9), 1209-1224. https://doi.org/10.1002/acp.1266
Shibata, K., Sasaki, Y., Bang, J. W., Walsh, E. G., Machizawa, M. G., Tamaki, M., Chang, L.-H., & Watanabe, T. (2017). Overlearning hyperstabilizes a skill by rapidly making neurochemical processing inhibitory-dominant. Nature Neuroscience, 20(3), 470-475. https://doi.org/10.1038/nn.4490
Tatel, C. E., & Ackerman, P. L. (2025). Procedural skill retention and decay: A meta-analytic review. Psychological Bulletin, 151(6), 696-736. https://doi.org/10.1037/bul0000481