Abstract
Learning Curve is a type of learning: the quantitative relationship between how much a task has been practiced and how well it is performed. Charted first by Ebbinghaus and named by Snoddy, it became a formal law when Newell and Rosenbloom argued that response time falls as a power function of practice. That claim is now contested: fits to individual learners often favor an exponential function, and the smooth group curve can be an artifact of averaging over abrupt, learner-specific transitions. This article defines the curve as a measure, traces its history from memory research to industrial cost accounting, sets out the power-law-versus-exponential debate, and distinguishes the phases of skill acquisition that generate it. Three interactive demonstrations fit competing functions, run an industrial doubling curve, and show a smooth average emerging from step-like individual data.
Keywords: learning curve, power law of practice, skill acquisition, automatization, practice
- A learning curve plots a performance measure — time, error rate, or cost — against the amount of practice, and its shape is the object of study, not just its downward slope. - The classic power law of practice says response time declines as a power function of trials; it is linear on log–log axes and implies ever-slowing improvement. - Fits to individual learners frequently favor an exponential law, and a smooth group power curve can be a statistical artifact of averaging over learners who improve abruptly at different times. - Skill acquisition passes through distinct phases — a slow cognitive phase, a faster associative phase, and an autonomous phase — which is why real curves show plateaus, dips, and leaps rather than one smooth descent. - The same mathematics recurs far outside the laboratory: Wright’s industrial learning curve describes how unit cost falls by a fixed fraction each time cumulative output doubles.
What a Learning Curve Is
A learning curve is a graph of a performance measure against the amount of practice that produced it. The horizontal axis counts experience — trials, repetitions, hours, or cumulative units produced — and the vertical axis records how well the task is done: the time taken to complete it, the proportion of errors, or the cost per unit. Because performance almost always improves with practice, the curve typically descends (for time, error, or cost) or rises (for accuracy or output), and the rate at which it does so is what the measure captures.
The learning curve is therefore best understood as an operational measure of learning rather than a theory of it. Two learners who both reach the same final skill can trace very different paths to it: one improving quickly then leveling off, another crawling through an early plateau before a sudden gain. The curve makes those differences visible and quantifiable, which is precisely why it has been adopted as a diagnostic tool in fields as far apart as memory research, motor skill, industrial engineering, and medical training.
Two features recur across domains. First, improvement is almost never linear: early practice buys large gains and later practice buys progressively smaller ones, a pattern of diminishing returns. Second, the curve usually approaches an asymptote — a performance floor (for time or error) or ceiling (for accuracy) that further practice cannot breach. The mathematical debate over the learning curve is, at heart, an argument about the exact functional form of that diminishing-returns approach to the asymptote.
A Short History of the Curve
The empirical study of the curve began with Hermann Ebbinghaus, who in the 1880s measured his own memory using nonsense syllables and the method of savings: relearning a list took fewer trials than learning it fresh, and the savings declined in an orderly way with the delay since first study. The resulting forgetting and relearning curves were the first quantitative descriptions of how learning changes with experience.
The term learning curve entered psychology through motor-skill research. George Snoddy (1926) charted mirror-tracing performance over repeated attempts and used the curve to distinguish two components of practice, which he called adaptation and facilitation. By the mid-twentieth century the curve was a standard tool for describing acquisition in tasks from typing to telegraphy.
The curve also had an independent life outside the laboratory. Studying airframe manufacture, T. P. Wright (1936) observed that the labor required to build an airplane fell by a roughly constant percentage each time the cumulative number built doubled — an industrial learning curve now central to cost estimation and operations management. That the same functional form describes a factory’s output and a single person’s reaction time is one of the more striking regularities in the study of practice.
The modern theoretical treatment arrived when Allen Newell and Paul Rosenbloom (1981) surveyed dozens of datasets and proposed that practice curves obey a single mathematical law — the power law — which the next section sets out.
The Power Law of Practice
Newell and Rosenbloom’s (1981) power law of practice states that the time T to perform a task declines as a power function of the number of practice trials N:
T = a + b · N−c
Here a is the asymptote (the irreducible floor that infinite practice would approach), b scales the amount of improvement available, and c is the learning rate — the exponent that governs how fast the curve bends toward its floor. The signature of a power function is that it becomes a straight line when both axes are plotted logarithmically: log(T − a) is linear in log(N). Newell and Rosenbloom found this log–log linearity across an unusually wide range of tasks and argued it reflected a common mechanism — the progressive chunking of task components into larger, faster-to-execute units.
The power law makes a strong and testable prediction: the relative rate of improvement is constant. Each doubling of practice yields the same proportional reduction in time, no matter how much practice has already accumulated. This is what gives the curve its characteristic shape of steep early gains giving way to a long, shallow tail — and it is why experts continue to improve, imperceptibly, over years of additional practice.
Two later theories supplied mechanisms for why a power law might emerge. John Anderson’s (1982) ACT theory attributed the speedup to knowledge compilation, in which slow, interpretive rule-following is gradually replaced by fast, task-specific procedures. Gordon Logan’s (1988) instance theory offered a different account: with practice, a learner accumulates stored memory traces of past solutions, and performance shifts from slow computation to fast retrieval of the best-matching instance — a race whose expected winning time falls as a power function of the number of stored instances.
The Exponential-Law Challenge
The power law’s status as an empirical regularity was overturned by a methodological argument. Andrew Heathcote, Scott Brown, and Douglas Mewhort (2000) pointed out that the datasets Newell and Rosenbloom relied on were almost all averaged over subjects. When they refitted practice data at the level of the individual learner and individual condition, an exponential law —
T = a + b · e−c · N
— fit better than the power law for the large majority of cases. The two functions look deceptively similar over a limited range of practice, which is why the distinction had gone unnoticed, but they differ in a theoretically crucial way: the exponential law has a constant relative learning rate per trial, whereas the power law’s relative rate slows as practice accumulates.
Why, then, does averaged data look so cleanly power-shaped? Because averaging exponential curves with different rates produces a curve that is itself well-approximated by a power function. The smooth group power law can be an artifact of aggregation rather than a property of any individual learner — a specific instance of the general hazard that a group average need not resemble any member of the group. Evans, Brown, Mewhort, and Heathcote (2018) refined this analysis, showing that once individual differences in learning rate and asymptote are modeled explicitly, the exponential form is favored and the apparent power law dissolves.
The most radical version of the argument comes from Gallistel, Fairhurst, and Balsam (2004), who examined acquisition trial by trial in individual animals and found that learning is frequently abrupt: performance stays near baseline and then steps up suddenly, at a trial that varies from subject to subject. A smooth, gradual group curve is precisely what averaging over such step functions produces, even though no individual ever traced a smooth curve. On this view the group learning curve may describe the distribution of onset times across learners rather than the dynamics within any one of them.
Phases of Skill Acquisition
If a single learner’s curve is not one smooth power function, what is it? A long tradition in cognitive psychology holds that skill acquisition passes through qualitatively distinct phases, each with its own dynamics, and that the observed curve is the concatenation of them. The canonical three-stage scheme is Paul Fitts and Michael Posner’s (1967): a cognitive stage of understanding the task, an associative stage of refining and connecting its components, and an autonomous stage of fast, low-effort execution. Anderson’s (1982) ACT account adopts the same three labels and supplies a mechanism for the transitions: an early cognitive phase, in which the learner works through explicit, verbalizable instructions; an associative phase, in which those instructions are compiled into smoother procedures and errors are pruned; and an autonomous phase, in which performance is fast, effortless, and resistant to interference.
Tenison and Anderson (2016) gave this phase account a quantitative test, using hidden-state models to identify when individual learners transition between a computation-based strategy and a retrieval-based one on the same problems. They found that the overall power-law-like speedup decomposes into two things: a within-phase speedup and a shift in the mix of strategies the learner is using — consistent with Logan’s instance theory operating within a staged architecture. The learning curve, on this synthesis, is smooth only because it sums over a discrete, strategy-switching process underneath.
This phase structure explains why real curves are lumpy. Gray and Lindstedt (2017) documented plateaus, dips, and leaps in extended practice: periods of no visible improvement, temporary regressions in which performance gets worse, and sudden jumps. They argued these are not noise but the visible trace of a learner abandoning a good-enough strategy, suffering a transient cost while a new one is assembled, and then leaping ahead once it works. Improvement, in other words, sometimes requires getting worse first — a pattern no monotonic power or exponential function can capture.
Worked Example
The industrial learning curve makes the mathematics concrete. Wright’s (1936) rule is stated as a learning rate: an 80% learning curve means that every time cumulative output doubles, the time (or cost) per unit falls to 80% of its previous value. Formally, the time for unit N is TN = T1 · Nb, where b = log2(learning rate).
Take a first-unit time of T1 = 100 hours and an 80% curve. The exponent is b = log2(0.80) = −0.322. Then:
- Unit 1: 100 × 1−0.322 = 100.0 hours - Unit 2: 100 × 2−0.322 = 80.0 hours (80% of unit 1) - Unit 4: 100 × 4−0.322 = 64.0 hours (80% of unit 2) - Unit 8: 100 × 8−0.322 = 51.2 hours (80% of unit 4) - Unit 16: 100 × 16−0.322 = 40.96 hours
Each doubling — from 1 to 2, 2 to 4, 4 to 8 — multiplies the per-unit time by exactly 0.80, which is the defining property of a power-law curve: a constant proportional gain per doubling of experience. A 70% curve (a steeper learner) would drop to 70% per doubling and reach unit 8 at just 34.3 hours; a 95% curve (a shallow learner) would still need 85.7 hours. The interactive industrial-curve demonstration below computes the unit and cumulative totals for any first-unit cost and learning rate.
| Form | Equation | Relative learning rate | Best fit at |
|---|---|---|---|
| Power law | T = a + b·N−c | Slows as practice accumulates | Averaged (group) data |
| Exponential law | T = a + b·e−c·N | Constant per trial | Individual-learner data |
| Step function | Abrupt onset at trial N0 | Zero, then a jump | Single-subject acquisition |
| Industrial (Wright) | TN = T1·Nlog2(LR) | Constant per doubling | Cumulative production cost |
Why an averaged curve can misrepresent the individual
Note. Each dashed gold line is one learner who improves abruptly at a different trial; the solid navy line is their average, which rises smoothly and resembles a power curve no individual traced. Original schematic.
Discussion
The learning curve is a rare object in cognitive psychology: a measure so robust that it recurs across memory, motor skill, problem solving, and industrial production, yet whose exact form remains genuinely contested. The resolution that has emerged is not that one law is right and the others wrong, but that the scale of analysis determines what is seen. At the level of a factory’s cumulative output or a large group of learners, a power law describes the data well and supports useful forecasting. At the level of a single learner, an exponential law usually fits better, and at the finest grain — trial by trial in one individual — acquisition can be abrupt and step-like, with the smooth curve emerging only under aggregation.
This has a practical consequence that reaches well beyond theory. In the health professions, learning curves are increasingly used to certify competence — how many supervised procedures must a trainee perform before their error rate falls to an acceptable floor? Pusic, Boutis, Hatala, and Cook (2015) warned that fitting a group curve and reading a single “number needed to reach competence” off it can badly mislead, precisely because individual trainees follow different curves with different asymptotes. A curve that is an averaging artifact cannot say when this trainee is ready. The measurement question and the theoretical question turn out to be the same question.
The learning curve also connects to its mirror image, the forgetting curve: what practice builds, disuse erodes, and the two curves together describe the full temporal dynamics of a skill or memory. An account of how quickly performance improves is incomplete without an account of how quickly it decays, which is why models of skill retention increasingly fit acquisition and forgetting jointly rather than in isolation.
Current Directions
Recent work has moved decisively toward modeling learning curves at the individual level with explicit statistical machinery rather than fitting an average and hoping it generalizes. Evans and colleagues’ (2018) hierarchical treatment, which estimates each learner’s own rate and asymptote while pooling information across learners, is representative of this shift: it recovers the exponential form that individual data favor while explaining why aggregate data look power-shaped. The same tools are being applied to education and clinical training, where the goal is a personalized curve that can flag when a specific learner has plateaued or is ready to advance.
A second active thread concerns the lumpiness of real practice. Gray and Lindstedt’s (2017) analysis of plateaus, dips, and leaps reframes non-monotonic performance not as measurement noise but as evidence of strategy discovery — the observable signature of a learner reorganizing how they do the task. This has renewed interest in detecting when and why strategy shifts occur, using the hidden-state and change-point methods that Tenison and Anderson (2016) brought to skill acquisition. Understanding the mechanics of a plateau — and how to break one — is now as central to the field as characterizing the smooth descent that surrounds it.
Common Misconceptions
- “A steep learning curve means something is hard to learn.”
- In common speech “steep learning curve” means difficult, but on the power law’s own definition the slope is the rate of improvement, so a steep curve shows rapid gains per unit of practice — fast learning. A genuinely hard-to-master skill produces a shallow curve (Newell & Rosenbloom, 1981).
- “The learning curve is always a smooth power function.”
- The smooth power curve is largely a property of averaged data. Individual learners more often fit an exponential law, and at the single-subject level acquisition can be abrupt and step-like (Heathcote et al., 2000).
- “A plateau means learning has stopped.”
- Plateaus are often periods of covert reorganization, during which a learner abandons one strategy and assembles a better one. Performance may even dip before it leaps, so flat or worsening output need not mean no learning is occurring (Gray & Lindstedt, 2017).
- “A group learning curve reveals how any single learner will improve.”
- Because averaging can create a curve that no member of the group traced, reading an individual’s expected trajectory — or a “number needed to reach competence” — off a group curve can be seriously misleading (Pusic et al., 2015).
Glossary
- Associative phase.
- The middle stage of skill acquisition, in which explicit instructions are compiled into smoother procedures and errors are pruned.
- Asymptote.
- The performance floor (for time, error, or cost) or ceiling (for accuracy) that a learning curve approaches but never reaches; the a term in the power and exponential laws.
- Automatization.
- The process by which a practiced task comes to be performed quickly and with little effort or attention, often by a shift from computation to memory retrieval.
- Averaging artifact.
- A feature of aggregated data — such as a smooth power curve — that arises from combining individuals and is present in no individual’s data.
- Chunking.
- The grouping of task components into larger units that can be executed as one; Newell and Rosenbloom’s proposed mechanism for the power law.
- Cognitive phase.
- The earliest stage of skill acquisition, in which the learner works through explicit, verbalizable instructions and performance is slow and error-prone.
- Diminishing returns.
- The pattern whereby early practice yields large gains and later practice yields progressively smaller ones.
- Exponential law of practice.
- The claim that performance improves as an exponential function of practice, with a constant relative learning rate per trial; favored by individual-level fits.
- Forgetting curve.
- The complement of the learning curve: the decline in retained performance as a function of time since practice.
- Industrial learning curve.
- Wright’s observation that unit production time or cost falls by a fixed fraction each time cumulative output doubles.
- Instance theory.
- Logan’s account in which practice accumulates stored memory traces, and performance shifts from computation to retrieval of the best-matching instance.
- Learning rate.
- The parameter (c, or the percentage in an industrial curve) that governs how quickly a learning curve bends toward its asymptote.
- Plateau.
- A stretch of practice over which performance shows little or no visible improvement, often masking covert strategy reorganization.
- Power law of practice.
- The claim that performance time declines as a power function of practice, appearing linear on log–log axes; historically the dominant description of the learning curve.
- Savings.
- Ebbinghaus’s measure of retained learning: the reduction in trials needed to relearn material compared with learning it fresh.
- Skill acquisition.
- The process of becoming proficient at a task through practice, generally described as passing through cognitive, associative, and autonomous phases.
Key Researchers
John R. Anderson (b. 1947). Architect of the ACT-R cognitive architecture; his knowledge-compilation account explains the power-law speedup as a shift from interpretive rule-following to task-specific procedures. Wikipedia - Wikidata - CMU Faculty - Google Scholar
Hermann Ebbinghaus (1850–1909). Founder of the experimental study of memory; his savings method produced the first quantitative learning and forgetting curves. Wikipedia - Wikidata
Andrew Heathcote (b. 1962). Co-author of the exponential-law critique showing that individual-level practice data favor an exponential over a power function. ORCID - Wikidata - Newcastle Faculty - Google Scholar
Gordon D. Logan (b. 1949). Proposed instance theory, giving a memory-retrieval mechanism for why practice speeds performance in a power-law fashion. Wikipedia - Wikidata - Vanderbilt Faculty - Google Scholar
Allen Newell (1927–1992). With Rosenbloom, formalized the power law of practice and proposed chunking as its mechanism. Wikipedia - Wikidata
Paul S. Rosenbloom (b. 1954). Co-formulated the power law of practice and the chunking account of skill acquisition. Wikidata - USC Viterbi Faculty - Google Scholar
Frequently Asked Questions
What is a learning curve?
A learning curve is a graph of how a performance measure — time, error rate, or cost — changes with the amount of practice. It quantifies the rate at which someone or something improves with experience.
Does a “steep learning curve” mean something is hard to learn?
No — that is the everyday usage, but technically it is backwards. A steep curve shows rapid improvement per unit of practice, meaning the skill is being learned quickly. Difficult skills produce shallow curves.
What is the power law of practice?
It is the claim that the time to perform a task declines as a power function of the number of practice trials, which appears as a straight line on log–log axes. It implies that each doubling of practice yields the same proportional improvement (Newell & Rosenbloom, 1981).
Is the power law actually correct?
It fits averaged group data well but has been challenged. When practice data are analyzed at the level of individual learners, an exponential law usually fits better, and the smooth group power curve can be an artifact of averaging over individuals (Heathcote et al., 2000).
Why do learning curves have plateaus?
Plateaus are often periods during which a learner is covertly reorganizing their strategy. Performance may stay flat or even temporarily worsen before jumping ahead once a new, more efficient approach is in place (Gray & Lindstedt, 2017).
What are the phases of skill acquisition?
A common account distinguishes a cognitive phase (working through explicit instructions), an associative phase (compiling those into smoother procedures), and an autonomous phase (fast, effortless performance) (Anderson, 1982).
What is an industrial learning curve?
It is Wright’s observation that unit production cost falls by a fixed percentage each time cumulative output doubles. An “80% curve,” for example, means each doubling of units cuts per-unit time to 80% of its previous value (Wright, 1936).
How is the learning curve related to the forgetting curve?
They are complementary. The learning curve describes how performance improves with practice; the forgetting curve describes how it decays with disuse. Together they capture the full temporal dynamics of a skill or memory.
References
Anderson, J. R. (1982). Acquisition of cognitive skill. Psychological Review, 89(4), 369–406. https://doi.org/10.1037/0033-295X.89.4.369
Evans, N. J., Brown, S. D., Mewhort, D. J. K., & Heathcote, A. (2018). Refining the law of practice. Psychological Review, 125(4), 592–605. https://doi.org/10.1037/rev0000105
Fitts, P. M., & Posner, M. I. (1967). Human performance. Brooks/Cole.
Gallistel, C. R., Fairhurst, S., & Balsam, P. (2004). The learning curve: Implications of a quantitative analysis. Proceedings of the National Academy of Sciences, 101(36), 13124–13131. https://doi.org/10.1073/pnas.0404965101
Gray, W. D., & Lindstedt, J. K. (2017). Plateaus, dips, and leaps: Where to look for inventions and discoveries during skilled performance. Cognitive Science, 41(7), 1838–1870. https://doi.org/10.1111/cogs.12412
Heathcote, A., Brown, S., & Mewhort, D. J. K. (2000). The power law repealed: The case for an exponential law of practice. Psychonomic Bulletin & Review, 7(2), 185–207. https://doi.org/10.3758/BF03212979
Logan, G. D. (1988). Toward an instance theory of automatization. Psychological Review, 95(4), 492–527. https://doi.org/10.1037/0033-295X.95.4.492
Newell, A., & Rosenbloom, P. S. (1981). Mechanisms of skill acquisition and the law of practice. In J. R. Anderson (Ed.), Cognitive skills and their acquisition (pp. 1–55). Erlbaum.
Pusic, M. V., Boutis, K., Hatala, R., & Cook, D. A. (2015). Learning curves in health professions education. Academic Medicine, 90(8), 1034–1042. https://doi.org/10.1097/ACM.0000000000000681
Snoddy, G. S. (1926). Learning and stability: A psychophysiological analysis of a case of motor learning with clinical applications. Journal of Applied Psychology, 10(1), 1–36. https://doi.org/10.1037/h0075814
Tenison, C., & Anderson, J. R. (2016). Modeling the distinct phases of skill acquisition. Journal of Experimental Psychology: Learning, Memory, and Cognition, 42(5), 749–767. https://doi.org/10.1037/xlm0000204
Wright, T. P. (1936). Factors affecting the cost of airplanes. Journal of the Aeronautical Sciences, 3(4), 122–128. https://doi.org/10.2514/8.155