Abstract
Child language is a form of language development: the developing linguistic system of a child, studied as a snapshot at successive ages rather than as the process of acquiring it. This article describes how that system is measured — through mean length of utterance and Brown’s stages — and how it grows across morphology, the lexicon, and syntax. It reviews the classic evidence that children command productive grammatical rules, from Berko’s wug test to the U-shaped course of overregularization, and the referential and expressive styles that distinguish the earliest vocabularies. It surveys the corpus and parent-report methods, from CHILDES to the MacArthur-Bates inventories, that turned child language into a shared, quantitative science. Three interactive demonstrations explore mean length of utterance and Brown’s stages, the order in which grammatical morphemes are acquired, and the rise and fall of overregularization.
Keywords: child language, mean length of utterance, grammatical morphemes, overregularization, CHILDES
Child language is the language a child produces and comprehends at a given point in development — the phonology, vocabulary, and grammar of the child considered as a system in its own right, not yet the adult system and no longer nothing. Where language development names the process of acquisition, child language names its product at each stage: what the two-year-old’s grammar actually is, described on its own terms (Brown, 1973; Clark, 2009). The distinction is the one the National Library of Medicine draws in its Medical Subject Headings, which files child language beneath language development as the language “spoken or understood by children at a particular maturational stage.”
Treating the child’s language as a system has a long payoff: it reveals that early speech is not a degraded copy of adult speech but rule-governed in its own right. A child who says two foots has not simply failed to learn feet; she has applied a plural rule so reliably that it overrides a form she once produced correctly. Reading child language this way — as a grammar to be described rather than a set of errors to be corrected — is the founding move of the modern field (Brown, 1973).
The productive system studied here rests on perceptual foundations laid in the first year. Infants tune from universal to native-language listeners, losing the ability to discriminate non-native contrasts as they commit to their own language (Werker & Tees, 1984; Kuhl, 2004), and they segment word-like units from continuous speech by tracking statistical regularities (Saffran et al., 1996). By the time children are combining words, comprehension already runs well ahead of production (Bergelson & Swingley, 2012). What follows is the growth of the child’s own productive grammar and lexicon.
- Child language is the child’s developing linguistic system described as a snapshot at a given stage, the product of acquisition rather than the process; MeSH files it beneath language development.
- Mean length of utterance (MLU), the average number of morphemes per utterance, indexes early grammatical growth and defines Brown’s five stages better than chronological age does.
- Berko’s wug test showed that even preschoolers apply productive morphological rules to invented words, proving that grammar is inferred rather than memorized word by word.
- Overregularization (goed, foots) follows a U-shaped course — correct, then rule-generalized error, then correct again — and is the signature of an emerging rule, though its measured rate is low.
- Shared corpora (CHILDES) and parent-report inventories (the MacArthur-Bates CDI, aggregated in Wordbank) turned child language into a reproducible, cross-linguistic science.
Figure 1
Brown’s Five Stages Indexed by Mean Length of Utterance
Measuring the System: MLU and Brown’s Stages
The central problem in describing child language is that chronological age is a poor yardstick: two children of the same age can differ enormously in grammatical maturity (Fenson et al., 1994). Roger Brown’s solution, from his longitudinal study of three children he called Adam, Eve, and Sarah, was to index development by the mean length of utterance — the average number of morphemes, not words, per utterance (Brown, 1973). Counting morphemes rather than words is what makes MLU sensitive to grammar: dogs counts as two morphemes (dog plus plural -s) and jumped as two (jump plus past -ed), so a child who has begun to inflect words scores higher than one producing the same number of bare stems.
Brown divided early acquisition into five stages defined by MLU bands (Table 1). Stage I (MLU 1.0–2.0) is the period of single words and the first telegraphic two-word combinations — more milk, mommy sock — stripped of grammatical morphemes. Grammatical morphemes appear in Stage II, sentence modalities such as questions and negation in Stage III, clause embedding in Stage IV, and clause coordination in Stage V. The value of the scale is that it captures the sequence of grammatical achievements in a single, replicable number, and the ordering of stages is far more stable across children than the ages at which they are reached.
| Stage | MLU | Approx. age | Grammatical hallmark |
|---|---|---|---|
| I | 1.0–2.0 | 15–30 mo | Single words and telegraphic two-word combinations |
| II | 2.0–2.5 | 28–36 mo | Grammatical morphemes appear (-ing, plural, in/on) |
| III | 2.5–3.0 | 36–42 mo | Sentence modalities: questions and negation |
| IV | 3.0–3.75 | 40–46 mo | Embedding of one clause within another |
| V | 3.75–4.5 | 42–52+ mo | Coordination of clauses |
The first demonstration makes MLU concrete: assemble a short sample of child utterances, each annotated with its morpheme count, and watch the mean length of utterance and the corresponding Brown stage update as the sample changes.
A fixed five-utterance sample. Toggle each bound morpheme to see how the inflections — not the words — drive the mean length of utterance.
Bound morphemes count separately, so stripping the three inflections drops the total from 14 to 11 and the MLU from 2.8 to 2.2 — moving the child from Stage III to Stage II. Stage boundaries follow Brown (1973); the age-to-stage mapping is deliberately omitted because it is loose.
Morphological Development
If child language is rule-governed, the strongest proof is that children extend rules to words they have never heard. Jean Berko’s 1958 wug test supplied exactly that proof. Shown a cartoon creature and told “This is a wug,” then shown two and prompted “Now there are two ___,” preschoolers reliably supplied wugs — a word that appears in no input and could not have been memorized (Berko, 1958). Because wug is novel, the plural could only have come from an internalized rule, and the same held for past tense, possessives, and derivations. The wug test remains the canonical demonstration that morphology is productive from the earliest stages.
Brown’s longitudinal data added a second regularity: the fourteen grammatical morphemes he tracked — the present progressive -ing, the prepositions in and on, the plural, the possessive, articles, past tense, third-person agreement, and the forms of the copula and auxiliary be — are acquired in a strikingly consistent order across children (Brown, 1973). The progressive -ing comes early; the contractible auxiliary comes late; and the sequence is largely the same from child to child, even though the ages vary. The ordering is not explained by frequency in the input alone: it tracks the grammatical and semantic complexity of each morpheme. The second demonstration steps through this acquisition sequence, pairing each morpheme with an example and its place in the order.
Children master these grammatical morphemes in a strikingly consistent order, regardless of how fast they progress. Step through the sequence.
Currently at rank 1: Present progressive -ing. The progressive -ing and the prepositions in and on come early; the contractible auxiliary comes last. Order is consistent across children even when rate is not (Brown, 1973).
The most telling evidence that children build rules is that they overapply them. Having produced irregular past tenses such as went and came correctly — presumably by rote — children later begin to say goed and comed, errors they could not have heard, and only still later recover the correct forms. This U-shaped course is the classic signature of a rule being extracted and then constrained (Marcus et al., 1992). Marcus and colleagues, analyzing overregularization across thousands of utterances, made the phenomenon precise: overregularization is real and diagnostic, but its rate is low — a median of around 2.5% of irregular past-tense uses — and it persists over years rather than appearing as a sudden collapse. The dip in correct performance is genuine but shallow, a correction to the folk image of a dramatic regression. The third demonstration traces this rise and fall of correct irregular production across development.
A child first says went correctly, then — on discovering the -ed rule — sometimes says goed, before recovering. Accuracy on irregular past-tense forms dips and returns.
“went” produced correctly, as an unanalyzed whole
The dip is real but shallow: across large samples the median overregularization rate was only about 2.5%, so errors like goed are the striking minority, not the rule (Marcus et al., 1992).
The Early Lexicon
The child’s growing vocabulary is not merely smaller than an adult’s; it is organized differently and grows in characteristic ways. Katherine Nelson’s study of the first fifty words revealed that children differ in style: some build a referential vocabulary dominated by object names, while others favor an expressive vocabulary rich in social and personal-interaction phrases (Nelson, 1973). Across children and, more strongly, across many languages, early nouns tend to be over-represented relative to verbs and function words — the noun bias — though its universality is debated and depends on the structure of the language being learned.
Early word meanings are also systematically off in ways that reveal the child’s categories. Children overextend, using dog for all four-legged animals or moon for round objects, before their semantic boundaries settle to adult conventions (Clark, 2009). These are not random errors: they follow the child’s perceptual and functional groupings, and they retreat as the lexicon fills in and words come to contrast with one another. Comprehension outpaces production throughout, so a child who says few words may nonetheless understand many (Bergelson & Swingley, 2012), and the pace of lexical growth is shaped by the quantity and quality of the language the child hears (Hoff, 2006; Fernald et al., 2013). Large parent-report samples show that the “typical” vocabulary at any age is really a wide band, with the fastest and slowest learners differing by hundreds of words (Fenson et al., 1994; Frank et al., 2021).
Methods: How Child Language Is Studied
Child language became a cumulative science when its evidence was made shareable. The single most important instrument is the Child Language Data Exchange System (CHILDES), the corpus and toolset Brian MacWhinney and Catherine Snow founded to pool transcribed child speech in a common format, with programs for computing measures such as MLU automatically (MacWhinney, 2000). Before CHILDES, transcripts sat in individual filing cabinets; after it, any researcher could reanalyze another’s data or test a hypothesis across dozens of children and many languages. Careful description of the input the child hears — the distinctive register of child-directed speech first characterized in detail by Snow — is part of the same tradition of taking naturalistic samples seriously (Snow, 1972).
The complementary instrument is the parent-report inventory. The MacArthur-Bates Communicative Development Inventories ask caregivers to check off the words and constructions their child produces, yielding norms from very large samples at low cost (Fenson et al., 1994). Aggregating these inventories across tens of thousands of children and dozens of languages, the open Wordbank database now lets researchers chart vocabulary trajectories and cross-linguistic regularities with a precision no single laboratory could reach (Frank et al., 2017). A sobering caveat runs through all of this: the evidence base is heavily skewed toward English and a handful of other well-studied languages, so claims about “the child” rest on a narrow and unrepresentative sample of the world’s languages (Kidd & Garcia, 2022).
Worked Example
Because MLU is the field’s workhorse measure, it is worth computing by hand. The rule is Brown’s: count the total number of morphemes across a sample of utterances and divide by the number of utterances. Bound morphemes count separately — the plural -s, the progressive -ing, the past -ed — but irregular forms and compounds count as one. Take the five-utterance sample in Table 2.
| Utterance | Morphemes counted | Count |
|---|---|---|
| Mommy sock | mommy + sock | 2 |
| Two shoes | two + shoe + -s | 3 |
| Doggie running | doggie + run + -ing | 3 |
| I want juice | I + want + juice | 3 |
| The balls | the + ball + -s | 3 |
| Total | 5 utterances | 14 |
The sample contains 14 morphemes across 5 utterances, so MLU = 14 ÷ 5 = 2.8. An MLU of 2.8 falls in the 2.5–3.0 band, placing this child in Brown’s Stage III — consistent with the emergence of questions and negation, and roughly the language of a three-to-three-and-a-half-year-old, though the age is only a loose guide. Notice how the bound morphemes drive the result: had the child produced two shoe, doggie run, and the ball without inflections, the total would fall to 11 and the MLU to 2.2, dropping the child into Stage II. The measure is doing exactly what Brown intended — registering grammatical growth that a word count would miss. The first demonstration recomputes this calculation for different samples.
Discussion
The study of child language reframed a developmental puzzle as a problem in grammar. By insisting that the child’s language is a system to be described rather than a set of adult targets missed, Brown and his contemporaries made it possible to ask precise questions: what is the child’s rule, how is it ordered relative to other rules, and what does its overapplication reveal (Brown, 1973; Berko, 1958)? The answers — productive morphology from the start, a consistent order of acquisition, U-shaped overregularization — are among the most robust findings in developmental science, and they set the explananda that every theory of acquisition must account for.
Those findings also feed the field’s central theoretical debate without settling it. That children extend rules to novel words and overregularize is compatible both with accounts positing innate grammatical machinery and with usage-based accounts on which rules are generalizations abstracted from stored exemplars. The consistent morpheme order, and evidence that ultimate attainment depends on the age at which acquisition begins (Newport, 1990), have been read as reflecting a maturationally constrained system; the low, extended rate of overregularization and its sensitivity to a word’s frequency have been read as favoring gradual, input-driven learning (Marcus et al., 1992). The description of child language is neutral between these readings, which is precisely why it is valuable: it is the shared evidence over which the theories contend.
Current Directions
Two shifts define the field’s active front. The first is a turn toward individual differences. The classic descriptions charted the average child, but children vary widely in rate and style, and researchers increasingly treat that variation as the signal to be explained — through differences in input, memory, and processing — rather than as noise around a universal path (Kidd et al., 2018; Kidd & Donnelly, 2020). Large open corpora and inventories make this possible, letting analyses hold estimates of “typical” development to a far higher evidential standard (Frank et al., 2017).
The second is a reckoning with the narrowness of the evidence base. The overwhelming majority of what is known about child language comes from English and a few other richly studied languages, and a systematic audit found the field’s samples to be strikingly unrepresentative of the world’s linguistic and cultural diversity (Kidd & Garcia, 2022). Correcting this — through corpora and inventories in under-studied languages — is now a central methodological priority, because regularities such as the noun bias or a given morpheme order cannot be claimed as universal until they are tested against the full range of human languages.
Common Misconceptions
- Errors like “goed” and “foots” mean a child’s language is going backwards.
- The opposite: overregularization is positive evidence that the child has extracted a productive rule and is applying it, which is why it appears after correct rote forms and follows a U-shaped course (Marcus et al., 1992).
- Vocabulary size is the best single measure of a child’s language level.
- Grammar develops partly independently of the lexicon; mean length of utterance, which counts morphemes, was devised precisely because it indexes grammatical growth that a word count misses (Brown, 1973).
- All children learn their first words in the same way.
- Children differ in style from the outset — some build an object-naming referential vocabulary, others a socially oriented expressive one — and they vary widely in rate (Nelson, 1973; Fenson et al., 1994).
Glossary
- Brown’s stages.
- Five stages of early grammatical development (I–V) defined by bands of mean length of utterance rather than by age, from telegraphic two-word speech through the coordination of clauses.
- Child-directed speech.
- The distinctive register caregivers use with young children — higher pitch, slower tempo, exaggerated intonation — part of the input whose systematic study began with Snow (1972).
- CHILDES.
- The Child Language Data Exchange System, a shared archive of transcribed child speech in a common format with tools for automatic analysis, part of the larger TalkBank system.
- Expressive style.
- An early-vocabulary profile weighted toward social and personal-interaction phrases rather than object names; contrasted with the referential style.
- Grammatical morpheme.
- A bound or free morpheme carrying grammatical rather than lexical meaning — the plural -s, the progressive -ing, articles, the copula — the class whose ordered acquisition Brown charted.
- MacArthur-Bates Communicative Development Inventories.
- Parent-report checklists of the words and constructions a child produces or understands, used to obtain vocabulary norms cheaply from very large samples.
- Mean length of utterance (MLU).
- The average number of morphemes per utterance in a language sample; the standard index of early grammatical development, computed as total morphemes divided by number of utterances.
- Morphology.
- The system of word structure — how stems combine with prefixes, suffixes, and inflections — the domain in which children demonstrate productive rules earliest and most clearly.
- Noun bias.
- The tendency for nouns, especially object names, to be over-represented in early vocabularies relative to verbs and function words; robust in many languages but debated in its universality.
- Overextension.
- The use of a word beyond its adult range — dog for all four-legged animals — following the child’s perceptual or functional categories before meanings settle to convention.
- Overregularization.
- The application of a regular rule to an irregular form — goed, foots — taken as evidence of an extracted rule; real but low in rate and U-shaped over development.
- Referential style.
- An early-vocabulary profile dominated by object names, associated with a focus on labelling things; contrasted with the expressive style.
- Telegraphic speech.
- Early multi-word speech that omits grammatical morphemes and function words, retaining mainly content words — mommy sock, more milk — characteristic of Brown’s Stage I.
- U-shaped development.
- A course in which performance is initially correct, declines as a rule is over-generalized, then recovers — the classic pattern for irregular past-tense forms.
- Wug test.
- Berko’s (1958) task in which children supply the plural or other inflection of an invented word (a wug), demonstrating that morphological rules are productive rather than memorized.
Key Researchers
Roger Brown (1925-1997). Late professor of psychology at Harvard University; his longitudinal study of Adam, Eve, and Sarah established mean length of utterance, Brown’s stages, and the ordered acquisition of grammatical morphemes in A First Language (1973). Wikipedia - Wikidata
Eve V. Clark (b. 1942). Professor of linguistics at Stanford University; her work on early word meaning, overextension, and the principles of contrast and conventionality anchors the study of the child’s lexicon. ORCID - Wikipedia - Wikidata - Google Scholar - Faculty Page
Jean Berko Gleason (b. 1931). Professor emerita of psychology at Boston University; creator of the 1958 wug test, the founding demonstration that children apply productive morphological rules to novel words. Wikipedia - Wikidata - Google Scholar - Faculty Page
Erika Hoff (b. 1951). Professor of psychology at Florida Atlantic University; her review of how social context shapes language and her research on bilingual and dual-language development are widely cited. ORCID - Wikipedia - Wikidata - Google Scholar - Faculty Page
Evan Kidd (contemporary). Researcher at the Australian National University and the Max Planck Institute for Psycholinguistics, Nijmegen; a leader of work on individual differences in acquisition and on the representativeness of the child-language evidence base. ORCID - Google Scholar - Faculty Page
Brian MacWhinney (b. 1945). Professor of psychology at Carnegie Mellon University; co-founder of the CHILDES corpus and the TalkBank system, and developer of the Competition Model of acquisition. ORCID - Wikipedia - Wikidata - Google Scholar - Faculty Page
Katherine Nelson (1930-2018). Late professor of psychology at the CUNY Graduate Center; documented the referential and expressive styles and the composition of the earliest vocabularies, and later founded work on event representation in memory. Wikipedia - Wikidata
Catherine E. Snow (b. 1945). Professor at the Harvard Graduate School of Education; her founding work on child-directed speech and on the language foundations of later literacy shaped the study of the input to child language. ORCID - Wikipedia - Wikidata - Google Scholar - Faculty Page
Michael Tomasello (b. 1950). Professor of psychology at Duke University and formerly co-director of the Max Planck Institute for Evolutionary Anthropology; the leading proponent of the usage-based, social-pragmatic account of how children construct grammar. Wikipedia - Wikidata - Google Scholar - Faculty Page
Frequently Asked Questions
What is child language?
Child language is the developing linguistic system of a child at a given stage — its phonology, vocabulary, and grammar considered in their own right rather than as an incomplete version of the adult system. In the MeSH classification it is filed beneath language development (Brown, 1973; Clark, 2009).
How is child language different from language development?
Language development is the process of acquiring language; child language is the product of that process at each point, the grammar and lexicon the child actually has at a given stage. The two are closely related, and MeSH files child language as a subtype of language development (Clark, 2009).
What is mean length of utterance?
Mean length of utterance, or MLU, is the average number of morphemes per utterance in a language sample. Because it counts morphemes rather than words, it registers grammatical growth, and it defines Brown’s five stages of early development (Brown, 1973).
What are Brown’s stages?
Five stages of early grammatical development defined by bands of MLU rather than by age, running from telegraphic two-word speech in Stage I through the coordination of clauses in Stage V. The order is consistent across children even though the ages vary (Brown, 1973).
What is the wug test?
A task in which a child is shown a novel creature called a wug and asked to complete “Now there are two ___.” Preschoolers reliably answer “wugs,” showing that they apply a productive plural rule to a word they could not have memorized (Berko, 1958).
Why do children say things like “goed” and “foots”?
These overregularizations are evidence that the child has extracted a regular rule and is applying it even to irregular forms. They follow a U-shaped course and, although diagnostic, occur at a low rate (Marcus et al., 1992).
How do researchers study child language?
Chiefly through shared transcript corpora such as CHILDES, which pool child speech in a common analyzable format, and through parent-report inventories such as the MacArthur-Bates CDI, aggregated in the open Wordbank database (MacWhinney, 2000; Frank et al., 2017).
Do all languages show the same patterns of child language?
Not necessarily, and this is a live concern. Most of what is known comes from English and a few other languages, and audits show the evidence base is unrepresentative of the world’s linguistic diversity, so regularities cannot be assumed universal until tested more broadly (Kidd & Garcia, 2022).
References
Bergelson, E., & Swingley, D. (2012). At 6–9 months, human infants know the meanings of many common nouns. Proceedings of the National Academy of Sciences, 109(9), 3253-3258. https://doi.org/10.1073/pnas.1113380109
Berko, J. (1958). The child’s learning of English morphology. Word, 14(2–3), 150-177. https://doi.org/10.1080/00437956.1958.11659661
Brown, R. (1973). A first language: The early stages. Harvard University Press.
Clark, E. V. (2009). First language acquisition (2nd ed.). Cambridge University Press.
Fenson, L., Dale, P. S., Reznick, J. S., Bates, E., Thal, D. J., & Pethick, S. J. (1994). Variability in early communicative development. Monographs of the Society for Research in Child Development, 59(5), 1-185. https://doi.org/10.2307/1166093
Fernald, A., Marchman, V. A., & Weisleder, A. (2013). SES differences in language processing skill and vocabulary are evident at 18 months. Developmental Science, 16(2), 234-248. https://doi.org/10.1111/desc.12019
Frank, M. C., Braginsky, M., Yurovsky, D., & Marchman, V. A. (2017). Wordbank: An open repository for developmental vocabulary data. Journal of Child Language, 44(3), 677-694. https://doi.org/10.1017/S0305000916000209
Frank, M. C., Braginsky, M., Yurovsky, D., & Marchman, V. A. (2021). Variability and consistency in early language learning: The Wordbank project. MIT Press.
Hoff, E. (2006). How social contexts support and shape language development. Developmental Review, 26(1), 55-88. https://doi.org/10.1016/j.dr.2005.11.002
Kidd, E., Donnelly, S., & Christiansen, M. H. (2018). Individual differences in language acquisition and processing. Trends in Cognitive Sciences, 22(2), 154-169. https://doi.org/10.1016/j.tics.2017.11.006
Kidd, E., & Donnelly, S. (2020). Individual differences in first language acquisition. Annual Review of Linguistics, 6, 319-340. https://doi.org/10.1146/annurev-linguistics-011619-030326
Kidd, E., & Garcia, R. (2022). How diverse is child language acquisition research? First Language, 42(6), 703-735. https://doi.org/10.1177/01427237211066405
Kuhl, P. K. (2004). Early language acquisition: Cracking the speech code. Nature Reviews Neuroscience, 5(11), 831-843. https://doi.org/10.1038/nrn1533
MacWhinney, B. (2000). The CHILDES project: Tools for analyzing talk (3rd ed., Vol. 1). Lawrence Erlbaum Associates.
Marcus, G. F., Pinker, S., Ullman, M., Hollander, M., Rosen, T. J., & Xu, F. (1992). Overregularization in language acquisition. Monographs of the Society for Research in Child Development, 57(4), 1-182. https://doi.org/10.2307/1166115
Nelson, K. (1973). Structure and strategy in learning to talk. Monographs of the Society for Research in Child Development, 38(1/2), 1-135. https://doi.org/10.2307/1165788
Newport, E. L. (1990). Maturational constraints on language learning. Cognitive Science, 14(1), 11-28. https://doi.org/10.1207/s15516709cog1401_2
Saffran, J. R., Aslin, R. N., & Newport, E. L. (1996). Statistical learning by 8-month-old infants. Science, 274(5294), 1926-1928. https://doi.org/10.1126/science.274.5294.1926
Snow, C. E. (1972). Mothers’ speech to children learning language. Child Development, 43(2), 549-565. https://doi.org/10.2307/1127555
Werker, J. F., & Tees, R. C. (1984). Cross-language speech perception: Evidence for perceptual reorganization during the first year of life. Infant Behavior and Development, 7(1), 49-63. https://doi.org/10.1016/S0163-6383(84)80022-3