Abstract

Psycholinguistics is the scientific study of the mental representations and processes that let people comprehend, produce, and acquire language. It treats language not as an abstract formal system but as a cognitive capacity implemented in real time by a memory-limited, prediction-prone mind. Its central findings are that processing is incremental, interpreting each word before the next arrives; that comprehension is often good enough rather than complete; and that the mind actively predicts upcoming input. Its methods range from measuring reading eye movements and speech errors to recording event-related brain potentials such as the N400. Modern work increasingly compares these human signatures against the internal states of large language models.

Keywords: language comprehension, sentence processing, speech production, prediction, mental lexicon

Psycholinguistics occupies the intersection of psychology and linguistics, asking how the abstract structures linguists describe are actually built and unpacked by the mind during the fractions of a second of ordinary speech and reading. The field acquired its modern shape when Noam Chomsky's review of B. F. Skinner's Verbal Behavior argued that language could not be a set of conditioned habits, because speakers routinely produce and understand sentences they have never encountered (Chomsky, 1959). That argument moved the explanatory burden from stimulus and response onto internal representations and computations, and the study of those computations is what psycholinguistics has pursued ever since.

Key Takeaways
  • Psycholinguistics studies the real-time mental processes of comprehending, producing, and acquiring language, not the formal grammar itself.
  • Language processing is incremental: the mind interprets each word as it arrives rather than waiting for a phrase to complete.
  • Spoken words are recognized by progressively narrowing a cohort of candidates as the acoustic signal unfolds.
  • Comprehension is frequently good enough rather than exhaustive, and the mind predicts upcoming words, as indexed by the N400 brain potential.
  • Its evidence comes from eye movements, speech errors, and electrophysiology, and increasingly from comparison with large language models.

## What Psycholinguistics Is

Psycholinguistics is the branch of cognitive science concerned with the psychological mechanisms underlying language use. Where a grammar specifies what the well-formed sentences of a language are, psycholinguistics asks how a person computes them: how a continuous acoustic stream is segmented into words, how words are retrieved from a store of tens of thousands, how a string of words is assembled into a structured meaning, and how the reverse path turns an intention into articulated speech. Three problems organize the field: comprehension, production, and acquisition. Each is studied as an information-processing problem with measurable time course, capacity limits, and error signatures.

The defining commitment is that these processes run in real time under sharp constraints. A listener cannot wait until an utterance ends to begin interpreting it; working memory could not hold the unanalyzed material (Working Memory is the limited store implicated here). Comprehension is therefore incremental and, as later sections show, often predictive. This makes timing the field's central dependent measure: reaction times, reading times, eye fixations, and the millisecond-resolved electrical response of the brain are all ways of watching a representation being built.

The third problem, acquisition, asks how a learner assembles this machinery in the first place. A foundational demonstration showed that eight-month-old infants segment a continuous stream of nonsense syllables into word-like units after only two minutes of exposure, tracking the transitional probabilities between adjacent syllables — evidence that part of language learning rests on domain-general statistical learning rather than innate linguistic rules alone (Saffran et al., 1996). The result reframed the nativist-empiricist debate that Chomsky's critique had sharpened: powerful distributional learning and structured prior knowledge need not be mutually exclusive, and the same predictive, probability-tracking machinery that this article traces in adult comprehension appears to be at work in the infant learning the words in the first place.

## Types of Psycholinguistics

The U.S. National Library of Medicine's Medical Subject Headings (MeSH) index Psycholinguistics as a descriptor with two narrower descriptors beneath it, listed in Table 1. MeSH is an indexing classification — a controlled vocabulary for cataloguing the literature — not a theory of how the language faculty divides into parts, so this list reflects how publications are tagged rather than a claim about natural joints in the mind. The two subtypes are not mutually exclusive, and neither exhausts the field; the substantive divisions psycholinguists actually work with are the comprehension, production, and acquisition problems described above. The parent descriptors under which MeSH files psycholinguistics are Behavioral Sciences, Psychological Phenomena, and Linguistics.

SubtypeIn brief
Neurolinguistic ProgrammingAn approach that claims a link between neurological processes, language, and learned behavioral patterns; classified here as a language-related descriptor, though it is not part of mainstream experimental psycholinguistics.
Semantic DifferentialA rating technique that measures the connotative meaning of words and concepts along bipolar adjective scales, such as good–bad or strong–weak.

Table 1. Direct subtypes of Psycholinguistics in the MeSH classification (tree F02.694).

## Spoken Word Recognition

Recognizing a spoken word is harder than it feels, because speech arrives as a continuous signal with no reliable silences between words and unfolds over time rather than all at once. The influential cohort model proposed that the first fragment of a word activates in parallel every word in the mental lexicon consistent with it — the cohort — and that this set is winnowed as more of the signal arrives, until only one candidate remains (Marslen-Wilson, 1987). The moment at which the input first becomes consistent with only one word is the uniqueness point, and recognition is typically timed to it rather than to the word's acoustic offset. A word like elephant can be identified before its final syllable, because no other word begins eleph-.

Lexical access is also strikingly automatic and, at least briefly, context-blind. Using cross-modal priming, listeners hearing an ambiguous word such as bug in a sentence biased toward insects nonetheless showed momentary activation of the unrelated spy meaning, indicating that both senses are retrieved before context selects one (Swinney, 1979). Recognition is thus best understood as fast, parallel activation followed by rapid selection, not a serial search — a pattern closely related to Speech Perception and to Recognition more broadly.

Figure 1

The Levels of the Language Comprehension Stream

The language comprehension stream from acoustic input to discourse meaning Five stacked levels connected by upward feedforward arrows and downward predictive feedback arrows: acoustic-phonetic input, phonological form, lexical access, syntactic structure, and semantic and discourse meaning. Semantic / discourse meaning Syntactic structure Lexical access (mental lexicon) Phonological form Acoustic-phonetic input feedforward prediction
Note. Green arrows show bottom-up feedforward flow as the signal is analyzed; dashed red arrows show top-down predictive feedback, by which higher levels pre-activate representations at lower ones. The levels interact continuously rather than strictly in sequence. Original schematic.

The demonstration below makes the cohort dynamic concrete: as each phoneme of a spoken word is added, the set of lexical candidates consistent with the input shrinks toward the uniqueness point.

Demo 1

Cohort narrowing

Choose a spoken target, then reveal it one phoneme at a time. Watch how many words remain compatible with what has been heard so far.

  captain
candlecandycancelcaptaincapturecarpetcartoonbeaconbeakerbeetle
Phonemes heard: 0 of 7. Cohort size: 10. No input yet: every word in the lexicon is a candidate.
Note. As each phoneme of the target is added, the cohort of words consistent with the input shrinks. The green word marks the uniqueness point — the moment a single candidate survives, often before the word is complete. Orthographic prefixes stand in for the phoneme sequence. Original interactive model after Marslen-Wilson (1987).

## Sentence Processing and Parsing

Understanding a sentence requires assigning it a structure — deciding which words group together and what relations hold among them. Because structure is not marked explicitly in the input, the parser must make commitments before the evidence is complete, and this is most visible when it commits wrongly. In a garden-path sentence such as The horse raced past the barn fell, readers initially treat raced as the main verb and are forced into a costly reanalysis when fell arrives and reveals that raced past the barn was a reduced relative clause. Eye-movement recordings show longer fixations and regressive saccades precisely at the disambiguating word (Frazier & Rayner, 1982).

Two families of theory explain this. The garden-path model holds that an autonomous syntactic module first builds the structurally simplest analysis — following principles such as minimal attachment — and consults meaning only to revise (Frazier & Rayner, 1982). Constraint-based accounts instead treat parsing as the parallel, probabilistic integration of many cues at once — syntactic, lexical, and contextual — with ambiguity resolved by the weight of evidence rather than by a structure-first rule (MacDonald et al., 1994). Eye tracking during natural reading has been the primary arena for adjudicating these views, revealing the fine time course of comprehension and its sensitivity to word frequency and predictability (Rayner, 1998).

A decisive method for showing how quickly context enters is the visual world paradigm, in which listeners' eye movements to objects in a scene are recorded as they hear a sentence. Listeners fixate the likely referent of a phrase before it is fully spoken, and even fixate a plausible object at the verb: hearing The boy will eat… they look to the cake before the noun is uttered, showing that the verb's meaning is used predictively (Altmann & Kamide, 1999; Tanenhaus et al., 1995). Yet comprehension is not always thorough. The good-enough framework documents that readers often settle for shallow interpretations that are adequate for the task but demonstrably incomplete, misanalyzing sentences they could parse fully if pressed (Ferreira & Patson, 2007). The demonstration below lets the reader step through a garden-path sentence and see where reanalysis becomes necessary.

Demo 2

Walking into the garden path

Step through the sentence one word at a time. The bar shows the modeled reading time at each word; watch what happens when the disambiguating word arrives.

Thehorseracedpastthebarnfell
220 ms
Word 1 of 7The: Begin a noun phrase.
Note. Reading times are illustrative fixed values (ms). In the reduced relative clause the parser commits early to a main-verb reading and pays a large reanalysis cost at 'fell'; marking the clause overtly ('that was') removes the spike. Original interactive model after Frazier & Rayner (1982).

## Language Production

Production runs the comprehension problem in reverse: from an intended message to a sequence of articulatory gestures, in roughly the same fraction of a second. The dominant framework decomposes this into ordered stages — conceptualization of the message, grammatical encoding that selects words (lemmas) and builds syntax, phonological encoding that specifies sound form, and articulation (Levelt et al., 1999). The strongest evidence for such stages is the patterning of speech errors. Slips are not random: sounds exchange with sounds and words with words of the same category, as in a fart smeller for a smart feller, implying distinct levels at which units are selected and sequenced.

The leading mechanistic account is the spreading-activation model, in which activation flows through a network of conceptual, word, and phoneme nodes, and the most active unit at each level is selected (Dell, 1986). Its interactive character — activation feeding back between levels — explains why errors tend to be both phonologically similar to and real words of the language, a bias a strictly feedforward system would not produce. Production and comprehension were long studied apart, but influential work argues they share representations, with the production system used to generate predictions during comprehension (Pickering & Garrod, 2013).

## The Brain and the Prediction Debate

Electrophysiology gave psycholinguistics a continuous, millisecond-resolved window onto comprehension. The N400, a negative event-related potential peaking about 400 ms after a word, was discovered when semantically anomalous endings — He spread the warm bread with socks — evoked a large deflection absent for expected endings (Kutas & Hillyard, 1980). The N400 is now understood to index the ease of semantic access, scaling inversely with how expected a word is in its context. A later, distinct positivity around 600 ms, the P600, is associated with syntactic reanalysis and integration (Hagoort, 2019).

The N400's sensitivity to expectancy reframed comprehension as actively predictive: the brain pre-activates likely upcoming words, and the response to a word depends on how well it was anticipated (Federmeier, 2007). How fine-grained that prediction is became a live controversy. A large multi-laboratory replication tested whether readers pre-activate specific words to the grain of their initial phonemes and found the effect far weaker and less reliable than an influential earlier study had suggested, tempering strong claims about lexical prediction (Nieuwland et al., 2018). The debate turns on definition as much as data: prediction can mean anything from graded pre-activation of meaning to commitment to a specific form, and progress required distinguishing these senses explicitly (Kuperberg & Jaeger, 2016). Anatomically, comprehension is supported by a bilateral-but-left-dominant network organized into a dorsal stream mapping sound to articulation and a ventral stream mapping sound to meaning (Hickok & Poeppel, 2007). The demonstration below models how a word's predictability drives the amplitude of its N400.

Demo 3

From cloze probability to the N400

Set how predictable a word is in its context (its cloze probability). Less predictable words carry more surprisal and evoke a larger N400. The presets reproduce the worked example in the text.

400 msbaseline
Surprisal = −log₂(0.60) = 0.74 bits. Modeled N400 amplitude = 0.74 µV of negativity. A highly expected word (large p) is read faster and evokes a small N400; an unexpected word (small p) evokes a large one.
Note. Surprisal S = −log₂(p) in bits; modeled N400 amplitude = 1.0 µV × S (an illustrative slope). The trace is an ERP-style Gaussian negativity at 400 ms, plotted downward, whose depth scales with amplitude. Original interactive model after Kutas & Hillyard (1980) and Federmeier (2007).
LevelCore processSignature measure
PhonologicalSegmenting the signal, recognizing wordsCohort narrowing; uniqueness point
LexicalRetrieving word form and meaningPriming; N400 amplitude
SyntacticAssigning structure to word stringsGarden-path reading times; P600
Semantic / discourseIntegrating meaning; predictingAnticipatory eye movements; N400
ProductionMessage to articulationSpeech-error patterns; naming latency

Table 2. Levels of language processing and their characteristic experimental signatures.

## Worked Example

A convenient way to quantify predictability is surprisal, defined as the negative base-2 logarithm of a word's probability in its context, measured in bits: surprisal = −log₂(p). The steady empirical finding is that both reading time and N400 amplitude rise roughly linearly with surprisal, so a low-probability word is both slower to read and evokes a larger N400.

Consider the sentence frame The children went outside to…. Suppose a cloze study — in which many people fill the blank — yields the completion play with probability 0.60 and the completion read with probability 0.05. The surprisal of play is −log₂(0.60) = 0.74 bits, whereas the surprisal of read is −log₂(0.05) = 4.32 bits. The difference is 3.58 bits. If a simple linear model sets modeled N400 amplitude at 1.0 microvolt of negativity per bit of surprisal (an illustrative slope), the two words differ by about 3.6 microvolts of N400 — the expected word producing a much smaller response than the unexpected one, exactly the pattern the N400PredictionDemo generates. This is why cloze probability, gathered from an offline questionnaire, predicts an online brain response measured in milliseconds: both are driven by the same underlying quantity, the word's probability given its context.

## Discussion

The recurring theme across comprehension, production, and their neural signatures is that the language system trades completeness for speed. It commits early and incrementally, predicts what is coming, and accepts good-enough analyses — a design that is efficient under real-time and memory constraints but that produces the garden paths, misinterpretations, and speech errors the field uses as evidence. This reframes classic questions. The old modularity debate — whether an encapsulated syntactic processor operates before meaning and context intrude — has largely given way to interactive, probabilistic models in which multiple constraints combine continuously, though the degree and locus of that interaction remain contested (MacDonald et al., 1994).

Prediction has become the organizing construct, but it is a graded and contested one. The evidence that the system anticipates upcoming meaning is robust; the claim that it routinely pre-activates specific word forms is weaker and was overstated in some early reports (Nieuwland et al., 2018). Precisely because prediction spans several distinct operations, theoretical progress depends on saying which one is meant (Kuperberg & Jaeger, 2016). These issues connect psycholinguistics to broader accounts of the mind as a prediction engine, including Predictive Coding, and to the memory systems — Semantic Memory and Working Memory — on which language processing depends.

## Current Directions

The most active current front is the comparison of human language processing with large language models (LLMs) — neural networks trained only to predict the next word. When such models are used to predict human brain activity, the best-performing models account for a substantial share of the variance in neural responses to language, and their fit improves with their next-word prediction ability, suggesting that predictive processing may be a shared computational principle (Schrimpf et al., 2021). Direct intracranial recordings echo this: word-by-word, the contextual embeddings of deep language models track the brain's responses, and both show evidence of predicting words before they are heard (Goldstein et al., 2022). Analyses of naturalistic listening further reveal that the brain computes a hierarchy of predictions at several linguistic levels simultaneously, from phonemes to syntax to meaning (Heilbron et al., 2022).

These convergences sharpen an old question about what the language system is for. Individual-subject neuroimaging has established that the brain's language network is a functionally specialized system, largely distinct from the circuits supporting general reasoning, and one review argues on this basis that language is primarily a tool for communication rather than the medium of thought (Fedorenko et al., 2024). Whether the striking model-brain alignments reflect genuinely shared mechanisms or merely shared statistics of language is the question the next decade of work is set to resolve.

## Common Misconceptions

Psycholinguistics is about learning foreign languages.
Second-language learning is one topic within it, but the field is far broader: it studies the mental processes of comprehending, producing, and acquiring language in general, including in fluent native speakers. Its founding argument concerned the nature of the knowledge underlying ordinary language use, not classroom instruction (Chomsky, 1959).
Comprehension builds a complete, accurate analysis of every sentence.
People routinely construct interpretations that are merely good enough for the moment, missing details they could recover if required. Readers asked about sentences they have just read reliably endorse interpretations their own parse should have ruled out (Ferreira & Patson, 2007). The misconception survives because comprehension feels effortless and complete from the inside.
The brain predicts the exact next word.
Prediction in comprehension is real but graded and largely about meaning, not a firm bet on a specific word form. A large multi-laboratory replication found only weak, unreliable evidence that readers pre-activate a specific word to the grain of its initial sound (Nieuwland et al., 2018). The strong version spread because an early, striking result was widely cited before it was tested at scale.

## Glossary

Cloze probability.
The proportion of people who complete a sentence frame with a given word; a standard measure of how expected that word is in context.
Cohort model.
An account of spoken word recognition in which the onset of a word activates all consistent candidates in parallel, and the set is narrowed as more input arrives.
Garden-path sentence.
A sentence whose early words invite a structural analysis that a later word forces the reader to abandon and rebuild, revealing the parser's commitments.
Good-enough processing.
The tendency to build shallow interpretations that are adequate for the task rather than complete and fully accurate analyses.
Incrementality.
The property that comprehension proceeds word by word, interpreting each input as it arrives rather than buffering until a unit completes.
Lexical access.
The retrieval of a word's stored form and meaning from the mental lexicon when it is heard or read.
Mental lexicon.
The mind's store of known words, holding their sound, spelling, meaning, and grammatical properties.
Minimal attachment.
A parsing principle by which the processor prefers the syntactically simplest structure, adding incoming words with the fewest new nodes.
N400.
A negative brain potential peaking near 400 ms after a word whose amplitude reflects the ease of accessing its meaning, larger for unexpected words.
P600.
A positive brain potential around 600 ms associated with syntactic reanalysis and the integration of structure.
Parsing.
The process of assigning a grammatical structure to a string of words during comprehension.
Phoneme.
The smallest unit of speech sound that can distinguish one word from another in a language.
Prediction.
In comprehension, the pre-activation of likely upcoming linguistic material before it is encountered, ranging from graded meaning to specific form.
Spreading activation.
A mechanism in which activation flows through a network of connected representations, used to model retrieval in both production and the lexicon.
Statistical learning.
Learning the structure of input by tracking distributional regularities such as the transitional probabilities between adjacent elements, implicated in how infants segment speech into words.
Surprisal.
The negative logarithm of a word's probability in context; a measure of its unexpectedness that predicts reading time and N400 amplitude.
Syntactic ambiguity.
A property of a word string that permits more than one grammatical structure, and hence more than one interpretation.
Uniqueness point.
The moment in a spoken word at which the accumulated sound is consistent with only one lexical candidate, allowing recognition.
Visual world paradigm.
A method that records eye movements to objects in a scene during listening, revealing the moment-by-moment time course of interpretation.

## Key Researchers

Noam Chomsky (b. 1928). Laureate Professor of Linguistics at the University of Arizona and Institute Professor Emeritus at MIT; his review of Skinner's Verbal Behavior helped launch the cognitive study of language. Faculty Page - Google Scholar

Gary S. Dell (b. 1950). Professor Emeritus of Psychology at the University of Illinois Urbana-Champaign; formulated the spreading-activation theory of language production. Faculty Page

Kara D. Federmeier. Professor of Psychology at the University of Illinois Urbana-Champaign; used event-related potentials to link the N400 to prediction in language comprehension. Faculty Page - ORCID

Evelina Fedorenko. Professor in Brain and Cognitive Sciences at MIT; maps the brain's functionally specialized language network with individual-subject imaging. Faculty Page - ORCID

Fernanda Ferreira (b. 1960). Distinguished Professor of Psychology at the University of California, Davis; developed the good-enough theory of language comprehension. Faculty Page - ORCID

Peter Hagoort (b. 1954). Director at the Max Planck Institute for Psycholinguistics and Radboud University's Donders Institute; advanced the neurobiology of language and the P600. Faculty Page - ORCID

Gina R. Kuperberg. Professor of Psychology at Tufts University and researcher at Massachusetts General Hospital; developed predictive-coding accounts distinguishing the N400 and P600. Faculty Page - ORCID

Marta Kutas. Distinguished Professor Emeritus of Cognitive Science at the University of California, San Diego; co-discovered the N400 potential. Faculty Page - ORCID

Willem J. M. Levelt (b. 1938). Director Emeritus of the Max Planck Institute for Psycholinguistics; built the standard staged model of speech production. Faculty Page - Google Scholar

William Marslen-Wilson. Emeritus Honorary Professor of Language and Cognition at the University of Cambridge; developed the cohort model of spoken word recognition. Faculty Page - ORCID

Keith Rayner (1943-2015). Atkinson Family Chair of Psychology at the University of California, San Diego; established eye tracking as the central method for studying reading. Google Scholar

## Frequently Asked Questions

What is psycholinguistics?
Psycholinguistics is the scientific study of the mental processes by which people comprehend, produce, and acquire language. It treats language as a real-time cognitive capacity rather than an abstract formal system, and it uses timing measures such as reading times and brain potentials to study it (Chomsky, 1959).

How is psycholinguistics different from linguistics?
Linguistics describes the structure of languages, specifying which sentences are grammatical and why. Psycholinguistics asks how a mind builds and unpacks those structures in real time, under memory and speed constraints, making it a branch of cognitive psychology as much as of language study (Levelt et al., 1999).

How do people recognize spoken words so quickly?
According to the cohort model, the first sounds of a word activate every candidate consistent with them, and this set is narrowed as more of the signal arrives until one word remains. Recognition often occurs at the uniqueness point, before the word is fully spoken (Marslen-Wilson, 1987).

What is a garden-path sentence?
It is a sentence whose opening words lead the reader toward one structure that a later word makes impossible, forcing reanalysis. Eye movements show longer fixations and regressions at the disambiguating word, revealing the parser's earlier commitment (Frazier & Rayner, 1982).

What does the N400 measure?
The N400 is a brain potential peaking around 400 ms after a word whose size reflects how easily the word's meaning is accessed in context. Unexpected or anomalous words evoke a larger N400 than expected ones (Kutas & Hillyard, 1980).

Does the mind really predict upcoming words?
Comprehension is predictive in that the mind pre-activates likely upcoming meaning, as shown by anticipatory eye movements and expectancy effects. However, evidence that it predicts a specific word form is weaker, and a large replication tempered strong claims of lexical prediction (Nieuwland et al., 2018).

Is comprehension always complete and accurate?
No. People often build good-enough interpretations that suffice for the task while missing details a full analysis would capture. This shallow processing is a normal feature of efficient comprehension rather than a failure (Ferreira & Patson, 2007).

Why compare the brain to large language models?
Large language models are trained only to predict the next word, and the better they predict, the better their internal states account for human brain responses to language. This alignment suggests predictive processing may be a shared principle, though whether it reflects common mechanisms is unresolved (Schrimpf et al., 2021).

## References

Altmann, G. T. M., & Kamide, Y. (1999). Incremental interpretation at verbs: Restricting the domain of subsequent reference. Cognition, 73(3), 247-264. https://doi.org/10.1016/S0010-0277(99)00059-1

Chomsky, N. (1959). A review of B. F. Skinner's Verbal Behavior. Language, 35(1), 26-58. https://doi.org/10.2307/411334

Dell, G. S. (1986). A spreading-activation theory of retrieval in sentence production. Psychological Review, 93(3), 283-321. https://doi.org/10.1037/0033-295X.93.3.283

Federmeier, K. D. (2007). Thinking ahead: The role and roots of prediction in language comprehension. Psychophysiology, 44(4), 491-505. https://doi.org/10.1111/j.1469-8986.2007.00531.x

Fedorenko, E., Ivanova, A. A., & Regev, T. I. (2024). The language network as a natural kind within the broader landscape of the human brain. Nature Reviews Neuroscience, 25(5), 289-312. https://doi.org/10.1038/s41583-024-00802-4

Ferreira, F., & Patson, N. D. (2007). The 'good enough' approach to language comprehension. Language and Linguistics Compass, 1(1-2), 71-83. https://doi.org/10.1111/j.1749-818X.2007.00007.x

Frazier, L., & Rayner, K. (1982). Making and correcting errors during sentence comprehension: Eye movements in the analysis of structurally ambiguous sentences. Cognitive Psychology, 14(2), 178-210. https://doi.org/10.1016/0010-0285(82)90008-1

Goldstein, A., Zada, Z., Buchnik, E., Schain, M., Price, A., Aubrey, B., Nastase, S. A., Feder, A., Emanuel, D., Cohen, A., Jansen, A., Gazula, H., Choe, G., Rao, A., Kim, C., Casto, C., Fanda, L., Doyle, W., Friedman, D., … Hasson, U. (2022). Shared computational principles for language processing in humans and deep language models. Nature Neuroscience, 25(3), 369-380. https://doi.org/10.1038/s41593-022-01026-4

Hagoort, P. (2019). The neurobiology of language beyond single-word processing. Science, 366(6461), 55-58. https://doi.org/10.1126/science.aax0289

Heilbron, M., Armeni, K., Schoffelen, J.-M., Hagoort, P., & de Lange, F. P. (2022). A hierarchy of linguistic predictions during natural language comprehension. Proceedings of the National Academy of Sciences, 119(32), e2201968119. https://doi.org/10.1073/pnas.2201968119

Hickok, G., & Poeppel, D. (2007). The cortical organization of speech processing. Nature Reviews Neuroscience, 8(5), 393-402. https://doi.org/10.1038/nrn2113

Kuperberg, G. R., & Jaeger, T. F. (2016). What do we mean by prediction in language comprehension? Language, Cognition and Neuroscience, 31(1), 32-59. https://doi.org/10.1080/23273798.2015.1102299

Kutas, M., & Hillyard, S. A. (1980). Reading senseless sentences: Brain potentials reflect semantic incongruity. Science, 207(4427), 203-205. https://doi.org/10.1126/science.7350657

Levelt, W. J. M., Roelofs, A., & Meyer, A. S. (1999). A theory of lexical access in speech production. Behavioral and Brain Sciences, 22(1), 1-38. https://doi.org/10.1017/S0140525X99001776

MacDonald, M. C., Pearlmutter, N. J., & Seidenberg, M. S. (1994). The lexical nature of syntactic ambiguity resolution. Psychological Review, 101(4), 676-703. https://doi.org/10.1037/0033-295X.101.4.676

Marslen-Wilson, W. D. (1987). Functional parallelism in spoken word-recognition. Cognition, 25(1-2), 71-102. https://doi.org/10.1016/0010-0277(87)90005-9

Nieuwland, M. S., Politzer-Ahles, S., Heyselaar, E., Segaert, K., Darley, E., Kazanina, N., Von Grebmer Zu Wolfsthurn, S., Bartolozzi, F., Kogan, V., Ito, A., Mézière, D., Barr, D. J., Rousselet, G. A., Ferguson, H. J., Busch-Moreno, S., Fu, X., Tuomainen, J., Kulakova, E., Husband, E. M., … Huettig, F. (2018). Large-scale replication study reveals a limit on probabilistic prediction in language comprehension. eLife, 7, e33468. https://doi.org/10.7554/eLife.33468

Pickering, M. J., & Garrod, S. (2013). An integrated theory of language production and comprehension. Behavioral and Brain Sciences, 36(4), 329-347. https://doi.org/10.1017/S0140525X12001495

Rayner, K. (1998). Eye movements in reading and information processing: 20 years of research. Psychological Bulletin, 124(3), 372-422. https://doi.org/10.1037/0033-2909.124.3.372

Saffran, J. R., Aslin, R. N., & Newport, E. L. (1996). Statistical learning by 8-month-old infants. Science, 274(5294), 1926-1928. https://doi.org/10.1126/science.274.5294.1926

Schrimpf, M., Blank, I. A., Tuckute, G., Kauf, C., Hosseini, E. A., Kanwisher, N., Tenenbaum, J. B., & Fedorenko, E. (2021). The neural architecture of language: Integrative modeling converges on predictive processing. Proceedings of the National Academy of Sciences, 118(45), e2105646118. https://doi.org/10.1073/pnas.2105646118

Swinney, D. A. (1979). Lexical access during sentence comprehension: (Re)consideration of context effects. Journal of Verbal Learning and Verbal Behavior, 18(6), 645-659. https://doi.org/10.1016/S0022-5371(79)90355-4

Tanenhaus, M. K., Spivey-Knowlton, M. J., Eberhard, K. M., & Sedivy, J. C. (1995). Integration of visual and linguistic information in spoken language comprehension. Science, 268(5217), 1632-1634. https://doi.org/10.1126/science.7777863