Abstract

A neuropsychological test is a type of psychological test that measures cognitive, sensory, and motor performance under standardized conditions to infer the integrity of the brain systems that support behaviour. This article treats the neuropsychological test as an instrument of inference from behaviour to brain: the two traditions of fixed and flexible batteries that organize the field, the cognitive domains its subtests sample, and the psychometric machinery of demographically corrected norms and diagnostic validity by which its scores are read. It follows the field from Luria's qualitative syndrome analysis and the Halstead-Reitan quantitative battery, through the executive-function measures that made discrete cognitive deficits testable, to the computerized and co-normed instruments now reshaping practice. Three interactive demonstrations model Stroop interference, the conversion of a raw score into a demographically corrected norm, and the sensitivity-specificity trade-off of a diagnostic cutoff.

Keywords: executive function, demographically corrected norms, diagnostic validity

A neuropsychological test is an instrument built to make an inference that cannot be made directly: from how a person performs a standardized task to the condition of the brain systems that performance depends on. It shares the machinery of every other psychological test, sampling behaviour under controlled conditions to estimate an unobservable attribute, but the attribute it targets is neither a trait nor an attainment. It is the functional integrity of cognition itself, and the reason the estimate matters is that a specific pattern of test failures can localize dysfunction, track its course, or distinguish one clinical picture from another. That diagnostic purpose is what gives the neuropsychological test its distinctive demands: its scores must be interpreted against who the person was before, not against a single population average, and its errors carry the weight of a clinical decision. The century of work behind the modern battery is therefore a century of two problems held together, a taxonomy of which cognitive functions there are to test, and a psychometrics rigorous enough to turn a score into a defensible statement about a brain.

Key Takeaways
  • A neuropsychological test measures cognitive, sensory, and motor performance to infer the integrity of the brain systems behind it, distinguishing it from a general psychological test by its diagnostic, brain-referenced purpose.
  • The field is organized by two traditions: the fixed battery, which administers the same complete set of tests to everyone, and the flexible battery, which selects tests to answer a specific clinical question.
  • Its subtests are grouped by cognitive domain, and executive function in particular is measured by classic conflict, set-shifting, and sequencing tasks whose interpretation rests on the separability of distinct control processes.
  • A raw score is meaningless until it is referred to a normative sample, and demographically corrected norms adjust for age, education, and other factors that shift expected performance independent of any impairment.
  • Because a test is used to classify people as impaired or unimpaired, its value depends on sensitivity and specificity, which trade off against each other as the diagnostic cutoff moves.

What a Neuropsychological Test Is

The defining feature of a neuropsychological test is the target of its inference. A general psychological test may estimate an aptitude, a personality trait, or an attitude; a neuropsychological test estimates the functional state of the brain, using performance on a cognitive, perceptual, or motor task as an index of the neural systems that task recruits. This changes how the instrument is built and read. Its content is chosen for sensitivity to brain dysfunction rather than to any curriculum or construct of individual differences, and its scores are interpreted not against a single grand mean but against an estimate of the individual's own premorbid functioning, the level at which that person would be expected to perform were the brain intact. A score that is unremarkable for the population can signal marked decline in a formerly high-functioning person, and a score that looks low can be entirely normal given a person's age and education. The reference standard for a survey of what tests clinicians actually rely on shows how stable this core toolkit has remained, and how much interpretive weight rests on the norms rather than the raw performance (Rabin et al., 2016).

Two broad strategies organize how the tests are assembled and given, and the contrast between them structures the whole field. Table 1 sets them out.

Table 1. The fixed and flexible battery traditions contrasted.
Dimension Fixed battery Flexible battery
Test selectionA standard complete set given to every examinee.Tests chosen to answer a specific referral question.
InterpretationQuantitative, actuarial, against battery-wide norms.Hypothesis-driven, qualitative and quantitative.
StrengthComprehensive coverage; co-normed comparability.Efficient; adapts to the patient and question.
WeaknessLong, costly; may test what is not in question.Norms drawn from different samples; examiner-dependent.
ExemplarHalstead-Reitan; Luria-Nebraska.The Boston process approach; most modern practice.

The contrast is real but not absolute, and modern practice has largely converged on a middle path. Most clinicians now work flexibly, selecting from a shared pool of well-normed tests to pursue the referral question, while drawing on the fixed tradition's insistence on quantitative norms and its concern that a battery cover the major domains rather than only the ones a first hypothesis suggests. The standard textbook of the field codifies exactly this hybrid, pairing a flexible, hypothesis-testing philosophy with a rigorous, norm-referenced quantitative core (Lezak et al., 2012).

Types of Neuropsychological Tests

MeSH places the neuropsychological test beneath psychological tests and indexes ten narrower descriptors under it, listed in Table 2. The classification is an indexing convenience for retrieval rather than a theory of test kinds: its categories mix instruments defined by a cognitive domain (language, memory and learning, mental navigation), by a single named test (Stroop, Trail Making, Wisconsin Card Sorting, Bender-Gestalt), by a complete battery (Luria-Nebraska), and by clinical purpose (mental status and dementia, psychiatric status rating). These bases of division are not mutually exclusive, and a given instrument can fall under more than one. Only descriptors with a dedicated article are linked.

Table 2. Direct subtypes of Neuropsychological Tests in the MeSH classification (tree F04.711.513).
Subtype In brief
Bender-Gestalt TestA visuomotor test requiring the copying of geometric figures, used as a screen for perceptual-motor and constructional impairment.
Language TestsMeasures of naming, comprehension, repetition, and fluency used to characterize aphasia and other language disorders.
Luria-Nebraska Neuropsychological BatteryA standardized fixed battery operationalizing Luria's qualitative examination into quantitatively scored items.
Memory and Learning TestsTests of encoding, retention, and retrieval across verbal and visual material, central to detecting amnestic syndromes.
Mental Navigation TestsTasks assessing spatial orientation and route-finding, probing the topographical memory systems of the medial temporal lobe.
Mental Status and Dementia TestsBrief global screens such as the MMSE and MoCA that sample several domains to flag cognitive impairment.
Psychiatric Status Rating ScalesStructured scales quantifying the severity of psychiatric symptoms, often paired with cognitive testing.
Stroop TestA colour-word interference task measuring the inhibition of a prepotent reading response, a core executive test.
Trail Making TestA timed connect-the-sequence task whose two parts index processing speed and set-shifting.
Wisconsin Card Sorting TestA rule-induction task measuring set-shifting and perseveration under changing, unstated sorting rules.

The ten indexed subtypes understate the working diversity of the field, which in practice draws on hundreds of instruments organized less by these labels than by the cognitive domain each is thought to tax. That domain structure, and the executive tests that dominate its most-studied region, is the subject of the next sections.

Origins: From Luria to the Quantified Battery

Modern neuropsychological testing grew from two contrasting responses to the same problem, how to turn the clinical examination of a brain-injured patient into a repeatable procedure. The first was qualitative. Alexander Luria, working from the study of localized wartime brain lesions, built an examination that probed functional systems one at a time, reading the specific pattern of preserved and failed subtasks as a syndrome that pointed to a particular brain region. His method was flexible and interpretive, and its power lay in the examiner's analysis rather than in a total score. It was later standardized into the fixed, quantitatively scored Luria-Nebraska battery, which preserved his content while making it comparable across patients and examiners.

The second response was quantitative from the outset. In the American tradition that became the Halstead-Reitan battery, a fixed set of tests was given to every patient and interpreted actuarially against norms, on the principle that consistent measurement and empirical cutoffs, rather than clinical impression, should carry the diagnostic load. It was within this tradition that Ralph Reitan validated the Trail Making Test as an indicator of organic brain damage, showing that a simple timed sequencing task discriminated patients with brain lesions from those without (Reitan, 1958). The tension between Luria's interpretive flexibility and Halstead-Reitan's actuarial rigour is the historical source of the two traditions in Table 1, and the eventual synthesis of the two, a flexible selection of rigorously normed tests, is the framework most contemporary practice inherits.

Cognitive Domains and the Executive Tests

Neuropsychological tests are organized by the cognitive domain each is designed to tax: attention, processing speed, language, visuospatial ability, learning and memory, and executive function, among others. A battery aims to sample the major domains so that a profile of relative strengths and weaknesses, rather than any single score, becomes the diagnostic object. The domain that has drawn the most theoretical and clinical attention is executive function, the set of control processes that coordinate and regulate other cognitive operations, and it is measured by a small family of classic tasks. A comprehensive review of executive-function instruments catalogues both their reach and a recurring problem: many are impure, so that a failure can reflect the target control process or a breakdown in a supporting function such as speed or language (Chan et al., 2008).

Figure 1. The cognitive domains a neuropsychological battery samples, with representative tests.
Domain map of a neuropsychological battery A central node labelled Neuropsychological Battery connects to six cognitive domains: attention, processing speed, language, visuospatial ability, learning and memory, and executive function. Each domain lists one or two representative tests, and executive function is highlighted as the most studied domain. Neuropsychological Battery Attention digit span, cancellation Processing Speed Trail Making A, coding Language naming, fluency Visuospatial figure copy, drawing Learning and Memory word-list recall Executive Function Stroop, WCST, Trails B

Three tasks anchor the executive battery. The Stroop Test requires naming the ink colour of a colour word while ignoring the word itself, so that reading the word, an automatic and prepotent response, conflicts with the colour-naming goal; the slowing on incongruent trials, first quantified by J. R. Stroop, indexes the inhibition of that prepotent response (Stroop, 1935). The Wisconsin Card Sorting Test asks the examinee to sort cards by a rule that is never stated and changes without warning, so that measured perseveration, continuing to sort by a rule that has stopped working, indexes the capacity to shift set (Grant & Berg, 1948). The Trail Making Test in its second part requires alternating between numbers and letters, taxing the same set-shifting capacity under time pressure (Reitan, 1958).

The first demonstration isolates the mechanism behind the Stroop effect. It shows naming time under congruent and incongruent conditions and lets the reader vary how automatic word reading is, so that the interference, the extra time to name a colour when the word fights it, grows precisely as reading becomes more automatic and harder to suppress.

Stroop interference: why an automatic habit costs you time

02505007501000600congruent774incongruent+174 msnaming time (ms)
Congruent RT600 ms
Incongruent RT774 ms
Interference174 ms

The rust segment is the Stroop cost: extra time spent overriding the automatic pull to read the word instead of naming its ink colour.

When reading is highly automatic the word is processed whether or not you intend it, so a colour that conflicts with the word must be forced through against that habit, adding 174 ms here. Stronger top-down control suppresses the reading response and shrinks the cost; at zero automaticity, a non-reader, there is no conflict and no interference at all.

Whether these tasks measure one executive capacity or several is not a rhetorical question but an empirical one, and its answer reshaped how the tests are interpreted. In a latent-variable analysis, Akira Miyake and colleagues showed that three commonly measured executive functions, shifting between tasks or mental sets, updating and monitoring working memory, and inhibiting prepotent responses, are clearly separable, yet remain moderately correlated, sharing some common underlying capacity (Miyake et al., 2000). The finding matters for testing directly: it means the Wisconsin Card Sorting Test and the Stroop are not interchangeable measures of a single executive faculty but partly distinct probes of separable control processes, and that a full executive assessment needs more than one.

Norms and the Interpretation of a Score

A raw neuropsychological score, a number of words recalled or seconds taken, carries no meaning until it is placed against a normative sample, a reference group whose distribution of scores defines what is expected. The raw score is converted to a standardized metric, most often a z-score, a T-score with a mean of 50 and a standard deviation of 10, or a percentile, that states how far the performance sits from the reference mean in units of the reference spread. The compendium that clinicians use to administer and score the major tests exists precisely to supply these norms, along with the administration procedures that make a score comparable to them (Strauss et al., 2006).

The choice of reference group is not incidental, because performance on most cognitive tests varies systematically with demographic factors that have nothing to do with brain injury. Older adults are slower, and years of education raise performance across many domains, so a single population norm will over-diagnose impairment in an older person with little schooling and miss it in a young, highly educated one. Demographically corrected norms address this by referring each score to a group matched on age, education, and sometimes sex and cultural background, so that the comparison isolates deviation from the person's own expected level. Robert Heaton's normative systems made these corrections standard for the major batteries. The Trail Making Test illustrates the size of the effect: its normative data are stratified by age and education because both shift completion times substantially in healthy people (Tombaugh, 2004).

The second demonstration makes this conversion manipulable. It takes a raw score and refers it to a norm whose expected mean shifts with the age and education the reader sets, then reports the resulting z-score, T-score, and percentile, and the impairment classification they imply, so that the same raw performance can move between normal and impaired as the reference group changes.

From raw score to norm: the same performance, two reference groups

-3-2-10123scorestandard deviations from the reference mean (z) →
Expected mean30.0
z-score-1.33
T-score36.7
Percentile9
ClassificationMildly impaired

Hold the raw score at 22 and raise education from 12 to 16 years: the norm expects more, so the same performance falls from the 9th percentile to the 3rd, and from mildly impaired to impaired.

The raw score never changes what the patient did; it changes only what was expected of them. A number normal against one reference group is a deficit against another, which is why the choice of norm, matched on age and education, is itself a clinical decision.

Diagnostic Validity and the Cutoff

Because a neuropsychological test is ultimately used to sort people into categories, impaired or not, this condition or that, its worth rests on diagnostic validity: how well its scores separate the groups it is meant to distinguish. Two quantities capture this at any chosen cutoff. Sensitivity is the proportion of truly impaired people the test correctly flags, and specificity is the proportion of truly unimpaired people it correctly clears. The two trade off against each other along a single dial, the cutoff: lowering the threshold to catch more genuine cases inevitably flags more healthy people as well, raising sensitivity while lowering specificity, and raising the threshold does the reverse. There is no cutoff that maximizes both at once whenever the score distributions of the two groups overlap, which they always do.

The brief global screens used to detect dementia show why the trade-off is consequential. The Mini-Mental State Examination samples orientation, registration, attention, recall, and language in a few minutes (Folstein et al., 1975), and the later Montreal Cognitive Assessment was designed to be more sensitive to the milder deficits of mild cognitive impairment that the older screen often missed (Nasreddine et al., 2005). Choosing a screen, and a cutoff on it, is choosing a point on the sensitivity-specificity curve, and the right point depends on the cost of a missed case against the cost of a false alarm.

The third demonstration makes that trade-off visible. It draws the score distributions of an unimpaired and an impaired group, lets the reader slide the diagnostic cutoff between them, and reports the sensitivity and specificity that result, along with the positive predictive value at a base rate the reader also sets, so that the abstract tension between catching cases and sparing the healthy becomes a single moving line.

The diagnostic cutoff: catching cases against sparing the healthy

203040506070cutoffflag ←→ clearUnimpairedImpairedtest score (T) →
Sensitivity76%
Specificity79%
Pos. predictive value47%

Slide the cutoff right to catch more real cases (higher sensitivity) at the price of clearing fewer healthy people (lower specificity). Drop the base rate and the same test's positive results become mostly false alarms.

Sensitivity and specificity are properties of the test and cutoff alone, but the positive predictive value, the chance that a flagged person is truly impaired, depends heavily on how common impairment is: at a base rate of 20% this cutoff yields a 47% predictive value, which is why a good test can still mislead in a low-prevalence setting.

Worked Example

Consider a patient who recalls 22 words on a verbal learning test. The question a clinician must answer is not whether 22 is high or low in the abstract, but whether it is low for this person, and the answer depends entirely on the norm the score is referred to. Suppose the test has a normative standard deviation of 6 words.

Take first an age-only norm. For the patient's age band the expected mean is 30 words. The standardized score is z = (22 − 30) / 6 = −8 / 6 = −1.33, which converts to a T-score of T = 50 + 10 × (−1.33) = 36.7 and, reading from the normal distribution, to about the 9th percentile: the patient performs at or above only 9 percent of the age-matched reference group. On a common convention that flags scores more than one standard deviation below the mean (T below 40) as borderline and more than 1.5 below (T below 35) as impaired, this score lands as mildly impaired.

Now apply a demographically corrected norm. The patient completed 16 years of education, and because schooling raises expected performance on this test, the education-matched expected mean rises to 33 words while the spread stays at 6. The same raw score is now z = (22 − 33) / 6 = −11 / 6 = −1.83, a T-score of T = 50 + 10 × (−1.83) = 31.7, and about the 3rd percentile. The identical performance of 22 words has moved from the 9th percentile to the 3rd, and from mildly impaired to impaired, purely because the reference group changed to one that expected more of a highly educated person. This is the whole point of demographic correction, and the reason the second demonstration lets the norm shift under a fixed raw score: a neuropsychological score is a statement about a person relative to who they were expected to be, and a real deficit in a high-functioning patient is exactly what a population norm is most likely to hide.

Discussion

The neuropsychological test occupies a peculiar position among psychological instruments. Its scientific value and its clinical hazards flow from the same source, that it licenses an inference from behaviour to the brain, and that inference is only ever probabilistic. A test score is sensitive to far more than the neural system it is meant to index, including effort, mood, fatigue, language, culture, and the impurity of the task itself, and the discipline of neuropsychology is in large part the discipline of not over-reading a number. The field's most searching self-criticism concerns exactly this gap between what the tests measure and what they are used to conclude: many established instruments have limited ecological validity, predicting performance in the clinic better than function in daily life, and were normed on samples that do not represent the patients now tested against them (Howieson, 2019).

Two developments define the field's trajectory. The first is the long, largely successful project of putting interpretation on a firm psychometric footing, through demographically corrected and co-normed batteries that let a clinician compare a patient's scores across domains on a common metric (Casaletto & Heaton, 2017). The second is the recognition that the classic paper-and-pencil tests, however well normed, are limited instruments: coarse in their timing, vulnerable to examiner variation, and insensitive to the subtle, everyday failures that matter most to patients. How to build the next generation of tests, more precise, more ecologically valid, and better grounded in cognitive and neural theory, is the field's central open problem, and the one its current research addresses most directly.

Current Directions

Three lines of work are reshaping neuropsychological testing. The first is computerization. Touchscreen and computer-administered batteries such as CANTAB, developed by Barbara Sahakian, Trevor Robbins, and colleagues, deliver stimuli with millisecond precision, score responses automatically, and use largely non-verbal tasks that reduce the confound of language and culture, making testing more standardized and more portable across sites and trials (Sahakian & Owen, 1992). The gain is precision and comparability; the open questions concern whether the computerized versions measure the same constructs as their paper ancestors and carry over their hard-won norms.

The second is the drive toward co-normed, psychometrically modern instruments. The trend, charted in a review of the field's past and future, is away from a patchwork of tests normed on incompatible samples and toward integrated batteries whose subtests share a single normative frame, so that a profile across domains can be read without the noise of mismatched reference groups (Casaletto & Heaton, 2017). The third is a more fundamental rethinking of what a neuropsychological test should be. A programmatic account of the tests of the future argues for grounding measurement in formal cognitive and psychometric models, exploiting item-response theory and adaptive testing, and building instruments whose scores map onto latent constructs rather than onto the idiosyncrasies of a particular task (Bilder & Reise, 2019). Across all three, a mature but ageing measurement technology is being asked to become more precise, more valid to everyday function, and better tied to a theory of the mind it was built to probe.

Key Researchers

Robert M. Bilder (1956-2025). Professor of psychiatry and biobehavioral sciences at the University of California, Los Angeles; argued for a next generation of neuropsychological tests grounded in formal cognitive and psychometric models, item-response theory, and adaptive testing rather than in the idiosyncrasies of legacy tasks. Google Scholar - Faculty Page

Robert K. Heaton. Professor of psychiatry at the University of California, San Diego; produced the demographically corrected normative systems for the major neuropsychological tests and charted the field's shift toward co-normed, computerized assessment. Google Scholar - Faculty Page

Muriel Deutsch Lezak (1927-2021). Neuropsychologist at Oregon Health & Science University; authored the field's standard reference textbook and championed the flexible, hypothesis-driven approach to test selection over rigid fixed batteries. Wikipedia

Alexander Romanovich Luria (1902-1977). Soviet neuropsychologist at Moscow State University; founder of modern neuropsychology, whose qualitative syndrome analysis of localized brain lesions became the basis of the Luria-Nebraska battery. Wikipedia

Akira Miyake. Professor of psychology at the University of Colorado Boulder; his latent-variable analysis showed that shifting, updating, and inhibition are separable but correlated executive functions, reframing how executive tests are interpreted. Google Scholar - Faculty Page

Laura A. Rabin. Professor of psychology at Brooklyn College, City University of New York; surveyed the test-usage practices of clinical neuropsychologists across North America, documenting which instruments the field actually relies on and how stable that core has remained. ORCID - Google Scholar - Faculty Page

Ralph M. Reitan (1922-2014). Neuropsychologist at the University of Arizona; developed the Halstead-Reitan battery and validated the Trail Making Test as an indicator of organic brain damage, establishing the fixed-battery tradition of quantitative assessment. Wikipedia

Barbara J. Sahakian. Professor of clinical neuropsychology at the University of Cambridge; co-developed CANTAB, the touchscreen computerized battery that pioneered language-independent, automatically scored testing of memory and executive function. ORCID - Wikipedia - Faculty Page

Glossary

Battery.
A set of neuropsychological tests administered together to sample multiple cognitive domains, interpreted as a profile of relative strengths and weaknesses rather than as a single score.
Ceiling effect.
A clustering of scores near the maximum possible, which prevents a test from discriminating among higher-functioning examinees and can mask mild impairment.
Cognitive domain.
A broad class of mental function, such as attention, language, memory, or executive function, that a group of tests is designed to tax and that organizes how a battery is assembled.
Demographically corrected norms.
Reference standards that adjust the expected score for age, education, and sometimes sex or cultural background, so that a comparison isolates deviation from a person's own expected level rather than from a single population mean.
Diagnostic validity.
The degree to which a test's scores correctly separate the clinical groups it is meant to distinguish, quantified principally by sensitivity and specificity at a chosen cutoff.
Ecological validity.
The extent to which test performance predicts real-world functioning in everyday life, as opposed to performance measured only under standardized clinic conditions.
Executive function.
The set of higher-order control processes, including inhibition, set-shifting, and working-memory updating, that coordinate and regulate other cognitive operations toward a goal.
Fixed battery.
An assessment approach in which the same complete, standardized set of tests is administered to every examinee and interpreted actuarially against battery-wide norms, exemplified by the Halstead-Reitan battery.
Flexible battery.
An assessment approach in which tests are selected to answer a specific referral question, guided by clinical hypotheses, exemplified by the Boston process approach and most modern practice.
Perseveration.
The continued application of a response or rule after it has ceased to be appropriate, measured on the Wisconsin Card Sorting Test as a failure to shift set.
Premorbid functioning.
The level of cognitive ability a person had before the onset of injury or illness, estimated so that current scores can be judged as decline rather than against a population average.
Processing speed.
The rate at which simple cognitive operations can be performed, a domain sampled by timed tasks such as the first part of the Trail Making Test and highly sensitive to age and diffuse brain change.
Sensitivity.
The proportion of truly impaired individuals whom a test correctly identifies as impaired at a given cutoff; it rises as specificity falls when the cutoff is lowered.
Specificity.
The proportion of truly unimpaired individuals whom a test correctly clears at a given cutoff; it rises as sensitivity falls when the cutoff is raised.
Stroop interference.
The slowing of colour naming when the printed word names a conflicting colour, reflecting the effort of inhibiting an automatic reading response; a core index of cognitive control.
T-score.
A standardized score with a mean of 50 and a standard deviation of 10, widely used in neuropsychology to express how far a raw score falls from its normative mean.

Frequently Asked Questions

What is a neuropsychological test?
It is a standardized test of cognitive, sensory, or motor performance used to infer the functional integrity of the brain systems that support that performance. Unlike a general psychological test, its purpose is diagnostic and brain-referenced, and its scores are interpreted against what a person would be expected to do rather than against a single population average (Lezak et al., 2012).

How does a neuropsychological test differ from a general psychological test?
The difference is the target of inference. A general psychological test estimates an aptitude, trait, or attitude, whereas a neuropsychological test estimates the state of the brain from performance on a cognitive task. That aim makes it sensitive to brain dysfunction by design and requires that its scores be read against age, education, and estimated premorbid level.

What is the difference between a fixed and a flexible battery?
A fixed battery gives the same complete set of tests to everyone and interprets them actuarially against common norms, as in the Halstead-Reitan battery. A flexible battery selects tests to answer a particular referral question. Most modern practice blends the two, choosing from a shared pool of well-normed tests while keeping the fixed tradition's quantitative rigour (Rabin et al., 2016).

Which tests measure executive function?
The classic executive tests are the Stroop Test, which measures inhibition of an automatic reading response (Stroop, 1935); the Wisconsin Card Sorting Test, which measures set-shifting and perseveration (Grant & Berg, 1948); and the Trail Making Test, whose second part taxes set-shifting under time pressure (Reitan, 1958). Because these processes are separable, more than one test is needed for a full picture (Miyake et al., 2000).

Why are demographically corrected norms important?
Performance on most cognitive tests varies with age and education independent of any injury, so a single population norm over-diagnoses impairment in older or less-schooled people and misses it in young, highly educated ones. Demographically corrected norms refer each score to a matched group, isolating true deviation from a person's expected level (Tombaugh, 2004).

What do sensitivity and specificity mean for a test?
Sensitivity is the proportion of truly impaired people a test correctly flags, and specificity is the proportion of truly unimpaired people it correctly clears. They trade off as the cutoff moves: catching more real cases also flags more healthy people. Choosing a cutoff means choosing a point on that trade-off (Nasreddine et al., 2005).

What is the Mini-Mental State Examination?
It is a brief global screen that samples orientation, registration, attention, recall, and language in a few minutes to flag cognitive impairment (Folstein et al., 1975). The later Montreal Cognitive Assessment was designed to be more sensitive to the milder deficits of mild cognitive impairment that the older screen often missed (Nasreddine et al., 2005).

How is neuropsychological testing changing?
It is moving toward computerized, touchscreen batteries that score responses with high precision and reduce language confounds (Sahakian & Owen, 1992), toward co-normed batteries that share a single normative frame (Casaletto & Heaton, 2017), and toward instruments grounded in formal cognitive and psychometric models rather than legacy tasks (Bilder & Reise, 2019).

References

Bilder, R. M., & Reise, S. P. (2019). Neuropsychological tests of the future: How do we get there from here? The Clinical Neuropsychologist, 33(2), 220-245. https://doi.org/10.1080/13854046.2018.1521993

Casaletto, K. B., & Heaton, R. K. (2017). Neuropsychological assessment: Past and future. Journal of the International Neuropsychological Society, 23(9-10), 778-790. https://doi.org/10.1017/S1355617717001060

Chan, R. C. K., Shum, D., Toulopoulou, T., & Chen, E. Y. H. (2008). Assessment of executive functions: Review of instruments and identification of critical issues. Archives of Clinical Neuropsychology, 23(2), 201-216. https://doi.org/10.1016/j.acn.2007.08.010

Folstein, M. F., Folstein, S. E., & McHugh, P. R. (1975). Mini-mental state: A practical method for grading the cognitive state of patients for the clinician. Journal of Psychiatric Research, 12(3), 189-198. https://doi.org/10.1016/0022-3956(75)90026-6

Grant, D. A., & Berg, E. A. (1948). A behavioral analysis of degree of reinforcement and ease of shifting to new responses in a Weigl-type card-sorting problem. Journal of Experimental Psychology, 38(4), 404-411. https://doi.org/10.1037/h0059831

Howieson, D. (2019). Current limitations of neuropsychological tests and assessment procedures. The Clinical Neuropsychologist, 33(2), 200-208. https://doi.org/10.1080/13854046.2018.1552762

Lezak, M. D., Howieson, D. B., Bigler, E. D., & Tranel, D. (2012). Neuropsychological assessment (5th ed.). Oxford University Press.

Miyake, A., Friedman, N. P., Emerson, M. J., Witzki, A. H., Howerter, A., & Wager, T. D. (2000). The unity and diversity of executive functions and their contributions to complex frontal lobe tasks: A latent variable analysis. Cognitive Psychology, 41(1), 49-100. https://doi.org/10.1006/cogp.1999.0734

Nasreddine, Z. S., Phillips, N. A., Bedirian, V., Charbonneau, S., Whitehead, V., Collin, I., Cummings, J. L., & Chertkow, H. (2005). The Montreal Cognitive Assessment, MoCA: A brief screening tool for mild cognitive impairment. Journal of the American Geriatrics Society, 53(4), 695-699. https://doi.org/10.1111/j.1532-5415.2005.53221.x

Rabin, L. A., Paolillo, E., & Barr, W. B. (2016). Stability in test-usage practices of clinical neuropsychologists in the United States and Canada over a 10-year period: A follow-up survey of INS and NAN members. Archives of Clinical Neuropsychology, 31(3), 206-230. https://doi.org/10.1093/arclin/acw007

Reitan, R. M. (1958). Validity of the Trail Making Test as an indicator of organic brain damage. Perceptual and Motor Skills, 8(3), 271-276. https://doi.org/10.2466/pms.1958.8.3.271

Sahakian, B. J., & Owen, A. M. (1992). Computerized assessment in neuropsychiatry using CANTAB: Discussion paper. Journal of the Royal Society of Medicine, 85(7), 399-402. https://pubmed.ncbi.nlm.nih.gov/1629849/

Stroop, J. R. (1935). Studies of interference in serial verbal reactions. Journal of Experimental Psychology, 18(6), 643-662. https://doi.org/10.1037/h0054651

Strauss, E., Sherman, E. M. S., & Spreen, O. (2006). A compendium of neuropsychological tests: Administration, norms, and commentary (3rd ed.). Oxford University Press.

Tombaugh, T. N. (2004). Trail Making Test A and B: Normative data stratified by age and education. Archives of Clinical Neuropsychology, 19(2), 203-214. https://doi.org/10.1016/S0887-6177(03)00039-8