Abstract

The Bender-Gestalt Test, which MeSH classifies under personality tests, is a brief visual-motor task in which a person copies nine geometric figures adapted from Gestalt psychology. Lauretta Bender introduced it in 1938 as a clinical measure of perceptual-motor maturation and organic impairment, and it became one of the most widely administered instruments in clinical and school psychology. Scoring systems developed by Koppitz for children and by Pascal, Suttell, Hutt, Briskin, and Lacks for adults convert reproduction errors such as rotations, distortions, integration failures, and perseveration into quantified indices of visual-motor integration. The second edition adds a recall trial and supplementary motor and perception tests under a single global scoring system. This article describes the figures, the major scoring methods, the constructs the test measures, and its contemporary diagnostic use, alongside the psychometric limits that constrain its interpretation.

Keywords: visual-motor integration, gestalt, neuropsychological assessment, developmental scoring

Key takeaways
  • The test asks a person to copy nine figures; how faithfully the copies preserve shape, angle, and spatial relation indexes visual-motor integration.
  • Separate scoring systems serve different purposes: Koppitz developmental scores for children, and the Pascal-Suttell, Hutt-Briskin, and Lacks error counts for screening adult brain dysfunction.
  • The 2003 second edition standardized administration, added a recall trial, and replaced competing error systems with one global scoring scale.
  • As a screen the test trades sensitivity against specificity; a copy score localizes nothing and must be read beside history and converging measures.

What the Bender-Gestalt Test Is

The Bender-Gestalt Test is a copying task. An examiner presents nine cards, one at a time, each bearing an abstract line drawing, and asks the person to reproduce each drawing on a blank sheet. Lauretta Bender assembled the instrument in 1938 from figures the Gestalt psychologist Max Wertheimer had used to demonstrate the laws of perceptual grouping, arguing that the way a person reorganizes these forms on paper reveals the maturation of the visual-motor function and its disruption by injury or disease (Bender, 1938). The figures themselves descend directly from Wertheimer's 1923 studies of how the visual field segregates into wholes (Wertheimer, 1923).

The test occupies an unusual position in assessment. MeSH indexes it under both neuropsychological tests and personality tests, a dual placement that reflects its history rather than a settled view of what it measures: early clinicians read the copies psychodynamically, while later work treated the same reproductions as a narrow probe of perceptual-motor skill. Because it is quick to give, needs almost no language, and travels easily across cultures, it remained for decades among the instruments clinicians reported using most often (Piotrowski, 1995). Surveys of practice through the end of the twentieth century placed it consistently among the most frequently administered psychological tests in the United States (Camara, Nathan, & Puente, 2000).

The Nine Gestalt Figures

The stimulus set is a design figure plus eight numbered figures. Each isolates a different demand on form perception and motor output: a dot-and-circle configuration, intersecting sinusoids, columns of dots that must be counted and aligned, open and closed curves that meet at a point, and a hexagon overlapped by a second hexagon. What makes them diagnostic is that each embodies a Gestalt grouping principle, so a faithful copy requires the person to perceive the intended whole and then hold that organization steady while the hand executes it.

Figure 1

Three representative figures and two of the errors scoring systems count

Three representative Bender-Gestalt figures and common reproduction errors A row of three panels. The first shows a row of dots reproduced with correct grouping. The second shows two overlapping hexagons copied with a rotation error. The third shows a wavy line copied with perseveration, extra cycles added. Correct grouping Rotation error Perseveration
Note. The left panel shows a correctly grouped row of dots; the centre panel a pair of overlapping hexagons copied with a rotation of the whole configuration; the right panel a wavy line copied with perseveration, the continuation of a pattern beyond its model. The figure is an original schematic drawn for this article.

The interactive figure below sets an intact reproduction beside the characteristic distortions clinicians code. Each control introduces one error class so that its effect on the copy can be seen in isolation.

The figures and the errors that get scored
Intact copy

The two hexagons overlap at the correct angle and the row holds five evenly spaced dots. This is the reproduction the scoring key treats as correct.

Copying even these simple forms draws on more than the eye and the hand. The person must segregate each figure from its background, an act of figure-ground organization; must preserve the Gestalt, the perceived whole that the parts compose; and must translate that percept into a graphomotor plan, a form of constructional praxis. Failure at any stage leaves a signature in the drawing, and the scoring systems are catalogs of those signatures.

Scoring Systems

The raw drawings mean little until a scoring system turns them into numbers, and the test's history is largely a history of competing systems built for different populations and purposes.

For children, the dominant method is the Koppitz Developmental Scoring System. Elizabeth Münsterberg Koppitz scored each reproduction for the presence of specific errors and showed that the total error count falls steeply with age, so that a child's score can be read against age norms as an index of perceptual-motor maturation (Koppitz, 1958). Her 1964 manual formalized this developmental scoring into thirty scorable items covering distortion of shape, rotation, integration failure, and perseveration (Koppitz, 1964). The system remains the reference method for reading children's Bender reproductions against developmental expectations.

For adults, the goal shifted from maturation to the detection of organic brain dysfunction, and three error systems competed. Pascal and Suttell built the first quantified adult method, scoring deviations from an ideal reproduction and summing them into a standardized index (Pascal & Suttell, 1951). The Hutt-Briskin system reduced scoring to a small set of discriminating error categories, and Patricia Lacks refined that set into the widely used Lacks method, sharpening the Hutt-Briskin errors into an efficient screen for brain impairment (Lacks, 1999). All three treat error classes such as rotation, integration error, and perseveration as evidence of compromised visual-motor integration.

Table 1. The major Bender-Gestalt scoring systems and their purposes.
Scoring systemPopulationWhat it scoresPrimary purpose
Koppitz Developmental Scoring SystemChildrenThirty developmental error itemsPerceptual-motor maturation against age norms
Pascal-SuttellAdultsDeviations from an ideal reproductionStandardized adult quality index
Hutt-Briskin systemAdultsSmall set of discriminating errorsScreen for organic brain dysfunction
LacksAdultsRefined Hutt-Briskin error setEfficient brain-dysfunction screen
Global scoring (Bender-Gestalt II)Ages 4 and upSingle global copy and recall scaleStandardized assessment across ages
Koppitz developmental scoring: reading a child against age norms
errorsage
Age-7 mean is about 8 errors. This child scored 8: Within age expectation.

Koppitz scoring counts specific errors and reads the total against age norms; because error counts fall steeply with age, the same raw score means very different things at five and at eleven. The curve here is an illustrative teaching approximation, not the published norm table.

The proliferation of incompatible systems was itself a problem: a score meant nothing without naming the method that produced it. The 2003 second edition, the Bender-Gestalt II, addressed this directly, standardizing administration, adding a recall trial and supplementary motor and perception tests, and replacing the competing error counts with a single global scoring scale (Brannigan & Decker, 2003).

What It Measures

At its narrowest, the test measures visual-motor integration, the coordination of visual perception with the fine motor output that reproduces what is seen. This construct develops through childhood and can be degraded by conditions that disturb perception, motor control, or the mapping between them, which is why the same drawing task indexes maturation in a child and impairment in an adult. Decker's psychometric work on the second edition traced how visual-motor processes grow through childhood and decline in later life, giving the copy score a developmental interpretation across the lifespan (Decker, 2008). Follow-up work separated the visual, motor, and integrative contributions to children's copies, showing that the single score aggregates partly dissociable skills (Decker, Englund, Carboni, & Brooks, 2011).

What the test does not do is localize. A low copy score signals that some part of the perceive-plan-execute chain is disturbed, but not where or why, and elevated error counts appear across intellectual disability, dementia, and focal injury alike. The screening question it can answer is probabilistic: given a base rate of impairment and the test's sensitivity and specificity, how much does a failing score raise the probability that dysfunction is present. The demonstration below makes that trade-off explicit.

The screen as a probability: predictive value at a given base rate
Impaired
Not impaired
Screen positive
49
32
Screen negative
11
108
PPV
60%
NPV
91%
Accuracy
79%

Across 200 referred adults, the same sensitivity and specificity yield very different predictive value as the base rate moves. At the default settings a positive screen is correct only about three times in five, which is why a failing Bender-Gestalt score is a reason to assess further rather than a diagnosis.

Read this way, the test is a screen, not a diagnosis. Its value lies in flagging cases that warrant fuller neuropsychological assessment, and its interpretation depends entirely on the sensitivity and specificity of the scoring system in the population at hand.

Worked Example

Consider a memory clinic that screens referred adults with the Lacks scoring system before deciding who proceeds to a full neuropsychological workup. Suppose the screen has a sensitivity of 0.82 and a specificity of 0.77 in this setting, and that 30 percent of the 200 adults referred in a year genuinely have brain dysfunction.

Of the 200, 60 have dysfunction and 140 do not. Applying the sensitivity, the screen flags 60 times 0.82, or about 49 of the 60 true cases, missing roughly 11. Applying the specificity, it correctly clears 140 times 0.77, or about 108 of the 140 unaffected adults, and falsely flags the remaining 32. The screen therefore produces about 49 true positives, 32 false positives, 108 true negatives, and 11 false negatives.

The positive predictive value is 49 divided by (49 plus 32), which is 49 of 81, or about 60 percent: a failing score leaves a meaningful chance the person is unimpaired. The negative predictive value is 108 divided by (108 plus 11), which is 108 of 119, or about 91 percent: a passing score is more trustworthy. Overall accuracy is (49 plus 108) of 200, or 78 percent. The lesson is the one every screen teaches. Even with respectable sensitivity and specificity, a positive result at this base rate is right only about three times in five, so the Bender-Gestalt outcome is a reason to look further, never a verdict on its own.

Discussion

The Bender-Gestalt Test has been criticized for the very features that made it popular. Its brevity and low language demand come at the cost of construct breadth: a single copy score cannot distinguish a perceptual deficit from a motor one, and the test's early psychodynamic interpretations, in which drawing style was read for personality, were never psychometrically defensible. As neuropsychology developed instruments designed to localize and to fractionate cognition, the Bender's role narrowed from a broad clinical probe to a specific screen of visual-motor integration.

Its persistence nonetheless has a rational basis. The task is nearly free of the language and cultural loading that burden many cognitive tests, which is why researchers continue to develop local norms for new populations (Tahmasebi, Alizadeh, Rezaei, & Salehi, 2016). Used as its evidence base supports, as a quick, low-cost flag for perceptual-motor disturbance that then routes a person to fuller assessment, it remains defensible. Used as its critics rightly warn against, as a stand-alone diagnosis or a localizing sign, it overreaches. The distinction is the whole of responsible practice with the instrument.

Current Directions

Contemporary work has moved in two directions. The first is cross-cultural standardization of the second edition. A Turkish standardization study established the psychometric properties of the Bender-Gestalt II under its global scoring system, extending the norms beyond the original American sample and testing whether the single global scale holds across languages (Korkmaz, Karakoç Demirkaya, & Özdemir, 2023). Norming projects in other populations pursue the same goal of making the copy score interpretable outside its country of origin (Tahmasebi et al., 2016).

The second direction is targeted clinical differentiation. Rather than asking whether the test detects impairment in general, recent studies ask whether it can help separate specific conditions. One line of work reports that Bender-Gestalt performance aids the clinical diagnosis of dementia with Lewy bodies in patients who already show mild cognitive impairment, exploiting the visuospatial and constructional demands that this dementia disrupts early (Murayama, Ota, & Iseki, 2024). A 2025 systematic review consolidated the modern evidence base, mapping where the test retains diagnostic value and where its reliability remains contested (Lafhal, Ait Ali, El Alaoui, & Ahami, 2025). Together these strands describe a mature instrument being repositioned from a general screen toward defined, evidence-tested roles.

Key Researchers

Lauretta Bender (1897-1987). The child neuropsychiatrist who created the Visual-Motor Gestalt Test in 1938, adapting Wertheimer's grouping figures into a clinical copying task; see Wikipedia.

Scott L. Decker (University of South Carolina). Co-author of the Bender-Gestalt II and author of psychometric work on visual-motor development and decline across the lifespan; ORCID 0000-0002-8085-7965.

Hicham Lafhal (Ibn Tofail University). Lead author of the 2025 systematic review consolidating the contemporary evidence base on the test; ORCID 0009-0006-4245-7043.

Norio Murayama (Showa Women's University). Contemporary researcher demonstrating the test's utility in the clinical diagnosis of dementia with Lewy bodies; ORCID 0009-0001-1586-4835.

Max Wertheimer (1880-1943). Founder of Gestalt psychology, whose 1923 figures illustrating the laws of perceptual grouping became the stimulus designs of the test; see Wikipedia.

The developmental and adult scoring systems that made the test usable were built by researchers who predate persistent scholarly identifiers. Elizabeth Münsterberg Koppitz (1918-1983) created the children's developmental scoring system; Gertrude Pascal and Barbara Suttell built the first quantified adult method; and Patricia Lacks (1941-2016) refined the Hutt-Briskin errors into the Lacks screen. Gary G. Brannigan (1947-2025) co-authored the second edition. Their contributions are cited throughout this article by way of their published work.

Glossary

Bender-Gestalt II
The 2003 second edition of the test, which standardized administration, added a recall trial and supplementary motor and perception tests, and replaced competing error systems with one global scoring scale.

Constructional praxis
The capacity to translate a perceived form into an organized motor plan that reproduces it; the ability the copying task most directly taxes.

Developmental scoring
Any method that reads a child's error count against age norms, treating the score as an index of perceptual-motor maturation rather than of impairment.

Distortion of shape
A scorable error in which the reproduced figure loses its essential form, such as a circle rendered as an ellipse or angles substituted for curves.

Figure-ground organization
The perceptual segregation of a figure from its background, the first step in apprehending a Gestalt figure before it can be copied.

Gestalt
A perceived whole whose organization is more than the sum of its parts; the grouping principles the figures embody derive from Gestalt psychology.

Hutt-Briskin system
An adult scoring method that reduced the reproduction to a small set of discriminating error categories, later refined into the Lacks screen for brain dysfunction.

Integration error
A failure to join the parts of a figure into a coherent whole, such as a gap between elements that should meet or overlapping forms drawn as separate.

Koppitz Developmental Scoring System
Elizabeth Münsterberg Koppitz's thirty-item method for scoring children's reproductions against age norms, the standard developmental scoring approach.

Perseveration
The continuation of a pattern beyond its model, such as adding extra cycles to a wavy line or extra dots to a row.

Rotation error
A reproduction in which the whole figure or a major element is turned appreciably from its original orientation while its internal form is preserved.

Sensitivity
The proportion of genuinely impaired people a screen correctly flags; one of the two rates that fix a screen's predictive value at a given base rate.

Specificity
The proportion of unimpaired people a screen correctly clears; traded against sensitivity in setting a screening cutoff.

Visual-motor integration
The coordination of visual perception with fine motor output, the core construct the copy score is taken to index.

Frequently Asked Questions

What does the Bender-Gestalt Test actually measure?
At its narrowest it measures visual-motor integration, the coordination of what a person sees with the hand movements that reproduce it. Older uses read the drawings for personality, but that interpretation was never psychometrically sound.

How is the test administered?
An examiner presents nine cards one at a time, each showing an abstract figure, and asks the person to copy each onto a blank sheet. The second edition adds a recall trial in which the person redraws the figures from memory.

Who developed the test?
Lauretta Bender assembled it in 1938 from figures that Max Wertheimer had used to illustrate Gestalt grouping principles. Bender treated the way a person reorganizes those figures on paper as a window on perceptual-motor maturation.

Why are there several different scoring systems?
Different populations and purposes drove different methods: Koppitz developmental scoring for children, and the Pascal-Suttell, Hutt-Briskin, and Lacks systems for screening adult brain dysfunction. A score is meaningless without naming the system that produced it.

Can the test diagnose brain damage?
No. It is a screen, not a diagnosis. A failing score raises the probability that some perceptual-motor disturbance is present but localizes nothing, so it should route a person to fuller assessment rather than settle a question.

What is the Koppitz Developmental Scoring System?
It is Elizabeth Münsterberg Koppitz's thirty-item method for scoring children's copies against age norms. Because error counts fall with age, a child's total can be read as an index of perceptual-motor development.

What changed in the Bender-Gestalt II?
The 2003 edition standardized administration, added a recall trial and supplementary motor and perception tests, and replaced the competing error systems with a single global scoring scale, so that scores became comparable across examiners.

Is the test still used today?
Yes, though in a narrower role than in its mid-century heyday. Recent work develops local norms for new populations and tests whether it can help differentiate specific conditions such as dementia with Lewy bodies.

References

Bender, L. (1938). A visual motor gestalt test and its clinical use (Research Monographs No. 3). American Orthopsychiatric Association.

Brannigan, G. G., & Decker, S. L. (2003). Bender Visual-Motor Gestalt Test, Second Edition (Bender-Gestalt II). Riverside Publishing.

Camara, W. J., Nathan, J. S., & Puente, A. E. (2000). Psychological test usage: Implications in professional psychology. Professional Psychology: Research and Practice, 31(2), 141-154. https://doi.org/10.1037/0735-7028.31.2.141

Decker, S. L. (2008). Measuring growth and decline in visual-motor processes with the Bender-Gestalt Second Edition. Journal of Psychoeducational Assessment, 26(1), 3-15. https://doi.org/10.1177/0734282907300685

Decker, S. L., Englund, J. A., Carboni, J. A., & Brooks, J. H. (2011). Cognitive and developmental influences in visual-motor integration skills in young children. Psychological Assessment, 23(4), 1010-1016. https://doi.org/10.1037/a0024079

Koppitz, E. M. (1958). The Bender Gestalt Test and learning disturbances in young children. Journal of Clinical Psychology, 14(3), 292-295. https://doi.org/10.1002/1097-4679(195807)14:3%3C292::AID-JCLP2270140321%3E3.0.CO;2-O

Koppitz, E. M. (1964). The Bender Gestalt Test for young children. Grune & Stratton.

Lacks, P. (1999). Bender Gestalt screening for brain dysfunction (2nd ed.). Wiley.

Lafhal, H., Ait Ali, D., El Alaoui, F. E., & Ahami, A. O. T. (2025). The Bender-Gestalt Test: A systematic review. Cureus, 17(3), e81122. https://doi.org/10.7759/cureus.81122

Korkmaz, S., Karakoç Demirkaya, S., & Özdemir, P. G. (2023). Bender-Gestalt II Test: Psychometric properties with the global scoring system on a Turkish standardization sample. Child Neuropsychology, 29(4), 607-627. https://doi.org/10.1080/09297049.2022.2104237

Murayama, N., Ota, K., & Iseki, E. (2024). The Bender Gestalt Test is useful for clinically diagnosing dementia with Lewy bodies in patients with mild cognitive impairment. Applied Neuropsychology: Adult, 31(6), 1296-1301. https://doi.org/10.1080/23279095.2022.2122059

Pascal, G. R., & Suttell, B. J. (1951). The Bender-Gestalt test: Quantification and validity for adults. Grune & Stratton.

Piotrowski, C. (1995). A review of the clinical and research use of the Bender-Gestalt Test. Perceptual and Motor Skills, 81(3, Pt. 2), 1272-1274. https://doi.org/10.2466/pms.1995.81.3f.1272

Tahmasebi, S., Alizadeh, H., Rezaei, S., & Salehi, M. (2016). Normalizing the Bender Visual-Motor Gestalt Test for 4- to 7-year-old children of Tehran, Iran. Journal of Rehabilitation, 17(1), 18-29. https://doi.org/10.20286/jrehab-170118

Wertheimer, M. (1923). Untersuchungen zur Lehre von der Gestalt. II. Psychologische Forschung, 4(1), 301-350. https://doi.org/10.1007/BF00410640