Abstract
The Stanford-Binet is a type of intelligence test: the American revision of the Binet-Simon scale, an individually administered battery that carried the intelligence quotient into psychological testing and has been renormed across five editions since 1916. This article treats it as a measurement instrument: how Terman's ratio IQ gave way to the deviation IQ at the 1960 Form L-M, how the fifth edition rebuilt the test as a point scale with five Cattell-Horn-Carroll factors measured across verbal and nonverbal routings, and how basal and ceiling rules let one instrument span ages two to eighty-five and over. Three interactive demonstrations model the ratio-to-deviation shift, the five-factor profile behind a full-scale score, and the adaptive basal-ceiling administration.
Keywords: intelligence quotient, deviation IQ, basal and ceiling
The Stanford-Binet Intelligence Scales are the oldest continuously published individual intelligence test, the instrument on which Lewis Terman's 1916 Stanford revision of the Binet-Simon scale first put the intelligence quotient into American practice. It is a member of the broad family of intelligence tests, sharing their machinery of standardized administration and norm-referenced scoring, and MeSH indexes it directly beneath that category. What distinguishes the Stanford-Binet is a century of revision that turned a single age-scale yielding a mental-age ratio into a modern point scale reporting a deviation IQ over five theory-driven factors. Following the test across its five editions is therefore a compact history of how the field learned to define, scale, and structure the score an intelligence test reports.
- The Stanford-Binet is the American revision of the Binet-Simon scale, an individually administered intelligence test first published by Terman in 1916 and now in its fifth edition (2003).
- It introduced the intelligence quotient into American testing as a mental-age ratio, then abandoned that ratio for the deviation IQ at the 1960 Form L-M, tracking the field's redefinition of the score.
- The fifth edition (SB5) is a point scale organized around five Cattell-Horn-Carroll factors, each measured through both a verbal and a nonverbal routing.
- Adaptive administration uses routing subtests plus basal and ceiling rules to give each examinee only the items near their ability, so one instrument spans ages two to over eighty-five.
- Like every norm-referenced test the Stanford-Binet must be periodically renormed against the Flynn effect, and its scores carry the same reliability, validity, and fairness obligations as any intelligence measure.
What the Stanford-Binet Is
The Stanford-Binet is an individually administered, norm-referenced intelligence test: a trained examiner works one-to-one with a single examinee, presenting a standardized series of cognitive tasks under fixed conditions and converting the raw performance into an index scaled against a representative national sample. It is the American descendant of the first practical intelligence scale, and MeSH files it as a narrower descriptor directly under intelligence tests, alongside the Wechsler scales, the two individually administered batteries that have defined the field for a century. Its headline output is a full-scale IQ, an estimate of general cognitive ability, reported on the modern deviation scale with a mean of 100 and a standard deviation of 15.
Two features mark the Stanford-Binet within that family. The first is longevity: no other individual intelligence test has been in continuous publication since 1916, so its five editions form an unusually clean record of how test theory changed. The second is its coverage of the full age range in one instrument. Where the Wechsler tradition splits into separate adult and child batteries, a single Stanford-Binet edition is normed from early childhood through late adulthood, a reach made possible by the adaptive administration described below. As with any such instrument, the quality of a Stanford-Binet score is judged by its reliability, the consistency of the measurement, and its validity, the support for the interpretations placed on it, obligations codified in the profession's testing standards (AERA et al., 2014).
From the Binet-Simon Scale to Five Editions
The lineage begins in Paris. In 1905 Alfred Binet and Théodore Simon, commissioned to identify schoolchildren needing special instruction, built a graded series of tasks ordered by the age at which a typical child could pass them (Binet & Simon, 1905). Their scale yielded a mental age, the age level of the hardest tasks a child could reliably complete, and it established the defining logic the Stanford-Binet inherited: a score is a comparison with a developmental norm rather than a count of correct answers.
The scale became American, and quotient-based, at Stanford. Lewis Terman's 1916 revision, the Stanford-Binet, adopted the intelligence quotient, mental age divided by chronological age and multiplied by 100, so that a child whose mental age matched their chronological age scored exactly 100 (Terman, 1916). Terman and Maud Merrill revised the test in 1937 into two parallel forms, L and M, and again in 1960, when the two forms were merged into a single Form L-M and, decisively, the ratio IQ was replaced by the deviation IQ that David Wechsler had introduced two years earlier (Wechsler, 1958). A fourth edition in 1986 moved toward a hierarchical factor model, and the current fifth edition (SB5), published by Gale Roid in 2003, rebuilt the instrument as a point scale over five factors (Roid, 2003). The published account of that revision history, edition by edition, is the test's own technical documentation (Becker, 2003). Figure 1 sets out the five editions and the defining change at each.
Figure 1
Five Editions of the Stanford-Binet, 1916-2003
The Ratio IQ and Why It Was Replaced
The original Stanford-Binet score was a ratio IQ, and its defect is the clearest lesson in the test's history. Dividing mental age by chronological age works while both climb together in childhood, but mental age stops rising in the late teens as cognitive development plateaus, whereas chronological age never stops. For a person of constant ability the ratio therefore falls with age, and the same numerical IQ means different things at different ages, which makes the ratio unusable for adults. The 1960 Form L-M solved this by adopting the deviation IQ, which expresses a person's standing in standard-deviation units relative only to their own age group, scaled to a mean of 100 and a standard deviation of 15 (Wechsler, 1958). The deviation score severed intelligence measurement from the mental-age ratio, and every Stanford-Binet edition since has used it.
The first demonstration makes the contrast manipulable. It follows a single developmental rate across the lifespan and plots the two scores side by side: a ratio IQ that tracks the deviation IQ in childhood but drifts downward once mental age reaches its plateau, against a deviation IQ that holds constant because it is always referred to the examinee's own age peers.
Why the ratio IQ was abandoned: it drifts, the deviation IQ does not
At chronological age 10, this child’s mental age is 12.0, so the ratio IQ is 120. The deviation IQ is 120, unchanged at any age. While mental age still climbs the two agree, but once it plateaus near sixteen the ratio falls for a person whose standing never dropped. The Stanford-Binet carried the ratio IQ from 1916 until Form L-M adopted the deviation IQ in 1960 for exactly this reason.
The Five-Factor Structure of the SB5
The fifth edition is organized by theory rather than by age level. Its architecture is the Cattell-Horn-Carroll (CHC) taxonomy, the consolidated model that arranges narrow abilities under broad group factors under a general factor at the apex (McGrew, 2009). CHC in turn rests on Raymond Cattell's 1963 separation of fluid intelligence, the capacity to reason with novel material, from crystallized intelligence, the store of accumulated knowledge (Cattell, 1963), joined to John Carroll's exhaustive re-analysis of more than four hundred factor-analytic datasets, whose three-stratum theory ordered abilities into narrow, broad, and general strata and supplied the empirical backbone McGrew merged with the Cattell-Horn model to form CHC (Carroll, 1993). Behind the whole hierarchy stands Charles Spearman's 1904 observation that all cognitive tasks correlate positively, from which a general factor g is inferred (Spearman, 1904). The SB5 measures five factors chosen from the CHC broad abilities: Fluid Reasoning, Knowledge, Quantitative Reasoning, Visual-Spatial Processing, and Working Memory. Each factor is assessed twice, through a verbal routing and a nonverbal routing, so the test can yield a Verbal IQ, a Nonverbal IQ, and the five factor indexes alongside the full-scale score (Roid, 2003). Table 1 sets out the five factors against the CHC broad abilities they draw on and what each is meant to tap.
| Factor | CHC broad ability | What it measures |
|---|---|---|
| Fluid Reasoning | Fluid intelligence (Gf) | Reasoning with novel material — completing patterns, inferring rules, and solving problems that acquired knowledge does not answer directly. |
| Knowledge | Crystallized intelligence (Gc) | The store of accumulated information and vocabulary a person has acquired through experience and schooling. |
| Quantitative Reasoning | Quantitative knowledge (Gq) | Reasoning with numbers and quantitative relationships, from counting and number concepts to word problems. |
| Visual-Spatial Processing | Visual processing (Gv) | Analyzing spatial relationships and patterns — assembling forms, reproducing designs, and tracking positions and orientations. |
| Working Memory | Short-term memory (Gsm) | Holding and transforming information in mind over the short term, as in repeating and reordering sequences. |
Whether the data support five distinct factors has been examined directly. A confirmatory factor analysis of the SB5 standardization sample tested the published five-factor model against simpler alternatives and found the general factor dominant, with the evidence for five cleanly separable factors weaker than the scoring structure implies (DiStefano & Dombrowski, 2006) — a reminder that a factor score printed on a report is a scoring convention, not a proof that the factor is a separable thing. The second demonstration builds the composite from its parts: the reader sets the five factor indexes and watches the full-scale score and the profile spread respond.
The fifth edition’s five-factor profile: the same Full-Scale IQ, different shapes
The five factor indexes average to a Full-Scale IQ of 106, but they scatter across 25 points from lowest to highest. A flat profile and a jagged one can share the same composite, which is why the fifth edition reports the factor indexes as well as the Full-Scale IQ: the shape carries clinical information the single number hides. (The composite here is a simple equal weighting for illustration; the published SB5 norms the composite from subtest scores.)
Adaptive Administration: Routing, Basal, and Ceiling
A test that spans ages two to over eighty-five cannot give every examinee every item; it would be impossibly long and mostly uninformative, since easy items tell nothing about a strong examinee and hard items nothing about a weak one. The Stanford-Binet handles this with adaptive administration. A short routing subtest first locates the examinee's approximate ability, so testing can start at an appropriate level rather than at the bottom (Roid & Barram, 2004). From that start the examiner works outward under two stopping rules. The basal is the level below which items are easy enough to be assumed passed and credited without administration; the ceiling is the level above which items are hard enough to be assumed failed. Only the block of items between basal and ceiling is actually given, and the raw score combines the administered items with the credited basal.
This is what lets one instrument cover the whole age range efficiently and measure each person near the region where their ability is most finely discriminated. It also connects to a familiar measurement hazard: a test with too low a ceiling suffers a ceiling effect, compressing the scores of the most able examinees who run out of hard enough items, which is one reason each edition extends the difficulty range. The third demonstration shows the mechanism directly: as the reader moves the examinee's ability level, the administered window of items slides up and down the difficulty ladder while the levels outside it are inferred rather than tested.
Adaptive administration: basal and ceiling bound a short block of items
The routing subtest starts the examinee near level 12. Testing continues down to a basal of level 11 (items easy enough to be passed and credited without administration) and up to a ceiling of level 14 (items hard enough to be failed). Only 4 of the 20 levels are actually given; the rest are inferred. This basal-ceiling logic is what lets one instrument span ages two to over eighty-five without asking anyone the whole item bank.
Reliability, Validity, and What the Score Predicts
The Stanford-Binet's full-scale IQ is highly reliable, and its standing rests, like any intelligence test's, on what the score forecasts. General cognitive ability measured in childhood predicts later educational achievement strongly (Deary et al., 2007), and across the lifespan it forecasts occupational attainment, health, and longevity, among the more robust predictive relationships in psychology (Deary, Penke, & Johnson, 2010). Because the full-scale score chiefly estimates g, the SB5 was deliberately built with a strong general-factor apex, which is part of why its composite carries so much predictive information.
Prediction is not explanation, and the authoritative task-force review that followed a period of public controversy is the model for how to state what is known: the tests are reliable and predictively valid, their scores are substantially heritable within groups, and the causes of average differences between groups were not established by the evidence available (Neisser et al., 1996). The professional standards make documenting a test's validity evidence and the fairness of its use an obligation of the examiner rather than an optional extra (AERA et al., 2014). None of this is special pleading for the Stanford-Binet; it is the frame within which any responsible use of its scores sits.
Worked Example
Follow the three demonstrations through one coherent case, checking that the arithmetic on the page matches the arithmetic in the demos.
Start with the ratio-versus-deviation contrast. Take a child whose development runs at 1.2 times the typical pace, and let mental age plateau at 16, as mental-age norms do in the late teens. At chronological age 10 the mental age is 1.2 × 10 = 12, so the ratio IQ is (12 / 10) × 100 = 120, and the deviation IQ, which simply expresses the same 1.2 rate as a standing, is 1.2 × 100 = 120. The two agree in childhood. Now follow the same child to chronological age 20. Mental age is capped at the plateau, 1.2 × 16 = 19.2, so the ratio IQ is now (19.2 / 20) × 100 = 96, while the deviation IQ is unchanged at 120. A 24-point ratio drop for a person whose ability never fell is exactly the artefact that retired the ratio IQ at Form L-M.
Now the factor profile. Suppose the SB5 returns five factor indexes: Fluid Reasoning 120, Knowledge 105, Quantitative Reasoning 110, Visual-Spatial Processing 95, and Working Memory 100. Averaging them as an equal-weight illustration gives (120 + 105 + 110 + 95 + 100) / 5 = 530 / 5 = 106, and the profile spread from the lowest to the highest factor is 120 − 95 = 25 points. The spread matters clinically: a 25-point gap flags an uneven profile that a single full-scale number would hide, which is why the factor indexes are reported alongside it. (The published SB5 composite is derived from subtest norms, not a plain mean; the equal-weight average here is only to make the demo's logic transparent.)
Finally the basal-ceiling administration. Place an examinee at ability level 12 on the demonstration's 20-level ladder, where an item at level L is passed exactly when ability is at least L. Descending from level 12, the first two consecutive passed levels set the basal at level 11; ascending, the first two consecutive failed levels set the ceiling at level 14. Only levels 11 through 14 are actually administered, ceiling minus basal plus one, 14 − 11 + 1 = 4 of the 20 levels; the rest are credited or skipped. Testing four levels instead of twenty, near the region where the examinee's ability is best discriminated, is the efficiency that lets one instrument span the whole age range.
Discussion
The Stanford-Binet's history is, in miniature, the history of how psychology learned to score an intelligence test. It introduced the intelligence quotient to American practice, discovered the fatal flaw in the mental-age ratio, adopted the deviation IQ that fixed it, and finally reorganized itself around an explicit theory of cognitive abilities. Each move was a genuine improvement in measurement, and the current edition is a well-normed, highly reliable instrument whose full-scale score predicts consequential outcomes as dependably as any in the field.
The tensions that remain are the ones common to all intelligence testing rather than peculiar to this test. The five-factor scoring structure is more confidently printed than the factor-analytic evidence for five separable factors warrants, with the general factor doing most of the work (DiStefano & Dombrowski, 2006), which bears on the broader question of whether g is a causal entity or a statistical summary of overlapping processes (Kovacs & Conway, 2016). Its norms drift upward with the Flynn effect and must be periodically refreshed (Flynn, 1987), and its scores, though substantially heritable, are demonstrably movable by schooling (Ritchie & Tucker-Drob, 2018). The measured score is informative and consequential, and it is not the whole of a person's intelligence, a distinction the field's most careful statements have always kept in view.
Current Directions
Two active lines of work bear directly on how a Stanford-Binet score should be understood. The first is genomic. Polygenic scores derived from genome-wide association studies now predict a meaningful share of the variance in intelligence-test performance directly from DNA, and the research frontier is joining these scores to brain imaging to trace the path from genetic variation through neural structure to measured ability (Plomin & von Stumm, 2018; Deary, Cox, & Hill, 2022). This promises a biological account of what the full-scale score partly reflects, while sharpening old ethical questions about prediction from the genome.
The second is the continuing analysis of environmental malleability. The demonstrated causal effect of education on measured intelligence is now a fixed point any complete account of the score must accommodate (Ritchie & Tucker-Drob, 2018), and it sits alongside the theoretical reassessment of the general factor that process-based accounts have reopened (Kovacs & Conway, 2016). For a test whose fifth edition was built around a strong g apex and five CHC factors, both questions are live: what exactly the composite measures, and how far the number can be moved. A sixth edition, whenever it comes, will have to answer them with fresh norms against a still-rising population mean.
Key Researchers
Alfred Binet (1857-1911). French psychologist at the Sorbonne; with Théodore Simon built the 1905 Binet-Simon scale, the age-graded ancestor of the Stanford-Binet, whose method of ordering tasks by the age at which a typical child passes them Terman carried directly into the American revision. Wikipedia
Raymond B. Cattell (1905-1998). Psychologist at the University of Illinois; distinguished fluid from crystallized intelligence, part of the Cattell-Horn-Carroll framework on which the SB5's five-factor structure was built. Wikipedia
Ian J. Deary. Professor of differential psychology at the University of Edinburgh; established the lifespan predictive reach of childhood intelligence scores for education, health, and mortality, the external-validity evidence that underwrites the interpretation of any Stanford-Binet score. ORCID - Wikipedia
James R. Flynn (1934-2020). Political scientist at the University of Otago; documented the secular rise in IQ scores, the Flynn effect, which is why the Stanford-Binet has been renormed at each revision and why old-norm scores on any edition run high. Wikipedia
Kevin S. McGrew. Director of the Institute for Applied Psychometrics; consolidated the Cattell-Horn-Carroll taxonomy that gave the SB5 its five explicitly CHC-aligned factor indexes and their verbal and nonverbal routings. Google Scholar - Wikipedia
Maud A. Merrill (1888-1978). Psychologist at Stanford University; co-authored the 1937 Terman-Merrill revision and led the 1960 Form L-M merger of its two parallel forms, extending the test's age range and modernizing its norms. Wikipedia
Robert Plomin. Behavioural geneticist at King's College London; led the move from twin-study heritability to the molecular genetics of intelligence, building the polygenic scores that now predict a share of the variance a test like the Stanford-Binet measures directly from DNA. ORCID - Faculty Page - Wikipedia
Stuart J. Ritchie. Psychologist at King's College London; meta-analyzed natural experiments to show that schooling causally raises measured intelligence, direct evidence that a Stanford-Binet score is partly an environmental product rather than a fixed capacity. ORCID - Google Scholar - Faculty Page
Gale H. Roid. Author of the Stanford-Binet Intelligence Scales, Fifth Edition (2003); rebuilt the test as a point scale with five CHC-aligned factors measured across verbal and nonverbal routings, ending the age-scale format the test had used since 1916. Faculty Page
Théodore Simon (1873-1961). French psychiatrist; Binet's collaborator on the 1905 scale and its revisions, whose standardized administration and age-graded item ordering became the template the Stanford-Binet preserved through every edition. Wikipedia
Charles Spearman (1863-1945). Psychologist at University College London; discovered the positive manifold and the general factor g, which the Stanford-Binet's full-scale IQ has always been read chiefly as an estimate of, and around whose strong apex the SB5 was explicitly designed. Wikipedia
Lewis M. Terman (1877-1956). Psychologist at Stanford University; produced the 1916 Stanford revision of the Binet-Simon scale, the Stanford-Binet, which imported the intelligence quotient into American testing and gave the test its name and its home institution. Wikipedia
David Wechsler (1896-1981). Clinical psychologist at Bellevue Psychiatric Hospital; introduced the deviation IQ, which the Stanford-Binet adopted at its third revision, Form L-M in 1960, abandoning the mental-age ratio that Terman's original had made famous. Wikipedia
Glossary
- Basal.
- The level below which items on an adaptive test are easy enough to be assumed passed and credited without administration, marking the lower edge of the block of items actually given.
- Binet-Simon scale.
- The 1905 French scale of age-graded tasks built by Binet and Simon, the first practical intelligence test and the direct ancestor of which the Stanford-Binet is the American revision.
- Cattell-Horn-Carroll theory.
- The consolidated taxonomy of cognitive abilities arranging narrow abilities under broad group factors under a general factor; the framework around which the SB5's five factor indexes are organized.
- Ceiling effect.
- The compression of scores that occurs when a test lacks items hard enough to distinguish among its ablest examinees, so their differences go unmeasured; a reason each edition extends its difficulty range.
- Ceiling.
- The level above which items on an adaptive test are hard enough to be assumed failed, marking the upper edge of the administered block; too low a ceiling produces a ceiling effect.
- Crystallized intelligence.
- The store of knowledge and skill accumulated through past learning; in Cattell's distinction, the residue of prior learning as opposed to the capacity to reason with novel material, tapped by the SB5 Knowledge factor.
- Deviation IQ.
- An intelligence score defined as a person's standing relative to same-age peers, scaled to a mean of 100 and a standard deviation of 15, which the Stanford-Binet adopted at the 1960 Form L-M in place of the mental-age ratio.
- Fluid intelligence.
- The capacity to reason and solve problems with novel material independent of acquired knowledge; the ability the SB5 Fluid Reasoning factor is built to measure.
- Form L-M.
- The 1960 third revision of the Stanford-Binet that merged the 1937 parallel forms L and M into one and replaced the ratio IQ with the deviation IQ.
- Full-scale IQ.
- The composite index a Stanford-Binet administration reports, an estimate of general cognitive ability drawn from all factors and both routings, scaled to a mean of 100 and a standard deviation of 15.
- General intelligence (g).
- The single common factor Spearman inferred from the positive correlations among all cognitive tasks; the construct a Stanford-Binet full-scale score chiefly estimates and around whose strong apex the SB5 was designed.
- Intelligence quotient (IQ).
- The standardized index the Stanford-Binet reports; originally mental age divided by chronological age times 100, now defined as a deviation score relative to an age norm.
- Mental age.
- The age level of the most difficult tasks a person can reliably pass, the developmental standing the Binet-Simon scale yielded and that the original Stanford-Binet ratio IQ used as its numerator.
- Nonverbal routing.
- The SB5 administration path that measures each of the five factors with tasks minimizing spoken language, yielding a Nonverbal IQ alongside the verbal path.
- Norm-referenced score.
- A score interpreted by comparison with the distribution of scores in a representative standardization sample rather than against an absolute standard; the form of scoring intrinsic to the Stanford-Binet.
- Ratio IQ.
- The original Stanford-Binet index, mental age divided by chronological age times 100; abandoned because mental age plateaus in adulthood while chronological age does not, so the ratio falls with age for constant ability.
- Routing subtest.
- A short initial subtest that locates an examinee's approximate ability so the Stanford-Binet can begin at an appropriate difficulty rather than at the bottom of the scale.
- Standardization.
- The fixing of administration, scoring, and interpretation procedures and the establishment of population norms, so that Stanford-Binet scores are comparable across examinees and meaningful against a reference distribution.
- Three-stratum theory.
- Carroll's factor-analytic model ordering cognitive abilities into three strata — many narrow abilities, a handful of broad group factors, and a single general factor — which, merged with the Cattell-Horn fluid-crystallized distinction, forms the Cattell-Horn-Carroll taxonomy behind the SB5's five factors.
Frequently Asked Questions
What is the Stanford-Binet test?
It is an individually administered, norm-referenced intelligence test, the American revision of the Binet-Simon scale first published by Lewis Terman at Stanford in 1916 and now in its fifth edition. It estimates general cognitive ability across the full age range and reports a full-scale IQ on a scale with a mean of 100 and a standard deviation of 15 (Terman, 1916; Roid, 2003).
How is the Stanford-Binet different from the Wechsler scales?
Both are individually administered intelligence tests indexed together in MeSH, and both report a deviation IQ. The chief difference is coverage: a single Stanford-Binet edition is normed across the whole age range from early childhood to late adulthood, whereas the Wechsler tradition splits into separate adult and child batteries (Roid, 2003).
What does the fifth edition (SB5) measure?
The SB5 is a point scale organized around five Cattell-Horn-Carroll factors: Fluid Reasoning, Knowledge, Quantitative Reasoning, Visual-Spatial Processing, and Working Memory. Each factor is measured through both a verbal and a nonverbal routing, so the test yields a Verbal IQ, a Nonverbal IQ, five factor indexes, and a full-scale IQ (Roid, 2003; McGrew, 2009).
Why did the Stanford-Binet abandon the ratio IQ?
The ratio IQ divides mental age by chronological age, but mental age plateaus in the late teens while chronological age keeps rising, so the ratio falls with age for a person of constant ability and cannot be used with adults. The 1960 Form L-M replaced it with the deviation IQ, which compares a person only with their own age group (Wechsler, 1958).
What are basal and ceiling rules?
They bound the block of items actually administered on an adaptive test. The basal is the level below which items are assumed passed and credited without administration; the ceiling is the level above which items are assumed failed. Only the items between them are given, which lets one instrument span a huge age range efficiently (Roid & Barram, 2004).
Is the Stanford-Binet affected by the Flynn effect?
Yes. Like every norm-referenced intelligence test its norms go stale as population performance rises about two to three points a decade, so a person of average ability scores above 100 against outdated norms. This is why the Stanford-Binet has been renormed at each revision (Flynn, 1987).
Does the SB5 really measure five separate abilities?
The test reports five factor indexes, but a confirmatory factor analysis of the standardization sample found the general factor dominant and the evidence for five cleanly separable factors weaker than the scoring structure implies (DiStefano & Dombrowski, 2006). The factor scores are a useful scoring convention rather than proof of five distinct entities.
What do Stanford-Binet scores predict?
The full-scale IQ chiefly estimates general ability, which predicts educational achievement, occupational attainment, health, and longevity across the lifespan (Deary et al., 2007; Deary, Penke, & Johnson, 2010). The scores are substantially heritable yet demonstrably raised by schooling (Ritchie & Tucker-Drob, 2018).
References
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
Becker, K. A. (2003). History of the Stanford-Binet Intelligence Scales: Content and psychometrics (Stanford-Binet Intelligence Scales, Fifth Edition, Assessment Service Bulletin No. 1). Riverside Publishing. https://www.hmhco.com/~/media/sites/home/hmh-assessments/clinical/stanford-binet/pdf/sb5_asb_1.pdf
Binet, A., & Simon, T. (1905). Methodes nouvelles pour le diagnostic du niveau intellectuel des anormaux. L'Annee Psychologique, 11, 191-244. https://doi.org/10.3406/psy.1904.3675
Carroll, J. B. (1993). Human cognitive abilities: A survey of factor-analytic studies. Cambridge University Press.
Cattell, R. B. (1963). Theory of fluid and crystallized intelligence: A critical experiment. Journal of Educational Psychology, 54(1), 1-22. https://doi.org/10.1037/h0046743
Deary, I. J., Strand, S., Smith, P., & Fernandes, C. (2007). Intelligence and educational achievement. Intelligence, 35(1), 13-21. https://doi.org/10.1016/j.intell.2006.02.001
Deary, I. J., Penke, L., & Johnson, W. (2010). The neuroscience of human intelligence differences. Nature Reviews Neuroscience, 11(3), 201-211. https://doi.org/10.1038/nrn2793
Deary, I. J., Cox, S. R., & Hill, W. D. (2022). Genetic variation, brain, and intelligence differences. Molecular Psychiatry, 27(1), 335-353. https://doi.org/10.1038/s41380-021-01027-y
DiStefano, C., & Dombrowski, S. C. (2006). Investigating the theoretical structure of the Stanford-Binet-Fifth Edition. Journal of Psychoeducational Assessment, 24(2), 123-136. https://doi.org/10.1177/0734282905285244
Flynn, J. R. (1987). Massive IQ gains in 14 nations: What IQ tests really measure. Psychological Bulletin, 101(2), 171-191. https://doi.org/10.1037/0033-2909.101.2.171
Kovacs, K., & Conway, A. R. A. (2016). Process overlap theory: A unified account of the general factor of intelligence. Psychological Inquiry, 27(3), 151-177. https://doi.org/10.1080/1047840X.2016.1153946
McGrew, K. S. (2009). CHC theory and the human cognitive abilities project: Standing on the shoulders of the giants of psychometric intelligence research. Intelligence, 37(1), 1-10. https://doi.org/10.1016/j.intell.2008.08.004
Neisser, U., Boodoo, G., Bouchard, T. J., Boykin, A. W., Brody, N., Ceci, S. J., Halpern, D. F., Loehlin, J. C., Perloff, R., Sternberg, R. J., & Urbina, S. (1996). Intelligence: Knowns and unknowns. American Psychologist, 51(2), 77-101. https://doi.org/10.1037/0003-066X.51.2.77
Plomin, R., & von Stumm, S. (2018). The new genetics of intelligence. Nature Reviews Genetics, 19(3), 148-159. https://doi.org/10.1038/nrg.2017.104
Ritchie, S. J., & Tucker-Drob, E. M. (2018). How much does education improve intelligence? A meta-analysis. Psychological Science, 29(8), 1358-1369. https://doi.org/10.1177/0956797618774253
Roid, G. H. (2003). Stanford-Binet Intelligence Scales, Fifth Edition: Technical manual. Riverside Publishing.
Roid, G. H., & Barram, R. A. (2004). Essentials of Stanford-Binet Intelligence Scales (SB5) assessment. John Wiley & Sons.
Spearman, C. (1904). General intelligence, objectively determined and measured. The American Journal of Psychology, 15(2), 201-293. https://doi.org/10.2307/1412107
Terman, L. M. (1916). The measurement of intelligence: An explanation of and a complete guide for the use of the Stanford revision and extension of the Binet-Simon Intelligence Scale. Houghton Mifflin.
Wechsler, D. (1958). The measurement and appraisal of adult intelligence (4th ed.). Williams & Wilkins. https://doi.org/10.1037/11167-000