Abstract
An intelligence test is the most encompassing form of aptitude test: a standardized instrument that estimates general cognitive ability from a broad sample of reasoning, knowledge, memory, and processing tasks, and reports the result as a single index scaled against a normative population. This article treats the intelligence test as a measurement problem: what construct it targets, how the intelligence quotient came to be defined as a deviation score rather than a mental-age ratio, how the factor structure of measured ability is organized, and what the scores predict. It follows the instrument from Binet and Simon's 1905 scale, through Wechsler's deviation IQ, to the general factor, the Cattell-Horn-Carroll taxonomy, the Flynn effect, and the modern genetics of ability. Three interactive demonstrations model the deviation IQ, the positive manifold behind a general factor, and the renorming problem the Flynn effect creates.
Keywords: intelligence quotient, general intelligence, Flynn effect
An intelligence test is built to place a person on a single dimension of general cognitive ability by sampling their performance across many different mental tasks and combining the results into one score. It is a species of aptitude test, sharing the machinery of standardized administration and norm-referenced scoring, but it is the broadest member of that family: where a specific aptitude test targets one narrow ability, an intelligence test deliberately spreads its items across verbal, quantitative, spatial, and memory domains so that the composite reflects what those domains have in common rather than any one of them. That commonality is the whole object of measurement, and the century of work behind the modern intelligence test is a sustained argument about what it is, how to scale it, whether it is one thing or many, and what its measurement licenses anyone to conclude.
- An intelligence test is the broadest form of aptitude test, estimating general cognitive ability from a wide sample of tasks and reporting a single norm-referenced index.
- The modern IQ is a deviation score, fixing a person's standing relative to same-age peers on a scale with a mean of 100 and a standard deviation of 15, having replaced the older mental-age ratio.
- Scores across all cognitive tasks correlate positively, the positive manifold, from which a general factor g is inferred; the Cattell-Horn-Carroll model organizes this into narrow, broad, and general strata.
- Test scores have risen roughly two to three points a decade across the twentieth century, the Flynn effect, which forces periodic renorming and complicates any claim that the tests measure a fixed trait.
- Intelligence-test scores are substantially heritable yet demonstrably shaped by environment, including schooling, and they predict educational, occupational, and health outcomes while remaining the subject of enduring debate over fairness and interpretation.
What an Intelligence Test Is
An intelligence test differs from a narrow ability test in the breadth of what it samples and the interpretation of what it yields. Its subtests are chosen to span distinct cognitive domains precisely so that the variance they share can be isolated from the variance peculiar to each, and the headline score is an estimate of that shared variance: a summary of how a person performs across mental tasks in general rather than at any single one. Because the target is a relative standing rather than an absolute quantity, an intelligence test is inherently norm-referenced. A raw score means nothing until it is compared with the distribution of scores in a representative standardization sample, and the reported index is a statement about where the examinee falls in that distribution, not about how many items were answered correctly.
Two properties follow from this design. First, an intelligence test must be standardized: administered, scored, and interpreted under fixed conditions, because a norm-referenced score is only meaningful if the examinee was tested the same way the norming sample was. Second, its quality is judged by the twin criteria of reliability, the consistency of the score across items, occasions, and forms, and validity, the degree to which the score supports the interpretations placed on it, including its correlation with external criteria (AERA et al., 2014). An intelligence test that is broad but unreliable, or reliable but predicts nothing, fails on its own terms. What makes the instrument distinctive is that the construct it estimates is not defined by any curriculum or task but emerges statistically from the pattern of correlations among its parts.
Types of Intelligence Tests
MeSH places the intelligence test beneath aptitude tests and indexes two narrower descriptors under it, the Stanford-Binet Test and the Wechsler Scales, the two individually administered batteries that have defined the field for a century. The classification is an indexing convenience for retrieval rather than an exhaustive typology: these two named families sit alongside group-administered tests, nonverbal and culture-reduced tests such as the Raven's Progressive Matrices, and infant and neuropsychological scales, and the categories are not mutually exclusive. Table 1 lists the two directly indexed subtypes; neither has a dedicated article yet, so neither is linked.
| Subtype | In brief |
|---|---|
| Stanford-Binet Test | The American revision of the Binet-Simon scale, an individually administered battery that introduced the intelligence quotient and remains a standard for the full age range. |
| Wechsler Scales | The family of adult and child batteries (WAIS, WISC) that introduced the deviation IQ and report a full-scale score alongside index scores for distinct ability domains. |
The two indexed families do not exhaust the instrument in practice. Group-administered intelligence tests trade the diagnostic richness of an individual session for the ability to test many people at once, as the US Army's Alpha and Beta tests first did in 1917. Nonverbal and culture-reduced tests such as the Raven's Progressive Matrices attempt to measure reasoning with minimal reliance on language or acquired knowledge, and their stability across cultures and cohorts has made them a workhorse of research on fluid ability (Raven, 2000). What unites all of these with the two named subtypes is the design goal of a broad, norm-referenced estimate of general ability, however the sampling is arranged.
Origins: From Mental Age to the Deviation IQ
The practical intelligence test begins in 1905, when Alfred Binet and Théodore Simon, commissioned to identify Parisian schoolchildren who needed special instruction, built a graded series of tasks ordered by the age at which a typical child could pass them (Binet & Simon, 1905). Their scale yielded a mental age: the age level of the hardest tasks a child could reliably complete, so that a child who passed the items typical of eight-year-olds had a mental age of eight regardless of chronological age. This was a measurement of standing relative to a developmental norm, and it introduced the defining logic of the intelligence test, that a score is a comparison with a reference population rather than a count.
The mental age became a quotient in America. Lewis Terman's 1916 Stanford revision of the Binet-Simon scale, the Stanford-Binet, adopted the intelligence quotient, mental age divided by chronological age and multiplied by 100, so that a child whose mental age matched their chronological age scored exactly 100 (Terman, 1916). The ratio IQ was intuitive and enormously influential, but it had a fatal defect: mental age stops rising in adulthood while chronological age does not, so the ratio falls with age for a person of constant ability and the same numerical IQ means different things at different ages. David Wechsler solved this in 1958 by redefining the IQ as a deviation score: a person's standing expressed in standard-deviation units relative to the distribution of their own age group, scaled to a mean of 100 and a standard deviation of 15 (Wechsler, 1958). The deviation IQ severed intelligence measurement from the mental-age ratio entirely, and it is the definition every major test now uses. Figure 1 shows how a raw performance is converted into this scale.
Figure 1
The Deviation IQ Scale
The first demonstration makes this conversion manipulable. It draws the normal distribution of ability, lets the reader move a raw performance up and down the scale, and reports the corresponding deviation IQ and percentile rank, so that the relationship between a position on the curve and an IQ score becomes concrete.
The deviation IQ: from a position on the curve to a score and a percentile
A performance 1.0 standard deviations from the age-group mean is a deviation IQ of 115, placing the examinee at the 84th percentile — the shaded proportion of the population scoring at or below them. The IQ is a restatement of a standing on the normal curve: it counts nobody’s correct answers, only their position relative to peers.
The Structure of Measured Intelligence
Behind the single IQ score lies a century-old question: is intelligence one thing or many? The empirical starting point is Charles Spearman's 1904 observation that performances on unrelated cognitive tasks are all positively correlated, a pattern he explained by a single general factor, g, common to every task, plus a factor specific to each (Spearman, 1904). This positive manifold, the fact that doing well on one mental test predicts doing well on almost any other, is why a broad composite carries so much information and is the construct an intelligence test's full-scale score chiefly estimates. The second demonstration builds this from the correlations up.
The positive manifold: how a general factor emerges from correlated subtests
| Mean subtest correlation | 0.47 |
| Variance of composite from g | 82% |
| Manifold present? | yes |
Each subtest loads on the common factor; the correlation between any two is the product of their loadings. Set the strength to zero and every off-diagonal correlation vanishes — no manifold, no g.
With the common factor at 1.00, distinct subtests correlate 0.47 on average and their composite draws 82% of its variance from what they share. The all-positive off-diagonal is the positive manifold; the shared variance it implies is what a full-scale IQ chiefly estimates. Whether that shared factor is a single ability or an overlap of many processes is the open theoretical question.
A single factor, however, does not capture the whole structure. Raymond Cattell distinguished fluid from crystallized intelligence in 1963, separating the capacity to reason with novel material from the store of knowledge accumulated through past learning (Cattell, 1963), and later work resolved measured ability into several broad factors between the narrow subtests and the general apex. John Carroll's survey of 460 datasets and the Cattell-Horn tradition were fused into the Cattell-Horn-Carroll (CHC) taxonomy, consolidated by Kevin McGrew, which arranges narrow abilities under about eight broad group factors under a general factor, and which now organizes and interprets the subtests of most contemporary intelligence batteries (McGrew, 2009). A modern test's full-scale score approximates g, its index scores approximate the broad factors, and its subtests sample the narrow abilities. Whether g is a real causal entity or a statistical summary of overlapping cognitive processes remains contested, a question recent process-based accounts have reopened (Kovacs & Conway, 2016).
What Intelligence Tests Predict
An intelligence test earns its standing from what its scores forecast. Childhood intelligence-test scores predict later educational achievement strongly (Deary et al., 2007), and general cognitive ability forecasts occupational attainment, job performance, health, and longevity across the lifespan, correlations that are among the more robust in psychology even where their size is debated (Deary, Penke, & Johnson, 2010). The neuroscience of these differences, from brain size and white-matter integrity to processing efficiency, is an active attempt to ground the predictive power of the score in biology (Deary, Penke, & Johnson, 2010).
Prediction, however, is not explanation, and a score that forecasts an outcome does not by itself reveal what it measures or why it works. The authoritative 1996 task-force report commissioned after a period of public controversy laid out what was securely known and what was not: that the tests are reliable and predictively valid, that scores are substantially heritable within groups, and that the causes of average differences between groups were not established by the evidence then available (Neisser et al., 1996). That careful separation of the well-supported from the speculative remains the responsible frame, and the professional standards make the documentation of a test's validity evidence, and the fairness of its use, an obligation of the test user rather than an optional extra (AERA et al., 2014).
The Flynn Effect and Renorming
The most striking complication for the idea that intelligence tests measure a fixed trait is that the scores keep rising. James Flynn documented that raw test performance climbed massively across the twentieth century in every nation with adequate data, so that norms established in one decade left later cohorts scoring well above the intended mean of 100 (Flynn, 1987). A meta-analysis of the accumulated evidence puts the gain at roughly two to three IQ points a decade, concentrated more on fluid than on crystallized measures (Trahan et al., 2014). Because the tests are norm-referenced, a rising population mean means a test's norms go stale: a person of genuinely average ability scores above 100 when measured against outdated norms, so tests must be periodically renormed to restore the mean to 100. The third demonstration shows how this drift inflates a fixed ability into a rising apparent IQ as norms age.
The Flynn effect: stale norms inflate a fixed ability into a rising IQ
Against norms 20 years old, a person of genuinely average ability is measured at IQ 104.6 — an inflation of 4.6 points, because the population mean has drifted up by about 2.3 points a decade since the norms were fixed. Renorming resets the age to zero and the mean to 100. The older a test’s norms, the more it flatters, which is why publishers renorm on a schedule and the standards require current norms.
The Flynn effect is theoretically awkward because genes cannot have changed over three generations, so the gains must be environmental, yet the same scores are highly heritable within any one cohort. The resolution is that the between-cohort change and the within-cohort variation have different causes, and the gains themselves point to environmental drivers such as better nutrition, more schooling, smaller families, and a more cognitively demanding visual and technical culture (Nisbett et al., 2012). The effect is a standing reminder that an IQ score is a standing relative to a moving population, not an absolute measurement of a stable quantity.
Heritability and Environment
Intelligence-test scores are among the most heritable of behavioural traits, with heritability estimates rising from around 0.2 in early childhood to 0.6 or higher in adulthood, and molecular genetics has begun to identify the many common variants of small effect that underlie this, building polygenic scores that predict test performance directly from DNA (Plomin & von Stumm, 2018). High heritability is easily misread as fixity, but it is neither: it is a population statistic about the sources of variation in a particular environment, not a statement about how malleable any individual's ability is. The clearest evidence of malleability is education itself. A meta-analysis of natural experiments found that an additional year of schooling raises measured intelligence by a few IQ points, an effect that persists across the lifespan (Ritchie & Tucker-Drob, 2018). Heritability and environmental influence are not competitors dividing a fixed pie; both are large, and the most recent genetic work aims to integrate them by tracing how genetic variation, brain structure, and cognitive ability are related rather than treating the score as a readout of either alone (Deary, Cox, & Hill, 2022).
Worked Example
Consider how the same underlying performance is scored under the two historical definitions of the IQ, and how the Flynn effect corrupts the answer if norms are left to age.
Take the ratio IQ first. A child of chronological age 10 performs at the level typical of 12-year-olds, giving a mental age of 12. The ratio IQ is (mental age / chronological age) × 100 = (12 / 10) × 100 = 120. Now follow the same child to age 20, and suppose their mental age has reached the adult plateau of 18 (mental-age norms flatten in the late teens). The ratio IQ is now (18 / 20) × 100 = 90 — a 30-point drop for a person whose ability never declined, purely because the denominator kept growing while the numerator stopped. This artefact is exactly why the ratio IQ was abandoned.
Now the deviation IQ, which has no such flaw because it compares a person only with their own age group. Suppose an adult scores one standard deviation above the mean of their age peers, z = +1.0. The deviation IQ is 100 + 15z = 100 + 15(1.0) = 115. Their percentile rank is the proportion of the normal distribution below z = 1.0, which is Φ(1.0) = 0.8413, the 84th percentile. An examinee two standard deviations above the mean, z = +2.0, scores 100 + 15(2.0) = 130, at the Φ(2.0) = 0.9772, or roughly 98th, percentile. The score is a fixed statement of standing regardless of age.
Finally the Flynn correction. Take the meta-analytic gain of 2.3 IQ points per decade, or 0.23 points per year (Trahan et al., 2014). A test was last normed 20 years ago. A person of genuinely average ability today, tested against those stale norms, is compared with a population whose mean was 2.3 × 2 = 4.6 points lower than today's, so their measured score is about 100 + 4.6 = 104.6 rather than the 100 their true standing warrants. The inflation is modest over 20 years but grows with the age of the norms, and in a high-stakes context, such as a cut score for a diagnosis or a program, a systematic four-to-five-point inflation is enough to change decisions — which is why the standards require current norms and why publishers renorm on a schedule.
Discussion
The intelligence test is at once one of psychology's most successful measurement technologies and one of its most contested. Its success is not in doubt: the tests are highly reliable, their scores load on a general factor that shows up in every dataset, and those scores predict a wide range of consequential outcomes with a consistency few psychological measures match. The deviation IQ is a clean, well-understood scale, and the CHC taxonomy gives a principled account of what the subtests sample. On the terms it sets for itself, the instrument works.
The controversies live at the boundary between the measurement and its interpretation. The Flynn effect shows that the scale measures a standing relative to a population that is itself moving, so the tests cannot be read as an absolute gauge of a fixed quantity (Flynn, 1987). Substantial heritability coexists with a real and lasting effect of schooling, so no simple nature-or-nurture reading survives contact with the evidence (Ritchie & Tucker-Drob, 2018). And the question of whether the general factor is a causal entity or a statistical summary of overlapping processes bears directly on what a full-scale IQ is a measurement of (Kovacs & Conway, 2016). None of these unsettles the predictive record; all of them caution against the over-reading that has dogged the instrument since Terman. The measured score is informative and consequential, and it is not the same thing as the whole of human intelligence — a distinction the field's most careful statements have always insisted on (Neisser et al., 1996).
Current Directions
Three lines of work are reshaping intelligence testing. The first is genomic. Polygenic scores derived from genome-wide association studies now predict a meaningful share of the variance in test performance directly from DNA, and the research frontier is integrating these scores with brain imaging to trace the pathway from genetic variation through neural structure to measured ability (Plomin & von Stumm, 2018; Deary, Cox, & Hill, 2022). This promises a biological account of the score but also sharpens old ethical questions about prediction from the genome.
The second is theoretical: a renewed challenge to the causal interpretation of g. Process overlap theory and related accounts argue that the positive manifold arises not from a single underlying ability but from the fact that many domain-general executive processes are tapped in common across tests, so that g is an emergent statistical byproduct rather than a thing the test measures (Kovacs & Conway, 2016). If this is right, the full-scale IQ remains predictively useful while ceasing to name a unitary trait. The third is the continuing analysis of environmental malleability, with the demonstrated causal effect of education on measured intelligence now a fixed point that any complete theory of the score must accommodate (Ritchie & Tucker-Drob, 2018). Across all three, the mature psychometrics of intelligence testing is being asked to say more precisely what the score is, where it comes from, and how far it can be moved.
Key Researchers
Alfred Binet (1857-1911). French psychologist at the Sorbonne; with Théodore Simon built the 1905 Binet-Simon scale, the first practical intelligence test, a graded series of age-ordered tasks designed to identify children needing special instruction. Wikipedia
Raymond B. Cattell (1905-1998). Psychologist at the University of Illinois; distinguished fluid from crystallized intelligence, the theoretical split between the capacity to reason with novel material and the residue of past learning that structures modern test interpretation. Wikipedia
Ian J. Deary. Professor of differential psychology at the University of Edinburgh; established the lifespan predictive reach of childhood intelligence scores for education, health, and mortality, and mapped the neuroscience and genetics of intelligence differences. ORCID - Wikipedia
James R. Flynn (1934-2020). Political scientist at the University of Otago; documented the massive secular rise in IQ scores across 14 nations, the Flynn effect, forcing the field to separate what tests measure from the trait and to renorm periodically. Wikipedia
Kevin S. McGrew. Director of the Institute for Applied Psychometrics; consolidated the Cattell-Horn-Carroll taxonomy of narrow, broad, and general cognitive abilities that current intelligence batteries use to organize and interpret their subtests. Google Scholar - Wikipedia
Robert Plomin. Behavioural geneticist at King's College London; led the move from twin-study heritability to the molecular genetics of intelligence, building the polygenic scores that predict test performance directly from DNA. ORCID - Faculty Page - Wikipedia
Stuart J. Ritchie. Psychologist at King's College London; meta-analyzed natural experiments to show that an additional year of schooling causally raises measured intelligence by a few IQ points, direct evidence that test scores are partly an environmental product. ORCID - Google Scholar - Faculty Page
Théodore Simon (1873-1961). French psychiatrist; Binet's collaborator on the 1905 scale and its revisions, whose standardized administration and age-graded item ordering became the template for every later individual intelligence test. Wikipedia
Charles Spearman (1863-1945). Psychologist at University College London; discovered the positive manifold and the general factor g, the construct an intelligence test's full-scale score chiefly estimates and the reason a single composite predicts so widely. Wikipedia
Robert J. Sternberg. Professor of psychology at Cornell University; the leading critic of the single-factor view from within the field, arguing that conventional tests sample a narrow analytic slice of ability and neglect creative and practical intelligence. ORCID - Google Scholar - Faculty Page
Lewis M. Terman (1877-1956). Psychologist at Stanford University; produced the 1916 Stanford-Binet, which imported the intelligence quotient into American testing and set the pattern for mass ability measurement. Wikipedia
David Wechsler (1896-1981). Clinical psychologist at Bellevue Psychiatric Hospital; built the Wechsler-Bellevue, WAIS, and WISC and replaced the mental-age ratio with the deviation IQ, the scale essentially every modern intelligence test now uses. Wikipedia
Glossary
- Cattell-Horn-Carroll theory.
- The consolidated taxonomy of cognitive abilities that arranges narrow abilities under broad group factors under a general factor, and that organizes the subtests of most modern intelligence batteries.
- Ceiling effect.
- The compression of scores that occurs when a test lacks items hard enough to distinguish among its most able examinees, so that high performers all reach the maximum and their differences go unmeasured.
- Crystallized intelligence.
- The store of knowledge and skill accumulated through past learning and experience; in Cattell's distinction, the residue of prior learning as opposed to the capacity to reason with novel material.
- Deviation IQ.
- An intelligence score defined as a person's standing relative to same-age peers, scaled to a mean of 100 and a standard deviation of 15, which replaced the mental-age ratio and is used by every major modern test.
- Fluid intelligence.
- The capacity to reason and solve problems with novel material independent of acquired knowledge; the aspect of ability on which the secular Flynn-effect gains are largest.
- Flynn effect.
- The substantial rise in raw intelligence-test performance across the twentieth century, roughly two to three IQ points a decade, which forces periodic renorming to keep the population mean at 100.
- General intelligence (g).
- The single common factor Spearman inferred from the positive correlations among all cognitive tasks; the construct an intelligence test's full-scale score chiefly estimates and the main source of its broad predictive power.
- Heritability.
- The proportion of variation in a trait within a population that is attributable to genetic differences in a given environment; for intelligence scores it rises with age from about 0.2 in childhood to 0.6 or higher in adulthood.
- Intelligence quotient (IQ).
- The standardized index an intelligence test reports; originally mental age divided by chronological age times 100, now defined as a deviation score relative to an age norm.
- Mental age.
- The age level of the most difficult tasks a person can reliably pass, the developmental standing that Binet and Simon's scale yielded and that the ratio IQ used as its numerator.
- Norm-referenced score.
- A score interpreted by comparison with the distribution of scores in a representative standardization sample rather than against an absolute standard; the form of scoring intrinsic to intelligence testing.
- Positive manifold.
- The empirical fact that scores on almost all cognitive tasks are positively correlated, so that able performance on one predicts able performance on others; the observation from which general intelligence was inferred.
- Predictive validity.
- The correlation between a test score and a later criterion such as educational or occupational success; a principal form of validity evidence for an intelligence test.
- Reliability.
- The consistency of a test score across items, occasions, and alternate forms; a precondition for validity, since a score that varies randomly cannot support any stable interpretation.
- Standardization.
- The fixing of administration, scoring, and interpretation procedures, and the establishment of population norms, so that scores are comparable across examinees and meaningful against a reference distribution.
- Wechsler Adult Intelligence Scale.
- The individually administered adult battery, one of the Wechsler scales, that introduced the deviation IQ and reports a full-scale score alongside index scores for distinct ability domains.
Frequently Asked Questions
What is an intelligence test?
It is a standardized instrument that estimates a person's general cognitive ability by sampling performance across a broad range of reasoning, knowledge, memory, and processing tasks, and reporting a single norm-referenced index. It is the broadest form of aptitude test, judged by its reliability and by the outcomes its scores predict (AERA et al., 2014).
What is IQ, and how is it calculated?
The intelligence quotient was originally a person's mental age divided by their chronological age, times 100. Since Wechsler, it is a deviation score: a standing relative to same-age peers, scaled to a mean of 100 and a standard deviation of 15, so that an IQ of 115 marks the 84th percentile (Wechsler, 1958).
How is an intelligence test different from an aptitude test?
An intelligence test is a type of aptitude test, the broadest one. Where a specific aptitude test targets a single narrow ability, an intelligence test deliberately samples many domains so its composite reflects the general ability they share. All are norm-referenced and judged by prediction (AERA et al., 2014).
What is the Flynn effect?
It is the finding that raw intelligence-test scores rose substantially across the twentieth century, roughly two to three points a decade, in every nation studied (Flynn, 1987; Trahan et al., 2014). Because the tests are norm-referenced, it forces periodic renorming to keep the mean at 100.
Are intelligence-test scores stable over a lifetime?
Rank order is fairly stable from childhood onward, and childhood scores predict adult education, health, and longevity (Deary et al., 2007; Deary, Penke, & Johnson, 2010). But the scores are not fixed: schooling causally raises them, and the population itself drifts upward over generations.
Is intelligence inherited or shaped by environment?
Both, substantially. Heritability rises with age to 0.6 or more, and polygenic scores now predict test performance from DNA (Plomin & von Stumm, 2018), yet an extra year of schooling causally raises measured IQ by a few points (Ritchie & Tucker-Drob, 2018). Heritability is a population statistic, not a measure of fixity.
Do intelligence tests measure one thing or many?
Both, at different levels. Scores across tasks correlate positively, supporting a general factor g (Spearman, 1904), while the Cattell-Horn-Carroll model resolves ability into broad and narrow factors beneath it (McGrew, 2009). Whether g is a real entity or a statistical byproduct is debated (Kovacs & Conway, 2016).
Are intelligence tests culturally biased?
The tests draw on knowledge and habits of mind that vary across environments, and average group differences and their causes remain contested, with the well-supported findings carefully separated from the speculative in the field's authoritative reviews (Neisser et al., 1996; Nisbett et al., 2012). Nonverbal and culture-reduced tests attempt to lessen the dependence on acquired content (Raven, 2000).
References
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
Binet, A., & Simon, T. (1905). Methodes nouvelles pour le diagnostic du niveau intellectuel des anormaux. L'Annee Psychologique, 11, 191-244. https://doi.org/10.3406/psy.1904.3675
Cattell, R. B. (1963). Theory of fluid and crystallized intelligence: A critical experiment. Journal of Educational Psychology, 54(1), 1-22. https://doi.org/10.1037/h0046743
Deary, I. J., Strand, S., Smith, P., & Fernandes, C. (2007). Intelligence and educational achievement. Intelligence, 35(1), 13-21. https://doi.org/10.1016/j.intell.2006.02.001
Deary, I. J., Penke, L., & Johnson, W. (2010). The neuroscience of human intelligence differences. Nature Reviews Neuroscience, 11(3), 201-211. https://doi.org/10.1038/nrn2793
Deary, I. J., Cox, S. R., & Hill, W. D. (2022). Genetic variation, brain, and intelligence differences. Molecular Psychiatry, 27(1), 335-353. https://doi.org/10.1038/s41380-021-01027-y
Flynn, J. R. (1987). Massive IQ gains in 14 nations: What IQ tests really measure. Psychological Bulletin, 101(2), 171-191. https://doi.org/10.1037/0033-2909.101.2.171
Kovacs, K., & Conway, A. R. A. (2016). Process overlap theory: A unified account of the general factor of intelligence. Psychological Inquiry, 27(3), 151-177. https://doi.org/10.1080/1047840X.2016.1153946
McGrew, K. S. (2009). CHC theory and the human cognitive abilities project: Standing on the shoulders of the giants of psychometric intelligence research. Intelligence, 37(1), 1-10. https://doi.org/10.1016/j.intell.2008.08.004
Neisser, U., Boodoo, G., Bouchard, T. J., Boykin, A. W., Brody, N., Ceci, S. J., Halpern, D. F., Loehlin, J. C., Perloff, R., Sternberg, R. J., & Urbina, S. (1996). Intelligence: Knowns and unknowns. American Psychologist, 51(2), 77-101. https://doi.org/10.1037/0003-066X.51.2.77
Nisbett, R. E., Aronson, J., Blair, C., Dickens, W., Flynn, J., Halpern, D. F., & Turkheimer, E. (2012). Intelligence: New findings and theoretical developments. American Psychologist, 67(2), 130-159. https://doi.org/10.1037/a0026699
Plomin, R., & von Stumm, S. (2018). The new genetics of intelligence. Nature Reviews Genetics, 19(3), 148-159. https://doi.org/10.1038/nrg.2017.104
Raven, J. (2000). The Raven's Progressive Matrices: Change and stability over culture and time. Cognitive Psychology, 41(1), 1-48. https://doi.org/10.1006/cogp.1999.0735
Ritchie, S. J., & Tucker-Drob, E. M. (2018). How much does education improve intelligence? A meta-analysis. Psychological Science, 29(8), 1358-1369. https://doi.org/10.1177/0956797618774253
Spearman, C. (1904). General intelligence, objectively determined and measured. The American Journal of Psychology, 15(2), 201-293. https://doi.org/10.2307/1412107
Terman, L. M. (1916). The measurement of intelligence: An explanation of and a complete guide for the use of the Stanford revision and extension of the Binet-Simon Intelligence Scale. Houghton Mifflin.
Trahan, L. H., Stuebing, K. K., Fletcher, J. M., & Hiscock, M. (2014). The Flynn effect: A meta-analysis. Psychological Bulletin, 140(5), 1332-1360. https://doi.org/10.1037/a0037173
Wechsler, D. (1958). The measurement and appraisal of adult intelligence (4th ed.). Williams & Wilkins. https://doi.org/10.1037/11167-000