Abstract
An aptitude test, which MeSH classifies under psychological tests, is a standardized instrument that measures a person's capacity to acquire a skill, using present performance to forecast future learning rather than to certify past attainment. This article treats the aptitude test as a prediction device: what distinguishes it from an achievement test, the factor-analytic theories of ability it rests on, and the evidence that its scores forecast training, job, and academic performance. It follows the field from Binet's first predictive scale and Spearman's general factor, through Thurstone's primary abilities and Cattell's fluid-crystallized distinction, to Carroll's three-stratum synthesis and the modern debate over how strongly, and how fairly, such tests predict. Three interactive demonstrations model predictive validity, aptitude-treatment interaction, and the way ability predictors shift as a skill is practised.
Keywords: predictive validity, general intelligence, range restriction
An aptitude test is an instrument built to answer a question about the future: not what a person has already learned, but how readily they will learn something they have not yet been taught. It shares its machinery with every other psychological test, sampling behaviour under standardized conditions to estimate an unobservable attribute, but its defining feature is the direction of the inference. An achievement test looks backward, certifying attainment; an aptitude test looks forward, forecasting performance in a task the examinee has not yet attempted. That forward inference is the whole reason the instrument exists and the whole basis on which it is judged, because an aptitude test earns its keep only insofar as its scores correlate with the outcome it claims to predict. The century of work behind the modern aptitude battery is therefore a century of two intertwined projects: a theory of what human abilities there are to measure, and a body of evidence about how well measuring them forecasts what people go on to do.
- An aptitude test measures the capacity to acquire a skill and is validated by how well its scores forecast future performance, distinguishing it from an achievement test, which certifies past learning.
- Its interpretation rests on factor theories of ability, from Spearman's general factor g through Thurstone's primary abilities to the Cattell-Horn-Carroll three-stratum model that organizes modern batteries.
- Predictive validity is a correlation, and its practical value depends on the selection ratio and base rate; a modest coefficient can still improve decisions substantially.
- Validity coefficients observed in an already-selected group are attenuated by range restriction, and correcting for it is both standard practice and, more recently, a contested source of inflated estimates.
- The predictors of performance are not fixed: general ability dominates early learning while narrow abilities take over with practice, and the balance of general versus specific ability remains an active debate.
What an Aptitude Test Is
The defining move of an aptitude test is prospective inference. Two examinees may sit an identical set of items, yet the same responses become an aptitude measurement or an achievement measurement depending only on the inference drawn from them: whether the score is read as evidence of what has been mastered, or as a forecast of what can be mastered. Because the target is a capacity that has not yet been exercised, an aptitude test is validated differently from an achievement test. Its content need not match any curriculum, and its quality is settled not by how faithfully it samples a body of taught material but by predictive validity, the correlation between the score obtained now and the performance observed later. Table 1 sets out the contrast that organizes the whole field.
| Dimension | Aptitude test | Achievement test |
|---|---|---|
| Inference | Capacity to acquire a skill not yet taught. | Mastery of material already taught. |
| Time orientation | Prospective; forecasts future performance. | Retrospective; certifies present attainment. |
| Content | Broad, not tied to a specific curriculum. | Bound to a defined body of instruction. |
| Validated by | Correlation with a later criterion. | Coverage of the content domain. |
| Typical use | Selection, placement, guidance. | Certification, grading, progress. |
The contrast is real but not absolute. Every aptitude test draws on knowledge the examinee has already acquired, since there is no way to sample a capacity except through behaviour that reflects prior learning, and a well-designed achievement test predicts future learning to some degree. The distinction is one of intended inference, not of watertight categories of item, which is why the same verbal-reasoning task can serve either purpose. What makes an instrument an aptitude test is that its scores are put to a predictive use and are held to a predictive standard.
Types of Aptitude Tests
MeSH indexes the aptitude test under a single narrower descriptor, intelligence tests, and places the category itself beneath psychological tests. The classification is an indexing convenience for retrieval rather than a theory of test kinds: an intelligence test is one especially broad aptitude measure, but aptitude testing also spans narrow single-ability tests and multi-aptitude batteries that report separate scores, and these are not mutually exclusive with the descriptor below. Table 2 lists the direct subtype; only those with a dedicated article are linked.
| Subtype | In brief |
|---|---|
| Intelligence Tests | Broad-band measures of general cognitive ability, the most encompassing aptitude tests, yielding an overall index such as the deviation IQ alongside subtest profiles. |
The single indexed subtype understates the diversity of aptitude instruments in practice. Alongside the broad intelligence test sit specific aptitude tests, each targeting one narrow ability such as mechanical reasoning, spatial visualization, or clerical speed, and multi-aptitude batteries such as the Differential Aptitude Tests or the military's ASVAB, which administer several specific tests together and report a profile of separate scores. Whether that profile adds anything beyond the single general score it also yields is one of the field's long-running empirical questions, taken up below.
Origins: From Binet to the Differential Battery
The practical aptitude test begins in 1905, when Alfred Binet and Théodore Simon, commissioned to identify Parisian schoolchildren who would struggle in ordinary instruction, built a graded series of tasks ordered by the age at which a typical child could pass them. The instrument was predictive by design: its purpose was to forecast scholastic difficulty before it occurred, and it did so by sampling reasoning and judgment rather than schoolroom knowledge. This was the prototype of the aptitude test, and every later ability battery inherits its logic of graded standardized tasks read as a forecast.
The theoretical scaffolding arrived a year earlier. In 1904 Charles Spearman observed that performances across unrelated cognitive tasks are all positively correlated, a pattern he explained by a single general factor, g, common to every task, plus factors specific to each (Spearman, 1904). The positive manifold, the empirical fact that able performance on one mental task predicts able performance on almost any other, is the reason a single composite score forecasts so widely, and it is the construct most aptitude batteries are built to capture. Louis Leon Thurstone then complicated the picture: applying his own multiple-factor analysis, he argued in 1938 that intelligence resolves into several primary mental abilities, verbal, numerical, spatial, and others, rather than one overarching factor (Thurstone, 1938). Thurstone's work is the psychometric charter for the differential battery, which reports separate ability scores rather than a single IQ.
Two mid-century developments deepened the theory of what aptitude tests measure. Raymond Cattell distinguished fluid from crystallized intelligence in 1963, separating the capacity to reason with novel material from the store of knowledge and skill accumulated through past learning (Cattell, 1963). The distinction matters directly to aptitude testing, since fluid reasoning is closer to a pure capacity to learn while crystallized ability is the residue of prior learning. John Carroll's 1993 survey then synthesized 460 factor-analytic datasets into a three-stratum model, arranging specific abilities under broad group factors under a general factor at the apex (Carroll, 1993). Fused with Cattell and Horn's work, this became the Cattell-Horn-Carroll (CHC) taxonomy, consolidated by Kevin McGrew, which now organizes and interprets the subtests of most contemporary aptitude and intelligence batteries (McGrew, 2009). Figure 1 sketches this hierarchy, the framework around which a modern aptitude battery's subtests are arranged.
Figure 1
The Three-Stratum Hierarchy of Cognitive Abilities
Predictive Validity
An aptitude test lives or dies by its predictive validity: the correlation between the score obtained now and the criterion measured later, whether training grades, job performance, or academic success. Because that correlation is the instrument's reason for existing, the practical question is not merely whether it is non-zero but what a coefficient of a given size buys in a real decision. The answer depends on two further quantities that have nothing to do with the test: the selection ratio, the proportion of applicants admitted, and the base rate, the proportion who would succeed anyway. A validity coefficient of 0.4 that looks unimpressive as a correlation can still move the success rate among those selected substantially when the selection ratio is low. This three-way relationship among validity, selection ratio, and base rate was first tabulated by Harold Taylor and J. T. Russell, whose 1939 tables give the expected proportion of successes among selected applicants for any combination of the three, and which remain the standard tool for translating a bare validity coefficient into the practical yield of a selection decision (Taylor & Russell, 1939).
The first demonstration makes that relationship manipulable. It draws a cloud of applicants whose aptitude scores correlate with a later criterion at a validity the reader sets, imposes a selection cut score and a success threshold, and reports the proportion of selected applicants who succeed, so that a bare correlation becomes the practical accuracy of a decision.
Predictive validity: from a correlation to the accuracy of a selection
| Base success rate | 51% |
| Selected who succeed | 69% |
| Selection ratio | 27% |
| Improvement | +18 pts |
Teal points are selected successes, rust are selected failures, grey are not selected. The lower the selection ratio, the more a fixed validity lifts the success rate above the base rate.
At a validity of 0.40 and a cut of z = 0.5, selecting on aptitude raises the success rate from 51% at the base rate to 69% among those chosen. A zero coefficient leaves the two equal: an aptitude test with no predictive validity adds nothing.
The evidence that aptitude tests do predict is among the most robust in applied psychology. Synthesizing eighty-five years of research, Frank Schmidt and John Hunter concluded that general mental ability is the single best predictor of job and training performance across occupations, and that its validity generalizes rather than shifting unpredictably from setting to setting (Schmidt & Hunter, 1998). General ability also forecasts occupational attainment and later performance over the span of a career (Schmidt & Hunter, 2004). The same holds in education: the Scholastic Assessment Test is so heavily g-loaded that it functions substantially as an aptitude measure (Frey & Detterman, 2004), childhood intelligence-test scores predict later educational achievement strongly (Deary et al., 2007), and standardized admissions tests predict graduate and professional success while retaining validity even after undergraduate grades are accounted for (Kuncel & Hezlett, 2007).
Range Restriction and Its Correction
A validity coefficient is almost always computed on people who have already been selected, and this quietly deflates it. If only high-scoring applicants are admitted, the range of aptitude scores among those whose performance is later measured is compressed, and a correlation computed on a restricted range is smaller than the correlation that would hold in the full applicant pool. Range restriction is therefore not a nuisance to be ignored but a systematic downward bias, and for most of the field's history the response was to correct for it, estimating what the validity would have been in the unrestricted population using the ratio of the unrestricted to the restricted standard deviation.
The correction is real and its logic is sound, but its application has become contentious. In a 2022 reanalysis, Paul Sackett and colleagues argued that the range-restriction corrections applied in influential meta-analyses had been based on artifact assumptions that were too large, inflating the corrected validity estimates for cognitive-ability tests, and that the true operational validities are meaningfully lower than the widely cited figures (Sackett et al., 2022). The reappraisal does not overturn the finding that aptitude tests predict, but it revises downward how strongly, and it is the most consequential recent correction to the selection literature. The professional standards treat the choice and documentation of such corrections as part of the validity evidence a test user is obliged to supply (AERA et al., 2014).
Aptitude, Instruction, and Practice
If an aptitude is a capacity to learn, then the natural test of the concept is whether it predicts how much someone benefits from instruction, and this reframes aptitude as readiness to profit from a given treatment rather than a fixed rank. Lee Cronbach and Richard Snow developed this into the theory of aptitude-treatment interaction (ATI): the best instructional method is not the same for everyone but depends on the learner's aptitude profile, so that two teaching methods can reverse their effectiveness across the range of a relevant aptitude (Cronbach & Snow, 1977). Where a simple main effect says one method is better, an interaction says the ranking of methods crosses over as aptitude changes.
The second demonstration shows this crossover directly. It plots predicted performance under two instructional treatments as a function of a learner's aptitude, lets the reader vary the strength of the interaction, and marks where the lines cross, so that the abstract claim that the better treatment depends on aptitude becomes a visible reversal.
Aptitude-treatment interaction: the better method depends on the learner
At interaction strength 0 the lines are nearly parallel and one treatment is better for everyone. As the strength rises the slopes diverge and the lines cross near z = -0.2: below the crossover the high-structure method wins, above it the low-structure method does. For the marked learner at z = 0.0, treatment A is predicted to work better (50 vs 48). This crossover is what it means for aptitude and treatment to interact.
A related discovery concerns what happens to the predictors of performance as a skill is practised. Phillip Ackerman showed that the ability that best predicts performance changes over the course of skill acquisition: general ability dominates early, when a task is novel and demands reasoning, but its predictive weight declines with practice as narrow perceptual-speed and psychomotor abilities take over on a task that has become routine (Ackerman, 1988). This refines what an aptitude test predicts and when, and it warns against assuming a single aptitude forecasts performance uniformly across the whole trajectory from novice to expert.
The third demonstration traces this shift. As the reader increases the number of practice trials, it shows the correlation of performance with general ability falling and the correlation with a narrow speed ability rising, so that the changing composition of what predicts skilled performance becomes a moving picture rather than a single snapshot.
Skill acquisition: the predictor of performance changes with practice
At 1 practice trials, performance correlates 0.63 with general ability and 0.12 with a narrow speed ability, so the dominant predictor is general ability. Early in learning a novel task loads on reasoning and general ability; with practice the task becomes routine and narrow perceptual-speed and psychomotor abilities take over. What an aptitude test predicts therefore depends on where along the learning curve the criterion is measured.
Worked Example
Consider an aptitude test used to select trainees, with an observed validity of 0.30 against training performance in a group of people who were themselves selected on a related measure. Two questions arise: how much does the score actually forecast, and how much of its apparent weakness is an artifact of range restriction?
Take the forecast first, in standardized units where both the aptitude score and the criterion have a standard deviation of 1. A validity of r = 0.30 means the test explains r² = 0.09, or 9%, of the variance in training performance, and the best prediction of an applicant's criterion score is r times their aptitude score: an applicant one standard deviation above the mean is predicted to perform 0.30 standard deviations above the mean. The uncertainty around that forecast is the standard error of estimate, √(1 − r²) = √(1 − 0.09) = √0.91 = 0.95 standard deviations, only slightly below the original spread of 1. A modest validity forecasts loosely for any one person, which is exactly why its practical value shows up in the aggregate hit rate the first demonstration reports, not in the precision of a single prediction.
Now correct for range restriction. Suppose the standard deviation of aptitude scores among the selected trainees is 0.60 of the standard deviation in the full applicant pool, so the ratio U = 1 / 0.60 = 1.667. Thorndike's Case II formula estimates the unrestricted validity as R = (r × U) / √(1 − r² + r²U²). Substituting, the numerator is 0.30 × 1.667 = 0.500, and the denominator is √(1 − 0.09 + 0.09 × 2.778) = √(0.91 + 0.250) = √1.160 = 1.077, giving R = 0.500 / 1.077 = 0.46. The correction lifts the observed 0.30 to an estimated 0.46 in the unrestricted pool, a substantial upward revision, and it is precisely the size of this routine correction that Sackett and colleagues argued had been overstated, on the ground that the assumed unrestricted variability was too large (Sackett et al., 2022). The worked numbers show why the debate matters: the same test is either a weak or a moderate predictor depending entirely on an artifact assumption the raw coefficient does not reveal.
Discussion
The aptitude test occupies an unusual position among psychological instruments: its scientific credentials and its social controversies both flow from the same fact, that it predicts. The predictive validity of general cognitive ability for training, job, and academic performance is one of the better-replicated findings in applied psychology, and it is the reason aptitude testing survives every wave of criticism. Yet prediction is also what makes the instrument consequential, because a test that forecasts who will succeed is inevitably used to decide who gets the chance to try, and the stakes of that use are what the professional standards on fairness and appropriate interpretation exist to govern (AERA et al., 2014).
Two tensions keep the field unsettled. The first is quantitative and recent: how strong the predictive validities actually are depends on corrections for range restriction and measurement error whose assumptions the Sackett reanalysis has reopened, so that even the headline numbers are now provisional (Sackett et al., 2022). The second is theoretical and old: whether the general factor is the right unit for prediction, or whether the specific abilities that differential batteries measure carry incremental information that a single g discards. The evidence that predictors shift across skill acquisition suggests the answer is not simple (Ackerman, 1988), and the question of what aptitude testing should measure, and how its results should be used, remains as much a matter of values as of psychometrics.
Current Directions
Three lines of work are reshaping the field. The first is the reappraisal of validity magnitudes already described: the correction of long-standing range-restriction assumptions has forced a re-examination of how well ability tests predict, and the revised, lower operational validities are propagating through the selection literature (Sackett et al., 2022). This is less a retreat than a recalibration, but it changes the numbers that decades of practice were built on.
The second is a revival of the specific-versus-general debate. A body of recent work argues that narrow abilities predict outcomes over and above general mental ability more than the dominant g-centric view allowed, so that the differential battery's separate scores may carry real incremental validity after all (Kell & Lang, 2017). Direct comparisons of the relative importance of general and specific abilities for career success suggest the balance is closer than the strong g position holds (Lang & Kell, 2020), reopening a question Thurstone and Spearman first joined a century ago. The third is the ongoing scrutiny of admissions testing in particular, where the predictive case for aptitude measures meets sustained debate over access and fairness, and where several institutions have moved to make such tests optional even as the evidence on their incremental validity is re-examined (Zwick, 2019). Across all three, the mature predictive technology of aptitude testing is being asked harder questions about how much it predicts, what exactly does the predicting, and to what social ends the prediction is put.
Key Researchers
Phillip L. Ackerman. Professor of psychology at the Georgia Institute of Technology; showed that the ability predictors of performance shift across skill acquisition, general ability dominating early learning while narrow perceptual-speed and psychomotor abilities take over with practice. ORCID - Google Scholar - Faculty Page
Alfred Binet (1857-1911). French psychologist at the Sorbonne; with Théodore Simon built the 1905 Binet-Simon scale, the prototype aptitude test, a graded series of tasks whose purpose was to forecast which children would struggle in ordinary instruction. Wikipedia
John B. Carroll (1916-2003). Psychologist at the University of North Carolina at Chapel Hill; synthesized 460 factor-analytic datasets into the three-stratum model of specific abilities, broad group factors, and general ability that modern aptitude batteries are organized around. Wikipedia
Raymond B. Cattell (1905-1998). Psychologist at the University of Illinois; distinguished fluid from crystallized intelligence, the key theoretical split for aptitude testing between the capacity to reason with novel material and the residue of past learning. Wikipedia
Lee J. Cronbach (1916-2001). Psychometrician at Stanford University; with Richard Snow framed aptitude-treatment interaction, the finding that the best instruction depends on the learner's aptitude profile, recasting aptitude as readiness to benefit from a treatment. Wikipedia
Nathan R. Kuncel. Professor of psychology at the University of Minnesota; marshalled the evidence that standardized admissions tests predict graduate and professional success and retain validity net of undergraduate grades, a central defence of aptitude testing in higher-education selection. Google Scholar - Faculty Page
Kevin S. McGrew. Director of the Institute for Applied Psychometrics; consolidated the Cattell-Horn-Carroll taxonomy of broad and narrow cognitive abilities that current aptitude and intelligence batteries use to organize and interpret their subtests. Google Scholar - Wikipedia
Paul R. Sackett. Professor of psychology at the University of Minnesota; led the 2022 reanalysis that corrected long-standing range-restriction assumptions and revised the validity of ability tests for selection downward from earlier meta-analytic estimates. ORCID - Faculty Page
Frank L. Schmidt (1944-2021). Industrial-organizational psychologist at the University of Iowa; with John Hunter developed psychometric meta-analysis and showed across eighty-five years of data that general mental ability is the single best predictor of job and training performance. Wikipedia
Charles Spearman (1863-1945). Psychologist at University College London; discovered the positive manifold and the general factor g, the construct most aptitude batteries are built to measure and the reason a single composite predicts across so many domains. Wikipedia
Louis Leon Thurstone (1887-1955). Psychometrician at the University of Chicago; introduced multiple-factor analysis and the seven primary mental abilities, the psychometric basis for differential aptitude batteries that report separate verbal, numerical, and spatial scores. Wikipedia
Glossary
- Achievement test.
- A test designed to measure what a person has already learned from a defined body of instruction, contrasted with an aptitude test by its retrospective, curriculum-bound inference.
- Aptitude-treatment interaction.
- The finding that the most effective instructional method depends on the learner's aptitude, so that two treatments can reverse their relative effectiveness across the range of a relevant ability.
- Aptitude.
- A relatively enduring capacity to acquire a skill or body of knowledge not yet taught; the attribute an aptitude test is built to forecast performance from.
- Base rate.
- The proportion of a group that would meet the success criterion without any selection; together with the selection ratio it determines the practical payoff of a given predictive validity.
- Cattell-Horn-Carroll theory.
- The consolidated taxonomy of cognitive abilities that arranges narrow abilities under broad group factors under a general factor, and that organizes the subtests of most modern aptitude and intelligence batteries.
- Crystallized intelligence.
- The store of knowledge and skill accumulated through past learning and experience; in Cattell's distinction, the residue of prior learning as opposed to the capacity to reason with novel material.
- Differential aptitude battery.
- A set of specific aptitude tests administered together and reported as a profile of separate ability scores rather than a single composite, grounded in Thurstone's primary mental abilities.
- Fluid intelligence.
- The capacity to reason and solve problems with novel material independent of acquired knowledge; in Cattell's distinction, the aspect of ability closest to a pure capacity to learn.
- General intelligence (g).
- The single common factor Spearman inferred from the positive correlations among all cognitive tasks; the construct most aptitude batteries are built to measure and the main source of their broad predictive power.
- Incremental validity.
- The predictive value a measure adds over and above the predictors already in use; the central question in asking whether specific abilities forecast outcomes beyond general mental ability.
- Positive manifold.
- The empirical fact that scores on almost all cognitive tasks are positively correlated, so that able performance on one predicts able performance on others; the observation from which general intelligence was inferred.
- Predictive validity.
- The correlation between a test score obtained now and a criterion measured later; the form of validity by which an aptitude test is principally judged, since its purpose is to forecast future performance.
- Primary mental abilities.
- The several broad, relatively independent factors, verbal, numerical, spatial, and others, that Thurstone extracted by multiple-factor analysis in place of a single general factor.
- Range restriction.
- The compression of score variance that occurs when a validity coefficient is computed only on an already-selected group, systematically deflating the observed correlation relative to the full applicant pool.
- Selection ratio.
- The proportion of applicants who are accepted; with the base rate, it governs how much a given predictive validity improves the success rate among those selected.
- Standard error of estimate.
- The standard deviation of the errors in predicting a criterion from a test score, equal to the criterion standard deviation times the square root of one minus the squared validity; it fixes how loosely a score forecasts an individual outcome.
- Three-stratum model.
- Carroll's hierarchical map of cognitive abilities, with numerous narrow abilities at the first stratum, about eight broad factors at the second, and a general factor at the third, derived from 460 factor-analytic datasets.
Frequently Asked Questions
What is an aptitude test?
It is a standardized test that measures a person's capacity to acquire a skill, using present performance to forecast how well they will do at a task they have not yet been trained on. Unlike an achievement test, which certifies past learning, its quality is judged by predictive validity, the correlation of its scores with a later criterion (AERA et al., 2014).
How is an aptitude test different from an achievement test?
The difference is the inference, not necessarily the items. An aptitude test looks forward, forecasting future performance, and need not match any curriculum; an achievement test looks backward, certifying mastery of a defined body of instruction. The same reasoning task can serve either purpose depending on whether its score is read as a forecast or as attainment.
Do aptitude tests actually predict performance?
Yes. General mental ability is among the best single predictors of job and training performance across occupations (Schmidt & Hunter, 1998), and standardized ability and admissions tests predict educational and graduate success, retaining validity even after prior grades are accounted for (Kuncel & Hezlett, 2007; Deary et al., 2007).
What is general intelligence, or g?
It is the single common factor Charles Spearman inferred in 1904 from the observation that performances on all cognitive tasks are positively correlated (Spearman, 1904). Because so many aptitude tests load heavily on g, a single composite score forecasts performance across a wide range of tasks.
What is range restriction, and why does it matter?
Range restriction is the compression of score variance that occurs when validity is computed on an already-selected group, which deflates the observed correlation. Correcting for it raises the estimate, but a 2022 reanalysis argued these corrections had been overstated, revising the validity of ability tests downward (Sackett et al., 2022).
Is a single general score enough, or do specific abilities add anything?
This is an open debate. The dominant view held that general mental ability carries most of the predictive weight, but recent work argues that narrow abilities add incremental validity beyond g (Kell & Lang, 2017; Lang & Kell, 2020), reopening a question that goes back to Thurstone's primary abilities.
What is aptitude-treatment interaction?
It is the finding, developed by Cronbach and Snow, that the best instructional method depends on the learner's aptitude, so that two teaching methods can reverse their effectiveness across the range of an ability (Cronbach & Snow, 1977). It reframes aptitude as readiness to benefit from a particular treatment rather than a fixed rank.
Does the same aptitude predict performance at every stage of learning?
No. Phillip Ackerman showed that general ability predicts performance most strongly early in skill acquisition, when a task is novel, while narrow perceptual-speed and psychomotor abilities take over as the task becomes practised (Ackerman, 1988). What an aptitude test predicts therefore depends on where in the learning curve the criterion is measured.
References
Ackerman, P. L. (1988). Determinants of individual differences during skill acquisition: Cognitive abilities and information processing. Journal of Experimental Psychology: General, 117(3), 288-318. https://doi.org/10.1037/0096-3445.117.3.288
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
Carroll, J. B. (1993). Human cognitive abilities: A survey of factor-analytic studies. Cambridge University Press.
Cattell, R. B. (1963). Theory of fluid and crystallized intelligence: A critical experiment. Journal of Educational Psychology, 54(1), 1-22. https://doi.org/10.1037/h0046743
Cronbach, L. J., & Snow, R. E. (1977). Aptitudes and instructional methods: A handbook for research on interactions. Irvington Publishers.
Deary, I. J., Strand, S., Smith, P., & Fernandes, C. (2007). Intelligence and educational achievement. Intelligence, 35(1), 13-21. https://doi.org/10.1016/j.intell.2006.02.001
Frey, M. C., & Detterman, D. K. (2004). Scholastic assessment or g? The relationship between the Scholastic Assessment Test and general cognitive ability. Psychological Science, 15(6), 373-378. https://doi.org/10.1111/j.0956-7976.2004.00687.x
Kell, H. J., & Lang, J. W. B. (2017). Specific abilities in the workplace: More important than g? Journal of Intelligence, 5(2), 13. https://doi.org/10.3390/jintelligence5020013
Kuncel, N. R., & Hezlett, S. A. (2007). Standardized tests predict graduate students' success. Science, 315(5815), 1080-1081. https://doi.org/10.1126/science.1136618
Lang, J. W. B., & Kell, H. J. (2020). General mental ability and specific abilities: Their relative importance for extrinsic career success. Journal of Applied Psychology, 105(9), 1047-1061. https://doi.org/10.1037/apl0000472
McGrew, K. S. (2009). CHC theory and the human cognitive abilities project: Standing on the shoulders of the giants of psychometric intelligence research. Intelligence, 37(1), 1-10. https://doi.org/10.1016/j.intell.2008.08.004
Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040-2068. https://doi.org/10.1037/apl0000994
Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262-274. https://doi.org/10.1037/0033-2909.124.2.262
Schmidt, F. L., & Hunter, J. E. (2004). General mental ability in the world of work: Occupational attainment and job performance. Journal of Personality and Social Psychology, 86(1), 162-173. https://doi.org/10.1037/0022-3514.86.1.162
Spearman, C. (1904). General intelligence, objectively determined and measured. The American Journal of Psychology, 15(2), 201-293. https://doi.org/10.2307/1412107
Taylor, H. C., & Russell, J. T. (1939). The relationship of validity coefficients to the practical effectiveness of tests in selection: Discussion and tables. Journal of Applied Psychology, 23(5), 565-578. https://doi.org/10.1037/h0057079
Thurstone, L. L. (1938). Primary mental abilities (Psychometric Monographs No. 1). University of Chicago Press.
Zwick, R. (2019). Assessment in American higher education: The role of admissions tests. The ANNALS of the American Academy of Political and Social Science, 683(1), 130-148. https://doi.org/10.1177/0002716219843469