Abstract

Psychometrics is the science of measuring psychological attributes: the theory and technique of building tests, quantifying the traits and abilities they estimate, and establishing that the numbers mean what they claim. This article treats it as the methodological core of psychological testing, tracing its two great theories of the test score: classical test theory, which splits an observed score into a true score and error and defines reliability as the share of variance that is not error, and item response theory, which models each response as a function of a latent trait and the item's properties. Around these sit the twin questions of reliability, from coefficient alpha to model-based omega, and validity, the century-long argument over what evidence licenses a test's interpretation. Three demonstrations model score precision, an item characteristic curve, and the attenuation of correlations by measurement error.

Keywords: psychometrics, reliability, validity, item response theory, factor analysis

Psychometrics is the branch of psychology concerned with the theory and technique of psychological measurement. Where a substantive theory asks what attention or memory is, psychometrics asks a prior and more general question: what it means to assign a number to a psychological attribute at all, and what has to be true of a test before the number it yields can be trusted. Its object is not any single trait but the measuring instrument itself, the test, and the chain of inference that runs from a person's responses to a claim about the quantity of some attribute they possess. The discipline supplies the tools every psychological test depends on, the models that relate responses to traits, the coefficients that quantify a score's precision, and the standards that decide whether a measure is any good (Lord & Novick, 1968).

Key Takeaways
  • Psychometrics is the science of psychological measurement: the theory and method of constructing tests, estimating the latent attributes they target, and justifying the interpretation of their scores.
  • Classical test theory models an observed score as a true score plus random error and defines reliability as the proportion of score variance that is not error.
  • Item response theory recasts measurement at the level of the item, modeling the probability of a response as a function of a person's latent trait and the item's difficulty and discrimination.
  • Reliability is necessary but not sufficient for validity; the modern view treats validity as a property of score interpretations and uses, supported by an explicit argument, not a fixed property of the test.
  • Measurement error attenuates observed correlations, so unreliable measures systematically understate the true relationships between constructs.

What Psychometrics Is

Psychometrics occupies an unusual position in psychology: it is a science whose subject matter is other sciences' instruments. Every claim that people differ in working-memory capacity, that a therapy reduces depression, or that one applicant is better suited to a job than another rests on a measurement, and psychometrics is the discipline that studies whether those measurements are sound. Its central problem is that psychological attributes are latent: unlike length or mass, a trait such as verbal ability or neuroticism cannot be read off directly but must be inferred from behavior it is presumed to cause. A test is a device for provoking a sample of that behavior under standardized conditions, and a psychometric model is a formal statement of how the unobserved attribute relates to the observed responses.

Because the attribute is never seen, two questions haunt every measurement and organize the whole field. The first is reliability: does the test give a stable, repeatable number, or is the score largely noise? The second is validity: granting that the number is stable, is it a measure of the intended attribute, or of something else? These questions are logically ordered, reliability bounding validity, and much of the technical apparatus of psychometrics exists to answer one or the other. The answers rest, in turn, on a model of what a test score is, and the field has produced two: classical test theory, which works with the total score, and item response theory, which works with the individual item.

Classical Test Theory

Classical test theory is the foundational model of the test score, given its definitive mathematical statement by Lord and Novick, who unified decades of scattered results into a single axiomatic framework (Lord & Novick, 1968). Its premise is disarmingly simple: an observed score X is the sum of a true score T, defined as the expected value of the observed score over infinitely many independent administrations, and a random error E that is uncorrelated with the true score. Because the error is random, it contributes variance but no systematic bias, and the variance of observed scores decomposes cleanly into true-score variance and error variance (Figure 1). Reliability is then defined as the ratio of true-score variance to total observed-score variance, a number between zero and one that states what fraction of the differences between people's scores reflects real differences in the attribute rather than measurement noise.

This definition has a direct practical consequence for the individual score. If reliability is known, the typical size of the error component can be recovered as the standard error of measurement, the standard deviation multiplied by the square root of one minus the reliability. The standard error converts an abstract reliability coefficient into a confidence band around a person's score, the interval within which their true score most plausibly lies, and it is the single most useful number for cautioning against over-reading a single test administration. The first demonstration lets the reader vary a scale's reliability and standard deviation and watch the resulting error band widen and narrow, and the Worked Example computes a case by hand.

Figure 1

The Classical Test-Score Model

Observed-score variance partitioned into true-score and error variance A horizontal bar represents total observed-score variance. It is divided into a large left portion labelled true-score variance and a smaller right portion labelled error variance. Below, an equation states that an observed score equals a true score plus error, and that reliability equals true-score variance divided by total observed-score variance. The larger the true-score portion relative to the whole, the higher the reliability. X = T + E observed score = true score + error True-score variance Error variance total observed-score variance Reliability = true-score variance ÷ total variance
Note. In classical test theory an observed score is a true score plus uncorrelated random error, so observed-score variance is the sum of true-score and error variance (Lord & Novick, 1968). Reliability is the proportion of the whole that is true-score variance; the standard error of measurement is the spread of the error portion. Original schematic.

Demonstration 1

Reliability, the true score, and the standard error

Observed-score varianceTrue 84%Error 16%Observed score = 100, with 95% confidence band6080100120140
True-score varianceError variance
SEM = 15 × √(1 − 0.84) = 6.00 points. An observed score of 100 carries a 95% confidence interval of about 88.2111.8 11.8).
A test score is part signal and part noise. Reliability is the share of score variance that is true-score variance; the rest is measurement error, whose spread — the standard error of measurement — sets the confidence band around any single observed score. Raising reliability shrinks the error share and narrows the band.

Reliability: From Alpha to Omega

The reliability coefficient is defined but not directly observable, because true scores are not observed, so psychometrics estimates it from a single administration through the internal consistency of a test's items. The overwhelmingly dominant estimator is Cronbach's coefficient alpha, which expresses reliability as a function of the number of items and their average intercorrelation and rises with both (Cronbach, 1951). Alpha's ubiquity is a mixed blessing. It is a lower bound on reliability that equals the true value only when the items are strictly parallel, an assumption real tests almost never meet, so a reported alpha is routinely misinterpreted as if it were reliability itself, or worse, as an index of unidimensionality, which it is not (Sijtsma, 2009).

Internal consistency is only one face of reliability, however. Because an observed score can also fluctuate across occasions, across raters, and across alternate forms, a single undifferentiated error term conceals which of these is doing the damage. Generalizability theory extends the classical model to decompose error into several distinct sources at once, treating each as a facet of the measurement design and estimating how dependably a score generalizes across each, so that a study can ask separately how much of the error is item sampling and how much is rater disagreement (Cronbach, Gleser, Nanda, & Rajaratnam, 1972). It is the framework that turns the vague notion of reliability across items, occasions, or raters into a set of separately estimable variance components.

The modern response has been to move past alpha toward model-based estimators. Coefficient omega derives reliability from an explicit factor-analytic model of the items rather than from alpha's restrictive parallel-test assumptions, and so estimates the quantity alpha was always meant to approximate without alpha's known downward bias (McNeish, 2018). A practical tutorial literature now walks researchers from alpha to omega, showing that the model-based coefficient is both more defensible and, with modern software, no harder to compute (Revelle & Condon, 2019). The shift matters because reliability is not merely a summary statistic: it caps the correlations a measure can enter into, which is where measurement error does its most insidious damage.

Demonstration 3

How measurement error hides a real relationship

0.000.250.500.751.00True correlation0.60Observed correlation0.43
r(observed) = 0.60 × √(0.81 × 0.64) = 0.43. The true relationship explains 36% of variance, but the attenuated one appears to explain only 19%. Dividing the observed value by 0.72 recovers the disattenuated 0.60.
Because error variance cannot correlate with anything, the correlation you observe between two measures is the true correlation between their constructs shrunk by the square root of the product of their reliabilities. Two strongly related constructs can look weakly related when measured with unreliable instruments; Spearman's correction runs the shrinkage backward.

The demonstration above makes the cost of unreliability visible. Because error variance cannot correlate with anything, the observed correlation between two measures is the true correlation between their constructs multiplied by the square root of the product of their reliabilities, an effect known as attenuation, first derived by Spearman (Spearman, 1904). Two constructs that are strongly related in reality will appear only weakly related if measured with unreliable instruments, and Spearman's correction for attenuation runs the formula backward to estimate what the correlation would be if both measures were perfectly reliable. The correction is a reminder that a disappointing empirical correlation may indict the measures rather than the theory.

Validity

Reliability guarantees a stable number; validity asks whether the number means what it is taken to mean, and it has been the most contested concept in the field. The modern framework began when Cronbach and Meehl introduced construct validity, arguing that a test measuring an unobservable attribute is validated not by any single criterion but by accumulating evidence that its scores behave as the surrounding theoretical network predicts (Cronbach & Meehl, 1955). Campbell and Fiske then supplied the era's most influential method, the multitrait-multimethod matrix, which demands that measures of the same construct by different methods converge while measures of different constructs diverge, separating the trait a test measures from the method it uses to measure it (Campbell & Fiske, 1959).

Messick unified the various types of validity, content, criterion, and construct, into a single overarching conception in which all validity is construct validity, and controversially folded the consequences of test use into the validity judgment, so that how a score is used bears on whether its interpretation is warranted (Messick, 1995). Kane recast this framework in a form practitioners could act on, treating validation as the evaluation of an explicit interpretive argument, a chain of inferences from response to score to interpretation to use, each link requiring its own evidence (Kane, 2013). Against this consensus that validity is a property of interpretations, Borsboom and colleagues pressed a dissenting realist claim: a test is valid, they argued, simply if the attribute it measures exists and causally produces variation in the test scores, relocating validity from the meaning of scores to the ontology of the attribute (Borsboom et al., 2004).

Factor Analysis and Latent Structure

Beneath both reliability and validity lies a question of structure: how many distinct attributes a set of items actually measures, and how they relate. Factor analysis, the statistical engine of psychometrics, answers it by explaining the correlations among many observed variables through a smaller number of latent factors. The technique grew directly out of Spearman's discovery that positive correlations among diverse cognitive tests could be explained by a single general factor, and it was Thurstone who generalized the method into multiple factor analysis, allowing several correlated factors rather than one and introducing the idea of rotation to simple structure, the search for a factor solution in which each variable loads clearly on as few factors as possible (Thurstone, 1931).

Simple structure is what makes a factor solution interpretable, and the choice of rotation remains a live methodological decision rather than a mechanical step. Factor analysis is now the tool a test constructor uses to decide whether a scale is unidimensional enough to sum into a single score, to locate the facets nested within a broad domain, and to estimate the model-based reliability coefficients that superseded alpha. It is also the historical bridge to item response theory: once the latent structure of a test is expressed as a formal measurement model, attention shifts naturally from the total score to the behavior of each item.

Item Response Theory

Item response theory rebuilds measurement from the item up. Instead of modeling the total score, it models the probability that a person will give a particular response to a particular item as a function of the person's location on a latent trait continuum and the item's own properties (Embretson & Reise, 2000). For a dichotomous item, this relationship is the item characteristic curve, an S-shaped logistic function whose difficulty parameter fixes the trait level at which a correct or endorsing response becomes more likely than not, and whose discrimination parameter sets how sharply the item distinguishes people just above and below that point. The second demonstration lets the reader manipulate both parameters and watch the curve shift and steepen.

This item-level formulation solves problems classical test theory cannot. Because person and item are placed on a common scale, a person's estimated trait level does not depend on which particular items they answered, and an item's properties do not depend on which particular sample took it, a property called invariance that underlies modern computerized adaptive testing. Item response theory also yields an information function that shows precisely where along the trait continuum a test measures well or poorly, replacing the single reliability coefficient of classical theory with a precision that varies by trait level. The framework has proved especially valuable in clinical and personality measurement, where bifactor models clarify when a set of items measuring several related facets can nonetheless be scored as a single dimension (Reise & Waller, 2009).

The two theories are best understood side by side (Table 1). They are not rivals to be adjudicated so much as tools suited to different problems: classical test theory remains the simpler and more robust choice for reporting a scale's overall precision, while item response theory earns its greater complexity wherever items must be compared, banked, or administered adaptively.

Table 1. Two theories of the test score compared.
Aspect Classical test theory Item response theory
Unit of analysis The total test score. The individual item response.
Central parameters Reliability and the standard error of measurement of the whole test. Each item's difficulty and discrimination, and each person's trait level.
Precision One reliability coefficient, treated as constant across the score range. An information function that varies with trait level, so precision is highest where items are targeted.
Sample and test dependence Item and person statistics depend on the particular sample and set of items. Person and item parameters are invariant across samples and item sets when the model fits.
Characteristic uses Reporting scale reliability, short fixed-form tests, everyday score interpretation. Computerized adaptive testing, item banking, test equating, clinical scale construction.

Demonstration 2

The item characteristic curve

0.000.250.500.751.00-3-2-10123b = 0.0latent trait θ →
At θ = 1.0, the probability of a keyed response is P = 0.82. Item information here is a²·P·(1−P) = 0.34, greatest when θ is near the difficulty b and the item is highly discriminating.
Item response theory models each item, not the total score. The probability of endorsing an item rises with the person's latent trait along an S-shaped curve. Difficulty (b) slides the curve along the trait axis — where it crosses 50 percent — and discrimination (a) sets how steeply the item separates people just above and below that point.

Worked Example

Suppose a cognitive-ability scale has a reliability of 0.84 and, in the population of interest, a standard deviation of 15 points. The standard error of measurement is the standard deviation times the square root of one minus the reliability: 15 times the square root of 1 minus 0.84, which is 15 times the square root of 0.16, or 15 times 0.4, giving 6.0 points. A person who scores 100 therefore carries a 95 percent confidence interval of roughly plus or minus 1.96 times 6.0, about 11.76 points, from a little under 88 to a little over 112. The width of that band, nearly a quarter of the score's standard deviation, is the practical meaning of a reliability that most researchers would regard as good, and it is a standing caution against treating a single observed score as an exact reading of the trait.

Now consider what that same imperfect reliability does to a correlation. Suppose the true correlation between this ability and a job-performance measure is 0.60, but performance is measured with a reliability of only 0.64 while the ability test's reliability is 0.81. The observed correlation is attenuated to the true value times the square root of the product of the two reliabilities: 0.60 times the square root of 0.81 times 0.64, which is 0.60 times the square root of 0.5184, or 0.60 times 0.72, giving 0.43. A genuine relationship of 0.60 presents as 0.43 purely because the instruments are unreliable. Running the correction backward, an observed 0.43 divided by 0.72 recovers the disattenuated 0.60. The attenuation demonstration recomputes this relationship as the reader varies the true correlation and the two reliabilities.

Discussion

Psychometrics has furnished psychology with something rare in the social sciences: a rigorous, mathematically explicit account of what its measurements are and how far they can be trusted. Classical test theory gave the field a decomposition of the score and a definition of reliability that remain in daily use a century later (Lord & Novick, 1968). Item response theory extended that account to the item and made possible adaptive tests that measure efficiently across the whole trait range (Embretson & Reise, 2000). And the long argument over validity, from construct validity through the interpretive-argument framework, forced psychologists to state precisely what a score is supposed to mean and what evidence would license that meaning (Cronbach & Meehl, 1955; Kane, 2013).

Yet the discipline's very rigor throws into relief how loosely its tools are often used in practice. A reliability coefficient designed as a lower bound is reported as though it were reliability itself; a validity concept built on accumulating evidence is satisfied with a single correlation; a factor model's assumptions are ignored in the rush to sum items into a score. The gap between the theory psychometrics has developed and the routine practice of the sciences that depend on it is the tension that animates the field's current, and notably self-critical, research front.

Current Directions

The most striking recent development is psychometrics turning its scrutiny on psychology's own measurement habits. Borsboom's widely read polemic argued that mainstream psychology had largely ignored a century of psychometric progress, continuing to sum items and report alpha while the measurement models that would justify those steps went unused (Borsboom, 2006). The replication crisis sharpened the point into an empirical program. A systematic audit of fifteen widely used self-report measures found that a large fraction failed basic tests of structural validity that their published use presupposed, hidden invalidity that ordinary reporting practices never surfaced (Hussey & Hughes, 2020). In parallel, a catalogue of questionable measurement practices, unreported scale modifications, unexamined factor structures, and validity simply assumed, argued that measurement is a neglected source of the field's unreliable findings and set out concrete standards for reporting validity evidence (Flake & Fried, 2020).

A more technical strand concerns the consequences of choosing the wrong measurement model. Fitting a reflective latent-variable model, in which the construct causes the items, to data that are actually formative, where the items constitute the construct, can distort estimated relationships more severely than random measurement error does, a result that raises the stakes on structural decisions that are often made by default (Rhemtulla et al., 2020). Together these lines mark a shift in the field's center of gravity, from developing ever more sophisticated models to insisting that the models actually be used, and that the validity of everyday measures be demonstrated rather than presumed.

Common Misconceptions

A high coefficient alpha shows that a scale measures a single thing.
Alpha is a function of the number of items and their average intercorrelation, and a long test can post a high alpha while measuring several distinct factors. Unidimensionality is a separate question that requires a factor model, not a reliability coefficient, to answer (Sijtsma, 2009).
A reliable test is therefore a valid one.
Reliability is only consistency; a scale can measure the wrong construct with great precision. Reliability bounds validity but never establishes it, which is why construct and criterion evidence must be gathered separately (Cronbach & Meehl, 1955).
Validity is a fixed property of a test.
On the dominant modern view, validity attaches to a particular interpretation and use of scores, not to the test as an object. A measure valid for one inference in one population may be invalid for another, so validation evaluates an argument about a specific use (Kane, 2013).

Glossary

Attenuation.
The reduction of an observed correlation below the true correlation between constructs, caused by measurement error in one or both measures; reversible in estimate by Spearman's correction for attenuation.
Classical test theory.
The model that treats an observed score as the sum of a true score and uncorrelated random error, and defines reliability as the ratio of true-score variance to observed-score variance.
Coefficient alpha.
The most common estimate of internal-consistency reliability, computed from the number of items and their average intercorrelation; a lower bound on reliability that equals it only for strictly parallel items.
Coefficient omega.
A model-based reliability estimate derived from a factor-analytic model of the items, avoiding the parallel-test assumptions that bias coefficient alpha downward.
Construct validity.
The degree to which a test measures the theoretical attribute it claims to, evidenced by scores behaving as the surrounding network of predicted relationships requires.
Convergent validity.
Agreement between measures of the same construct obtained by different methods; one half of the requirement imposed by a multitrait-multimethod matrix.
Criterion validity.
The extent to which a test score correlates with an external outcome it is meant to predict or to stand in for, whether measured concurrently or in the future.
Discriminant validity.
Divergence between measures of different constructs, even when obtained by the same method; the complement of convergent validity in evaluating a measure.
Factor analysis.
A statistical method that explains the correlations among many observed variables in terms of a smaller number of latent factors, used to determine how many attributes a set of items measures.
Generalizability theory.
An extension of classical test theory that partitions measurement error into several distinct sources, or facets, such as items, occasions, and raters, and estimates how dependably a score generalizes across each.
Item characteristic curve.
The S-shaped function relating a person's latent trait level to the probability of a given response to an item, governed by the item's difficulty and discrimination parameters.
Item response theory.
A family of measurement models that express the probability of an item response as a function of a latent trait and item parameters, placing persons and items on a common scale.
Latent variable.
An attribute that is not observed directly but is inferred from its presumed effects on measured responses; the object most psychometric models are built to estimate.
Multitrait-multimethod matrix.
A table of correlations among several traits each measured by several methods, used to separate convergent and discriminant validity from method variance.
Reliability.
The consistency of a test score across items, occasions, or raters, defined in classical test theory as the proportion of observed-score variance attributable to true-score variance.
Standard error of measurement.
The expected spread of an individual's observed scores around the true score, computed as the standard deviation times the square root of one minus the reliability; it sets the width of a score's confidence interval.
True score.
The expected value of a person's observed score over infinitely many independent administrations; the noise-free quantity a test aims to estimate in classical test theory.
Validity.
The degree to which evidence and theory support the intended interpretations and uses of test scores; on the modern view a property of interpretations rather than of the test itself.

Key Researchers

Denny Borsboom (b. 1973). Professor of Psychology at the University of Amsterdam; he reframed validity theory around the realist claim that a test is valid when the attribute it measures exists and causes variation in scores, and founded the network approach to psychological measurement. ORCID - Faculty Page - Google Scholar - Wikipedia - Wikidata

Lee J. Cronbach (1916-2001). Vida Jacks Professor of Education at Stanford University; he originated coefficient alpha, the standard index of internal-consistency reliability, and with Paul Meehl formalized the concept of construct validity. Wikipedia - Wikidata

Susan E. Embretson. Professor of Psychology at the Georgia Institute of Technology; she advanced item response theory for psychological research and pioneered the cognitive design of tests, in which item difficulty is modeled from the cognitive processes an item demands. ORCID - Faculty Page - Google Scholar

Jessica K. Flake. Associate Professor of Psychology at the University of British Columbia; she documented the prevalence of questionable measurement practices in psychology and set out standards for the validity evidence everyday research should report. ORCID - Faculty Page - Google Scholar

Frederic M. Lord (1912-2000). Distinguished psychometrician at the Educational Testing Service; he founded item response theory and, with Melvin Novick, wrote the canonical text that unified classical test theory with latent-trait models. Wikipedia - Wikidata

Samuel Messick (1931-1998). Distinguished psychometrician at the Educational Testing Service; he unified validity into a single construct-centred framework and added the consequences of test use as an explicit validity concern. Wikipedia - Wikidata

Steven P. Reise. Research Professor of Psychology at the University of California, Los Angeles; he extended item response theory to clinical and personality measurement, especially bifactor models that clarify when a multi-item scale can be scored as if it measured one thing. ORCID - Faculty Page - Google Scholar

William Revelle. Professor of Psychology at Northwestern University; he clarified the move from coefficient alpha to model-based omega and authored the R psych package that put those estimators into routine use. ORCID - Faculty Page - Google Scholar - Wikipedia - Wikidata

Charles Spearman (1863-1945). Grote Professor of Mind and Logic at University College London; he founded the mathematics of mental measurement, deriving the correction for attenuation and the reliability of the sum score, and originated the two-factor theory of intelligence that gave factor analysis its first subject. Wikipedia - Wikidata

L. L. Thurstone (1887-1955). Charles F. Grey Distinguished Service Professor at the University of Chicago; he developed multiple factor analysis and the concept of rotation to simple structure, and the law of comparative judgment that grounds psychological scaling. Wikipedia - Wikidata

Frequently Asked Questions

What is psychometrics?
Psychometrics is the science of psychological measurement, the theory and technique of constructing tests, estimating the latent attributes they target, and justifying the interpretation of their scores (Lord & Novick, 1968).

What is the difference between reliability and validity?
Reliability is the consistency of a score across items, occasions, or raters, while validity is whether the score measures the intended construct; reliability is necessary for validity but does not guarantee it (Cronbach & Meehl, 1955).

What is Cronbach's alpha and what are its limits?
Alpha estimates internal-consistency reliability from the number of items and their average intercorrelation, but it is only a lower bound on reliability and is not a measure of unidimensionality, so it is widely over-interpreted (Sijtsma, 2009).

How does item response theory differ from classical test theory?
Classical test theory models the total score, whereas item response theory models the probability of each response as a function of a latent trait and item parameters, placing persons and items on a common scale (Embretson & Reise, 2000).

What is the standard error of measurement?
It is the typical spread of a person's observed scores around their true score, computed as the standard deviation times the square root of one minus the reliability, and it sets the confidence interval around a score (Lord & Novick, 1968).

Why do unreliable measures weaken observed correlations?
Because measurement error cannot correlate with anything, an observed correlation is the true correlation multiplied by the square root of the product of the two measures' reliabilities, an attenuation that Spearman's correction reverses in estimate (Spearman, 1904).

Is validity a property of a test?
On the dominant modern view, validity is a property of a particular interpretation and use of scores, evaluated as an explicit argument, rather than a fixed property of the test as an object (Kane, 2013).

What are questionable measurement practices?
They are underreported or unexamined choices in how a measure is built and scored, such as silent scale modifications or assumed rather than demonstrated validity, that undermine the credibility of psychological findings (Flake & Fried, 2020).

References

Borsboom, D., Mellenbergh, G. J., & van Heerden, J. (2004). The concept of validity. Psychological Review, 111(4), 1061-1071. https://doi.org/10.1037/0033-295X.111.4.1061

Borsboom, D. (2006). The attack of the psychometricians. Psychometrika, 71(3), 425-440. https://doi.org/10.1007/s11336-006-1447-6

Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81-105. https://doi.org/10.1037/h0046016

Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297-334. https://doi.org/10.1007/BF02310555

Cronbach, L. J., Gleser, G. C., Nanda, H., & Rajaratnam, N. (1972). The dependability of behavioral measurements: Theory of generalizability for scores and profiles. John Wiley & Sons. ISBN 9780471188506.

Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302. https://doi.org/10.1037/h0040957

Embretson, S. E., & Reise, S. P. (2000). Item response theory for psychologists. Lawrence Erlbaum Associates. ISBN 0805828192.

Flake, J. K., & Fried, E. I. (2020). Measurement schmeasurement: Questionable measurement practices and how to avoid them. Advances in Methods and Practices in Psychological Science, 3(4), 456-465. https://doi.org/10.1177/2515245920952393

Hussey, I., & Hughes, S. (2020). Hidden invalidity among 15 commonly used measures in social and personality psychology. Advances in Methods and Practices in Psychological Science, 3(2), 166-184. https://doi.org/10.1177/2515245919882903

Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1-73. https://doi.org/10.1111/jedm.12000

Lord, F. M., & Novick, M. R. (1968). Statistical theories of mental test scores. Addison-Wesley. ISBN 978-0201043105.

McNeish, D. (2018). Thanks coefficient alpha, we'll take it from here. Psychological Methods, 23(3), 412-433. https://doi.org/10.1037/met0000144

Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749. https://doi.org/10.1037/0003-066X.50.9.741

Reise, S. P., & Waller, N. G. (2009). Item response theory and clinical measurement. Annual Review of Clinical Psychology, 5, 27-48. https://doi.org/10.1146/annurev.clinpsy.032408.153553

Revelle, W., & Condon, D. M. (2019). Reliability from alpha to omega: A tutorial. Psychological Assessment, 31(12), 1395-1411. https://doi.org/10.1037/pas0000754

Rhemtulla, M., van Bork, R., & Borsboom, D. (2020). Worse than measurement error: Consequences of inappropriate latent variable measurement models. Psychological Methods, 25(1), 30-45. https://doi.org/10.1037/met0000220

Sijtsma, K. (2009). On the use, the misuse, and the very limited usefulness of Cronbach's alpha. Psychometrika, 74(1), 107-120. https://doi.org/10.1007/s11336-008-9101-0

Spearman, C. (1904). The proof and measurement of association between two things. The American Journal of Psychology, 15(1), 72-101. https://doi.org/10.2307/1412159

Thurstone, L. L. (1931). Multiple factor analysis. Psychological Review, 38(5), 406-427. https://doi.org/10.1037/h0069792