Abstract

A psychological test, which MeSH classifies under behavioral disciplines and activities, is a standardized procedure for sampling behaviour to measure an unobservable attribute such as ability, personality, or symptom severity. This article treats the test as a scientific object rather than a catalogue of instruments: what turns a sample of behaviour into a measurement, and the three properties that decide whether the resulting number can be trusted. It follows the field from Binet's first practical scale and Spearman's general intelligence, through Cronbach and Meehl's account of reliability and construct validity, to the modern argument-based and item-response views that govern test construction. Particular attention is given to the reframing of validity as a property of score interpretations rather than of tests. Three interactive demonstrations model reliability, criterion validity, and the standardization of scores against a norm.

Keywords: reliability, validity, standardization

A psychological test is an instrument that turns a controlled sample of behaviour into a number standing for an attribute that cannot be observed directly. A person cannot be inspected for intelligence, conscientiousness, or depression the way a patient can be inspected for a fracture; the attribute is a construct, inferred from performance on tasks or answers to items chosen because they are supposed to depend on it. The whole discipline of test theory exists to answer one question about that inference: under what conditions is the number a trustworthy stand-in for the attribute? The answer has three parts, developed over a century and now codified in the professional standards the field holds itself to (AERA et al., 2014). A test's score must be reliable, reproducible rather than dominated by chance; it must be valid, supporting the interpretation placed on it; and it must be standardized, administered and scored uniformly and referred to a norm that gives a raw score meaning.

Key Takeaways
  • A psychological test measures an unobservable construct by sampling behaviour under standardized conditions; its output is a trustworthy number only insofar as it is reliable, valid, and referred to a norm.
  • Reliability is the reproducibility of a score; internal consistency is commonly indexed by Cronbach's alpha, and the Spearman-Brown relation shows reliability rises with test length.
  • Validity is not a property of a test but of a specific interpretation of its scores, unified by Cronbach and Meehl and Messick around the concept of construct validity and later recast as an argument to be defended.
  • Standardization refers a raw score to a norm group, converting it into a percentile or a standard score such as the deviation IQ, so that an individual is measured against a comparison population.
  • Recent measurement critiques argue that many psychological instruments are adopted without evidence that they measure their intended construct, reopening validity as an unfinished empirical problem.

What a Psychological Test Is

The defining move of a psychological test is indirect measurement. Because the target attribute is latent, the test presents a set of standardized stimuli, items, tasks, or questions, and treats the person's responses as observable indicators of the unobservable construct. This only works if the responses are collected the same way for everyone: the same items, the same instructions, the same time limits, the same scoring rules. Uniform administration is what makes two people's scores comparable, and it is the reason a test is a measurement rather than an interview. The classical account decomposes any observed score into a true score, the value that would be obtained on average over infinitely many independent administrations, plus an error term, and every property of a good test is a statement about how large that error is and what the true score means.

Three properties govern the quality of the number, and they are logically ordered: a score cannot be valid for any purpose if it is not first reliable, and a reliable, valid score is still uninterpretable for an individual without a norm. Table 1 sets out the three, each with the question it answers and the way it is estimated. The rest of the article takes them in turn, because the history of testing is essentially the progressive clarification of these three ideas, and because the modern controversies in the field are disagreements about what the second of them, validity, actually requires (Borsboom et al., 2004).

Table 1. The three properties that make a test score trustworthy.
Property Question it answers How it is estimated
ReliabilityWould the same person get the same score again?Internal consistency (alpha), test-retest correlation, parallel forms.
ValidityDoes the score support the interpretation placed on it?Construct evidence, criterion correlations, a validity argument.
StandardizationWhat does this raw score mean for this person?Norms from a reference sample; percentiles and standard scores.

The ordering is not merely pedagogical. Reliability caps validity: the correlation between a test and any criterion cannot exceed the square root of the product of their reliabilities, so an unreliable test is validity-limited before any question of meaning arises. This is why the field's earliest technical work was on reliability, and why an estimate of it is the first thing a test manual reports.

Types of Psychological Tests

MeSH indexes the psychological test under seven narrower descriptors, grouped by what a test is for rather than by any single measurement principle; the parent category is behavioral disciplines and activities. These subtypes are not mutually exclusive, and the classification is an indexing convenience for retrieval, not a theory of test kinds: a single instrument can be an aptitude test, be scored as a rating scale, and be evaluated by the methods of psychometrics at once. Table 2 lists the direct subtypes; only those with a dedicated article are linked.

Table 2. Direct subtypes of Psychological Tests in the MeSH classification (tree F04.711).
Subtype In brief
Aptitude TestsMeasure the capacity to acquire a skill, forecasting future performance rather than present attainment.
Behavior Rating ScaleQuantify the frequency or severity of observed behaviour through structured self- or informant report.
Ecological Momentary AssessmentSample experience and behaviour repeatedly in real time in the person's natural environment.
Neuropsychological TestsPerformance tasks that localize and quantify cognitive deficits linked to brain function.
Patient Health QuestionnaireA brief self-report screen for common mental disorders, best known as the depression module PHQ-9.
Personality TestsAssess stable traits and dispositions through self-report inventories or projective methods.
PsychometricsThe measurement theory itself, the reliability, validity, and scaling methods every test is built on and judged by.

These categories differ in how directly the behaviour is sampled, but they answer to the same three properties. The Patient Health Questionnaire is a useful illustration: its nine-item depression module, the PHQ-9, is short and freely distributed, yet its reliability and criterion validity against structured diagnostic interviews were documented before it entered routine use, which is what licenses its clinical interpretation (Kroenke et al., 2001).

Origins: From Binet to the Standardized Battery

The practical psychological test begins in 1905, when Alfred Binet and Théodore Simon, commissioned to identify Parisian schoolchildren needing special instruction, built a scale of graded tasks ordered by the age at which a typical child could pass them. Their innovation was the mental age: a child's score was the age level of the hardest tasks they could complete, and comparing it to chronological age gave a single index of developmental standing. This was the first instrument to measure cognitive ability by graded standardized tasks rather than by physical proxies, and it established the template every ability test still follows. When the ratio of mental to chronological age was multiplied by 100, it became the intelligence quotient, and Binet's scale, exported and revised, became the Stanford-Binet.

Alongside the applied work ran a theoretical discovery that gave testing its central idea. In 1904 Charles Spearman observed that a person's performances across unrelated cognitive tasks were all positively correlated, a pattern he explained by positing a single general factor, g, common to every task, plus factors specific to each (Spearman, 1904). To extract g he invented factor analysis, and in doing so gave test theory its foundational premise: that observed scores are surface indicators of underlying latent variables. Every later notion of a construct, and every factor model of personality or ability, descends from this move. David Wechsler then reshaped the ability test for adults, replacing the age-ratio quotient with the deviation IQ, a standard score fixing the population mean at 100 and the standard deviation at 15, and splitting the scale into verbal and performance indices; the Wechsler batteries remain the dominant clinical intelligence tests. The biological reality of the individual differences these tests capture is still an active research question, now pursued with neuroimaging and genetics (Deary et al., 2010).

Reliability

Reliability is the reproducibility of a score, the degree to which it reflects the true score rather than momentary error. It is estimated in several ways, each isolating a different source of error: test-retest reliability correlates the same test given twice and captures stability over time; parallel-forms reliability correlates two equivalent versions; and internal consistency asks whether the items of a single administration agree with one another, as they should if they measure one thing. The dominant internal-consistency index is coefficient alpha, introduced by Lee Cronbach in 1951 as a generalization of earlier split-half formulas (Cronbach, 1951). Alpha rises with two quantities: the average correlation among items and the number of items. The second dependence is the Spearman-Brown relation, which states that lengthening a test with equivalent items increases its reliability in a predictable way, and it is the practical lever a test builder pulls to reach a target reliability.

The first demonstration makes that lever manipulable. Holding constant the average inter-item correlation, it shows how the standardized alpha of a test climbs as items are added, and how the same increase shrinks the standard error of measurement on a fixed score scale, so that the abstract coefficient becomes a concrete gain in the precision of an individual's score.

Reliability explorer: test length, item correlation, and precision

0.00.20.40.60.70.80.91.0number of items (k)reliability (alpha)

Current test: alpha = 0.804. On a scale with SD = 15, the standard error of measurement is 6.6 points, so a 95% confidence interval spans ±13.0 points around an observed score. Adding items (moving right along the curve) raises alpha and shrinks the error; the gold lines mark alpha = 0.80 and 0.90.

Alpha's ubiquity has made its limits important. It is not, as often assumed, a measure of unidimensionality: a test can be multidimensional and still return a high alpha, and a set of items can be perfectly unidimensional yet return a modest alpha if the items are weakly correlated. Sijtsma showed formally that alpha is a lower bound to reliability whose gap from the true value is generally unknown, so that a reported alpha is neither necessary nor sufficient evidence that a scale is good (Sijtsma, 2009). The methodological response has been to prefer model-based coefficients such as omega, computed from a fitted factor model rather than assumed equal item weights, which recover reliability under weaker and more realistic conditions (McNeish, 2018). The practical upshot is that reliability remains foundational but its default estimator is no longer taken on trust.

Validity

Validity is the more demanding property, and the history of testing is in large part the history of clarifying what it means. The early view treated validity as a coefficient, the correlation of a test with a criterion it was meant to predict, and a test could have as many validities as it had criteria. A single scatterplot captures this criterion validity: the higher the correlation between a predictor score and the outcome it forecasts, the more accurately a cut score on the test sorts people who will succeed from those who will not. The utility of selection testing rests on exactly this relationship, and decades of accumulated evidence establish that well-constructed ability and work-sample tests predict job performance substantially (Schmidt & Hunter, 1998). Figure 1 contrasts reliability and validity as two independent properties of a test, using the familiar image of shots at a target: a test can be consistent without being accurate.

Figure 1

Reliability and Validity as Independent Properties

Three target boards showing that a test can be reliable without being valid Three dartboard-style targets, each with concentric rings and a central bullseye. The left target, labelled Neither reliable nor valid, has shots scattered widely across the whole board. The middle target, labelled Reliable but not valid, has shots tightly clustered together but far from the bullseye, in the upper-left ring. The right target, labelled Reliable and valid, has shots tightly clustered on the bullseye at the centre. The figure shows that consistency of shots is reliability, while hitting the centre is validity, and the two are independent. Neither scattered and off-centre Reliable, not valid consistent but off-centre Reliable and valid consistent and on-centre

Note. Consistency of the shots is reliability; hitting the bullseye is validity. The middle target is the cautionary case, a highly reliable score that measures the wrong thing. Positions are illustrative.

The decisive conceptual advance came in 1955, when Cronbach and Meehl argued that criterion correlations are not enough: when a test measures a theoretical construct with no single definitive criterion, its validity rests on the whole network of lawful relations the construct is expected to enter, the nomological network (Cronbach & Meehl, 1955). This made construct validity the master concept, subsuming the older content and criterion types. Samuel Messick carried the unification to its conclusion, arguing that validity is a single, integrated judgment about the adequacy of interpretations and actions based on scores, and that it must include the social consequences of test use (Messick, 1995). On this view the recurring phrase a valid test is a category error: validity is a property of an interpretation of a score for a purpose, never of the instrument itself.

The second demonstration makes criterion validity tangible. It draws a cloud of applicants whose predictor scores correlate with a later outcome at a validity coefficient the reader sets, imposes a selection cut score and a success threshold, and reports the proportion of selected applicants who go on to succeed, so that a bare correlation becomes the practical accuracy of a decision.

Criterion validity: from a correlation to the accuracy of a decision

success thresholdcutpredictor score →criterion outcome →
Base success rate54%
Selected who succeed76%
Selected (of 240)119
Improvement+21 pts

Teal points are selected successes, orange are selected failures, grey are not selected. As the validity coefficient rises, selecting above the cut yields a higher success rate than the base rate.

With a validity coefficient of 0.50 and a cut at z = 0.0, selecting applicants above the cut raises the success rate from 54% at the base rate to 76% among those selected. A zero coefficient leaves the two equal: prediction adds nothing.

Two later developments sharpen the modern picture. Michael Kane recast Messick's integrated judgment as an explicit argument: to validate a test is to state the chain of inferences from score to conclusion and marshal evidence for each link, so that validation becomes a structured, falsifiable case rather than an open-ended accumulation of correlations (Kane, 2013). And a strand of theoretical work has pushed back on the whole interpretive tradition, arguing that validity has a simpler ontological core: a test is valid when the attribute it targets exists and causally produces variation in the scores (Borsboom et al., 2004). Contemporary treatments fold both strands into practical guidance on constructing measures whose validity is built in rather than argued after the fact (Clark & Watson, 2019).

Standardization and Norms

A reliable, valid score is still meaningless for an individual until it is referred to a comparison. Answering 40 of 60 items correctly is neither good nor bad in itself; it acquires meaning only against the distribution of scores in a defined norm group. Standardization has two faces. The first is procedural: every examinee receives the same items, instructions, and scoring, so that differences in scores reflect the person rather than the testing conditions. The second is normative: the raw score is transformed by reference to a representative standardization sample into a percentile rank, the proportion of the norm group scoring below the person, or into a standard score that expresses the raw score as a distance from the norm mean in standard-deviation units. Wechsler's deviation IQ is the paradigm case, placing the mean at 100 and the standard deviation at 15 so that any score can be read directly as a standing in the population.

The third demonstration shows this transformation. It plots the normal norm distribution and lets the reader move a raw score across it, reading off the corresponding standard score and percentile rank, so that the arbitrary raw scale is converted into the population-referenced number a test report actually gives.

Standardization: a raw score becomes a percentile against a norm

557085100115130145standard score

A score of 115 lies 1.00 standard deviations from the norm mean, placing the person at the 84th percentile: about 84% of the norm group scores at or below this value. The same raw performance would carry a different percentile against a different norm group, which is why representative norming is a condition of valid interpretation.

Norms are the quiet foundation of fair interpretation, and they are also where testing meets its social obligations. A norm group that does not represent the population a person belongs to yields a misleading standing, and a score interpreted without regard to the conditions under which the norms were collected can mislead badly. The professional standards devote much of their length to exactly these requirements, treating representative norming, uniform administration, and documented fairness as conditions of valid use rather than optional refinements (AERA et al., 2014).

Worked Example

Consider a clinician using a 20-item ability scale reported with a Cronbach's alpha of 0.80, scaled like a deviation IQ to a mean of 100 and a standard deviation of 15. Two questions arise: how precise is an individual's score, and how much would the test have to grow to make it more precise?

Precision is given by the standard error of measurement, SEM = SD × √(1 − α) = 15 × √(1 − 0.80) = 15 × √0.20 = 15 × 0.447 = 6.7 points. A 95% confidence interval for a person's true score is their observed score plus or minus 1.96 × 6.7 = 13 points, so an obtained score of 100 carries a true-score interval of roughly 87 to 113. That band is wide enough to matter clinically, which motivates the second question.

To raise reliability from 0.80 to 0.90, the Spearman-Brown formula gives the required lengthening factor k = [α′(1 − α)] / [α(1 − α′)] = [0.90 × 0.20] / [0.80 × 0.10] = 0.18 / 0.08 = 2.25. The 20-item test must grow to 20 × 2.25 = 45 equivalent items. The pay-off is a tighter score: at α = 0.90 the SEM falls to 15 × √0.10 = 4.7 points and the 95% interval narrows to about ±9 points, an interval of roughly 91 to 109 around a score of 100. The same items that raise alpha from 0.80 to 0.90 correspond, at a mean inter-item correlation of about 0.17, to exactly this move from 20 to 45 items, which is the relationship the first demonstration lets the reader trace directly. The example shows why reliability is reported before anything else: it fixes, in the score's own units, how much of an individual's number to believe.

Discussion

The psychological test is one of the more consequential technologies psychology has produced, and its intellectual history is a slow tightening of standards. The trajectory from Spearman's factor to Kane's validity argument is a movement from treating the test as the object of evaluation to treating the interpretation of its scores as the object, with the instrument merely the occasion for an inference that must be justified. This is a genuine advance, but it has an uncomfortable corollary: if validity is a property of interpretations rather than instruments, then no test is ever validated once and for all, and a score valid for one decision may be invalid for another. The field's own standards embrace this, which is why they read less like a specification and more like an ethics of inference (AERA et al., 2014).

Two unresolved tensions temper any confidence. The first is that much of applied psychology uses instruments whose validity was never seriously established. A measurement-practices critique has documented how often scales are adopted, modified, and reported without evidence that they measure the intended construct, a pattern of questionable measurement practices that quietly undermines the literatures built on them (Flake & Fried, 2020). The second is deeper and concerns the constructs themselves: if the theoretical entities tests are meant to measure are weakly specified, then no amount of psychometric sophistication can secure validity, because there is no well-defined target to hit. This is one face of the broader theory crisis in psychology, the argument that the discipline's difficulty is less a shortage of data than a shortage of theories precise enough to be measured against (Eronen & Bringmann, 2021). Testing's technical apparatus is mature; what it measures is, for many constructs, still under-defined.

Current Directions

Three developments are reshaping how tests are built and judged. The first is the shift from classical test theory to item response theory, which models the probability of a given response as a function of the person's latent trait level and the properties of the individual item, rather than treating the total score as the unit of analysis. IRT lets item difficulty and discrimination be estimated on a common scale, supports adaptive testing that selects items to a person's ability in real time, and connects a test's measurement properties to the cognitive processes its items demand (Embretson & Reise, 2000). The second is the reform of reliability practice already underway, replacing the reflexive report of alpha with model-based coefficients and explicit dimensionality analysis (McNeish, 2018).

The third, and the most consequential, is the measurement-reform movement itself. Prompted by the replication crisis, methodologists have argued that unexamined measurement is a hidden source of irreproducibility, and have pressed for construct validation to be treated as a precondition of research rather than an afterthought, with pre-registered measurement plans and transparent scale documentation (Clark & Watson, 2019; Flake & Fried, 2020). Running beneath all three is a renewed engagement with the philosophical question of what a psychological attribute is, and whether the constructs psychology measures are the kind of thing that can bear the causal interpretation validity theory now asks of them (Borsboom et al., 2004; Eronen & Bringmann, 2021). The instruments are more sophisticated than ever; the open frontier is the theory the instruments are supposed to be measuring.

Key Researchers

Alfred Binet (1857-1911). French psychologist at the Sorbonne; with Théodore Simon built the 1905 Binet-Simon scale, the first practical intelligence test, introducing the mental-age metric and the measurement of ability by graded standardized tasks. Wikipedia

Denny Borsboom. Professor of psychology at the University of Amsterdam; reframed validity theory around the claim that a test is valid when its target attribute causally produces variation in scores, and pioneered the network approach to psychological constructs. ORCID - Faculty Page - Wikipedia

Lee J. Cronbach (1916-2001). Psychometrician at Stanford University; gave the field coefficient alpha in 1951 and, with Paul Meehl, the construct-validity framework in 1955, two of the most-cited ideas in measurement. Wikipedia

Ian J. Deary. Differential psychologist at the University of Edinburgh; his Lothian Birth Cohort studies use decades-old intelligence-test scores to establish the stability, biological correlates, and life outcomes of measured cognitive ability. ORCID - Faculty Page - Wikipedia

Susan E. Embretson. Psychometrician at the Georgia Institute of Technology; a leading item-response-theory methodologist who developed cognitive-psychometric models and automatic item generation linking a test's measurement properties to the mental processes its items demand. ORCID - Faculty Page

Kurt Kroenke. Professor of medicine at Indiana University and the Regenstrief Institute; developed the PHQ-9 and related brief self-report measures, showing how a short, freely available standardized scale can carry documented reliability and validity into routine clinical screening. ORCID - Faculty Page

Paul E. Meehl (1920-2003). Clinical psychologist at the University of Minnesota; co-authored the 1955 construct-validity framework and formalized the nomological network, the idea that a test's meaning is fixed by its lawful relations to other constructs. Wikipedia

Samuel Messick (1931-1998). Psychometrician at Educational Testing Service; unified the older validity trinity into a single construct-validity framework and added consequential validity, the argument that a test's social consequences are part of its validation. Wikipedia

Charles Spearman (1863-1945). Psychologist at University College London; founded factor analysis and the theory of general intelligence (g) in 1904, giving test theory its premise that observed scores are indicators of underlying latent factors. Wikipedia

Robert J. Sternberg. Professor of human development at Cornell University; triarchic theorist of intelligence and a persistent critic of conventional ability tests, arguing that standard measures capture analytic ability while neglecting creative and practical intelligence. ORCID - Faculty Page - Wikipedia

David Wechsler (1896-1981). Clinical psychologist at Bellevue Psychiatric Hospital; designed the Wechsler adult and child intelligence scales, replacing the mental-age quotient with the deviation IQ and separate verbal and performance indices, still the dominant clinical intelligence batteries. Wikipedia

Glossary

Construct validity.
The degree to which a test measures the theoretical attribute it is intended to, judged from the whole network of relations the construct is expected to enter; since 1955 the master concept subsuming other validity types.
Criterion validity.
The extent to which a test score correlates with an external outcome it is meant to predict or estimate, such as later job or academic performance.
Cronbach's alpha.
The most widely used index of internal-consistency reliability, rising with the average inter-item correlation and the number of items; formally a lower bound to a test's reliability.
Deviation IQ.
A standard score for intelligence expressing a person's standing relative to a norm group, conventionally with a mean of 100 and a standard deviation of 15, introduced by Wechsler in place of the mental-age quotient.
Factor analysis.
A statistical method, originated by Spearman, that explains the correlations among many observed variables by a smaller number of latent factors; the technical basis of construct measurement.
General intelligence (g).
The single common factor Spearman inferred from the positive correlations among all cognitive tasks; the most replicated finding in the study of mental ability.
Internal consistency.
A form of reliability estimated from a single administration, indexing how far the items of a test agree with one another as they should if they measure a common attribute.
Item response theory.
A modern measurement framework modelling the probability of a response as a function of a person's latent trait level and item properties such as difficulty and discrimination, rather than of the total score.
Norm.
The distribution of scores in a representative reference sample against which an individual raw score is compared to give it meaning.
Percentile rank.
The percentage of a norm group scoring at or below a given raw score; the most directly interpretable norm-referenced expression of an individual's standing.
Psychometrics.
The branch of psychology concerned with the theory and technique of measurement, including the reliability, validity, and scaling methods by which every test is constructed and evaluated.
Reliability.
The reproducibility of a test score, the degree to which it reflects a stable true score rather than measurement error; it sets an upper limit on validity.
Standard error of measurement.
The standard deviation of the errors around a true score, equal to the score standard deviation times the square root of one minus the reliability; it fixes the width of a confidence interval for an individual's score.
Standardization.
The dual requirement that a test be administered and scored uniformly and that its raw scores be referred to norms from a representative sample, so that individuals can be compared on a common scale.
Test-retest reliability.
The reliability estimated by correlating the scores of the same people on two separate administrations of a test, capturing the stability of the score over time.
Validity.
The degree to which evidence and theory support the interpretations and uses placed on a test's scores; a property of an interpretation for a purpose rather than of the instrument itself.

Frequently Asked Questions

What is a psychological test?
It is a standardized procedure that samples a person's behaviour, through items, tasks, or questions administered uniformly, in order to measure an unobservable psychological attribute such as ability, personality, or symptom severity. Its output is a trustworthy number only insofar as it is reliable, valid, and referred to a norm (AERA et al., 2014).

What is the difference between reliability and validity?
Reliability is the consistency of a score, whether the same person would score the same again; validity is whether the score supports the interpretation placed on it. A test can be highly reliable yet invalid, measuring something consistently other than what is claimed, so reliability is necessary but not sufficient for validity (Borsboom et al., 2004).

What is Cronbach's alpha?
Alpha is the most common index of internal-consistency reliability, increasing with the average correlation among a test's items and with the number of items (Cronbach, 1951). It is a lower bound to reliability rather than a measure of unidimensionality, and model-based coefficients are increasingly preferred (Sijtsma, 2009).

Can a test itself be valid?
Strictly, no. Modern validity theory holds that validity is a property of a specific interpretation and use of scores, not of the instrument, so the same test can yield valid inferences for one purpose and invalid ones for another (Messick, 1995; Kane, 2013).

What is standardization?
Standardization has two parts: administering and scoring a test identically for everyone, and referring the raw score to norms from a representative reference sample. Together they convert a raw score into a percentile or standard score that expresses an individual's standing in a comparison population (AERA et al., 2014).

What is general intelligence, or g?
It is the single common factor Charles Spearman inferred in 1904 from the observation that performances on all cognitive tasks are positively correlated (Spearman, 1904). Its biological basis remains an active research question pursued with neuroimaging and genetics (Deary et al., 2010).

How is item response theory different from classical test theory?
Classical theory analyses the total test score; item response theory models each response as a function of the person's latent trait level and the item's properties, placing persons and items on a common scale and enabling adaptive testing that tailors items to ability (Embretson & Reise, 2000).

Why do critics say psychology has a measurement problem?
Because many instruments are adopted and modified without evidence that they measure their intended construct, a pattern of questionable measurement practices that weakens the findings built on them and contributes to the replication crisis (Flake & Fried, 2020; Eronen & Bringmann, 2021).

References

American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.

Borsboom, D., Mellenbergh, G. J., & van Heerden, J. (2004). The concept of validity. Psychological Review, 111(4), 1061-1071. https://doi.org/10.1037/0033-295X.111.4.1061

Clark, L. A., & Watson, D. (2019). Constructing validity: New developments in creating objective measuring instruments. Psychological Assessment, 31(12), 1412-1427. https://doi.org/10.1037/pas0000626

Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297-334. https://doi.org/10.1007/BF02310555

Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302. https://doi.org/10.1037/h0040957

Deary, I. J., Penke, L., & Johnson, W. (2010). The neuroscience of human intelligence differences. Nature Reviews Neuroscience, 11(3), 201-211. https://doi.org/10.1038/nrn2793

Embretson, S. E., & Reise, S. P. (2000). Item response theory for psychologists. Lawrence Erlbaum Associates.

Eronen, M. I., & Bringmann, L. F. (2021). The theory crisis in psychology: How to move forward. Perspectives on Psychological Science, 16(4), 779-788. https://doi.org/10.1177/1745691620970586

Flake, J. K., & Fried, E. I. (2020). Measurement schmeasurement: Questionable measurement practices and how to avoid them. Advances in Methods and Practices in Psychological Science, 3(4), 456-465. https://doi.org/10.1177/2515245920952393

Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1-73. https://doi.org/10.1111/jedm.12000

Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606-613. https://doi.org/10.1046/j.1525-1497.2001.016009606.x

McNeish, D. (2018). Thanks coefficient alpha, we'll take it from here. Psychological Methods, 23(3), 412-433. https://doi.org/10.1037/met0000144

Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749. https://doi.org/10.1037/0003-066X.50.9.741

Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262-274. https://doi.org/10.1037/0033-2909.124.2.262

Sijtsma, K. (2009). On the use, the misuse, and the very limited usefulness of Cronbach's alpha. Psychometrika, 74(1), 107-120. https://doi.org/10.1007/s11336-008-9101-0

Spearman, C. (1904). General intelligence, objectively determined and measured. The American Journal of Psychology, 15(2), 201-293. https://doi.org/10.2307/1412107