Abstract

A personality inventory is a type of personality test: a structured self-report questionnaire whose items are answered on fixed scales and summed into scores on one or more trait dimensions. This article treats the inventory as measurement engineering rather than trait theory: how its items are selected, how it defends against the distortions self-report invites, and how a raw score is made interpretable. It sets out the three traditions of scale construction, rational, empirical criterion-keyed, and factor-analytic, and the different guarantees each buys. It examines the validity scales and the K-correction that adjust for socially desirable responding, the norm-referenced T-score metric on which a raw total is read, and the shift from discrete clinical types to continuous dimensional traits. Three demonstrations model criterion keying, the correction of response bias, and the conversion of a raw score into a normed T-score.

Keywords: personality inventory, criterion keying, response bias, validity scales, test norms

A personality inventory is a standardized questionnaire in which a person rates a fixed set of statements about themselves, and the answers are scored against one or more trait scales. It is the dominant form of the self-report personality test, and MeSH classifies it as one narrower class within that parent category. What makes the inventory a distinct object of study is not the trait theory it serves, which it shares with every other personality measure, but the machinery it interposes between a respondent's answers and a usable score. An inventory must decide which items to include out of the indefinitely many that could be written, it must protect the resulting score against a respondent who answers carelessly or self-flatteringly, and it must convert a raw total, meaningless on its own, into a number that locates the person against a reference population. These three problems, item selection, distortion control, and norming, define the inventory as an engineering achievement, and they are the subject of this article (Hopwood & Donnellan, 2010).

Key Takeaways
  • A personality inventory is a structured self-report questionnaire whose items are answered on fixed scales and summed into trait scores; it is the dominant class of self-report personality test.
  • Inventories are built by one of three methods, rational, empirical criterion-keying, or factor-analytic, each of which buys a different guarantee about what the resulting scale measures.
  • Because self-report invites distortion, inventories carry validity scales that detect careless, defensive, or socially desirable responding, and some apply a correction that adjusts trait scores for the detected bias.
  • A raw score is meaningless until it is referred to a norm; the standard metric is the T-score, with a population mean of 50 and standard deviation of 10, on which a conventional clinical cutoff sits at 65.
  • The history of the inventory runs from discrete clinical types toward continuous dimensional traits, a shift that now reorganizes both normal-range and clinical personality measurement.

What a Personality Inventory Is

A personality inventory operationalizes a trait as a sum of item responses. The respondent reads a series of statements, such as I make plans and stick to them or I am easily upset, and rates each on a fixed scale, dichotomous in the older instruments and typically a five-point agreement scale in modern ones. Items belonging to the same scale are summed, after reverse-scoring those worded in the opposite direction, and the total is the raw scale score. This is the defining structure of the inventory and the source of both its strengths and its vulnerabilities: it is efficient, fully standardized, and objectively scored, requiring no clinical judgment to administer, but it takes the respondent's self-description at face value and inherits whatever inaccuracy that description carries.

What the raw score means is a separate question from how it is computed, and it is the question that governs an inventory's worth. A scale score is a claim about a construct, an unobservable disposition posited to explain regularities in behavior, and the evidence that the score measures that construct rather than something else is the burden of construct validity, the framework Cronbach and Meehl established for reasoning about what a test score is entitled to claim (Cronbach & Meehl, 1955). For an inventory this burden is specific: the internal structure of the instrument, which items cluster into which scales and how strongly, must match the trait structure the scale purports to measure, and evaluating that correspondence is a distinct methodological problem that inventory construction has repeatedly had to confront (Hopwood & Donnellan, 2010). The remainder of this article follows the raw score through the three stages that make it interpretable: how its items were chosen, how it is defended against distortion, and how it is normed.

Figure 1
The inventory as a measurement pipeline.
The four stages that turn item responses into an interpretable inventory score. A left-to-right pipeline: an item pool is filtered by a construction method into a keyed scale; raw responses are summed; validity scales apply a correction for response bias; and the corrected raw score is referred to a norm to produce a T-score read against a clinical cutoff. Item pool construction method keys items Raw score sum of keyed, reverse-scored items Validity check bias detected; K-correction applied T-score referred to a norm; read at cutoff 65 From item responses to an interpretable score selection · scoring · distortion control · norming

Note. A construction method keys items from a pool; keyed responses are summed and reverse-scored into a raw total; validity scales detect response bias and a K-style correction adjusts the total; and the corrected raw score is referred to a normative distribution to yield a T-score, read against the conventional clinical cutoff of 65. Original schematic.

Types of Personality Inventory

Within the medical subject headings, Personality Inventory is itself the parent of five narrower classes of instrument (Table 1), each a named inventory or family that the indexing literature catalogues under the general heading. As with any MeSH classification, this is an indexing scheme for retrieving the assessment literature rather than a theory that partitions personality measurement at its natural joints; the categories are not mutually exclusive, and they mix historically important single instruments with whole method families. None of the five is yet a separate article on this site, so they are listed for orientation rather than linked; the descriptor above them all, Personality Tests, is the live parent to which the inventory belongs.

Table 1. Direct subtypes of Personality Inventory in the MeSH classification (tree F04.711.647.513).
Subtype In brief
Cattell Personality Factor Questionnaire The Sixteen Personality Factor Questionnaire (16PF), the prototypical factor-analytic inventory, measuring sixteen primary trait factors distilled from the personality lexicon.
Manifest Anxiety Scale A scale of items selected to measure chronic anxiety as a stable trait, historically drawn from the MMPI item pool.
Millon Clinical Multiaxial Inventory A clinical inventory aligning its scales with a theory of personality disorders and their relation to clinical syndromes.
MMPI The Minnesota Multiphasic Personality Inventory, the foundational empirically keyed clinical inventory and the source of the validity-scale tradition.
Test Anxiety Scale A focused inventory measuring the disposition to experience anxiety in evaluative testing situations.

The five subtypes already display the tension the rest of this article develops. The MMPI and the Millon inventory were built for clinical description and organized around discrete diagnostic groups; the 16PF was built by factor analysis to recover the dimensions of normal personality; the two anxiety scales isolate a single trait from a larger pool. These are not merely different instruments but products of different construction philosophies, and the philosophy an inventory follows determines what its scores can be trusted to mean.

Methods of Scale Construction

An inventory begins with a problem of selection: of the effectively unlimited statements that could be written about a person, which belong on a scale? Three strategies have answered this question, and they differ in what they take as the authority for including an item.

The rational or content method includes an item because its meaning, on inspection, matches the construct: a conscientiousness scale includes I keep my belongings neat and clean because the statement is manifestly about orderliness. This method is transparent and makes the scale's content valid by construction, but it rests entirely on the constructor's theory of the trait and offers no defense if that theory is wrong, and it is maximally vulnerable to distortion because the relevance that makes an item easy to write also makes its desirable answer easy to guess.

The empirical or criterion-keyed method abandons item content as the authority and substitutes an external criterion. An item earns its place on a scale if, and only if, members of a criterion group, patients with a particular diagnosis, say, answer it differently from a reference group, regardless of what the item appears to be about. This was the radical innovation of the Minnesota Multiphasic Personality Inventory, whose clinical scales were assembled by administering a large item pool to diagnostic groups and to normative visitors, and retaining for each scale the items that statistically separated the target group from the norm (Hathaway & McKinley, 1940). The method's strength is that it makes no assumption about why an item works, only that it does, so an item with no obvious connection to the construct is kept if it discriminates; its cost is that the resulting scale is heterogeneous in content and hard to interpret, and its keying is only as good as the criterion groups on which it was built. The first demonstration reproduces this logic, letting the reader set the endorsement rates of candidate items in a criterion and a reference group and see which items a discrimination threshold would key onto the scale.

The factor-analytic method takes the internal structure of responses as its authority. A large item set is administered to a broad sample, the correlations among items are factor-analyzed, and scales are formed from the items that load on each recovered factor, so that the scales correspond to the empirical dimensions along which people actually vary. Cattell's Sixteen Personality Factor Questionnaire was the founding instrument of this tradition, and the tradition culminated in the inventories built to measure the Five-Factor Model, whose factor structure was shown to replicate across instruments, observers, and cultures (McCrae & Costa, 1987). The factor-analytic inventory buys internal coherence, each scale is homogeneous and its items intercorrelate by design, and it grounds the choice of how many traits to measure in the data rather than in theory (McCrae & John, 1992). That the number is empirical rather than fixed is shown by the persistence of the question: lexical factor analyses in several languages support a six-factor solution, the HEXACO model, that adds an honesty-humility dimension to the familiar five, so even the dimensionality an inventory should target remains under active revision (Ashton & Lee, 2007). Its limitation is the mirror of criterion keying's: coherence is not the same as external validity, and a set of internally clean factors is still only a hypothesis about what predicts behavior until it is tested against outcomes.

Demonstration 1

Empirical criterion keying

ItemEndorsement % (◆ criterion / ◇ reference)I often feel that life is not…Δ54I am troubled by frequent hea…Δ24I sometimes tease animals.Δ32I feel sad most of the time.Δ46I enjoy reading about history.Δ3I believe I am being followed.Δ27I like to visit new places.Δ-5My sleep is fitful and distur…Δ26
Discrimination threshold (criterion − reference)Δ ≥ 25 pts
The scale keeps 5 of 8 candidate items. An item survives only if the criterion group endorses it at least 25 points more often than the reference group. Notice that a plausible-sounding item can be dropped for weak discrimination while an odd-sounding one is kept for strong discrimination: the empirical method keys on what works, not on what an item appears to mean.
Note. Each candidate item shows the percentage endorsing it in a criterion group (navy) and a reference group (gray). An item is keyed onto the scale (gold check) only when the endorsement difference meets the discrimination threshold, independent of the item's apparent content. Raising the threshold builds a shorter, more sharply discriminating scale. Endorsement rates are illustrative and fixed. Original interactive schematic.

Response Bias and Validity Scales

Every inventory that trusts a person to describe themselves is exposed to the ways that description can go wrong, and the inventory tradition is distinguished from other personality measures by the formal machinery it builds to detect and correct those failures. The distortions fall into two broad kinds. Acquiescence is a content-independent tendency to agree, or to disagree, with items regardless of what they assert, and it is countered structurally by balancing each scale with reverse-keyed items so that indiscriminate agreement cancels out. Socially desirable responding is the subtler and more studied threat: the tendency to answer in the direction that presents the self favorably. Paulhus showed that this is not one disposition but two separable components, an unconscious self-deceptive enhancement, an honestly held but inflated self-view, and a deliberate impression management directed at an audience, and that the two dissociate under instructions and situations in ways a single social-desirability score conflates (Paulhus, 1984). The distinction matters for an inventory because the two components call for different responses: impression management can be detected and screened, while self-deceptive enhancement is bound up with genuine self-regard and cannot be simply subtracted out.

The Minnesota inventory institutionalized the detection of these threats through dedicated validity scales administered alongside the clinical scales. Its L scale catches the naive respondent claiming implausible virtue, its F scale catches deviant or careless responding by counting rarely endorsed items, and its K scale measures a more sophisticated defensiveness, a guardedness that suppresses the endorsement of genuine problems. The critical insight was that a validity scale need not merely flag a distorted profile but can be used to correct the substantive scores. Meehl and Hathaway demonstrated that K functions as a statistical suppressor variable: adding a fraction of a respondent's K score to certain clinical scales removes defensiveness-related variance that otherwise masks real elevation, sharpening the discrimination between clinical and normal profiles (Meehl & Hathaway, 1946). This K-correction, applied as a fixed fractional weight that differs by scale, was among the first formal attempts to model and remove a response style rather than merely note its presence. The second demonstration makes the mechanism explicit, letting the reader distort a true profile through defensiveness and then apply a K-style correction to recover it.

Demonstration 2

Response bias and the K correction

30405060708090T = 65SmDPtSc◦ true ● observed + K correction
Defensiveness (raw K score)K = 12
With defensiveness at K = 12, a guarded respondent under-reports genuine elevations, pulling the observed scores toward the mean. The K correction adds a fixed fraction of K back onto each scale, according to its weight (0 for Depression, 0.5 for Somatic complaints, 1.0 for Psychasthenia and Schizophrenia), moving the observed markers back toward the true standing.
Note. Each row is a clinical scale plotted on the T-score metric. The hollow marker is the person's true standing; defensiveness suppresses the observed score toward the mean (leftward). Applying a K-style suppressor correction adds a fixed fraction of the detected defensiveness back onto each scale, per its K weight, moving the filled marker back toward truth. The dashed line is the T = 65 clinical cutoff. Illustrative values; the correction models Meehl and Hathaway's use of K as a suppressor variable. Original interactive schematic.

Norms and T-Score Interpretation

A raw scale score, the sum of keyed item responses, carries no meaning in isolation: a raw depression score of 24 is high or low only relative to how people in general score. The inventory therefore refers every raw score to a norm, the distribution of scores in a defined reference population, and reports not the raw total but its standing within that distribution. The standard metric for this is the T-score, a linear transformation of the raw score onto a scale with a mean of 50 and a standard deviation of 10, so that a T of 50 is exactly average, a T of 60 is one standard deviation above the mean, and a T of 70 is two. The transformation is $T = 50 + 10 \times (X - M)/SD$, where $X$ is the raw score and $M$ and $SD$ are the normative mean and standard deviation. On this metric the conventional threshold for clinical significance sits at a T of 65, roughly the 93rd percentile, above which a score is treated as elevated. The third demonstration renders this conversion, letting the reader move a raw score across the normative distribution and read off the corresponding z-score, T-score, and percentile against the clinical cutoff.

A normed score is only as trustworthy as the scale it summarizes is precise, and the precision of an inventory scale is indexed by its reliability, the proportion of score variance that is not measurement error. For the homogeneous scales that factor-analytic and rational construction produce, reliability is usually estimated by internal consistency, the average intercorrelation among a scale's items expressed as coefficient alpha (Cronbach, 1951). Reliability bounds interpretation directly: it fixes the standard error of measurement, and hence the width of the confidence band that must be placed around any observed T-score, as the worked example below makes concrete.

The T-score makes two different scales commensurable, so that a profile of elevations across many scales can be read at a glance, but it also imports the reference population into every interpretation. A T-score is a statement about standing in a particular norm group, and a score referred to the wrong norm is as misleading as a mismeasured raw total; the same raw answers yield different T-scores against a clinical and a general-population norm. This dependence is why norming is not a clerical step appended to construction but part of the measurement itself, and why an inventory's manual devotes as much attention to the composition of its normative sample as to the keying of its items. The interpretive gain, a single comparable metric on which a whole profile can be read, is bought with a permanent commitment to the population that defines it.

Demonstration 3

Raw score to normed T-score

T = 6520304050607080T-score metric
Raw score (norm: M = 16, SD = 4)X = 24
Raw 24 = z = 2.00 = T = 70.0, the 97.7th percentile. This is at or above the conventional T = 65 clinical cutoff and would be read as an elevation.
Note. A raw scale score is referred to the normative distribution (mean 16, SD 4 here) and converted to a T-score of mean 50 and SD 10. The shaded area is the percentile — the share of the reference population scoring at or below the value. The dashed line marks the T = 65 clinical cutoff. Move the raw score to see how the same total maps to a standing that only the norm makes meaningful. Original interactive schematic.

From Clinical Types to Dimensional Traits

The inventory was born clinical and typological. The Minnesota inventory's scales were built to separate diagnostic groups, and its logic was categorical: a profile placed a person nearer to or farther from a clinical type. The great shift of the following decades was from this categorical, type-based conception toward a dimensional one, in which a person is located at a continuous point on a small number of trait axes rather than sorted into a kind. The inventories built to measure the Five-Factor Model carried this dimensional conception into normal-range assessment, and their proponents argued explicitly that the same continuous trait framework that describes normal personality also serves clinical assessment better than discrete categories (Costa & McCrae, 1992).

The argument was pressed hardest against the classification of personality disorders, where the categorical model had long strained against evidence of continuity and comorbidity between supposedly distinct types. A body of work assembling the alternatives converged on the conclusion that personality pathology is better represented as extreme standing on the same dimensions that describe normal personality than as a set of discrete disorders (Widiger & Simonsen, 2005). This program produced its own inventories: a maladaptive-trait model and measure was constructed for the fifth edition of the diagnostic manual, defining pathological personality as elevation on a set of trait dimensions and building an inventory, the PID-5, to assess them (Krueger et al., 2012). The dimensional program has since been generalized beyond personality into a proposed reorganization of psychopathology as a whole, the Hierarchical Taxonomy of Psychopathology, which arranges disorders themselves into a dimensional hierarchy recovered from the empirical covariation of symptoms, and which treats the trait inventory as a primary instrument of measurement (Kotov et al., 2017). The inventory that began by sorting patients into clinical types now supplies the continuous coordinates of a dimensional model of mental disorder.

Worked Example

Consider an examinee whose raw score on a depression scale is 24, where the normative sample has a mean of 16 and a standard deviation of 4. The standardized score is $(24 - 16)/4 = 2.0$, two standard deviations above the mean. Converting to the T metric gives $T = 50 + 10 \times 2.0 = 70$, a score at roughly the 98th percentile and well above the conventional clinical cutoff of 65. The raw total of 24, uninterpretable on its own, becomes a clear statement of standing once referred to the norm.

The precision of that statement is bounded by the scale's reliability. With an internal-consistency reliability of 0.84, and the T-score standard deviation of 10 fixed by construction, the standard error of measurement is $10 \times \sqrt{1 - 0.84} = 10 \times \sqrt{0.16} = 10 \times 0.4 = 4$ T-points. A 95 percent confidence interval around the observed T of 70 is therefore roughly $70 \pm 1.96 \times 4$, about $\pm 8$ points, spanning 62 to 78, a reminder that even a clearly elevated score is an estimate with a band of uncertainty rather than a point.

Now consider how a validity-scale correction reshapes interpretation. Suppose a second examinee scores a raw 10 on a somatic-complaints scale whose norm has a mean of 12 and standard deviation of 4, giving an uncorrected $z = (10 - 12)/4 = -0.5$ and $T = 50 - 5 = 45$, an unremarkable, slightly below-average score. But this examinee also scores 12 on the K defensiveness scale, and the scale carries a K-correction weight of 0.5, so the corrected raw score is $10 + 0.5 \times 12 = 10 + 6 = 16$. Recomputing, $z = (16 - 12)/4 = 1.0$ and $T = 50 + 10 = 60$. The correction has moved the score from below average to the edge of elevation, because the examinee's high defensiveness implies that a raw score of 10 understated genuine complaints that guardedness had suppressed. The 15-point swing from 45 to 60 is the quantitative expression of Meehl and Hathaway's insight that a response style is not merely noise to be flagged but variance that can be modeled and removed (Meehl & Hathaway, 1946).

Discussion

The personality inventory is best understood not as a theory of personality but as a solution to three engineering problems that any self-report measure must solve. The construction methods answer the first, which items to trust, and the three traditions represent genuinely different bets: criterion keying trusts an external criterion and accepts uninterpretable scales in exchange for demonstrated discrimination, while factor analysis trusts the internal structure of responses and accepts that coherence is not yet validity (Hathaway & McKinley, 1940; McCrae & Costa, 1987). The validity-scale apparatus answers the second, how to defend a score against its own respondent, and the K-correction remains the clearest case of an inventory not merely flagging distortion but modeling and subtracting it (Meehl & Hathaway, 1946; Paulhus, 1984). The norm answers the third, how to make a raw total mean something, and the T-score's convenience is inseparable from its dependence on the population that defines it.

That inventory scores are stable enough to be worth measuring, yet not fixed, is itself an empirical finding rather than an assumption. Longitudinal evidence shows that trait levels change in orderly, population-wide ways across the life course even as the rank ordering of individuals is largely preserved, which is what licenses an inventory to treat a trait as an enduring quantity while acknowledging that the quantity is not immutable (Roberts et al., 2006). The deeper unresolved question is the one the construction methods first posed and the dimensional program has now inherited: at what level, and against which criterion, an inventory's scales should be validated, given that internal coherence, external discrimination, and predictive utility do not always point to the same instrument (Hopwood & Donnellan, 2010).

Current Directions

The most active current work treats the inventory less as a fixed instrument than as one realization of an underlying trait space, and asks how the many inventories in circulation relate to that space and to one another. A systematic evaluation of widely used trait scales found that most, though not all, could be located within the Five-Factor space, a result that both clarifies the reach of the dominant framework and identifies the constructs that escape it, and that bears directly on whether scores from different inventories can be treated as measuring the same thing (Bainbridge et al., 2022). In parallel, hierarchical inventories such as the Big Five Inventory-2 have refined the internal architecture of measurement, separating each broad domain into distinct facets so that the same responses can be reported at whichever level of breadth a use requires, and improving both the bandwidth and the fidelity of the resulting profile (Soto & John, 2017). The dimensional reorganization of clinical assessment continues to press inventories toward serving as the measurement layer of an empirically derived taxonomy rather than as freestanding tests, so that the questions now driving the field are less about any single instrument than about the structure the whole family of instruments is converging on (Roberts & Yoon, 2022).

Common Misconceptions

The items of a personality inventory are chosen because they obviously measure the trait.
Only rationally constructed scales work that way. The empirical criterion-keying method that built the MMPI retains an item purely because it statistically discriminates a criterion group from a reference group, regardless of whether its content appears related to the construct, so a keyed scale can be deliberately heterogeneous in content (Hathaway & McKinley, 1940).
A validity scale only flags a distorted profile for discarding.
A validity scale can do more than flag distortion; it can correct for it. The MMPI's K scale is used as a statistical suppressor, with a fraction of its score added to certain clinical scales to remove defensiveness-related variance and sharpen the separation of clinical from normal profiles (Meehl & Hathaway, 1946).
A personality inventory measures a fixed trait that does not change over a person's life.
Trait scores are stable enough to rank people consistently, but they are not immutable. A meta-analysis of longitudinal studies found orderly, population-wide mean-level change in traits across the life course, so an inventory measures an enduring quantity that still moves rather than a fixed one, and the same raw total is meaningful only once referred to the appropriate norm (Roberts et al., 2006).

Glossary

Acquiescence bias.
A content-independent tendency to agree, or disagree, with items regardless of what they assert; countered by balancing a scale with reverse-keyed items so the tendency cancels.
Construct validity.
The evidence that a scale measures the unobservable construct it claims to, rather than something else; the framework Cronbach and Meehl established for reasoning about what a test score is entitled to assert.
Criterion keying.
The empirical construction method that includes an item on a scale only if it statistically discriminates a criterion group from a reference group, independent of the item's apparent content.
Dimensional model.
A conception that locates a person at a continuous point on trait axes rather than sorting them into discrete types; increasingly the basis of both normal-range and clinical inventories.
Factor-analytic method.
A construction method that forms scales from the items loading on factors recovered by factor-analyzing item intercorrelations, so scales match the empirical dimensions of variation.
Impression management.
The deliberate component of socially desirable responding, in which answers are consciously shaded for an audience; distinguished by Paulhus from unconscious self-deceptive enhancement.
K-correction.
The MMPI procedure of adding a fixed fraction of a respondent's K defensiveness score to certain clinical scales, using K as a suppressor to remove response-style variance.
Norm.
The distribution of scores in a defined reference population against which a raw score is referred; the composition of the norm sample determines what a normed score means.
Personality inventory.
A structured self-report questionnaire whose items are answered on fixed scales and summed into scores on one or more trait dimensions; the dominant class of self-report personality test.
Rational method.
A construction method that includes an item because its content, on inspection, matches the target construct; transparent but dependent on the constructor's theory and vulnerable to distortion.
Reliability.
The proportion of a scale's score variance that is not measurement error; for a homogeneous scale it is commonly estimated by internal consistency (coefficient alpha) and it fixes the standard error of measurement bounding a T-score.
Reverse-keyed item.
An item worded so that agreement counts against the trait, scored in the opposite direction before summing; balancing a scale with reverse-keyed items cancels a respondent's content-independent tendency to agree.
Self-deceptive enhancement.
The unconscious component of socially desirable responding, an honestly held but inflated self-view; distinguished by Paulhus from deliberate impression management and not simply subtractable from a score.
Socially desirable responding.
The tendency to answer inventory items in the direction that presents the self favorably; shown to comprise separable self-deceptive and impression-management components.
Suppressor variable.
A measure that carries little valid variance of its own but improves a scale when partialled in, because it removes irrelevant variance from it; the MMPI's K scale acts as one when its score is added to certain clinical scales.
T-score.
A normed score linearly transformed to a mean of 50 and standard deviation of 10, on which a person's standing across scales is read and a clinical cutoff conventionally sits at 65.
Validity scale.
A scale embedded in an inventory to detect a response style, such as careless, defensive, or overly virtuous answering, rather than to measure a substantive trait.

Key Researchers

Raymond B. Cattell (1905-1998). Research professor at the University of Illinois; he applied factor analysis to trait structure and built the Sixteen Personality Factor Questionnaire, the founding factor-analytic inventory. Wikipedia - Wikidata

Paul T. Costa (b. 1942). Scientist Emeritus at the National Institute on Aging; with Robert McCrae he built the NEO Personality Inventory and argued for a dimensional, trait-based approach to clinical personality assessment. ORCID - Wikipedia - Wikidata

Starke R. Hathaway (1903-1984). Psychologist at the University of Minnesota; with J. C. McKinley he built the Minnesota Multiphasic Personality Inventory and established the empirical criterion-keying method of scale construction. Wikipedia - Wikidata

Roman Kotov (b. 1977). Professor of Psychiatry at Stony Brook University; he leads the Hierarchical Taxonomy of Psychopathology, the dimensional model that reorganizes clinical assessment around empirically derived trait structure. ORCID - Faculty Page - Google Scholar

Robert R. McCrae (b. 1949). Independent scientist, formerly of the National Institute on Aging; with Paul Costa he established the cross-instrument and cross-observer robustness of the five-factor trait structure that grounds modern factor-analytic inventories. ORCID - Wikipedia - Wikidata

J. Charnley McKinley (1891-1950). Neuropsychiatrist at the University of Minnesota; co-author of the MMPI, he contributed the clinical framing of its original diagnostic scales. Wikipedia - Wikidata

Paul E. Meehl (1920-2003). Regents' Professor of Psychology at the University of Minnesota; he introduced the K suppressor correction for the MMPI and, with Cronbach, formalized construct validity for psychological tests. Wikipedia - Wikidata

Christopher J. Soto (b. 1981). Professor of Psychology at Colby College; he led the development of the Big Five Inventory-2, a hierarchical inventory that separates each broad domain into distinct facets. ORCID - Faculty Page - Google Scholar

Frequently Asked Questions

What is a personality inventory?
A personality inventory is a structured self-report questionnaire in which a person rates a fixed set of statements about themselves, and the answers are summed into scores on one or more trait scales; it is the dominant class of self-report personality test (Hopwood & Donnellan, 2010).

How are the items of a personality inventory chosen?
By one of three methods: the rational method includes items whose content matches the trait, the empirical or criterion-keying method includes items that discriminate a criterion group from a reference group, and the factor-analytic method forms scales from items that load on recovered factors (Hathaway & McKinley, 1940).

What is criterion keying?
Criterion keying is the empirical construction method, pioneered in the MMPI, that retains an item on a scale only if members of a criterion group answer it differently from a reference group, regardless of whether the item's content appears related to the construct (Hathaway & McKinley, 1940).

What is a validity scale?
A validity scale is a set of items embedded in an inventory to detect a response style rather than a substantive trait, such as the MMPI's L, F, and K scales, which catch naive virtue-claiming, careless or deviant responding, and sophisticated defensiveness (Meehl & Hathaway, 1946).

What is the K-correction?
The K-correction is the MMPI procedure of adding a fixed fraction of a respondent's K defensiveness score to certain clinical scales, using K as a statistical suppressor to remove defensiveness-related variance and sharpen the separation of clinical from normal profiles (Meehl & Hathaway, 1946).

What is a T-score?
A T-score is a normed score obtained by transforming a raw score to a scale with a mean of 50 and standard deviation of 10, so a T of 60 is one standard deviation above average; a conventional clinical cutoff sits at a T of 65.

What is the difference between social desirability's two components?
Paulhus showed that socially desirable responding comprises self-deceptive enhancement, an honestly held but inflated self-view, and impression management, a deliberate shading of answers for an audience, and that the two dissociate under different instructions and situations (Paulhus, 1984).

How have personality inventories changed over time?
They have moved from discrete clinical types toward continuous dimensional traits, a shift that carried the inventory from the categorical MMPI scales to the trait dimensions of the Five-Factor inventories and the dimensional models now reorganizing psychopathology (Kotov et al., 2017).

References

Ashton, M. C., & Lee, K. (2007). Empirical, theoretical, and practical advantages of the HEXACO model of personality structure. Personality and Social Psychology Review, 11(2), 150-166. https://doi.org/10.1177/1088868306294907

Bainbridge, T. F., Ludeke, S. G., & Smillie, L. D. (2022). Evaluating the Big Five as an organizing framework for commonly used psychological trait scales. Journal of Personality and Social Psychology, 122(4), 749-777. https://doi.org/10.1037/pspp0000395

Costa, P. T., & McCrae, R. R. (1992). Normal personality assessment in clinical practice: The NEO Personality Inventory. Psychological Assessment, 4(1), 5-13. https://doi.org/10.1037/1040-3590.4.1.5

Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297-334. https://doi.org/10.1007/BF02310555

Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302. https://doi.org/10.1037/h0040957

Hathaway, S. R., & McKinley, J. C. (1940). A multiphasic personality schedule (Minnesota): I. Construction of the schedule. The Journal of Psychology, 10(2), 249-254. https://doi.org/10.1080/00223980.1940.9917000

Hopwood, C. J., & Donnellan, M. B. (2010). How should the internal structure of personality inventories be evaluated? Personality and Social Psychology Review, 14(3), 332-346. https://doi.org/10.1177/1088868310361240

Kotov, R., Krueger, R. F., Watson, D., Achenbach, T. M., Althoff, R. R., Bagby, R. M., . . . Zimmerman, M. (2017). The Hierarchical Taxonomy of Psychopathology (HiTOP): A dimensional alternative to traditional nosologies. Journal of Abnormal Psychology, 126(4), 454-477. https://doi.org/10.1037/abn0000258

Krueger, R. F., Derringer, J., Markon, K. E., Watson, D., & Skodol, A. E. (2012). Initial construction of a maladaptive personality trait model and inventory for DSM-5. Psychological Medicine, 42(9), 1879-1890. https://doi.org/10.1017/S0033291711002674

McCrae, R. R., & Costa, P. T. (1987). Validation of the five-factor model of personality across instruments and observers. Journal of Personality and Social Psychology, 52(1), 81-90. https://doi.org/10.1037/0022-3514.52.1.81

McCrae, R. R., & John, O. P. (1992). An introduction to the five-factor model and its applications. Journal of Personality, 60(2), 175-215. https://doi.org/10.1111/j.1467-6494.1992.tb00970.x

Meehl, P. E., & Hathaway, S. R. (1946). The K factor as a suppressor variable in the Minnesota Multiphasic Personality Inventory. Journal of Applied Psychology, 30(5), 525-564. https://doi.org/10.1037/h0053634

Paulhus, D. L. (1984). Two-component models of socially desirable responding. Journal of Personality and Social Psychology, 46(3), 598-609. https://doi.org/10.1037/0022-3514.46.3.598

Roberts, B. W., Walton, K. E., & Viechtbauer, W. (2006). Patterns of mean-level change in personality traits across the life course: A meta-analysis of longitudinal studies. Psychological Bulletin, 132(1), 1-25. https://doi.org/10.1037/0033-2909.132.1.1

Roberts, B. W., & Yoon, H. J. (2022). Personality psychology. Annual Review of Psychology, 73, 489-516. https://doi.org/10.1146/annurev-psych-020821-114927

Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology, 113(1), 117-143. https://doi.org/10.1037/pspp0000096

Widiger, T. A., & Simonsen, E. (2005). Alternative dimensional models of personality disorder: Finding a common ground. Journal of Personality Disorders, 19(2), 110-130. https://doi.org/10.1521/pedi.19.2.110.62628