Abstract
Personality assessment, which MeSH classifies under behavioral disciplines and activities, is the measurement of the enduring patterns of thought, feeling, and behavior that distinguish one person from another, catalogued as descriptor D010552. Its methods range from structured self-report inventories such as the NEO-PI-R and the Big Five Inventory to projective techniques such as the Rorschach, and every method is judged by the same psychometric standards of reliability and validity. This article traces the field from the lexical hypothesis through factor-analytic trait models to the Five-Factor Model that organizes most modern inventories, examines the person-situation critique that once threatened trait measurement, and works through the standard error of measurement that quantifies how much confidence a single test score deserves.
Keywords: personality assessment, Five-Factor Model, reliability and validity
What Personality Assessment Is
Personality assessment is the systematic measurement of the relatively stable patterns of thinking, feeling, and behaving that characterize an individual and differentiate that person from others. In MeSH the descriptor (D010552) sits under behavioral disciplines and activities, an indexing placement that groups it with the methods and practices of the behavioral sciences rather than with any single theory of personality. The object of measurement is the trait — a dimension along which people differ consistently across time and situation — and the aim is to locate a person on that dimension precisely enough to describe, predict, or make decisions about behavior.
What distinguishes assessment from casual impression is that it is psychometric: every instrument, whether a questionnaire or an inkblot, produces a score whose trustworthiness is evaluated by the same two criteria. Reliability asks whether the score is stable and internally consistent; validity asks whether it measures the construct it claims to and predicts what it should. A personality test is not credible because it feels insightful but because its scores survive these checks, a standard that separates the field's accepted instruments from the popular questionnaires that resemble them.
Personality assessment serves clinical, organizational, forensic, and research purposes, and the method chosen reflects the purpose. A clinician screening for maladaptive patterns, an employer predicting job performance, and a researcher mapping the structure of individual differences all measure personality, but they weigh the trade-offs among the available methods differently. Those methods, and the trait structure they have converged on, are the substance of the field.
Types of Personality Assessment
MeSH classifies personality assessment under the broad parent behavioral disciplines and activities and records a single narrower descriptor beneath it. The classification is an indexing scheme for the biomedical literature, not a theory of the field, so its one formal subtype should be read as the term NLM chose to index separately rather than as an exhaustive taxonomy of assessment methods; the working division of the field into self-report inventories, projective techniques, and behavioral methods, described in the next section, cuts across it.
| Subtype | In brief |
|---|---|
| Q-Sort | A ranking method in which the assessor sorts a fixed set of descriptive statements into a forced, quasi-normal distribution from least to most characteristic of the person, yielding an ipsative profile rather than a set of independent scale scores. |
Methods of Assessment
The field's methods fall into three broad families that differ in how directly they ask about personality. Self-report inventories — also called objective tests — present standardized items to which the respondent answers on a fixed scale, and the answers are scored by key rather than by clinical judgment. The dominant modern examples measure the Five-Factor Model: Costa and McCrae's Revised NEO Personality Inventory (NEO-PI-R) scores five broad domains and thirty facets from 240 items (#ref-costa-1992), while shorter instruments such as the Big Five Inventory and its revision trade facet detail for brevity (#ref-soto-2017). Their strength is objectivity and psychometric transparency; their weakness is dependence on the respondent's self-knowledge and willingness to answer candidly.
A modern inventory reports a person as five domain scores, each a T-score (mean 50, standard deviation 10) read against a norm group. Move each slider to build a profile and see how the score translates into a standing relative to the average. The dashed line marks the norm mean of 50.
The profile is illustrative: the descriptors follow the conventional T-score bands (about 45–55 average, 56 and up high), not any one publisher's norms. A real report interprets facets and validity scales alongside the five domains. Computed locally, not stored.
Projective techniques take the opposite approach, presenting ambiguous stimuli — inkblots, ambiguous scenes — on the theory that the respondent's interpretation reveals personality structure that a direct question would not elicit. The Rorschach inkblot method is the archetype. Projective methods have long been the field's most contested: because scoring depends heavily on the system used and the examiner, their psychometric credentials have been challenged repeatedly. A large meta-analysis of individual Rorschach variables found that validity is highly uneven, with some scores well supported and many others weak or unsupported, which is the empirically honest summary of the method's standing (#ref-mihura-2013).
Behavioral and observational methods infer personality from samples of actual behavior rather than from either self-report or projection — structured ratings, situational tests, and informant reports. They sidestep the candor problem but are costly and capture narrower slices of behavior. Across all three families, a landmark review of the evidence on psychological testing concluded that well-constructed tests have validity comparable to that of many accepted medical tests, and that no single method is sufficient on its own — a case for multimethod assessment rather than reliance on any one instrument (#ref-meyer-2001).
Reliability and Validity
Every assessment method is answerable to the same psychometric standards, and the more rigorous of the two is validity. Cronbach and Meehl's classic analysis established construct validity as the framework for personality measures: because a trait such as extraversion cannot be observed directly, a test's validity is established by accumulating evidence that its scores behave as the theory of the construct predicts — correlating with related measures (convergent validity), diverging from unrelated ones (discriminant validity), and predicting relevant outcomes (criterion validity) (#ref-cronbach-1955). No single correlation validates a test; validity is a cumulative argument. Campbell and Fiske gave that argument an operational form in the multitrait-multimethod matrix, which arrays several traits each measured by several methods so that convergent and discriminant validity can be read directly from the pattern of correlations — strong agreement among different methods measuring the same trait, weak correlation among measures of different traits (#ref-campbell-1959).
Reliability is the prerequisite. A score that fluctuates randomly from one administration to the next cannot measure anything stable, so instruments are evaluated for internal consistency and for stability over time. The most widely reported index of internal consistency is coefficient alpha, Cronbach's statistic summarizing how strongly a scale's items intercorrelate and, by extension, how much of a score reflects the common trait rather than item-specific noise (#ref-cronbach-1951). Reliability also sets a ceiling on validity: a test can correlate with an outcome no more strongly than the square root of the product of the two measures' reliabilities, which is why unreliable measurement caps what any instrument can predict. The practical expression of reliability is the standard error of measurement, the quantity that turns an abstract reliability coefficient into a concrete band of uncertainty around a person's score — the subject of the worked example below.
A score is a range, not a point. The standard error of measurement turns a scale's reliability into a band of uncertainty: SEM = SD × √(1 − reliability). Set the reliability, the score's standard deviation, and an observed T-score to read off the 95% confidence interval around the true score.
At the defaults (r = 0.90, SD = 10, score = 65) the SEM is 3.16 and the 95% interval is about [58.8, 71.2] — the figures worked through in the text. Lower the reliability and the band widens sharply: unreliable measurement is imprecise measurement. Computed locally, not stored.
Reliability is not fixed by the construct; it can be engineered. Adding items that tap the same trait raises a scale's internal consistency in a predictable way described by the Spearman-Brown relationship, which is why longer scales are generally more reliable than short ones and why abbreviating an inventory has a measurable psychometric cost. The demo below makes that trade-off explicit: it shows how a scale's reliability climbs as items are added and how quickly the returns diminish, the calculation a test developer runs when deciding how long an inventory needs to be.
Adding items that tap the same trait raises a scale's reliability in a predictable way. Set the reliability of a short base scale and its length, then choose a new length to read the predicted reliability off the curve. The returns diminish: each block of items buys less than the last.
The Spearman-Brown formula assumes the added items are of the same quality as the originals — an ideal a real revision only approximates. It shows why longer scales tend to be more reliable and why shortening an inventory has a measurable psychometric cost. Computed locally, not stored.
The Trait Structure Behind Assessment
What a personality inventory measures depends on a prior question: how many traits are there, and what are they? The modern answer grew out of the lexical hypothesis, Allport and Odbert's premise that the important ways people differ have become encoded in language, which led them to extract some 18,000 person-descriptive terms from an unabridged dictionary (#ref-allport-1936). The list was a catalogue, not a structure. Cattell reduced it by early factor analysis to sixteen source traits and built the 16 Personality Factor questionnaire on them, the first influential attempt to derive an inventory's scales from the covariation among trait terms rather than from theory alone (#ref-cattell-1943).
Cattell's sixteen factors did not replicate cleanly, and decades of reanalysis converged instead on five broad dimensions. Goldberg, who named them the Big Five, showed that the same five-factor structure recurred across different sets of trait adjectives and different samples (#ref-goldberg-1993), while McCrae and Costa demonstrated that the same five emerged from questionnaire items and, crucially, from both self-reports and observer ratings — evidence that the factors index something real about the person rather than an artifact of self-description (#ref-mccrae-1987). John and Srivastava's taxonomy consolidated the history, measurement, and competing conventions into the framework most inventories now assume (#ref-john-1999).
| Domain | High scorers tend to be |
|---|---|
| Openness to experience | Imaginative, curious, aesthetically sensitive, and open to unconventional ideas. |
| Conscientiousness | Organized, dependable, self-disciplined, and achievement-oriented. |
| Extraversion | Sociable, assertive, energetic, and prone to positive emotion. |
| Agreeableness | Cooperative, trusting, warm, and considerate of others. |
| Neuroticism | Prone to anxiety, emotional reactivity, and negative affect (its low pole is emotional stability). |
The Five-Factor Model matters to assessment because it gives disparate inventories a common coordinate system: a modern review found that the great majority of widely used personality scales can be located within the five-dimensional space, which lets scores from different instruments be compared and integrated (#ref-bainbridge-2022). It is the trait structure the profile demo above visualizes and the reason a score on almost any contemporary personality scale can be read as a position on one of these five axes.
Worked Example
Personality scores are usually reported as T-scores, standardized to a mean of 50 and a standard deviation of 10, so that a single scale can be interpreted against its norm group. The question a T-score raises is how much confidence a single number deserves, and the answer is given by the standard error of measurement (SEM), which converts a scale's reliability into a band of uncertainty around the observed score.
Suppose a respondent obtains a T-score of 65 on a conscientiousness scale whose reliability is r = 0.90, with the standard T-score deviation of SD = 10. The standard error of measurement is
SEM = SD × √(1 − r) = 10 × √(1 − 0.90) = 10 × √0.10 = 10 × 0.3162 = 3.16.
To place a 95% confidence interval around the observed score, multiply the SEM by 1.96 and add and subtract it:
margin = 1.96 × 3.16 = 6.20,
so the interval is 65 − 6.20 to 65 + 6.20, or about [58.8, 71.2].
The interpretation is the whole point of reporting a score psychometrically: the respondent's true conscientiousness score is not known to be exactly 65; it lies, with 95% confidence, somewhere between roughly 59 and 71. Two observations follow. First, even a highly reliable scale (r = 0.90 is good for a personality measure) leaves a band of several T-points, which is why responsible interpretation treats a score as a range rather than a point and is cautious about small differences between scales. Second, the band widens as reliability falls: at r = 0.75 the same SD gives SEM = 10 × √0.25 = 5.0 and a 95% margin of 9.8, nearly ten T-points in either direction. The reliability demo lets any reliability and observed score be entered so the resulting confidence band can be read off directly, making concrete why reliability is the precondition for trusting an assessment at all.
Discussion
Personality assessment's defining episode is the person-situation debate. In 1968 Mischel marshalled evidence that scores on personality traits predicted actual behavior only weakly — correlations rarely exceeding about 0.30 — and argued that behavior was governed more by situations than by broad, stable dispositions, a critique that questioned the very enterprise of trait measurement (#ref-mischel-1968). For a period it appeared that the field's central object might be an illusion of the observer rather than a property of the person.
The resolution reshaped the field rather than ending it. Aggregation was the key: single acts are noisy, but averaged across many occasions, behavior is predicted by traits far more strongly than Mischel's ceiling implied, and the person-by-situation interaction — how a given person behaves in a given kind of situation — proved to be a stable, measurable signature in its own right. The Five-Factor Model's cross-observer replication answered the deeper worry by showing that trait scores are not merely self-descriptions: independent observers place a person at the same location on the five axes, which is difficult to explain unless the traits track something real (#ref-mccrae-1987). Assessment emerged from the debate more careful about what a trait score claims — a probabilistic tendency aggregated over occasions, not a prediction of any single act — and that chastened understanding is the modern standard against which instruments are judged.
Current Directions
The contemporary agenda is less about discovering new dimensions than about establishing what assessment scores genuinely predict and how well those predictions replicate. The most direct expression of that agenda is the Life Outcomes of Personality Replication project, which tested whether the many published links between personality traits and consequential life outcomes hold up under pre-registered replication; a substantial majority did, but with effect sizes smaller than the original reports, a sobering calibration of how much a personality score can be expected to explain (#ref-soto-2019). The finding neither vindicates nor dismisses trait assessment; it sizes it honestly.
A parallel effort is consolidating the measurement instruments themselves. The revision of the Big Five Inventory (the BFI-2) illustrates the current design priorities: a hierarchical structure that reports both broad domains and narrower facets, improved discrimination among the scales, and careful attention to the wording effects that inflate or deflate scores (#ref-soto-2017). At the level of the field's structure, work locating the wider ecosystem of trait scales within the five-factor space is clarifying which instruments measure the same thing under different names and which capture genuinely distinct variance (#ref-bainbridge-2022). A comprehensive review of personality psychology frames the through-line: assessment is increasingly evaluated not by the elegance of its factor structure but by the size and replicability of the real-world outcomes its scores forecast, and by evidence that measured traits change in orderly ways across the lifespan (#ref-roberts-2022). The open questions are empirical — which scores predict which outcomes, for whom, and how stably — rather than taxonomic.
Common Misconceptions
- A personality test reveals a hidden true type.
- Validated personality assessment measures positions on continuous trait dimensions, not membership in discrete types. The popular type-sorting questionnaires that resemble clinical instruments generally lack the reliability and validity evidence that defines a credible measure (#ref-meyer-2001).
- A test score is an exact measurement of the trait.
- Every score carries measurement error. The standard error of measurement places a confidence band around an observed score, so a responsible interpretation treats the number as a range rather than a precise value (#ref-cronbach-1955).
- Projective tests such as the Rorschach are pseudoscience with no validity.
- The truth is more selective. Meta-analysis shows that some Rorschach scores are well supported while many others are weak, so the method's validity is variable and score-specific rather than uniformly absent or present (#ref-mihura-2013).
- Personality traits do not predict real behavior.
- This overstates Mischel's critique. Aggregated across many occasions, traits predict behavior and life outcomes reliably, though with modest effect sizes, and independent observers agree on where a person falls on the major dimensions (#ref-soto-2019; #ref-mccrae-1987).
Glossary
- Big Five.
- The five broad trait dimensions — openness, conscientiousness, extraversion, agreeableness, and neuroticism — recovered repeatedly from analyses of trait language and questionnaire items; Goldberg's name for the structure.
- Coefficient alpha.
- Cronbach's index of a scale's internal consistency, summarizing how strongly its items intercorrelate; the most commonly reported reliability statistic for personality scales.
- Construct validity.
- The cumulative evidence that a test measures the unobservable construct it claims to, assembled from convergent, discriminant, and criterion relationships; the framework Cronbach and Meehl established for personality measures.
- Convergent and discriminant validity.
- Two components of construct validity: a valid measure should correlate with other measures of the same trait (convergent) and correlate less with measures of different traits (discriminant).
- Criterion validity.
- The degree to which a test score predicts a relevant external outcome, such as job performance or a clinical diagnosis.
- Factor analysis.
- A statistical method that reduces the correlations among many trait items or terms to a smaller number of underlying dimensions; the technique that produced Cattell's sixteen factors and the Big Five.
- Five-Factor Model (FFM).
- The trait model organizing personality into five broad domains, operationalized by the NEO-PI-R; the coordinate system most modern inventories assume.
- Lexical hypothesis.
- The premise, from Allport and Odbert, that the socially important dimensions of personality have become encoded as single terms in natural language, so the structure of trait words reveals the structure of traits.
- Multitrait-multimethod matrix.
- Campbell and Fiske's design for assessing construct validity, arranging several traits each measured by several methods so that convergent and discriminant validity can be read from the correlation pattern.
- NEO-PI-R.
- The Revised NEO Personality Inventory of Costa and McCrae, a 240-item self-report measure scoring the five domains and thirty facets of the Five-Factor Model.
- Person-situation debate.
- The controversy Mischel opened in 1968 over whether behavior is governed by stable traits or by situations; resolved by aggregation, which restored the predictive power of traits across many occasions.
- Personality assessment.
- The systematic, psychometric measurement of the enduring patterns of thought, feeling, and behavior that distinguish one person from another.
- Projective technique.
- An assessment method presenting ambiguous stimuli, such as inkblots, on the theory that the respondent's interpretation discloses personality structure a direct question would not elicit.
- Q-sort.
- The single MeSH subtype of personality assessment: sorting a fixed set of descriptive statements into a forced quasi-normal distribution to yield an ipsative profile of a person.
- Reliability.
- The consistency of a measurement — its internal consistency and stability over time; a precondition for validity that also sets a ceiling on how strongly a test can predict anything.
- Self-report inventory.
- An objective personality test in which the respondent answers standardized items on a fixed scale, scored by key rather than by clinical judgment; the dominant modern method.
- Standard error of measurement (SEM).
- The standard deviation of the errors around a person's observed score, equal to SD × √(1 − reliability); it defines the confidence band within which the true score is likely to lie.
- T-score.
- A standardized score scaled to a mean of 50 and a standard deviation of 10, the metric in which most personality inventories report results against a norm group.
- Trait.
- A dimension along which people differ consistently across time and situation; the unit of measurement in personality assessment.
- Validity.
- The degree to which a test measures the construct it claims to and supports the inferences drawn from its scores; the more demanding of the two psychometric standards.
Key Researchers
Gordon W. Allport (1897-1967). With H. S. Odbert, extracted roughly 18,000 trait terms from an unabridged dictionary, founding the lexical hypothesis on which the modern trait taxonomies rest. Wikipedia - Wikidata
Raymond B. Cattell (1905-1998). Reduced the trait lexicon by factor analysis to sixteen source traits and built the 16 Personality Factor questionnaire, a founding instrument of factor-analytic assessment. Wikipedia - Wikidata
Paul T. Costa (b. 1942). With Robert McCrae, developed the Five-Factor Model and its measure, the NEO-PI-R, the most widely used self-report inventory in modern personality assessment. ORCID - Google Scholar - Wikipedia
Lewis R. Goldberg (1932-2026). Coined the term Big Five, established its factor-marker structure, and built the public-domain International Personality Item Pool that made large-scale assessment research feasible. Wikipedia - Google Scholar - Wikidata
Oliver P. John (b. 1959). Co-authored the canonical Big Five taxonomy and the Big Five Inventory, the short measure that made the five-factor structure usable across survey research. Google Scholar - Faculty - Wikipedia
Robert R. McCrae (b. 1949). With Paul Costa, developed the Five-Factor Model and the NEO-PI-R and demonstrated the model's validity across instruments and across self- and observer reports. ORCID - Google Scholar - Wikipedia
Brent W. Roberts (living). A leader in the study of personality development — how the traits assessment measures change and stabilize across the lifespan — and in the replication of personality science. ORCID - Google Scholar - Faculty
Christopher J. Soto (living). Developer of the BFI-2 and leader of the Life Outcomes of Personality Replication project testing which personality-outcome links reproduce under pre-registration. ORCID - Google Scholar - Faculty
Frequently Asked Questions
What is personality assessment? It is the systematic, psychometric measurement of the enduring patterns of thought, feeling, and behavior that distinguish one person from another. In MeSH it is descriptor D010552, filed under behavioral disciplines and activities.
What are the main methods of personality assessment? There are three broad families: self-report inventories such as the NEO-PI-R, in which a person answers standardized items; projective techniques such as the Rorschach, which use ambiguous stimuli; and behavioral or observational methods, which infer personality from samples of actual behavior or from informant reports.
What are reliability and validity? Reliability is the consistency of a score across items and over time; validity is the evidence that a test measures the construct it claims to and predicts what it should. Reliability is the precondition, and validity is the more demanding standard, because it must be built up from many separate lines of evidence.
What is the Five-Factor Model? It is the trait structure organizing personality into five broad domains: openness, conscientiousness, extraversion, agreeableness, and neuroticism. It gives different inventories a shared coordinate system, so scores from separate instruments can be compared and integrated.
Is the Rorschach inkblot test valid? Partly. A large meta-analysis found that the method's validity is uneven: some scores are well supported by evidence while many others are weak or unsupported, so its credentials are variable and specific to the score rather than uniform across the whole instrument.
What was the person-situation debate? It was the controversy Mischel opened in 1968 by showing that trait scores predicted single behaviors only weakly and arguing that situations mattered more than dispositions. It was resolved by aggregation: averaged across many occasions, traits predict behavior reliably, and the field emerged more precise about what a trait score claims.
What is a T-score? It is a standardized score scaled to a mean of 50 and a standard deviation of 10, the metric most personality inventories use so that a raw score can be interpreted against a norm group. A T-score of 60, for instance, is one standard deviation above the norm-group average.
How accurate is a single personality test score? A score should be read as a range, not a point. The standard error of measurement converts a scale's reliability into a confidence band, so even a highly reliable score spans several points, and responsible interpretation is cautious about small differences between scales or occasions.
References
Allport, G. W., & Odbert, H. S. (1936). Trait-names: A psycho-lexical study. Psychological Monographs, 47(1), i-171. https://doi.org/10.1037/h0093360
Bainbridge, T. F., Ludeke, S. G., & Smillie, L. D. (2022). Evaluating the Big Five as an organizing framework for commonly used psychological trait scales. Journal of Personality and Social Psychology, 122(4), 749-777. https://doi.org/10.1037/pspp0000395
Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81-105. https://doi.org/10.1037/h0046016
Cattell, R. B. (1943). The description of personality: Basic traits resolved into clusters. The Journal of Abnormal and Social Psychology, 38(4), 476-506. https://doi.org/10.1037/h0054116
Costa, P. T., Jr., & McCrae, R. R. (1992). Revised NEO Personality Inventory (NEO-PI-R) and NEO Five-Factor Inventory (NEO-FFI): Professional manual. Psychological Assessment Resources. OCLC 29360101.
Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297-334. https://doi.org/10.1007/BF02310555
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302. https://doi.org/10.1037/h0040957
Goldberg, L. R. (1993). The structure of phenotypic personality traits. American Psychologist, 48(1), 26-34. https://doi.org/10.1037/0003-066X.48.1.26
John, O. P., & Srivastava, S. (1999). The Big Five trait taxonomy: History, measurement, and theoretical perspectives. In L. A. Pervin & O. P. John (Eds.), Handbook of personality: Theory and research (2nd ed., pp. 102-138). Guilford Press. ISBN 9781572304833.
McCrae, R. R., & Costa, P. T., Jr. (1987). Validation of the five-factor model of personality across instruments and observers. Journal of Personality and Social Psychology, 52(1), 81-90. https://doi.org/10.1037/0022-3514.52.1.81
Meyer, G. J., Finn, S. E., Eyde, L. D., Kay, G. G., Moreland, K. L., Dies, R. R., Eisman, E. J., Kubiszyn, T. W., & Reed, G. M. (2001). Psychological testing and psychological assessment: A review of evidence and issues. American Psychologist, 56(2), 128-165. https://doi.org/10.1037/0003-066X.56.2.128
Mihura, J. L., Meyer, G. J., Dumitrascu, N., & Bombel, G. (2013). The validity of individual Rorschach variables: Systematic reviews and meta-analyses of the Comprehensive System. Psychological Bulletin, 139(3), 548-605. https://doi.org/10.1037/a0029406
Mischel, W. (1968). Personality and assessment. Wiley. OCLC 222767.
Roberts, B. W., & Yoon, H. J. (2022). Personality psychology. Annual Review of Psychology, 73, 489-516. https://doi.org/10.1146/annurev-psych-020821-114927
Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology, 113(1), 117-143. https://doi.org/10.1037/pspp0000096
Soto, C. J. (2019). How replicable are links between personality traits and consequential life outcomes? The Life Outcomes of Personality Replication project. Psychological Science, 30(5), 711-727. https://doi.org/10.1177/0956797619831612