Abstract

A personality test is a standardized instrument for measuring the enduring dispositions that distinguish one person from another. This article treats it as a psychological test whose object is the trait, tracing the two families the field developed: the self-report inventory, which scores answers to statements against trait scales, and the projective technique, which infers structure from responses to ambiguous stimuli. It sets out the framework organizing most modern inventories, the Five-Factor Model, and the psychometric standards any test must meet, internal-consistency reliability and construct validity established through a multitrait-multimethod matrix. It examines what tests predict, the person-situation debate over whether trait scores generalize across situations, and the machine-learning and item-level methods reshaping the field. Three demonstrations model a Big Five profile, how coefficient alpha depends on test length and item intercorrelation, and the meaning of a validity coefficient.

Keywords: personality test, five-factor model, reliability, construct validity, projective techniques

A personality test is a psychological test designed to measure the relatively stable patterns of thought, feeling, and behavior that characterize an individual across time and circumstance. It is one branch of the broader family of psychological tests, distinguished from ability and achievement measures by its object: not what a person can do at their best, but what a person is typically disposed to do. The enterprise rests on the assumption that such dispositions, the traits, exist, that they differ in degree between people, and that those differences can be quantified from a structured sample of behavior or self-description (Meyer et al., 2001). How to sample that behavior, how to score it, and how to know whether the resulting number means what it claims to mean are the questions that have organized a century of test construction.

Key Takeaways
  • A personality test measures enduring dispositions, the traits, that differ in degree between people and are inferred from a standardized sample of self-description or behavior.
  • Tests fall into two broad families: self-report inventories, which score answers to statements against trait scales, and projective techniques, which infer structure from responses to ambiguous stimuli.
  • Most modern inventories organize traits within the Five-Factor Model, five broad dimensions recovered repeatedly from both trait ratings and the natural-language lexicon of personality description.
  • A usable test must be both reliable, giving a consistent score, and valid, measuring the construct it claims to; reliability is necessary for validity but does not guarantee it.
  • Trait scores predict consequential life outcomes at modest but robust levels, a fact reconciled with the person-situation critique by aggregating behavior across many situations rather than predicting a single act.

What a Personality Test Is

A personality test operationalizes an abstract disposition as a measurable quantity. The disposition itself, extraversion or conscientiousness, is a construct: it cannot be observed directly but is posited to explain regularities in behavior. The test is a set of standardized observations, most often responses to a fixed set of items, from which the construct is estimated. The logic connecting the two was given its definitive statement by Cronbach and Meehl, who argued that a psychological test does not simply measure a trait but participates in a nomological network of expected relationships, and that the evidence for its validity is the degree to which its scores behave as the network predicts (Cronbach & Meehl, 1955). A personality test is therefore inseparable from a theory of the trait it measures; the number it yields is meaningful only against a web of expectations about what should and should not correlate with it.

Two features distinguish the personality test from other psychological tests. It has no correct answers: unlike an ability test, where a response is scored right or wrong, a personality item is scored only for the direction and degree of the disposition it indicates. And its data are, in the dominant tradition, self-reported, which makes the test vulnerable to distortions the respondent introduces, whether deliberate impression management or the subtler biases of acquiescence and socially desirable responding. Much of the methodology of personality testing is a response to these two facts: the absence of a scoring key and the fallibility of the respondent as a reporter on themselves.

Types of Personality Tests

Within the medical subject headings, Personality Tests is the parent of five narrower classes of instrument, and the standard classification (Table 1) reflects the different ways a test can sample the dispositions it aims to measure. These categories are an indexing classification, the scheme by which the assessment literature is catalogued, rather than a theory that carves personality measurement at its joints; they are not mutually exclusive, and a single published instrument may draw on more than one. A projective technique and a self-report inventory can target the same trait by opposite routes, and the tradition to which a test belongs reflects its assumptions about how personality is best revealed.

Table 1. Direct subtypes of Personality Tests in the MeSH classification (tree F04.711.647).
Subtype In brief
Bender-Gestalt Test A visuomotor task in which the respondent copies geometric designs; used mainly to screen for neuropsychological impairment and, in its projective use, for personality indicators.
Personality Inventory A structured self-report questionnaire whose items are answered on fixed scales and summed into trait scores, the form of the MMPI, the NEO-PI-R, and the Big Five inventories.
Projective Techniques Methods presenting ambiguous stimuli, such as inkblots or pictures, on the premise that the structure a respondent imposes reveals latent dispositions.
Semantic Differential A rating procedure that locates the connotative meaning a person attaches to a concept on a series of bipolar adjective scales.
Word Association Tests Methods that infer conflict or disposition from the content and latency of a respondent's replies to a series of stimulus words.

The two families that dominate contemporary practice, the personality inventory and the projective technique, embody opposite wagers about the respondent. The inventory trusts people to describe themselves accurately when asked directly; the projective technique assumes that the most important dispositions are not available to introspection and must be inferred indirectly. The remainder of this article concentrates on the inventory, which structures the modern field, returning to the projective tradition where the contrast is instructive.

The Structure of Personality: The Five-Factor Model

An inventory must decide which traits to measure, and for most of the twentieth century that decision was made by theoretical fiat, each author proposing their own list. The resolution came from the lexical hypothesis: the idea that the important individual differences have, over time, become encoded as single words in the natural language, so that the structure of personality can be recovered by factor-analyzing the vocabulary of trait description. Repeated analyses of trait-descriptive adjectives converged on five broad, replicable factors, first identified by Tupes and Christal in United States Air Force samples (Tupes & Christal, 1992) and later established as robust across methods and populations by Goldberg, who named them the Big Five (Goldberg, 1990). The five, conventionally labeled openness to experience, conscientiousness, extraversion, agreeableness, and neuroticism, were shown by McCrae and Costa to recover the same structure when the data came from questionnaires rather than adjective lists and from observer ratings rather than self-report, evidence that the factors are properties of the trait domain and not artifacts of one method (McCrae & Costa, 1987).

The Five-Factor Model is a hierarchical description: each broad domain subsumes narrower facets, and each facet is indicated by several items, so that a respondent's answers aggregate upward from item to facet to domain (Figure 1). Modern inventories build this hierarchy explicitly; the Big Five Inventory-2, for instance, measures each of the five domains through three facets and each facet through several items, trading a small increase in length for the ability to report both the broad domain and its component parts (Soto & John, 2017). The model is descriptive rather than final: the six-factor HEXACO model adds an honesty-humility dimension that the Big Five folds into agreeableness, and argues that six factors better capture the lexical evidence in several languages (Ashton & Lee, 2007). It is also dimensional rather than typological: it locates a person at a continuous point on each factor rather than sorting people into discrete kinds. This distinguishes it from the type-based tests the public most associates with personality assessment; when the best known of these, the Myers-Briggs Type Indicator, is examined against the five factors, its dichotomous types resolve into continuous dimensions cut arbitrarily at the middle, so that the same responses are better described as scores than as types (McCrae & Costa, 1989). The first demonstration renders a Big Five profile as five domain scores, letting the reader set each and see the percentile and verbal descriptor that a scored inventory would report.

Figure 1

The Hierarchical Structure of a Five-Factor Domain

A three-level hierarchy from items to facets to a personality domain A tree diagram with three tiers. At the top, a single dark box labelled Domain, Conscientiousness. Three lines descend to a middle tier of three boxes labelled Orderliness, Productiveness, and Responsibility, marked as facets. From each facet, three lines descend to small boxes in the bottom tier, marked as items. Scores aggregate upward: items sum into facets, and facets sum into the domain. Domain Facets Items Conscientiousness Orderliness Productiveness Responsibility
Note. A single domain, conscientiousness, decomposes into three facets, each measured by several items, following the Big Five Inventory-2 structure (Soto & John, 2017). A respondent's item answers sum into facet scores and facet scores into the domain, so the same responses can be reported at whichever level of breadth a use requires. Original schematic.

Set the Scores, Read the Profile

Scoring a Big Five Inventory

Drag each domain to a T-score. A score of 50 is the population average; every ten points is one standard deviation. The percentile and band on the right are recomputed from the normal distribution as you move each slider.

Openness: T 6288th · very high
Conscientiousness: T 5569th · high
Extraversion: T 4842th · average
Agreeableness: T 5050th · average
Neuroticism: T 4016th · very low
mean (T 50)OpennessConscientiousnessExtraversionAgreeablenessNeuroticism
Five domain scores on the T-score metric an inventory reports (mean 50, standard deviation 10). Each raw score is located in the norm distribution and converted to a percentile through the normal cumulative distribution, then given the verbal band a score report would print. This is the scoring step of a self-report inventory: a raw sum interpreted against norms. An illustrative scoring model on Gaussian norms; percentiles computed locally, not stored.

Reliability and Validity

A personality score is worth nothing unless it is both reliable, consistent on repetition, and valid, a measure of the intended construct. Reliability comes first because it bounds validity: a score swamped by measurement error cannot correlate with anything, including the trait it targets. The most widely reported index of reliability for an inventory is internal consistency, the degree to which the items of a scale measure the same thing, summarized by Cronbach's coefficient alpha (Cronbach, 1951). Alpha rises with the average correlation among items and with the number of items, which is why lengthening a homogeneous scale raises its reliability; the second demonstration makes this dependence explicit, and the Worked Example computes a case by hand.

Reliability is necessary but not sufficient. A scale can be perfectly consistent and still measure the wrong thing, so the harder question is construct validity: whether the score behaves as a measure of the trait should. Campbell and Fiske gave the field its enduring method for answering it, the multitrait-multimethod matrix, which requires that measures of the same trait by different methods agree (convergent validity) and that measures of different traits, even by the same method, diverge (discriminant validity) (Campbell & Fiske, 1959). The requirement is exacting: a self-report extraversion scale that correlates highly with an observer rating of extraversion but only modestly with a self-report of agreeableness has evidence for validity, whereas one that correlates more with other self-reports than with other measures of the same trait is measuring method, not trait. Reasoning about a test score as a decision about the presence of a trait, with its attendant hits and false alarms, connects personality measurement to signal detection theory, in which the separation of true signal from noise is made formal.

Lengthen the Test, Watch Alpha Rise

Coefficient Alpha by Length and Item Correlation

Set the number of items and their average correlation with one another. Alpha rises with both, which is why a longer scale of equally related items is more reliable. The 0.70 line marks the conventional minimum for research use.

Number of items (k)10
Average inter-item correlation0.30
0.00.51.00.70number of items
Coefficient alpha = 0.81 (good) for 10 items at an average inter-item correlation of 0.30.
Standardized coefficient alpha as a function of the number of items and their average intercorrelation, alpha = k times rbar over one plus rbar times (k minus one). The curve traces alpha against test length at the chosen item correlation; the marker is the current scale. The defaults reproduce the Worked Example (ten items at an average correlation of 0.30 give alpha 0.81; twenty give 0.90). Computed locally from the closed form, not stored.

Self-Report and the Projective Alternative

The self-report inventory is efficient, standardized, and cheap, but it inherits the fallibility of the self as witness. Respondents may not know themselves accurately, and even when they do they may not report accurately, shading answers toward social approval or endorsing items indiscriminately. Inventories defend against these threats with reverse-keyed items that catch indiscriminate agreement, with validity scales that detect distortion, and with forced-choice formats that pit equally desirable options against each other. The projective techniques take a more radical route around the problem: by presenting an ambiguous stimulus with no obviously desirable response, they aim to bypass conscious self-presentation entirely and reach dispositions the respondent could not report even if willing.

The scientific standing of the projective tradition has been sharply contested. A comprehensive review concluded that most projective instruments, and most of the scores derived from the Rorschach in particular, lacked the validity evidence their clinical use presupposed (Lilienfeld et al., 2000). A later systematic meta-analysis was more discriminating, finding that a minority of Rorschach variables, especially those indexing thought disorder and reality testing, had respectable validity while many others did not, so that the instrument is neither uniformly valid nor uniformly worthless (Mihura et al., 2013). Underlying the whole debate is a deeper challenge to the trait concept itself. Mischel marshaled evidence that trait scores predict behavior in any single situation only weakly, the correlation rarely exceeding about 0.30, and argued that behavior is controlled more by situations than by broad dispositions (Mischel, 1968). The person-situation debate that followed forced the field to specify exactly what a personality test can and cannot predict.

Predicting Life Outcomes

The resolution of the person-situation debate was a lesson in the level at which traits operate. A trait does not predict a single act well, but it predicts the aggregate of many acts across situations reliably, because the situational influences that dominate any one occasion average out over many. Assessed at that level, trait scores predict consequential outcomes at levels that rival or exceed the standard predictors of applied psychology. A landmark review compared the predictive validity of personality traits against socioeconomic status and cognitive ability for outcomes including mortality, divorce, and occupational attainment, and found that traits, conscientiousness above all, held their own or better across the board (Roberts et al., 2007). More broadly, the validity coefficients of well-constructed psychological tests are comparable to those of established medical tests, a comparison that reframes the modest-looking correlations of the field as unremarkable by the standards of applied prediction generally (Meyer et al., 2001).

What a validity coefficient of that size means in practice is easy to misjudge. A correlation of 0.30, which the person-situation critique treated as a ceiling on trait prediction, explains only nine percent of the variance in a single outcome, yet it shifts the odds of that outcome substantially between people at opposite ends of the trait, and its practical value accumulates across the many decisions to which it is applied (Roberts & Yoon, 2022). The third demonstration makes this concrete, letting the reader set a validity coefficient and see both the scatter it produces and the change in outcome odds it implies between high and low scorers.

Set the Validity, See the Spread

What a Validity Coefficient Buys

Set the validity coefficient, the correlation between the test score (horizontal) and the outcome (vertical). Read the variance it explains, the success rates it implies for high and low scorers, and watch the cloud of cases tighten toward a line as validity grows.

Validity coefficient (r)0.30
test score →
At r = 0.30 the test explains 9% of outcome variance. Success rate: 65% among high scorers versus 35% among low scorers.
A validity coefficient is the correlation between a test score and the outcome it predicts. Its square is the share of outcome variance explained; the binomial effect-size display re-expresses it as the success rate in the high-trait versus low-trait half. The scatter is a deterministic 200-point sample drawn from the chosen population correlation (seeded, so it never changes on reload). At r = 0.30 the test explains 9 percent of variance yet shifts success from 35 to 65 percent between the extremes. Illustrative sample; computed locally, not stored.

Worked Example

Consider a conscientiousness scale of 10 items whose average correlation with one another is 0.30. Its internal-consistency reliability, in the standardized form of coefficient alpha, follows from the number of items and their average intercorrelation. The standardized alpha is the number of items times the average inter-item correlation, divided by one plus the average correlation multiplied by one less than the number of items. Here that is 10 times 0.30, which is 3.0, divided by 1 plus 0.30 times 9, which is 1 plus 2.7, or 3.7. The quotient is 3.0 divided by 3.7, about 0.81, comfortably above the 0.70 conventionally treated as the minimum for research use.

The two levers are visible in the formula. Suppose the scale is lengthened to 20 items of the same average intercorrelation: alpha becomes 20 times 0.30, or 6.0, divided by 1 plus 0.30 times 19, which is 6.7, giving 6.0 divided by 6.7, about 0.90. Doubling the length raised reliability from 0.81 to 0.90 without improving a single item, because alpha rewards the accumulation of consistent evidence. Reliability in turn bounds the precision of an individual score. With a reliability of 0.81 and a scale standard deviation of 15 points, the standard error of measurement is 15 times the square root of 1 minus 0.81, which is 15 times the square root of 0.19, or 15 times 0.436, about 6.5 points. A person's observed score of 60 therefore carries a 95 percent confidence interval of roughly plus or minus 13 points, from about 47 to 73, a spread wide enough to caution against over-interpreting a single administration. The reliability demonstration recomputes alpha from the same formula as the reader varies the two inputs.

Discussion

The history of the personality test is the gradual disciplining of an intuitively appealing but methodologically treacherous idea, that a person can be summarized by a small set of numbers. The Five-Factor Model gave the field a shared and replicable answer to the question of which numbers, ending a century of competing proprietary lists (Goldberg, 1990). The psychometric apparatus of reliability and construct validity gave it standards for deciding whether a proposed measure of any of those numbers is sound (Cronbach & Meehl, 1955; Campbell & Fiske, 1959). And the person-situation debate, initially a threat to the whole enterprise, sharpened rather than destroyed it, by establishing that traits predict aggregated behavior rather than single acts, and that the modest single-act correlations Mischel emphasized are exactly what a proper theory of traits should expect (Mischel, 1968; Roberts et al., 2007).

Two tensions remain unresolved. The first is between the self-report inventory's efficiency and its dependence on an honest and self-aware respondent, a dependence the projective tradition rejected without producing a broadly validated alternative (Lilienfeld et al., 2000; Mihura et al., 2013). The second is between the descriptive success of a small number of broad factors and the recurring evidence that important variance lives below them, in the facets and the individual items that the domain scores average away. Whether the future of personality measurement lies in refining the five broad factors or in modeling personality at a finer grain is the question the field is now actively contesting (Roberts & Yoon, 2022).

Current Directions

The most active current work is pushing personality measurement below the level of the broad trait. A prominent argument holds that the Big Five domains, useful as they are for description, are too coarse for either prediction or explanation, and that the field should model personality at the level of individual items or narrow characteristics, each of which carries specific, sometimes causally interpretable, information that the domain average discards (Mõttus et al., 2020). Complementing this, machine-learning methods are being brought to bear on personality data, both to extract trait-relevant signal from behavioral traces such as digital records and language, and to model the nonlinear structure that linear factor models cannot represent; the promise is a measurement freed from the fixed questionnaire item, and the open problem is validating predictions that no longer rest on transparent self-report (Bleidorn & Hopwood, 2019). A third strand asks how far the Big Five actually serves as the common framework it is assumed to be: a systematic evaluation of widely used trait scales found that most, though not all, could be located within the five-factor space, clarifying both the reach of the model and the constructs that escape it (Bainbridge et al., 2022). Together these lines mark a shift from asking which broad traits exist to asking at what level personality is most usefully measured.

Common Misconceptions

A reliable personality test is therefore an accurate one.
Reliability and validity are distinct properties. A scale can yield a highly consistent score that nonetheless measures the wrong construct, which is why internal consistency must be supplemented by convergent and discriminant validity evidence before a score can be trusted to mean what it claims (Campbell & Fiske, 1959).
The Rorschach and other projective tests have no scientific validity.
The truth is more discriminating than the blanket dismissal. A systematic meta-analysis found that a subset of Rorschach scores, particularly those indexing thought disorder, have meaningful validity, while many other scores do not, so the instrument is neither uniformly valid nor uniformly worthless (Mihura et al., 2013).
Because behavior varies so much by situation, personality tests cannot predict it.
The weak prediction of a single act, which drove the person-situation critique, gives way to robust prediction once behavior is aggregated across many situations, because situational influences average out; traits predict consequential life outcomes at levels comparable to the standard predictors of applied psychology (Roberts et al., 2007).

Glossary

Acquiescence bias.
The tendency to agree with items regardless of content, a response style that inflates or distorts self-report scores and is countered by reverse-keyed items.
Big Five.
The five broad, replicable factors of personality, openness, conscientiousness, extraversion, agreeableness, and neuroticism, recovered from both trait ratings and the natural-language lexicon.
Coefficient alpha.
An index of internal-consistency reliability that summarizes how strongly a scale's items intercorrelate; it rises with the number of items and their average intercorrelation.
Construct validity.
The degree to which a test measures the theoretical construct it claims to, evidenced by the score behaving as the surrounding network of predictions requires.
Construct.
An abstract, unobservable disposition, such as a trait, that a test estimates from observable responses and that is defined by its place in a network of expected relationships.
Convergent validity.
Agreement between measures of the same trait obtained by different methods; one half of the requirement a multitrait-multimethod matrix imposes.
Discriminant validity.
Divergence between measures of different traits, even when obtained by the same method; the complement of convergent validity in evaluating a measure.
Facet.
A narrow trait nested beneath a broad Five-Factor domain; several facets compose a domain and several items compose a facet in a hierarchical inventory.
Five-Factor Model.
The dominant descriptive framework of personality, organizing traits into the five broad domains of the Big Five, each subsuming narrower facets.
Lexical hypothesis.
The proposal that the socially important personality differences have become encoded as single words in natural language, so that trait structure can be recovered from the vocabulary.
Person-situation debate.
The controversy, launched by Mischel, over whether behavior is governed more by stable traits or by situations, resolved by the finding that traits predict aggregated rather than single behavior.
Personality inventory.
A structured self-report questionnaire whose items are answered on fixed scales and summed into trait scores; the dominant form of the modern personality test.
Projective technique.
A method presenting an ambiguous stimulus, such as an inkblot, on the premise that the structure a respondent imposes reveals dispositions unavailable to direct self-report.
Reliability.
The consistency of a test score across items, occasions, or raters; a necessary condition for validity, since error variance cannot correlate with the target trait.
Standard error of measurement.
The expected spread of an individual's observed scores around the true score, computed from the scale standard deviation and its reliability; it sets the width of a score's confidence interval.
Trait.
A relatively stable disposition to think, feel, or behave in a characteristic way, differing in degree between people and constituting the object a personality test measures.
Validity coefficient.
The correlation between a test score and a criterion it is meant to predict; even modest values shift outcome odds meaningfully across the range of the trait.

Key Researchers

Michael C. Ashton. Professor of Psychology at Brock University; with Kibeom Lee he developed the six-factor HEXACO model, adding an honesty-humility dimension to the lexical description of personality structure. Faculty Page - Google Scholar

Wiebke Bleidorn (b. 1980). Professor of Differential Psychology and Psychological Assessment at the University of Zurich; she leads research on lifespan personality development and on machine-learning approaches to personality assessment. Faculty Page - ORCID - Wikipedia - Wikidata

Paul T. Costa (b. 1942). Scientist Emeritus at the National Institute on Aging; with Robert McCrae he built the Revised NEO Personality Inventory, the instrument that operationalized the Five-Factor Model for clinical and research use. ORCID - Wikipedia - Wikidata

Lee J. Cronbach (1916-2001). Vida Jacks Professor of Education at Stanford University; he originated coefficient alpha, the standard index of internal-consistency reliability, and with Paul Meehl formalized the concept of construct validity. Wikipedia - Wikidata

Lewis R. Goldberg (1932-2026). Senior Scientist at the Oregon Research Institute; he established the lexical Big Five factor structure, coined the term Big Five, and created the public-domain International Personality Item Pool. Google Scholar - Faculty Page - Wikipedia - Wikidata

Robert R. McCrae (b. 1949). Independent scientist, formerly of the National Institute on Aging; with Paul Costa he established the cross-instrument, cross-observer, and cross-cultural robustness of the five-factor trait structure. ORCID - Wikipedia - Wikidata

Walter Mischel (1930-2018). Robert Johnston Niven Professor of Humane Letters in Psychology at Columbia University; his 1968 monograph launched the person-situation debate by questioning how far trait scores predict behavior across situations. Wikipedia - Wikidata

Hermann Rorschach (1884-1922). Swiss psychiatrist; he created the inkblot test, the prototypical projective technique, which infers personality structure from a respondent's interpretation of ambiguous stimuli. Wikipedia - Wikidata

Christopher J. Soto. Professor of Psychology at Colby College; he led the development of the Big Five Inventory-2, a hierarchical Big Five measure with fifteen facets now widely used and cross-culturally validated. Faculty Page - ORCID

Frequently Asked Questions

What is a personality test?
A personality test is a standardized psychological test that measures the enduring dispositions, or traits, that differ in degree between people, inferring them from a structured sample of self-description or behavior (Meyer et al., 2001).

What are the Big Five personality traits?
The Big Five are five broad, replicable dimensions, openness, conscientiousness, extraversion, agreeableness, and neuroticism, recovered repeatedly from both trait ratings and the natural-language vocabulary of personality description (Goldberg, 1990).

What is the difference between a self-report inventory and a projective technique?
A self-report inventory asks a person to rate statements about themselves and sums the answers into trait scores, whereas a projective technique presents an ambiguous stimulus and infers dispositions from the structure the person imposes on it (Lilienfeld et al., 2000).

How is the reliability of a personality test measured?
The most common index is internal-consistency reliability, summarized by coefficient alpha, which reflects how strongly the items of a scale intercorrelate and rises with both the number of items and their average intercorrelation (Cronbach, 1951).

What does it mean for a personality test to be valid?
Validity means the test measures the construct it claims to, shown when scores converge with other measures of the same trait and diverge from measures of different traits, the pattern required by a multitrait-multimethod matrix (Campbell & Fiske, 1959).

Do personality tests actually predict behavior?
They predict aggregated behavior and consequential life outcomes at modest but robust levels, comparable to the standard predictors of applied psychology, even though they predict any single act only weakly (Roberts et al., 2007).

Is the Rorschach inkblot test scientifically valid?
A systematic meta-analysis found that some Rorschach scores, especially those indexing thought disorder, have meaningful validity while many others do not, so the instrument is neither uniformly valid nor uniformly without support (Mihura et al., 2013).

What was the person-situation debate?
It was the controversy launched by Mischel over whether behavior is governed more by stable traits or by situations, resolved by the finding that traits predict behavior aggregated across many situations rather than any single act (Mischel, 1968).

References

Ashton, M. C., & Lee, K. (2007). Empirical, theoretical, and practical advantages of the HEXACO model of personality structure. Personality and Social Psychology Review, 11(2), 150-166. https://doi.org/10.1177/1088868306294907

Bainbridge, T. F., Ludeke, S. G., & Smillie, L. D. (2022). Evaluating the Big Five as an organizing framework for commonly used psychological trait scales. Journal of Personality and Social Psychology, 122(4), 749-777. https://doi.org/10.1037/pspp0000395

Bleidorn, W., & Hopwood, C. J. (2019). Using machine learning to advance personality assessment and theory. Personality and Social Psychology Review, 23(2), 190-203. https://doi.org/10.1177/1088868318772990

Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81-105. https://doi.org/10.1037/h0046016

Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297-334. https://doi.org/10.1007/BF02310555

Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302. https://doi.org/10.1037/h0040957

Goldberg, L. R. (1990). An alternative “description of personality”: The Big-Five factor structure. Journal of Personality and Social Psychology, 59(6), 1216-1229. https://doi.org/10.1037/0022-3514.59.6.1216

Lilienfeld, S. O., Wood, J. M., & Garb, H. N. (2000). The scientific status of projective techniques. Psychological Science in the Public Interest, 1(2), 27-66. https://doi.org/10.1111/1529-1006.002

McCrae, R. R., & Costa, P. T. (1987). Validation of the five-factor model of personality across instruments and observers. Journal of Personality and Social Psychology, 52(1), 81-90. https://doi.org/10.1037/0022-3514.52.1.81

McCrae, R. R., & Costa, P. T. (1989). Reinterpreting the Myers-Briggs Type Indicator from the perspective of the five-factor model of personality. Journal of Personality, 57(1), 17-40. https://doi.org/10.1111/j.1467-6494.1989.tb00759.x

Meyer, G. J., Finn, S. E., Eyde, L. D., Kay, G. G., Moreland, K. L., Dies, R. R., Eisman, E. J., Kubiszyn, T. W., & Reed, G. M. (2001). Psychological testing and psychological assessment: A review of evidence and issues. American Psychologist, 56(2), 128-165. https://doi.org/10.1037/0003-066X.56.2.128

Mihura, J. L., Meyer, G. J., Dumitrascu, N., & Bombel, G. (2013). The validity of individual Rorschach variables: Systematic reviews and meta-analyses of the Comprehensive System. Psychological Bulletin, 139(3), 548-605. https://doi.org/10.1037/a0029406

Mischel, W. (1968). Personality and assessment. Wiley.

Mõttus, R., Wood, D., Condon, D. M., Back, M. D., Baumert, A., Costantini, G., Epskamp, S., Greiff, S., Johnson, W., Lukaszewski, A., Murray, A., Revelle, W., Wright, A. G. C., Yarkoni, T., Ziegler, M., & Zimmermann, J. (2020). Descriptive, predictive and explanatory personality research: Different goals, different approaches, but a shared need to move beyond the Big Few traits. European Journal of Personality, 34(6), 1175-1201. https://doi.org/10.1002/per.2311

Roberts, B. W., Kuncel, N. R., Shiner, R., Caspi, A., & Goldberg, L. R. (2007). The power of personality: The comparative validity of personality traits, socioeconomic status, and cognitive ability for predicting important life outcomes. Perspectives on Psychological Science, 2(4), 313-345. https://doi.org/10.1111/j.1745-6916.2007.00047.x

Roberts, B. W., & Yoon, H. J. (2022). Personality psychology. Annual Review of Psychology, 73, 489-516. https://doi.org/10.1146/annurev-psych-020821-114927

Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology, 113(1), 117-143. https://doi.org/10.1037/pspp0000096

Tupes, E. C., & Christal, R. E. (1992). Recurrent personality factors based on trait ratings. Journal of Personality, 60(2), 225-251. https://doi.org/10.1111/j.1467-6494.1992.tb00973.x