Abstract

The Minnesota Multiphasic Personality Inventory (MMPI) is a type of personality inventory: a broad self-report instrument whose true-or-false items are scored onto clinical scales that screen for patterns of psychopathology. It is the most widely used such instrument in history and the paradigm case of empirical criterion keying — items were placed on a scale not because they seemed to measure a construct but because they empirically separated a diagnostic group from a normative sample, whatever their surface content. This article treats the MMPI as a case study in that method: how contrasted-group keying builds a scale, why the instrument carries validity scales measuring how a person answered, how the K scale corrects clinical scores as a suppressor variable, and what the restructured editions changed. Its logic is the counterpoint to theory-driven test design.

Keywords: personality assessment, empirical criterion keying, validity scales

The Minnesota Multiphasic Personality Inventory, universally abbreviated MMPI, is the most heavily used and most heavily researched personality inventory ever built, and it owes that standing to a method rather than a theory. Where a theory-driven instrument specifies in advance what its scales should contain, the MMPI was assembled by a procedure that treated the question of what an item measures as an empirical one, to be settled by data rather than by inspection. Reading the instrument well means understanding that procedure, because nearly everything distinctive about the MMPI — the numbered scales, the validity indices, the K correction — is a consequence of building a test by contrasted-group keying rather than by design from a construct.

That method was Starke Hathaway and J. C. McKinley's. Working at the University of Minnesota in the late 1930s, they wanted a practical instrument that could screen psychiatric patients efficiently, and they were willing to let the data decide which items belonged on which scale (Hathaway & McKinley, 1940). The result was a test whose items are often opaque about what they measure, whose scales are named by number to discourage over-literal reading, and whose interpretation rests on decades of accumulated empirical correlates rather than on the face meaning of any single question. The MMPI is, in short, the great worked example of what psychologists came to call the empirical, or criterion-keying, approach to test construction.

Key Takeaways
  • The MMPI is a broad self-report inventory whose true-or-false items are scored onto clinical scales that screen for patterns of psychopathology, interpreted from empirical correlates rather than item content.
  • Its defining method is empirical criterion keying: an item joins a scale because it statistically separated a criterion diagnostic group from a normative sample, regardless of what the item appears to ask.
  • Because the scored datum is the response itself and not its literal content, the instrument carries validity scales (L, F, and K) that measure how candidly and consistently a person answered before the clinical scales are read.
  • The K scale is used as a suppressor variable: a fraction of it is added back to certain clinical scales to sharpen the discrimination between genuine and defended responding.
  • Successive editions — MMPI-2, the Restructured Form, and the MMPI-3 — re-normed the instrument and separated a pervasive demoralization factor from the distinctive core of each clinical scale.

What the MMPI Is

The MMPI is a self-report inventory in which a respondent answers several hundred true-or-false statements about symptoms, attitudes, and habits, and those answers are scored onto a profile of scales representing dimensions of psychopathology and personality (Hathaway & McKinley, 1943). The original clinical scales were numbered rather than named — 1 through 0 — precisely because their creators did not want a clinician to read a scale as a literal measure of the diagnostic label it was keyed against; a high score on the scale keyed to depression, for instance, carries a web of empirical meanings that outruns the everyday word. The output is a profile across scales, read as a configuration, not a single verdict.

Two features distinguish the MMPI from a questionnaire built around a construct. The first is that its scales were not written to measure a trait and then checked; they were selected item by item for their power to separate diagnostic groups, so an item's presence on a scale is a claim about its empirical behaviour, not its meaning. The second is that the instrument does not trust the respondent to be a candid or accurate reporter. Because it treats the act of endorsing an item as the datum, it must also measure the style of that endorsing, which is why a set of validity scales is woven into the same booklet as the clinical ones.

This makes the MMPI an instructive contrast with a theory-driven inventory such as the Millon Clinical Multiaxial Inventory, whose scales are deduced from a theory of personality and aligned to the diagnostic manual of the day. The MMPI's original scales were instead keyed against the psychiatric nosology of the 1940s, and the instrument has always been read as a set of dimensional empirical screens rather than as a direct map onto the categories of any one edition of the diagnostic manual (American Psychiatric Association, 2013). The two instruments answer the same clinical need by opposite philosophies of construction.

Building a Scale from Data: Empirical Criterion Keying

The method at the heart of the MMPI is contrasted-group, or criterion, keying, and its logic is deceptively simple. To build a scale for a given condition, the constructors assembled a criterion group of patients carrying that diagnosis and a comparison group of normative respondents, then administered a large pool of candidate items to both. Any item that a substantially different proportion of the two groups endorsed became a member of the scale, keyed in whichever direction the criterion group favoured. An item earned its place by discriminating, not by sounding relevant (Hathaway & McKinley, 1940). Items that read as obviously about the condition often failed to discriminate and were dropped; items with no evident bearing on it often separated the groups sharply and were kept.

The consequence is that MMPI items need not be, and frequently are not, transparent about what they measure. This was not an embarrassment to be tolerated but a principle to be defended, and Paul Meehl gave it its classic statement: a 'structured' personality test item is not a self-evident measure whose content must be taken at face value, but a piece of verbal behaviour whose diagnostic meaning is whatever its empirical correlates turn out to be (Meehl, 1945). On this view the answer of 'true' or 'false' is a response to be counted, and the question of what the response indicates is answered by the keying data, not by the semantics of the sentence. The method thus dissolves the intuition that a good test item must look like the thing it measures.

Keying scale by scale against separate criterion groups had a second consequence that shaped everything downstream: the resulting scales overlap and share items, and they carry along whatever the criterion groups had in common beyond their defining diagnosis (McKinley & Hathaway, 1944). Because clinical groups of every kind tend to be more distressed, demoralized, and forthcoming about trouble than a normative sample, items tapping general maladjustment were keyed onto many scales at once. That shared variance is the source of the MMPI's characteristic scale intercorrelations, and it is precisely what the later restructured editions set out to isolate and remove. The strength of empirical keying — that it captures whatever actually distinguishes a group — is inseparable from its cost, that some of what distinguishes every clinical group is the same thing.

Figure 1. Empirical criterion keying: an item joins a scale by how differently two groups endorse it, not by its content.
A candidate item endorsed at different rates by a criterion group and a normative group Two vertical bars show the percentage endorsing a candidate item. The criterion clinical group endorses it at seventy percent; the normative group endorses it at thirty percent. The forty-point difference between the bars is what places the item on the scale, keyed in the direction the criterion group favours. The item's surface content is marked irrelevant to the decision. The gap between groups keys the item, not what it says 100% 50% 0% 70% Criterion group 30% Normative group 40-point gap Item content is not consulted; the discrimination decides membership.

Note. The endorsement difference, not the wording, determines whether a candidate item is keyed onto the scale and in which direction. Illustrative values. Original schematic.

Measuring How a Person Answered: The Validity Scales

Because the MMPI scores the fact of a response rather than trusting its content, it must gauge the manner in which a respondent approached the whole booklet, and it does so with validity scales built into the instrument from the start (Hathaway & McKinley, 1943). Three carry the classic load. The L (Lie) scale collects items describing minor faults almost everyone will admit, so that denying them signals a naive attempt to look virtuous. The F (Infrequency) scale collects items very rarely endorsed in the normative sample, so that endorsing many of them signals unusual responding — severe disturbance, carelessness, random answering, or deliberate exaggeration. The K (Correction) scale is subtler, capturing a guarded, defensive test-taking attitude that depresses clinical scores without the obvious naivety the L scale catches.

Read together, the validity scales describe a respondent's response style before any clinical scale is interpreted, and their configuration is diagnostic in its own right. A profile high on L and K but low on F suggests defensive under-reporting — an effort to appear well-adjusted that may mask genuine elevations. The opposite configuration, F markedly high with K low, suggests over-reporting: a person endorsing far more rare symptoms than a genuine clinical presentation typically produces. Distinguishing genuine distress from symptom exaggeration is among the instrument's most consequential tasks, and it is the specific target of the over-reporting validity scales that the modern editions expanded and that a substantial meta-analytic literature has since evaluated (Ingram & Ternes, 2016).

The validity scales are why the MMPI can be used at all in settings where respondents have incentives to distort. In forensic, disability, and pre-employment evaluations, the manner of responding cannot be assumed candid, and the validity indices are typically the first scales a clinician reads, because a clinical profile is uninterpretable until the response style that produced it is established (Tylicki et al., 2022). The instrument's willingness to measure the respondent's approach to the test, and not merely their symptoms, is a direct descendant of the empirical premise that a response means whatever the data say it means — including data about honesty.

Table 1. The MMPI lineage: successive editions and what distinguished each from its predecessor.

Edition Introduced Distinguishing development
MMPI1943The original; empirically keyed clinical scales and the L, F, and K validity scales
MMPI-21989Restandardized on a large contemporary sample; introduced uniform T-scores and new content and validity scales
MMPI-2-RF2008Restructured Clinical scales isolating a shared demoralization factor from each scale's distinctive core
MMPI-32020Newly normed on a demographically current sample; updated scales and expanded over-reporting validity indices

Note. Item counts and full scale rosters differ across editions and are omitted here. Each revision re-normed the instrument while preserving the empirical-keying logic of the original. Original schematic.

Correcting the Score: K as a Suppressor Variable

The most technically interesting move in the MMPI's scoring is the use of the K scale not merely as a warning flag but as an active correction to the clinical scales. Hathaway and McKinley observed that a defensive test-taking attitude systematically lowers clinical scores, so two people with the same underlying disturbance can produce different profiles simply because one answered more guardedly. Meehl and Hathaway proposed treating K as a suppressor variable — a measure that carries no direct information about the clinical construct itself but captures the contaminating defensiveness, so that adding a portion of it back to a clinical score removes that contamination and sharpens the scale's discrimination (Meehl & Hathaway, 1946).

In practice a fixed fraction of the K raw score is added to five of the clinical scales, with the fraction chosen empirically for each scale to maximize its separation of criterion from normative respondents. Some scales receive the full K, others a half or a fifth, and several receive none, reflecting how much each was found to be depressed by defensiveness. The correction is mechanical and applied before the profile is plotted, so a clinician reads K-corrected scores without recomputing them, but its logic is entirely statistical: K is worth adding in the exact degree that it predicts the error a defensive style introduces. This is a genuinely subtle idea — a variable unrelated to the target is used to improve the target's measurement precisely because it is related to the error.

The K correction also exposes the seam in empirical test construction. Because the weights were estimated on particular samples to optimize a particular discrimination, they inherit the peculiarities of those samples, and the question of whether a correction tuned on one population transfers to another is exactly the kind of question the empirical method raises and cannot settle in advance. That concern, generalized from the K weights to the whole edifice of keyed scales, is what motivated the restructuring of the clinical scales two generations later, when the shared demoralization variance the original keying had spread across scales was deliberately extracted and modelled on its own (Simms et al., 2005).

The Inventory in Motion

The three demonstrations below make manipulable the parts of the MMPI that prose can only describe. The first performs empirical keying, deciding whether a candidate item joins a scale from how differently two groups endorse it. The second applies the K correction, adding a fraction of a suppressor score back to a clinical scale. The third reads a validity configuration, classifying a response style from the L, F, and K scores.

Demo 1 — Keying an item by contrasted groups
70%Criterion30%Normative+40 pts

Difference: +40 percentage points (threshold to key: 20). KEYED — scored in the "true" direction

The item reads as on-topic, but its wording is never consulted: only the 20-point gap between the groups decides membership. An off-topic item that separates the groups is kept; an on-topic item that does not is dropped. That is empirical criterion keying.

The empirical-keying demonstration makes the instrument's central method concrete. Setting the endorsement rate of a candidate item in the criterion group and in the normative group changes the gap between them, and the demonstration keys the item onto the scale only when that gap crosses the selection threshold — in whichever direction the criterion group favours. Deliberately mismatching an item's stated content with its discrimination shows the method's signature: content is never consulted, and an item that sounds irrelevant is kept whenever it separates the groups.

Demo 2 — The K correction as a suppressor
Raw= 35K add+15navy = raw clinical score · gold = suppressor contribution (1K)

Corrected score = 20 + 1 × 15 = 35. The suppressor adds 15 points.

The same K score adds many points to a full-K scale (Pt, Sc) and few to a fifth-K scale (Ma), because the fraction is tuned to how badly a defensive style was found to depress each scale. K carries no direct information about the disorder — it corrects the error that defensiveness introduces.

The K-correction demonstration exposes the suppressor logic. Choosing a clinical scale sets its empirically fixed K fraction, and adjusting the raw clinical score and the K score shows how many points the correction adds and where the corrected score lands. Watching a full-K scale and a fifth-K scale respond differently to the same K value shows that the correction is calibrated scale by scale to how much each is depressed by a defensive style.

Demo 3 — Reading the L, F, K validity configuration
6550L95F45K

Configuration: Over-reporting

F is far above the normative range while K is not elevated — the signature of endorsing many rare items without the defensiveness that would suppress them. Clinical elevations may reflect exaggeration.

The validity-scale demonstration reads the profile before the profile. Setting the L, F, and K scores selects a response-style configuration, and the demonstration classifies it — valid, defensive under-reporting, or symptom over-reporting — with the reasoning that a clinician applies. Moving F up while holding K low tips the profile into over-reporting; raising L and K together tips it into defensiveness, the two distortions the validity scales exist to catch.

Worked Example

Consider first the keying of a single item. Suppose a candidate item is endorsed true by 70% of a criterion group of patients and by 30% of the normative sample, and the constructor's rule is to key any item whose groups differ by at least 20 percentage points. The difference here is 70 − 30 = 40 points, well past the threshold, so the item is keyed onto the scale in the true direction — the direction the criterion group favoured — and it is kept regardless of whether its wording appears to concern the diagnosis. Had the two groups instead endorsed the item at 48% and 42%, the 6-point difference would fall short of the threshold and the item would be discarded, however plainly relevant its content seemed. The decision is made entirely by the gap.

Now apply the K correction. Suppose a respondent's raw score on the Psychasthenia scale is 20, and that scale receives the full K weight of 1.0. If the respondent's K raw score is 15, the correction adds 1.0 × 15 = 15 points, giving a K-corrected raw of 20 + 15 = 35. The same K score applied to a scale weighted at only 0.2 would add just 0.2 × 15 = 3 points. The suppressor's contribution is thus large where defensiveness was found to distort the scale badly and small where it was not, which is the whole point of tuning a separate fraction per scale.

Finally, read the validity configuration. Suppose the same protocol shows an F score at a T of 95 with K at a T of 45. F far above the normative range with K below it is the signature of over-reporting: the respondent has endorsed many rarely-endorsed items while showing none of the defensiveness that would suppress them. Before any clinical scale is interpreted, this configuration warns that the elevations to follow may reflect exaggeration rather than the level of genuine disturbance, and it is exactly why the validity scales are read first.

Discussion

The MMPI earns its place in cognitive psychology as the definitive worked example of empirical test construction — the doctrine that what a test item measures is an empirical question answered by the item's correlates, not a semantic one answered by its content. Hathaway and McKinley did not write scales to measure constructs and then validate them; they let contrasted-group discrimination select the items and accepted whatever scales resulted, however opaque. Meehl's defence of that approach reframed the personality-test item as a sample of verbal behaviour rather than a self-report to be taken at face value, a move whose influence reaches well beyond this one instrument into how psychologists think about measurement itself (Meehl, 1945). The validity scales and the K correction are the same premise carried through: if a response means whatever it predicts, then how a person responds is itself measurable and correctable.

The method's strengths and weaknesses are two sides of one coin. Empirical keying captures real, sometimes counterintuitive, discriminating power that a construct-first approach can miss, and it produced an instrument of remarkable durability and predictive reach. But it also builds in whatever the criterion groups happened to share, spreading a broad demoralization factor across scales and leaving the scales substantially intercorrelated and difficult to interpret in isolation. Auke Tellegen's restructuring of the clinical scales was a direct response: it extracted that pervasive demoralization dimension and rebuilt each clinical scale around its distinctive residual core, aiming for scales that were more discriminant and more transparent about what they measured (Simms et al., 2005). The Restructured Form that followed rebuilt the instrument on that foundation while preserving the validity-scale machinery the original had pioneered (Ben-Porath & Tellegen, 2008).

The instrument's limitations, like its strengths, are legible from its construction. Norms tied to a particular sample age with that sample, which is why the instrument has been re-standardized repeatedly (Butcher et al., 1989). Scales keyed against the psychiatric categories of the 1940s inherit the nosology of their era, and correction weights estimated on one population may not transfer cleanly to another. Each of these is a direct consequence of choosing to build a test from data rather than from theory, and each has been the object of continuing empirical work — the restandardizations that refresh the norms, the restructurings that disentangle the scales, and the validity research that keeps testing how well the instrument distinguishes genuine disturbance from distorted responding. A measure is best understood by seeing what its method of construction bought and what it necessarily cost.

Current Directions

The MMPI remains an intensely active research instrument, and the questions now asked of it cluster around two themes: detecting invalid responding and connecting the scales to contemporary models of psychopathology. The first theme has grown into a substantial quantitative literature on the over-reporting validity scales, whose ability to flag symptom exaggeration is central to the instrument's use in forensic and disability evaluation. Meta-analytic work has pooled that evidence to estimate how well the over-reporting scales separate genuine from feigned presentations and where their thresholds should sit (Ingram & Ternes, 2016), and validation of the newest edition's over-reporting scales in high-stakes forensic samples continues that line (Tylicki et al., 2022).

The second theme places the restructured scales within the dimensional models that increasingly organize clinical psychology. Rather than treating the scales as freestanding empirical predictors, current work maps them onto hierarchical structures of psychopathology, asking how the specific-problems and restructured scales locate within a broader taxonomy of internalizing, externalizing, and thought-disorder dimensions (Sellbom, 2017). Reviews of the Restructured Form have consolidated this reframing, presenting the instrument not as a legacy test but as a modern, structurally-informed assessment of personality and psychopathology whose empirical roots are compatible with a dimensional nosology (Sellbom, 2019). The MMPI's empirical method, in other words, is being rejoined to the theory-building it originally set aside.

Common Misconceptions

MMPI items measure what they appear to ask about.
Often they do not. Items are keyed onto scales by their power to separate diagnostic groups, so an item's scale membership reflects its empirical discrimination, not its surface content (Meehl, 1945).
The validity scales just flag invalid tests for exclusion.
They do more: the K scale is added back to clinical scores as a suppressor correction, and the validity configuration is itself interpreted as a response style before the clinical scales are read (Meehl & Hathaway, 1946).
The numbered clinical scales are single, pure measures of their diagnostic labels.
No. Empirical keying left the scales sharing a broad demoralization factor and substantially intercorrelated, which is exactly what the restructured editions set out to separate (Simms et al., 2005).
The MMPI is scored against the current DSM like a diagnostic checklist.
It is not. Its scales were keyed to the psychiatric nosology of the 1940s and are read as dimensional empirical screens, unlike an instrument built to align with a specific edition of the diagnostic manual (American Psychiatric Association, 2013).

Glossary

Clinical scale.
One of the MMPI's empirically keyed scales, originally numbered rather than named, that scores a dimension of psychopathology from a respondent's item endorsements.

Contrasted-group method.
The construction procedure of comparing a criterion group with a normative group and keeping items that separate them; the operational form of empirical keying.

Criterion group.
A sample of respondents defined by carrying the diagnosis or characteristic a scale is meant to detect, against which candidate items are contrasted.

Demoralization.
A pervasive factor of general distress and unhappiness shared across clinical groups, spread by empirical keying through many MMPI scales and isolated by the restructured editions.

Empirical criterion keying.
The method of selecting test items for a scale by their statistical power to distinguish a criterion group from a comparison sample, regardless of the items' content.

F (Infrequency) scale.
A validity scale of rarely-endorsed items; many endorsements signal unusual responding such as severe disturbance, carelessness, or over-reporting.

K (Correction) scale.
A validity scale capturing a defensive test-taking attitude, a fraction of which is added back to certain clinical scales as a suppressor correction.

L (Lie) scale.
A validity scale of minor faults most people admit; denying them signals a naive effort to appear unusually virtuous.

Over-reporting.
A response style in which a respondent endorses more or more severe symptoms than a genuine presentation typically produces, targeted by the F-family validity scales.

Profile.
The pattern of scores across the MMPI's scales, interpreted as a configuration rather than as isolated single-scale values.

Response style.
The manner in which a respondent approached the test — candid, defensive, exaggerating, or inconsistent — measured by the validity scales before the clinical scales are read.

Restructured Clinical (RC) scales.
Revised clinical scales that remove the shared demoralization factor and rebuild each scale around its distinctive core, introduced in the MMPI-2-RF.

Self-report inventory.
A questionnaire scored from a respondent's own answers about their behaviour, feelings, or attitudes, as opposed to observer ratings or performance tasks.

Suppressor variable.
A measure unrelated to the target construct but related to an error contaminating its measurement, added to a score to remove that error and sharpen prediction.

Uniform T-score.
A standardized score, introduced in the MMPI-2, adjusted so that a given T value marks the same percentile across the clinical scales.

Validity scale.
A scale that assesses how a respondent answered rather than what they reported, used to judge whether and how a profile can be interpreted.

Key Researchers

Yossef S. Ben-Porath. Professor of psychological sciences at Kent State University; lead author of the MMPI-2-RF and the MMPI-3 and a principal architect of the instrument's restructured validity and clinical scales. Faculty Page - Google Scholar

Starke R. Hathaway (1903-1984). Co-creator of the MMPI at the University of Minnesota; with J. C. McKinley he developed the empirical-keying method and built the original clinical and validity scales. Wikipedia

Paul E. Meehl (1920-2003). Minnesota clinical psychologist who articulated the theory behind the MMPI's empirical construction and, with Hathaway, introduced the K scale as a suppressor variable. Wikipedia

Martin Sellbom. Professor of clinical psychology at Monash University; a leading contemporary MMPI researcher mapping the restructured and specific-problems scales onto modern dimensional models of psychopathology. ORCID - Google Scholar - Faculty Page

Auke Tellegen (1930-2024). Personality psychologist at the University of Minnesota; co-author of the MMPI-2 restandardization and architect of the Restructured Clinical scales that separated demoralization from each scale's distinctive core. Wikipedia

Frequently Asked Questions

What does the MMPI measure?
It measures patterns of psychopathology and personality from a respondent's true-or-false self-report, scoring the answers onto empirically keyed clinical scales that are interpreted from their accumulated correlates rather than from item content (Hathaway & McKinley, 1943).

What is empirical criterion keying?
It is the method of selecting an item for a scale because it statistically separated a criterion diagnostic group from a normative sample, keeping the item regardless of what it appears to ask (Hathaway & McKinley, 1940).

Why are the MMPI clinical scales numbered instead of named?
The scales were numbered to discourage reading a score as a literal measure of its diagnostic label, since empirical keying gives each scale a web of correlates that outruns the everyday meaning of the diagnosis (McKinley & Hathaway, 1944).

What are the L, F, and K validity scales?
They measure response style rather than symptoms: L detects a naive effort to look virtuous, F detects unusual or over-reported responding, and K detects a defensive test-taking attitude, so a clinician can judge how a profile should be read (Hathaway & McKinley, 1943).

What is the K correction?
It is the practice of adding a fixed fraction of the K score back to certain clinical scales, treating K as a suppressor variable that captures defensiveness so that removing it sharpens the scales' discrimination (Meehl & Hathaway, 1946).

How does the MMPI differ from a theory-driven inventory?
It builds scales from data rather than from a construct: items are kept for their empirical discrimination, whereas a theory-driven instrument specifies its scales in advance from an account of the trait (Meehl, 1945).

What did the Restructured Form change?
The MMPI-2-RF rebuilt the clinical scales to remove a shared demoralization factor and rebuild each around its distinctive core, aiming for scales that are more discriminant and easier to interpret in isolation (Ben-Porath & Tellegen, 2008).

Is the MMPI still used and researched?
Yes. It remains among the most widely used clinical and forensic instruments, and current work examines its over-reporting validity scales and maps its restructured scales onto dimensional models of psychopathology (Sellbom, 2019).

References

American Psychiatric Association. (2013). Diagnostic and statistical manual of mental disorders (5th ed.). American Psychiatric Association. https://doi.org/10.1176/appi.books.9780890425596

Ben-Porath, Y. S., & Tellegen, A. (2008). MMPI-2-RF (Minnesota Multiphasic Personality Inventory-2 Restructured Form): Manual for administration, scoring, and interpretation. University of Minnesota Press.

Butcher, J. N., Dahlstrom, W. G., Graham, J. R., Tellegen, A., & Kaemmer, B. (1989). MMPI-2: Manual for administration and scoring. University of Minnesota Press.

Hathaway, S. R., & McKinley, J. C. (1940). A multiphasic personality schedule (Minnesota): I. Construction of the schedule. The Journal of Psychology, 10(2), 249-254. https://doi.org/10.1080/00223980.1940.9917000

Hathaway, S. R., & McKinley, J. C. (1943). The Minnesota Multiphasic Personality Inventory (manual). University of Minnesota Press.

Ingram, P. B., & Ternes, M. S. (2016). The detection of content-based invalid responding: A meta-analysis of the MMPI-2-Restructured Form's (MMPI-2-RF) over-reporting validity scales. The Clinical Neuropsychologist, 30(4), 473-496. https://doi.org/10.1080/13854046.2016.1187769

McKinley, J. C., & Hathaway, S. R. (1944). The Minnesota Multiphasic Personality Inventory: V. Hysteria, hypomania, and psychopathic deviate. Journal of Applied Psychology, 28(2), 153-174. https://doi.org/10.1037/h0059245

Meehl, P. E. (1945). The dynamics of 'structured' personality tests. Journal of Clinical Psychology, 1(4), 296-303. https://doi.org/10.1002/1097-4679(194510)1:4%3C296::AID-JCLP2270010410%3E3.0.CO;2-%23

Meehl, P. E., & Hathaway, S. R. (1946). The K factor as a suppressor variable in the Minnesota Multiphasic Personality Inventory. Journal of Applied Psychology, 30(5), 525-564. https://doi.org/10.1037/h0053634

Sellbom, M. (2017). Mapping the MMPI-2-RF specific problems scales onto extant psychopathology structures. Journal of Personality Assessment, 99(4), 341-350. https://doi.org/10.1080/00223891.2016.1206909

Sellbom, M. (2019). The MMPI-2-Restructured Form (MMPI-2-RF): Assessment of personality and psychopathology in the twenty-first century. Annual Review of Clinical Psychology, 15, 149-177. https://doi.org/10.1146/annurev-clinpsy-050718-095701

Simms, L. J., Casillas, A., Clark, L. A., Watson, D., & Doebbeling, B. N. (2005). Psychometric evaluation of the restructured clinical scales of the MMPI-2. Psychological Assessment, 17(3), 345-358. https://doi.org/10.1037/1040-3590.17.3.345

Tylicki, J. L., Gervais, R. O., & Ben-Porath, Y. S. (2022). Examination of the MMPI-3 over-reporting scales in a forensic disability sample. The Clinical Neuropsychologist, 36(7), 1878-1901. https://doi.org/10.1080/13854046.2020.1856414