Abstract

Rorschach Test is a type of ink blot tests: a projective personality measure in which a respondent reports what each of ten symmetrical inkblots might be, and the examiner codes every response for where it was seen, which perceptual feature drove it, how well the form fits, and what was seen. Hermann Rorschach introduced the ten blots in 1921; John Exner's Comprehensive System later standardized administration and scoring, making the test empirically researchable. This article treats the Rorschach as a case study in how a measure earns scientific standing: how the location-determinant-content coding works, how standardization stabilized reliability, how a variable-by-variable meta-analysis separated the scores that predict from the scores that do not, and how the empirically-normed R-PAS now carries the surviving variables forward.

Keywords: Rorschach inkblot test, Comprehensive System, projective assessment

The Rorschach is the founding instrument of the projective tradition and, for most of the twentieth century, its most contested. Ten symmetrical inkblots are shown one at a time, the respondent says what each might be, and the examiner records the answers and then inquires into where each percept was located and what about the blot made it look that way (Rorschach, 1942). What makes the test cognitively interesting is that it scores the perceptual process rather than the story: not what a person sees so much as how they impose structure on an ambiguous field. What makes it scientifically interesting is the long, public struggle to decide whether that process, once coded, measures anything real — a struggle that turned the Rorschach into the clearest worked example in psychology of the difference between a test being standardized, being reliable, and being valid.

Key Takeaways
  • The Rorschach is a projective personality test of ten symmetrical inkblots, scored on the perceptual process rather than the content of what is seen.
  • Each response is coded on three primary dimensions — location (where), determinant (what perceptual feature), and content (what object) — plus a form-quality judgement of perceptual accuracy.
  • Exner's Comprehensive System standardized administration, coding, and norms in the 1970s, converting a family of incompatible scoring schools into a single researchable instrument.
  • Reliability and validity are distinct: standardization and good inter-coder agreement made scores dependable, but whether a given score predicts anything is a separate, harder question.
  • A variable-by-variable meta-analysis found solid support for the structural, perception-based indices and little for many thematic ones; the R-PAS successor keeps the supported variables and adjusts scores for response complexity.

What the Rorschach Test Is

The Rorschach is a performance-based personality test: rather than asking a respondent to describe themselves, it sets a perceptual task and scores how they perform it. The materials are ten inkblots, printed one to a card, each bilaterally symmetrical — five achromatic, two with red detail, three fully chromatic. In the response phase the examiner hands over each card in turn and asks what it might be, writing the answers down verbatim. In the inquiry phase the examiner returns to each response and establishes two things: the location, meaning which part of the blot was used, and the determinant, meaning what feature of that area — its shape, its colour, its shading, or an attributed sense of movement — made it look that way (Rorschach, 1942).

The scored record is therefore not the list of objects a person names but the coding of how each percept was formed. A response of a bat and a response of a butterfly may be coded identically if both use the whole blot, are driven by form alone, and match the blot's contour well; two responses naming the same object may be coded very differently if one is a precise use of the whole blot and the other a vague reading of an odd corner. This is the projective hypothesis applied to perception: given a deliberately unstructured stimulus, the structure a person supplies is taken to reflect something about how they characteristically organize experience (Exner, 2003).

That commitment to scoring process over content is what places the Rorschach in cognitive rather than merely clinical territory. The coding is, in effect, a taxonomy of perceptual decisions — how much of the field to use, which of its properties to be governed by, how closely to hold the percept to the actual contour. Whether those decisions carry stable personality meaning is the empirical question the rest of the test's history is about; that they are perceptual-organization decisions, and can be coded reliably, is not in dispute.

Location, Determinant, and Content

The coding rests on three primary dimensions, and stating them precisely is the key to everything that follows. Location records where on the blot the percept was seen: the whole blot (W), a common and readily delineated detail (D), an unusual or small detail (Dd), or the white space inside or around the blot (S). Determinant records what perceptual feature drove the response: pure form (F) when only the shape is used, human movement (M) when the percept is seen as a person in action, animal or inanimate movement (FM, m), chromatic colour (C, CF, FC by how much form co-determines it), achromatic colour (C'), and the shading determinants of texture, vista, and diffuse shading (T, V, Y). Content records the class of object seen — whole and part human (H, Hd), whole and part animal (A, Ad), anatomy (An), and a long tail of further categories (Exner, 2003).

Cutting across these is form quality, a judgement of perceptual accuracy: how well the reported percept actually fits the contour of the area used, graded from superior through ordinary and unusual to minus, the last reserved for percepts that ignore or violate the blot's real shape. Form quality is the single most cognitively transparent code, because it scores the fit between a perception and the stimulus that occasioned it, and it is the code most directly tied to the structural indices that later validity work would single out (Mihura et al., 2013).

Figure 1. The Rorschach coding pipeline: a free response is coded on three dimensions plus form quality, then aggregated into a structural summary.
From ambiguous blot to structural summary An ambiguous inkblot elicits a free response, which is coded on location, determinant, and content, and judged for form quality; the codes from all responses are aggregated into a structural summary that is the basis of interpretation. Scoring the perceptual process Ambiguous blot free response: what might this be? Location (W / D / Dd / S) where it was seen Determinant (F / M / C ...) what feature drove it Content (H / A / An ...) what object was seen Form quality (+ / o / u / -) how well form fits Structural summary codes from all responses aggregated into ratios and totals (e.g. form-quality proportions, movement and colour balance) the numeric basis of interpretation

Note. Schematic of the coding logic; codes and abbreviations are illustrative of the Comprehensive System, not a complete scoring key.

The codes from every response are then aggregated into a structural summary: proportions, ratios, and totals such as the balance of form-driven to colour-driven responses or the fraction of percepts with adequate form quality. It is this aggregate, not any single answer, that the Comprehensive System interprets — a design that spreads each index across the whole record and is the reason the coding scheme, however elaborate, ultimately behaves like a set of summary scores (Exner, 2003).

Table 1. The primary Rorschach coding dimensions, with representative codes.

Dimension Representative codes What it records
LocationW, D, Dd, SWhich part of the blot was used — whole, common detail, unusual detail, or white space
DeterminantF, M, FM, m, C, CF, FC, T, V, YThe perceptual feature that drove the response — form, movement, colour, or shading
ContentH, Hd, A, Ad, AnThe class of object seen — human, animal, anatomy, and further categories
Form quality+, o, u, -How well the reported percept fits the actual contour of the area used

The Comprehensive System

For its first half-century the Rorschach was not one test but several. Rorschach died in 1922, the year after publication, leaving the method without a settled scoring scheme, and five major systems — Beck, Klopfer, Hertz, Piotrowski, Rapaport-Schafer — grew up around it, differing in what they coded, how they coded it, and how they administered the cards (Exner, 2003). A score meant something different depending on whose manual the examiner followed, which made the accumulating research literature nearly impossible to pool: two studies using the same test could not be compared because they were not, in the operational sense, using the same instrument.

John Exner's Comprehensive System, developed through the 1970s, was the response. Exner compared the five systems empirically and assembled a single scheme from the elements each had best evidence for, specifying administration, coding, and — crucially — a normative reference sample against which an individual record could be read (Exner, 2003). Standardization did for the Rorschach what fixing the response format did for its psychometric cousins: it turned a craft practised in incompatible dialects into one operationalized procedure, so that a coded record produced in one clinic meant the same thing as one produced in another, and a study's findings could in principle be replicated.

Standardization is not the same as validation, and the distinction is the hinge of the whole story. The Comprehensive System made the Rorschach reliable in the specific sense that trained coders applying it to the same record largely agree, and it supplied norms so that scores could be located against a reference distribution. What it could not do by itself was establish that any particular score measured the personality construct it was named for. It made the test researchable; whether the research would vindicate the scores was, in the 1970s, still entirely open (Meyer & Archer, 2001).

Reliability by Standardization

Two kinds of reliability matter for a coded projective test, and the Comprehensive System improved both. The first is inter-coder reliability: whether two examiners independently coding the same responses assign the same codes. Because the System specifies coding criteria explicitly, agreement for the well-defined structural codes is high, and later work formalized the requirement that acceptable field reliability be demonstrated rather than assumed (Meyer & Archer, 2001). Coding agreement is the reliability the Rorschach most clearly earns, because it depends only on the clarity of the rules, not on any claim about what the scores mean.

The second is temporal stability: whether a person's scores hold across a retest interval. A meta-analytic review of Rorschach retest data found that many summary scores show stability coefficients comparable to those of other personality measures, with the more trait-like indices more stable than the state-sensitive ones — a pattern that is exactly what a working instrument should show and that undercut the claim that Rorschach scores were merely noise (Gronnerod, 2003). Stability, like coding agreement, is a genuine property the test can be shown to have.

But reliability bounds validity without supplying it. A score can be coded consistently and remain stable over months and still predict nothing outside the test, because consistency and stability are properties of the measurement, not evidence about what is measured. The Comprehensive System won the reliability argument and thereby sharpened, rather than settled, the real dispute: granting that Rorschach scores are dependable numbers, are they dependable numbers about anything? (Meyer & Archer, 2001)

The Validity Controversy

That question produced one of the most public methodological disputes in modern psychology. Through the 1990s a series of critiques argued that the Comprehensive System's norms were unrepresentative — tending to make ordinary people look maladjusted — and that the evidence for most of its scores predicting external criteria was thin or absent (Wood et al., 1996). A broad review of projective techniques concluded that many widely used scores lacked the empirical support their clinical use implied, and pressed the field to hold the Rorschach to the same validity standards as any other test (Lilienfeld et al., 2000). The critique was pursued at book length, framing the Rorschach as a case study in how a clinical tradition can outrun its evidence (Wood et al., 2003).

Defenders answered on the same empirical ground rather than retreating from it. A comparative meta-analysis found that Rorschach validity coefficients were, on average, of about the same magnitude as those of the MMPI — neither negligible nor large, but real and comparable to an accepted self-report instrument (Hiller et al., 1999). Reviewers argued that the accumulated record, read as a whole rather than through its weakest studies, showed the test performing at the level of other assessment methods (Meyer & Archer, 2001), and the field's professional body issued a formal statement affirming that the instrument had adequate support for a defined range of uses while acknowledging the limits of others (Society for Personality Assessment, 2005).

The dispute was ultimately resolved not for the test as a whole but variable by variable. A systematic meta-analysis estimated the validity of each Comprehensive System variable separately against external criteria and found a clear pattern: the structural, perception-based indices — form quality and the cognitive and perceptual variables closest to it — had solid support, while many thematic and content-based scores had little or none (Mihura et al., 2013). This is the finding that made sense of both camps at once: the critics were right that many scores did not predict, and the defenders were right that some did, and the reason the global debate had been irresolvable was that a single test-level verdict was the wrong unit of analysis.

Rorschach Scoring in Motion

The three demonstrations below make the test's logic manipulable. The first builds the code for a single response, showing how location, determinant, form quality, and content combine into a scored record that ignores the object named. The second lays out the variable-by-variable validity verdict, showing why a single average across scores misrepresents a test whose variables genuinely differ. The third runs the response-complexity confound and the complexity adjustment its modern successor uses to remove it.

Demo 1 — Coding one response, with the object held fixed
Recorded codeW F o ALocationWDeterminantFForm qualityoContentAPerceptual organization8/9well-organized percept

Recorded code: W F o A. Content class A (animal) is held fixed — a bat and a butterfly code identically.

The object named does not enter the code’s meaning. Location (Whole blot), determinant (Pure form), and form quality (ordinary) are what the examiner records, so the very same animal yields a well-organized whole-blot, form-driven, good-fit percept or a poorly-organized space-driven, colour-swept, contour-violating one depending only on how the blot was used. That is the concrete meaning of scoring the perceptual process rather than the content.

The scoring demonstration assembles a single response's code. Choosing a location, a determinant, a form-quality grade, and a content class builds the code string the examiner would record — and holding the named object fixed while the perceptual choices vary shows that the same object yields entirely different codes depending on how the blot was used. Watching a whole-blot, form-driven, good-fit percept and an odd-detail, colour-driven, poor-fit percept of the very same object produce opposite records is the concrete meaning of scoring the process rather than the content.

Demo 2 — Validity, variable by variable
0.300.05020 varsr = 0.3040 varsr = 0.05threshold 0.15test-level average 0.13per-variable validity coefficient

Test-level average 0.13, pooling 60 variables; the threshold of 0.15 selects 20 of them.

The single average (gold) sits well below the supported variables and well above the unsupported ones, describing neither. A third of the variables are as valid as much of personality assessment; the rest carry almost no signal. Averaging over them destroys exactly the information a validity study should recover — which is why the dispute was settled variable by variable, not by any global coefficient.

The validity demonstration sorts a set of representative variables by their level of empirical support and lets a threshold slide across them. Setting the threshold shows how many variables clear it and how the naive test-level average — a single number pooling supported and unsupported scores — sits well below the best variables and well above the worst, misrepresenting both. The demonstration makes visible why the variable-by-variable meta-analysis, not any global coefficient, was what resolved the validity dispute.

Demo 3 — The response-complexity confound, and the adjustment
Raw countAdjusted (ref = 90)A6B8A9B6B scores higherA scores higheradjustment reverses the raw ordering

Raw: A 6, B 8. Adjusted (× 90 ÷ complexity): A 9, B 6.

Complexity partly reflects sheer productivity rather than the trait, so a raw count tracks how much a respondent said as much as what the score is meant to measure — and at the default settings the richer record outscores the sparser one on the raw count while the adjustment reverses it. Rescaling each count to a common reference complexity removes the part of the score that measured only quantity, which is the complexity-adjusted scoring R-PAS introduced.

The complexity demonstration runs the productivity confound the Rorschach has always faced and the adjustment its successor applies. Setting two respondents' raw count of some determinant and their overall response complexity shows the raw count tracking complexity as much as the trait, and even reversing the true ordering, while rescaling each score to a common reference complexity recovers the ordering the trait alone would give. This is the complexity-adjusted scoring the R-PAS system introduced, made concrete.

Worked Example

Begin with the complexity confound the third demonstration runs. Suppose two respondents are scored for some determinant. Respondent A gives a raw count of 6 at an overall response complexity of 60; respondent B gives a raw count of 8 at a complexity of 120, having produced a longer, richer record. On the raw counts B outscores A, 8 to 6, so the ranking favours B. But complexity partly reflects sheer productivity rather than the trait, so rescale each count to a common reference complexity of 90: A becomes 6 × 90 ÷ 60 = 9, and B becomes 8 × 90 ÷ 120 = 6. Adjusted for complexity, A outscores B, 9 to 6, reversing the raw ordering. The adjustment does with arithmetic what a fixed record does by design — it removes the part of the score that measured only how much the respondent said.

Now the reliability payoff of aggregating across items. The Spearman-Brown formula gives the reliability of a test of k comparable items each of reliability r as R = kr ÷ (1 + (k − 1)r). Take a single card's contribution as a modest r = 0.10. The ten Rorschach cards give R = (10 × 0.10) ÷ (1 + 9 × 0.10) = 1.0 ÷ 1.9 ≈ 0.53. To reach a conventional R ≈ 0.80 at the same per-item reliability would take about thirty-six comparable items: solving k × 0.10 ÷ (1 + (k − 1) × 0.10) = 0.80 gives k ≈ 36. This is the quantitative reason the Rorschach interprets aggregated summary scores rather than single responses, and why precise, agreed coding of every response matters so much when there are only ten cards to build from.

Finally, why a single validity number misleads. Suppose a review covers sixty variables, of which twenty genuinely predict an external criterion at r = 0.30 and forty do not, at r = 0.05. The test-level average is (20 × 0.30 + 40 × 0.05) ÷ 60 = (6 + 2) ÷ 60 ≈ 0.13 — a modest figure that makes the whole test look weakly valid, when in fact a third of its variables are as valid as much of personality assessment and the rest carry almost no signal. The average is an artefact of pooling; the variable-by-variable estimate recovers the real structure, which is exactly the correction the meta-analysis supplied (Mihura et al., 2013). (The illustrative counts here are chosen to show the arithmetic, not to report the review's exact figures.)

Discussion

The Rorschach earns its place in cognitive psychology less as a personality test than as the field's longest-running lesson in what it takes for a measure to be believed. Its history separates three things that lay opinion runs together. Standardization, achieved by Exner's Comprehensive System, made the test one procedure rather than five and so made its literature poolable (Exner, 2003). Reliability, both inter-coder agreement and temporal stability, was then demonstrable, because those are properties of clear rules and stable respondents (Gronnerod, 2003). Neither guaranteed validity, which had to be established, and it was established only when the question was asked at the right grain (Mihura et al., 2013).

That grain is the durable methodological point. The decades of test-level argument — is the Rorschach valid, yes or no — were irresolvable because the test is not the kind of thing that has a single validity. It is a bundle of scores that differ in what they measure and how well, and averaging over them destroys the very information a validity study should recover. The variable-by-variable meta-analysis resolved the dispute not by finding new data so much as by adopting the correct unit of analysis, and in doing so it vindicated part of each camp's claim while refuting the global versions of both (Hiller et al., 1999).

The modern settlement is neither the wholesale rejection the harshest critics urged nor the wholesale defence its advocates once mounted. The scores with meta-analytic support are retained and used; the scores without it are set aside; and the difference between the two is now something a clinician can look up rather than argue about. A test that spent fifty years as a symbol of everything unscientific in clinical psychology became, in the end, an unusually clean demonstration of how psychological measurement is supposed to be adjudicated (Society for Personality Assessment, 2005).

Current Directions

The clearest institutional expression of the modern settlement is the Rorschach Performance Assessment System, an empirically-grounded successor built to keep the variables that survived meta-analytic scrutiny and drop those that did not (Meyer et al., 2011). R-PAS makes two changes with direct measurement rationale: it selects variables by their evidence base rather than by tradition, and it addresses the productivity confound by adjusting scores for response complexity, so that a rich, high-complexity record and a sparse one are placed on a comparable footing. Early psychometric work on the system has focused on demonstrating that its raw and complexity-adjusted scores can be coded with acceptable interrater agreement (Pignolo et al., 2017), including in non-patient samples (Kivisalu et al., 2016) — the reliability groundwork any new scoring scheme must lay before validity claims can rest on it.

A second line is methodological, insisting that psychological tests be validated by formal, systematic synthesis rather than by accumulated clinical impression — the standard the Rorschach dispute was finally settled by, now proposed as the general rule (Mihura et al., 2019). A third is process-level, using methods that observe what respondents actually do while responding. Eye-tracking studies relating the visual complexity of a record to measurable cognitive engagement begin to test whether the perceptual variables index the perceptual work they are named for, moving the evidence from correlations with external criteria toward the cognitive mechanism of the response itself (Ales et al., 2020). The common direction is away from defending or dismissing the instrument whole and toward measuring, and mechanistically understanding, its specific scores.

Common Misconceptions

The Rorschach interprets the object a respondent names in the blots.
The named object carries little weight. Scoring is driven by the perceptual process — where on the blot the percept was located, what feature drove it, and how well its form fits — not by the symbolic meaning of a bat or a butterfly (Exner, 2003).
Because it is standardized, the Rorschach is valid.
Standardization and validity are different achievements. Exner's system made the test one reliable procedure, but whether a given score predicts an external criterion had to be established separately, and the answer differs by score (Meyer & Archer, 2001).
Meta-analysis showed the Rorschach is worthless.
It showed the opposite of a single verdict. Estimated variable by variable, the structural, perception-based indices had solid support while many thematic scores did not — a split, not a blanket dismissal (Mihura et al., 2013).
The ten inkblots are random and meaningless.
They are fixed, symmetrical stimuli, unchanged since 1921, chosen for their ambiguity; the test's standardization depends on every respondent facing exactly the same ten cards in the same order (Rorschach, 1942).

Glossary

Comprehensive System.
John Exner's standardized scheme for administering, coding, and norming the Rorschach, assembled in the 1970s from the five earlier systems and dominant in research and practice for decades.

Content.
The Rorschach coding dimension recording the class of object a respondent reports — human, animal, anatomy, and further categories — counted rather than symbolically interpreted.

Determinant.
The perceptual feature of the blot that drives a response — its form, colour, shading, or an attributed sense of movement; one of the three primary Rorschach coding dimensions.

Form quality.
A graded judgement of how well a reported percept fits the actual contour of the blot area used, from superior to minus; the code most directly indexing perceptual accuracy.

Inter-coder reliability.
The degree to which two examiners independently coding the same responses assign the same codes; the reliability the Comprehensive System most clearly secures.

Location.
The Rorschach coding dimension recording which part of the blot was used — the whole blot, a common detail, an unusual detail, or the white space.

Movement (M).
A determinant coded when a percept is seen as engaged in action, most importantly human movement; among the structural variables with the strongest empirical support.

Projective hypothesis.
The premise that responses to an unstructured stimulus are shaped by the perceiver's own dispositions rather than by the stimulus, so the response reveals the person.

R-PAS.
The Rorschach Performance Assessment System, an empirically-grounded successor to the Comprehensive System that retains the variables with meta-analytic support and adjusts scores for response complexity.

Reliability.
The consistency of a measurement across coders, occasions, or forms; a necessary but not sufficient condition for a score to be valid.

Response complexity.
The overall richness and quantity of a Rorschach record; a productivity factor that inflates raw counts and which R-PAS scores adjust for.

Structural summary.
The aggregate of ratios, proportions, and totals computed from the codes of all responses, and the numeric basis on which the Rorschach is interpreted.

Temporal stability.
The consistency of a person's scores across a retest interval; shown by meta-analysis to be comparable, for many Rorschach summary scores, to that of other personality measures.

Validity.
The degree to which a score measures the construct it is meant to; distinct from reliability and, for Rorschach variables, established one variable at a time.

Key Researchers

John E. Exner Jr. (1928-2006). American psychologist who integrated the five competing Rorschach schools into the Comprehensive System (1974), the standardized administration and scoring scheme that made the test empirically researchable and dominated clinical use for three decades. Wikipedia - Wikidata

Scott O. Lilienfeld (1960-2020). Clinical psychologist at Emory University whose 2000 review of projective techniques and later critiques set the evidentiary bar for what Rorschach scores may and may not claim, shaping the modern validity debate. Wikipedia - Wikidata

Gregory J. Meyer. Professor of psychology at the University of Toledo and lead developer of the Rorschach Performance Assessment System (R-PAS), the empirically-normed successor built to retain only the Comprehensive System variables that survived meta-analytic validation. ORCID - Google Scholar - Faculty Page

Hermann Rorschach (1884-1922). Swiss psychiatrist whose Psychodiagnostik (1921) introduced the ten standardized inkblots and a perception-based scoring of location, determinant, and content that bears his name. Wikipedia - Wikidata

James M. Wood. Psychologist at the University of Texas at El Paso whose 1996 critical examination of the Comprehensive System and 2003 book What's Wrong with the Rorschach? drove the reappraisal of the test's norms and validity. ORCID - Faculty Page

Frequently Asked Questions

What is the Rorschach test?
It is a projective personality test in which a respondent reports what each of ten symmetrical inkblots might be, and the examiner codes every response for where it was seen, what perceptual feature drove it, and what was seen, scoring the perceptual process rather than the content (Rorschach, 1942).

How is a Rorschach response scored?
Each response is coded on three primary dimensions (location, determinant, and content) plus a form-quality judgement of how well the percept fits the blot, and the codes from all responses are aggregated into a structural summary of ratios and totals (Exner, 2003).

What is the Comprehensive System?
It is the standardized scheme John Exner assembled in the 1970s from the five earlier Rorschach systems, specifying administration, coding, and norms so that a coded record means the same thing across examiners and studies (Exner, 2003).

Is the Rorschach reliable?
For the well-defined structural codes, yes: trained coders largely agree, and meta-analysis shows many summary scores are as stable over time as other personality measures. Reliability, however, does not by itself establish that a score measures anything (Gronnerod, 2003).

Is the Rorschach valid?
It depends on the variable. A systematic meta-analysis found solid support for the structural, perception-based indices and little for many thematic scores, so validity is a property of individual variables rather than of the test as a whole (Mihura et al., 2013).

Why was the Rorschach so controversial?
Critics argued its norms were unrepresentative and most of its scores unsupported, while defenders pointed to validity comparable to accepted self-report tests; the dispute was irresolvable at the test level because the test's variables genuinely differ in validity (Wood et al., 1996).

What is R-PAS?
The Rorschach Performance Assessment System is an empirically-grounded successor that keeps the variables with meta-analytic support, drops those without it, and adjusts scores for response complexity to remove a productivity confound (Meyer et al., 2011).

Does what I see in the inkblot reveal my personality?
Not in the way popular culture suggests. The object named carries little weight; what is scored is the perceptual process behind the response, and only the scores with demonstrated validity are interpreted (Mihura et al., 2019).

References

Ales, F., Giromini, L., & Zennaro, A. (2020). Complexity and cognitive engagement in the Rorschach task: An eye-tracking study. Journal of Personality Assessment, 102(4), 538-550. https://doi.org/10.1080/00223891.2019.1575227

Exner, J. E. (2003). The Rorschach: A comprehensive system, Volume 1: Basic foundations and principles of interpretation (4th ed.). John Wiley & Sons.

Gronnerod, C. (2003). Temporal stability in the Rorschach method: A meta-analytic review. Journal of Personality Assessment, 80(3), 272-293. https://doi.org/10.1207/S15327752JPA8003_06

Hiller, J. B., Rosenthal, R., Bornstein, R. F., Berry, D. T. R., & Brunell-Neuleib, S. (1999). A comparative meta-analysis of Rorschach and MMPI validity. Psychological Assessment, 11(3), 278-296. https://doi.org/10.1037/1040-3590.11.3.278

Kivisalu, T. M., Lewey, J. H., Shaffer, T. W., & Canfield, M. L. (2016). An investigation of interrater reliability for the Rorschach Performance Assessment System (R-PAS) in a nonpatient U.S. sample. Journal of Personality Assessment, 98(4), 382-390. https://doi.org/10.1080/00223891.2015.1118380

Lilienfeld, S. O., Wood, J. M., & Garb, H. N. (2000). The scientific status of projective techniques. Psychological Science in the Public Interest, 1(2), 27-66. https://doi.org/10.1111/1529-1006.002

Meyer, G. J., & Archer, R. P. (2001). The hard science of Rorschach research: What do we know and where do we go? Psychological Assessment, 13(4), 486-502. https://doi.org/10.1037/1040-3590.13.4.486

Meyer, G. J., Viglione, D. J., Mihura, J. L., Erard, R. E., & Erdberg, P. (2011). Rorschach Performance Assessment System: Administration, coding, interpretation, and technical manual. Rorschach Performance Assessment System.

Mihura, J. L., Meyer, G. J., Dumitrascu, N., & Bombel, G. (2013). The validity of individual Rorschach variables: Systematic reviews and meta-analyses of the Comprehensive System. Psychological Bulletin, 139(3), 548-605. https://doi.org/10.1037/a0029406

Mihura, J. L., Bombel, G., Dumitrascu, N., Roy, M., & Meadows, E. A. (2019). Why we need a formal systematic approach to validating psychological tests: The case of the Rorschach Comprehensive System. Journal of Personality Assessment, 101(4), 374-392. https://doi.org/10.1080/00223891.2018.1458315

Pignolo, C., Giromini, L., Ando, A., Ghirardello, D., Di Girolamo, M., Ales, F., & Zennaro, A. (2017). An interrater reliability study of Rorschach Performance Assessment System (R-PAS) raw and complexity-adjusted scores. Journal of Personality Assessment, 99(6), 619-625. https://doi.org/10.1080/00223891.2017.1296844

Rorschach, H. (1942). Psychodiagnostics: A diagnostic test based on perception (P. Lemkau & B. Kronenberg, Trans.). Hans Huber. (Original work published 1921)

Society for Personality Assessment. (2005). The status of the Rorschach in clinical and forensic practice: An official statement by the Board of Trustees of the Society for Personality Assessment. Journal of Personality Assessment, 85(2), 219-237. https://doi.org/10.1207/s15327752jpa8502_16

Wood, J. M., Nezworski, M. T., & Stejskal, W. J. (1996). The Comprehensive System for the Rorschach: A critical examination. Psychological Science, 7(1), 3-10. https://doi.org/10.1111/j.1467-9280.1996.tb00658.x

Wood, J. M., Nezworski, M. T., Lilienfeld, S. O., & Garb, H. N. (2003). What's wrong with the Rorschach? Science confronts the controversial inkblot test. Jossey-Bass.