Abstract

Psychological phenomena is the highest-level category under which psychology files the observable regularities of mind and behaviour that a theory is obliged to explain. The term is a container, not a construct: it names the class of things to be studied rather than any single thing. This article treats the category itself as the object of analysis, setting out the taxonomy MeSH imposes on the domain and examining the contested question of what kind of thing a phenomenon is — a natural kind, a constructed category, or a causal network. It reviews the three levels at which any phenomenon can be explained, then turns to how one is measured and how the field decides which reported phenomena are real. A worked example computes the statistical power of a replication attempt and the reproducibility rate it implies.

Keywords: psychological phenomena, construct validity, levels of analysis, replication

Psychology does not study one thing. It studies a sprawling and loosely bounded collection of regularities — that attention has limits, that memory is reconstructive, that judgements are anchored, that traits are stable — held together less by a shared subject matter than by a shared method of inquiry. Psychological phenomena is the name for that collection at its most inclusive. Because the category is a container, the interesting questions about it are not about its contents but about its structure: how the domain is carved into subtypes, what sort of entities its members are, at what level they are to be explained, and how a claim that some phenomenon exists is established or overturned. These are the questions on which the health of the science turns, and they are the subject of this article.

Key Takeaways
  • Psychological phenomena is an umbrella category — the class of regularities psychology must explain — rather than a single construct with its own mechanism.
  • What kind of thing a phenomenon is remains contested: essentialist natural kinds, historically constructed categories, and causal networks are the three leading answers, and each implies a different research strategy.
  • Marr's three levels — computational, algorithmic, and implementational — distinguish the questions one can ask about any phenomenon, and confusing them produces spurious disagreement.
  • A phenomenon reaches the evidence only through an operationalisation, so measurement validity, not just statistical significance, governs whether a finding means what it claims.
  • The replication and theory crises are disputes about which reported phenomena are real and how weak theories about them evade correction.

What Psychological Phenomena Are

A psychological phenomenon is a regularity in experience or behaviour that a psychological theory is expected to explain: an effect, a capacity, a disposition, or a process that recurs reliably enough to be named and studied. The category is defined by inclusion rather than by essence. It embraces the momentary (a masked prime shifting a response), the enduring (a personality trait), the normative (a well-functioning mind), and the pathological, unified only by their being states or processes of a psychological subject.

This breadth is not a defect to be tidied away; it reflects a genuine feature of the field. Unlike physics, which can point to a small set of fundamental quantities, psychology inherited its objects from ordinary language and clinical practice, and only later tried to make them precise (Danziger, 1997). The central terms — intelligence, motivation, attention, personality — were names before they were variables, and the work of turning them into measurable constructs is still incomplete. Treating psychological phenomena as a bounded natural domain, rather than as a working collection assembled by history and convenience, is the first mistake to avoid.

Types of Psychological Phenomena

MeSH places Psychological Phenomena at the root of tree F02, the top-level branch of its classification for the psychological sciences, with no parent above it. Its fifteen direct subtypes, listed in Table 1, partition the domain by function rather than by mechanism: a capacity (mental competency), a state (mental health), a class of processes (mental processes), a body of theory (psychological theory), an applied field (applied psychology), and so on. The divisions are orthogonal in principle but overlap in practice — psychomotor performance is also a mental process, resilience is also a facet of mental health — because the tree is an indexing classification built to route literature, not a causal taxonomy of the mind. A term's position records how the National Library of Medicine files work about it, not a claim that its subtypes are mutually exclusive kinds. Only subtypes with a live article on this site are linked.

Table 1. Direct subtypes of Psychological Phenomena in the MeSH classification (tree F02).
Subtype In brief
Mental Competency The capacity to understand information and make reasoned decisions with legal or clinical standing.
Mental Health A state of psychological well-being and effective functioning, not merely the absence of disorder.
Mental Processes The internal operations — attention, memory, reasoning — by which information is transformed.
Parapsychology The study of putative phenomena, such as telepathy, that lie outside established scientific explanation.
Personal Autonomy The capacity for self-governed, independent decision and action.
Psychological Posttraumatic Growth Positive psychological change arising from the struggle with adversity.
Psycholinguistics The mental processes underlying the comprehension, production, and acquisition of language.
Psychological Theory The formal frameworks proposed to explain and predict behaviour and mental life.
Applied Psychology The use of psychological principles to solve practical problems in work, health, and education.
Psychomotor Performance The coordinated cognitive and motor processes underlying skilled physical action.
Psychophysiology The study of how psychological states relate to bodily physiological activity.
Religion and Psychology The intersection of religious belief and practice with psychological functioning.
Psychological Resilience The capacity to adapt successfully and recover in the face of adversity.
Social Theory Frameworks explaining how social structures and interactions shape thought and behaviour.
Theory of Planned Behavior A model predicting intentional behaviour from attitudes, subjective norms, and perceived control.

The Ontology of a Psychological Phenomenon

Grant that a phenomenon such as depression or working memory is real; a harder question remains. What kind of thing is it? Three answers organise the debate, and the choice is not academic — it dictates how the phenomenon should be measured, modelled, and treated (Kendler et al., 2011).

The essentialist view treats a phenomenon as a natural kind with a hidden common cause: the observable symptoms or behaviours are surface indicators of an underlying entity, as a fever indicates an infection. On this account the right model is a latent variable that generates its indicators, and the research task is to find the cause behind them. The constructionist view denies the hidden essence: the category is a useful grouping fixed by history and practice, and its coherence is a fact about classification rather than about the world (Danziger, 1997). The network view offers a third possibility: the phenomenon is not a thing behind its components but the pattern of causal relations among them — insomnia causes fatigue, fatigue causes low mood, low mood worsens insomnia — so the syndrome is a self-sustaining system with no core to discover (Borsboom et al., 2019). This last view has driven a decade of network modelling in psychopathology and beyond (Robinaugh et al., 2020). The demonstration in the next section shows how a single set of correlations can be read under either the latent or the network interpretation.

Levels of Analysis

Even once its ontological status is fixed, a phenomenon can be explained at more than one level, and the levels answer different questions. Marr's canonical distinction separates the computational level (what problem the system solves and why), the algorithmic level (what representations and operations solve it), and the implementational level (how those operations are physically realised in the brain) (Marr, 1982). A complete account addresses all three, but a claim pitched at one level neither confirms nor refutes a claim at another: a computational theory of face recognition is not threatened by ignorance of its neural substrate, and a neural finding does not by itself adjudicate between competing algorithms.

Confusing the levels is a reliable source of spurious disagreement. Debates about whether a phenomenon is really cognitive or really neural often dissolve once it is noticed that the parties are describing the same system at different levels. The demonstration below lets the reader select a phenomenon and see the three levels stated side by side.

Computational levelWhat & why?
Map a retinal image to the identity of a known person, invariant to pose, lighting, and expression.
Algorithmic levelHow, in principle?
Encode a face as a point in a multidimensional 'face space' relative to an average face, then match to stored exemplars.
Implementational levelHow, in the brain?
Selective responses in the fusiform face area and a distributed occipitotemporal face-processing network.

The same phenomenon is fully described only across all three levels; a claim at one level neither confirms nor refutes a claim at another.

The levels framework also disciplines the ontology debate of the previous section: the latent-variable and network models are competing algorithmic-level accounts of the same computational problem, which is why they can be compared on how well each reproduces the observed pattern of associations rather than by appeal to which is more fundamental.

Measuring a Psychological Phenomenon

A phenomenon reaches the evidence only through an operationalisation — a concrete procedure that turns the construct into a number. This is where most of the slippage between claim and finding occurs. Two research traditions historically pulled in opposite directions: an experimental tradition that manipulates variables and a correlational tradition that measures individual differences, each studying the same phenomena by an incompatible method and rarely speaking to the other (Cronbach, 1957). Bridging them requires that a measure actually capture the construct it names, a property called construct validity — established, on the classic account, not by any single correlation but by the measure's place in a whole network of theoretical expectations (Cronbach & Meehl, 1955).

Validity is routinely assumed rather than demonstrated. A survey of measurement practice finds that researchers frequently invent or modify scales without reporting any evidence that they measure the intended construct, so a statistically robust effect may be an artefact of an invalid instrument (Flake & Fried, 2020). The problem compounds at the point of inference: a verbal hypothesis about a phenomenon is far broader than the single operationalisation used to test it, so a significant result licenses a much narrower conclusion than is usually drawn (Yarkoni, 2022). Whether the phenomenon is modelled as a latent common cause or as a network of interacting parts changes what a good measurement even looks like, as the demonstration makes concrete (Borsboom et al., 2021).

LX1X2X3
r(X1, X2)0.36r(X2, X3)0.36r(X1, X3)0.36

A common cause forces all three pairwise correlations to be equal — the signature of a latent variable.

Which Phenomena Are Real

The final question is the most consequential: given a reported effect, is the phenomenon real? For a decade psychology has answered this by attempting large-scale replication, and the results reset expectations across the field. A coordinated replication of one hundred published studies reproduced the original significant result in only about a third to a half of cases, depending on the criterion, forcing a distinction between phenomena that are robust and those that were statistical accidents (Open Science Collaboration, 2015). The response was to make direct replication a routine part of the science rather than a rare and unwelcome intrusion (Zwaan et al., 2018).

Replication is powerful because it converts an ontological question into a statistical one: a real phenomenon, studied with adequate power, reappears; a spurious one does not. But the inference is only as good as the power of the replication attempt, and much of the original literature was badly underpowered. The demonstration below, and the worked example that follows, make the link between statistical power and the expected reproducibility rate explicit.

Statistical power0.8043%Expected significant replications out of 100: 43

Underpowered: even a real phenomenon would fail to replicate often, so a low reproducibility rate is expected and uninformative.

Worked Example

Consider a phenomenon whose true effect, expressed as a standardised mean difference, is d = 0.4 — a modest but genuine effect. A replication team runs a two-group study with n = 40 participants per group. What is the probability that they obtain a significant result at the conventional two-tailed threshold of α = 0.05, and what does that imply for a set of such replications?

The non-centrality parameter for a two-sample t test is

δ = d × √(n/2) = 0.4 × √(40/2) = 0.4 × √20 ≈ 0.4 × 4.472 = 1.789

Using the normal approximation to the t distribution, the critical value at α = 0.05 (two-tailed) is z = 1.96. Statistical power is the probability that the observed statistic exceeds this critical value given the true effect:

Power = Φ(δ − 1.96) = Φ(1.789 − 1.96) = Φ(−0.171) ≈ 0.432

So a single replication has only about a 43% chance of success even though the phenomenon is real. Across a batch of, say, 100 such replications the expected number of significant results is

100 × 0.432 ≈ 43

A 43% reproducibility rate would, on this reasoning, be exactly what one should expect for a body of real effects studied at d = 0.4 with 40 participants per group — which is close to the rate the large replication projects actually observed. The lesson is that a low replication rate does not by itself prove the original phenomena were illusory; underpowered replication of genuine effects produces the same pattern. Figure 1 plots power against sample size for this effect, showing how far the field was from the 80% power that would make a failed replication genuinely informative.

Figure 1

Statistical power of a two-group replication as a function of per-group sample size, for a true effect of d = 0.4 at α = 0.05 (two-tailed, normal approximation).

Power curve for d = 0.4 Power rises with per-group sample size, passing the worked-example point of about 0.43 at n = 40 and reaching 0.80 only near n = 100. Per-group sample size n Power 0.80 n = 40, power ≈ 0.43
Note. Power is computed as Φ(d√(n/2) − 1.96). The 0.80 reference line is reached only near n = 100 per group, so the worked example's n = 40 study is substantially underpowered for a d = 0.4 effect.

Discussion

The four questions of this article — how the domain is classified, what kind of thing a phenomenon is, at what level it is explained, and how its reality is established — are not independent. The ontology one adopts fixes the appropriate measurement model; the measurement model determines what counts as evidence; and the evidential standard decides which phenomena survive. Much apparent disagreement in psychology is really a disagreement at one of these joints mistaken for a dispute about facts. A recurring diagnosis holds that the field's difficulties are downstream of weak theory: because theories of psychological phenomena rarely make risky, precise predictions, they are hard to falsify and easy to defend against inconvenient data (Meehl, 1978). Modern restatements agree, arguing that psychology accumulates isolated effects without the integrating theoretical framework that would let them constrain one another (Muthukrishna & Henrich, 2019).

The constructive reading is that the replication crisis was a symptom, not the disease. Improving statistical power and pre-registration addresses the reliability of individual findings, but the deeper problem is that a phenomenon defined loosely enough to be measured many incompatible ways cannot be decisively confirmed or refuted by any of them. Progress on which phenomena are real depends on progress in saying clearly what they are.

Current Directions

The most active response reframes psychology's central difficulty as a theory crisis rather than a replication crisis: the field lacks formal, quantitatively specified theories, and until phenomena are defined precisely enough to yield point predictions, more data will not settle disputes about them (Eronen & Bringmann, 2021). This has prompted calls to spend less effort testing under-specified hypotheses and more on the descriptive, exploratory work of characterising phenomena before modelling them (Scheel et al., 2021).

Two methodological programmes carry the agenda forward. Formal-theory approaches build explicit computational or mathematical models whose predictions can be compared and falsified, moving beyond verbal theories whose vagueness has been the recurring complaint. Network psychometrics, meanwhile, has matured into a standard toolkit for modelling phenomena as systems of interacting variables, with agreed estimation methods and accuracy diagnostics (Borsboom et al., 2021). Whether these programmes converge on a more cumulative science, or simply add rigour to a still-fragmented catalogue of effects, is the open question of the coming decade.

Common Misconceptions

A psychological phenomenon is a single well-defined entity.
Many are better understood as historically constructed categories or as networks of interacting components rather than as unitary things with a hidden essence (Kendler et al., 2011).
A statistically significant effect establishes that a phenomenon is real.
Significance from a single, possibly invalid operationalisation licenses a far narrower claim than the verbal phenomenon it is taken to support, and may not replicate (Yarkoni, 2022).
A failed replication proves the original phenomenon was illusory.
An underpowered replication of a genuine effect fails at a predictable rate; the reproducibility rate must be read against the power of the attempts, not taken at face value (Open Science Collaboration, 2015).

Glossary

Algorithmic level.
In Marr's framework, the level of analysis specifying the representations and operations by which a system solves its computational problem.
Computational level.
In Marr's framework, the level specifying what problem a system solves and why, independent of how it is solved.
Construct validity.
The degree to which a measurement procedure actually captures the construct it is intended to measure.
Construct.
A theoretical attribute — such as intelligence or anxiety — that is not observed directly but inferred from measurable indicators.
Essentialism.
The view that a phenomenon is a natural kind with a hidden common cause underlying its observable indicators.
Implementational level.
In Marr's framework, the level specifying how a system's operations are physically realised, as in neural hardware.
Latent variable.
An unobserved common cause posited to generate the correlations among a set of measured indicators.
Natural kind.
A category whose members share a real, mind-independent essence that supports scientific generalisation.
Network model.
A representation of a phenomenon as a system of directly interacting components rather than as indicators of a common cause.
Operationalisation.
The concrete procedure that turns an abstract construct into an observable, quantifiable measurement.
Psychological phenomenon.
A regularity in experience or behaviour that a psychological theory is expected to explain.
Replication.
A repeat of a study with a new sample to test whether its reported effect reappears.
Statistical power.
The probability that a study detects an effect of a given size when that effect is genuinely present.
Theory crisis.
The argument that psychology's difficulties stem from theories too vague to yield the risky predictions that would let data adjudicate them.

Key Researchers

Denny Borsboom (b. 1973). Developed the network theory of psychological phenomena, treating constructs as systems of interacting components rather than latent causes. ORCID - Google Scholar - Wikipedia - Wikidata

Lee J. Cronbach (1916-2001). Diagnosed psychology's split into experimental and correlational disciplines and shaped the theory of construct validity. Wikipedia - Wikidata

Kurt Danziger (1926-2026). Historian of psychology who showed that its central categories are historically constructed rather than natural kinds discovered in the world. Wikipedia - Wikidata

Eiko I. Fried. Works on the measurement and network modelling of psychological phenomena, documenting how routinely their operationalisations go unvalidated. ORCID - Google Scholar - Wikidata

Kenneth S. Kendler. Psychiatric epistemologist who analyses what kind of things psychological and psychiatric phenomena are, contrasting essentialist and mechanistic accounts. ORCID - Google Scholar - Wikipedia - Wikidata

Daniel Lakens. Methodologist who argues for more descriptive, exploratory research to establish which psychological phenomena are real before confirmatory testing. ORCID - Google Scholar - Wikipedia - Wikidata

David Marr (1945-1980). Introduced the three levels of analysis — computational, algorithmic, and implementational — at which any psychological phenomenon can be explained. Wikipedia - Wikidata

Paul E. Meehl (1920-2003). Critiqued the weak theory-testing practices that leave psychological phenomena poorly explained and resistant to falsification. Wikipedia - Wikidata

Frequently Asked Questions

What exactly counts as a psychological phenomenon? Any regularity in experience or behaviour that a psychological theory is expected to explain, such as an effect, a capacity, a disposition, or a process, that recurs reliably enough to be named and studied. It is an umbrella category rather than a single construct (Danziger, 1997).

How are psychological phenomena classified? MeSH places them at the root of its F02 tree and divides them into fifteen functional subtypes, from mental processes to applied psychology. The classification indexes literature rather than asserting mutually exclusive natural kinds.

Is a psychological phenomenon a real thing or just a label? It depends on the phenomenon and on one's theory of it. The essentialist, constructionist, and network views give different answers, and the choice determines how the phenomenon should be measured and modelled (Kendler et al., 2011).

What are Marr's levels of analysis? Three distinct levels at which any phenomenon can be explained: the computational (what problem is solved), the algorithmic (what representations and operations solve it), and the implementational (how it is physically realised) (Marr, 1982).

Why does measurement matter so much? Because a phenomenon reaches the evidence only through an operationalisation. If the measure lacks construct validity, even a statistically robust effect may not mean what it is taken to mean (Flake & Fried, 2020).

What was the replication crisis? A large replication project found that only about a third to a half of published psychological effects reproduced, prompting a distinction between robust phenomena and statistical accidents and making direct replication routine (Open Science Collaboration, 2015).

Does a failed replication mean the effect was fake? Not necessarily. An underpowered replication of a genuine effect fails at a predictable rate, so the reproducibility rate must be interpreted against the statistical power of the attempts (Zwaan et al., 2018).

What is the theory crisis? The argument that psychology's deeper problem is not unreliable data but vague theory: until phenomena are defined precisely enough to yield risky predictions, more data cannot decisively settle disputes about them (Eronen & Bringmann, 2021).

References

Borsboom, D., Cramer, A. O. J., & Kalis, A. (2019). Brain disorders? Not really: Why network structures block reductionism in psychopathology research. Behavioral and Brain Sciences, 42, e2. https://doi.org/10.1017/S0140525X17002266

Borsboom, D., Deserno, M. K., Rhemtulla, M., Epskamp, S., Fried, E. I., McNally, R. J., Robinaugh, D. J., Perugini, M., Dalege, J., Costantini, G., Isvoranu, A.-M., Wysocki, A. C., van Borkulo, C. D., van Bork, R., & Waldorp, L. J. (2021). Network analysis of multivariate data in psychological science. Nature Reviews Methods Primers, 1, 58. https://doi.org/10.1038/s43586-021-00055-w

Cronbach, L. J. (1957). The two disciplines of scientific psychology. American Psychologist, 12(11), 671-684. https://doi.org/10.1037/h0043943

Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302. https://doi.org/10.1037/h0040957

Danziger, K. (1997). Naming the mind: How psychology found its language. Sage Publications. ISBN 9780803977631.

Eronen, M. I., & Bringmann, L. F. (2021). The theory crisis in psychology: How to move forward. Perspectives on Psychological Science, 16(4), 779-788. https://doi.org/10.1177/1745691620970586

Flake, J. K., & Fried, E. I. (2020). Measurement schmeasurement: Questionable measurement practices and how to avoid them. Advances in Methods and Practices in Psychological Science, 3(4), 456-465. https://doi.org/10.1177/2515245920952393

Kendler, K. S., Zachar, P., & Craver, C. (2011). What kinds of things are psychiatric disorders? Psychological Medicine, 41(6), 1143-1150. https://doi.org/10.1017/S0033291710001844

Marr, D. (1982). Vision: A computational investigation into the human representation and processing of visual information. W. H. Freeman and Company. ISBN 9780716712848.

Meehl, P. E. (1978). Theoretical risks and tabular asterisks: Sir Karl, Sir Ronald, and the slow progress of soft psychology. Journal of Consulting and Clinical Psychology, 46(4), 806-834. https://doi.org/10.1037/0022-006X.46.4.806

Muthukrishna, M., & Henrich, J. (2019). A problem in theory. Nature Human Behaviour, 3(3), 221-229. https://doi.org/10.1038/s41562-018-0522-1

Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716

Robinaugh, D. J., Hoekstra, R. H. A., Toner, E. R., & Borsboom, D. (2020). The network approach to psychopathology: A review of the literature 2008-2018 and an agenda for future research. Psychological Medicine, 50(3), 353-366. https://doi.org/10.1017/S0033291719003404

Scheel, A. M., Tiokhin, L., Isager, P. M., & Lakens, D. (2021). Why hypothesis testers should spend less time testing hypotheses. Perspectives on Psychological Science, 16(4), 744-755. https://doi.org/10.1177/1745691620966795

Yarkoni, T. (2022). The generalizability crisis. Behavioral and Brain Sciences, 45, e1. https://doi.org/10.1017/S0140525X20001685

Zwaan, R. A., Etz, A., Lucas, R. E., & Donnellan, M. B. (2018). Making replication mainstream. Behavioral and Brain Sciences, 41, e120. https://doi.org/10.1017/S0140525X17001972