Abstract
A psychological theory is a system of concepts and propositions that explains how mental and behavioural phenomena are produced and predicts how they behave under new conditions. In the Medical Subject Headings vocabulary it is a kind of psychological phenomenon, descriptor D011582. This article distinguishes a theory from a description or a statistical model, surveys the families of theory the vocabulary enumerates, and sets out the standards by which theories are judged: falsifiability, construct validity and the nomological network, paradigms and research programmes, and levels of analysis. It closes on the theory crisis, the recognition that psychology's difficulty in replicating findings is downstream of a difficulty in building theories precise enough to be decisively tested. Three interactive demonstrations let the reader vary a prediction's boldness, assemble a nomological network, and steer a research programme between the progressive and the degenerating.
Keywords: psychological theory, falsifiability, construct validity, nomological network, theory crisis
A psychological theory is a structured set of statements that explains a class of mental or behavioural phenomena by specifying the entities involved, the relations among them, and the mechanisms through which they operate, in a form that yields testable predictions (Popper, 2002). It is more than a summary of what has been observed and more than an equation that fits a curve: a theory says why the pattern holds, and in doing so commits itself to consequences that could turn out to be false. In the Medical Subject Headings vocabulary, Psychological Theory is descriptor D011582, filed at tree position F02.739 as a narrower kind of Psychological Phenomena (D011579); the vocabulary further divides it into six narrower descriptors, an indexing taxonomy this article treats as a catalogue of historically influential schools rather than a claim about how theories should be carved up today. The construct sits at the foundation of the discipline because everything downstream depends on it: what a study measures, which hypotheses count as risky, and whether an accumulation of findings amounts to understanding are all decided by the quality of the theory behind them (Meehl, 1978; Muthukrishna & Henrich, 2019). The sections below set out what makes a system of ideas a theory, the families the vocabulary recognises, the criteria by which theories are appraised, and the modern argument that psychology has too few good ones.
- A theory explains and predicts; it is not a restatement of data or a model that merely fits. Its value lies in the consequences it forbids.
- Falsifiability is the classic criterion of demarcation: a scientific theory rules out possible observations, and the more it rules out, the more a surviving test corroborates it.
- Because psychological constructs cannot be observed directly, a theory earns construct validity by embedding each construct in a nomological network of lawful relations to other constructs and to measurements.
- Whole theoretical frameworks change less by isolated refutation than through paradigms and research programmes, which can be progressive or degenerating.
- The modern theory crisis holds that weak, verbally stated theories, not weak methods alone, are why psychological findings so often fail to replicate.
What a Psychological Theory Is
The defining work of a theory is explanation, and explanation is more than description. A curve fitted to reaction-time data describes how response time grows with the number of items to be searched; a theory of visual search explains that growth by positing a mechanism, a serial or parallel comparison process, from which the curve follows as a consequence. The difference matters because only the explanatory version makes commitments beyond the data in hand: it says what should happen when the display changes, when attention is loaded, when the observer is given more time, and each of those commitments is a way the theory could be wrong (Popper, 2002). A description that only summarises the observed cannot be wrong about anything it has not yet seen, and so it cannot be tested.
This is why the currency of a good theory is what it forbids. A theory that is compatible with every possible outcome explains nothing, because no observation could count against it; a theory that is compatible with only a narrow band of outcomes stakes a great deal on the world turning out one particular way (Meehl, 1978). Paul Meehl made this the centre of his critique of soft psychology: the standard practice of predicting merely that two groups will differ, or that a correlation will be non-zero, asks the theory to forbid almost nothing, because with enough participants some difference is nearly always found regardless of whether the theory is true. A theory that instead predicts a specific numerical value, or the precise shape of a function, exposes itself to refutation in a way that a directional prediction never does, and a specific prediction that survives is correspondingly more impressive. The first demonstration below makes this trade-off between boldness and safety concrete.
A theory is also more than a model. A model is a formal or computational object, a set of equations or a simulation, that can be fitted to data; a theory is the account of the world that the model is supposed to express. The distinction is easy to lose because a good theory is often stated as a model, but a model can fit well while expressing no coherent theory, and a theory can be illuminating while resisting formalisation (Guest & Martin, 2021). The healthiest relationship runs in one direction: the theory motivates the model, the model forces the theory to be explicit, and the discipline of making a verbal theory computational routinely reveals that it was never precise enough to generate the predictions its authors believed it made (Robinaugh et al., 2021).
Types of Psychological Theory
The Medical Subject Headings vocabulary treats Psychological Theory (D011582) as a parent heading with six narrower descriptors, listed in Table 1. These are best read as the historically dominant schools of psychological explanation rather than as a principled partition: the categories overlap, they were coined at different moments and for different purposes, and a modern theory of a given phenomenon will often draw on several at once or belong to none of them. The list is also an indexing classification, built to catalogue the literature, so its boundaries reflect bibliographic convenience as much as conceptual structure. Only one of the six, theory of mind, is at present a live topic on this site; the remainder are named here with their vocabulary glosses and will be linked as their articles are built.
Table 1
The Narrower Descriptors of Psychological Theory in MeSH (F02.739)
| School | Tree number | Core commitment |
|---|---|---|
| Behaviorism | F02.739.138 | Explains behaviour through learned associations between stimuli and responses, treating inner states as outside the scope of a scientific account. |
| Existentialism | F02.739.418 | Locates the springs of behaviour in the person's confrontation with freedom, meaning, and finitude rather than in mechanism. |
| Gestalt Theory | F02.739.527 | Holds that perception and thought are organised into wholes whose properties are not reducible to their parts. |
| Personal Construct Theory | F02.739.660 | Casts the person as a scientist who anticipates events through a private system of bipolar constructs. |
| Psychoanalytic Theory | F02.739.794 | Explains conduct through unconscious conflict among dynamic mental agencies shaped by early experience. |
| Theory of Mind | F02.739.897 | The capacity to attribute mental states to oneself and others, itself a topic of theory as well as a name for a school. |
Note. Glosses are tightened from each descriptor's MeSH scope note. The categories are not mutually exclusive and do not exhaust modern theory; cognitive psychology's own explanatory frameworks, such as information-processing and computational theories, are not among these historical headings at all.
Falsifiability and Demarcation
The question of what separates a scientific theory from one that only looks scientific was given its most influential answer by Karl Popper: a theory is scientific to the extent that it is falsifiable, that it forbids some observations and so could in principle be refuted (Popper, 2002). Popper's target was theories that seemed to explain everything, that could accommodate any outcome after the fact, and his diagnosis was that their apparent strength, their universal applicability, was in fact their fatal weakness: a claim consistent with every possible observation says nothing about which observations to expect. On this view the growth of knowledge is not the accumulation of confirmations but the survival of bold conjectures through severe attempts to refute them, and a theory that has survived tests it was likely to fail is corroborated, a status Popper carefully distinguished from proof.
The practical force of falsifiability in psychology is a matter of degree rather than a yes-or-no verdict, and this is where Meehl's contribution sharpens Popper's. A theory's testability depends on how much it forbids, which depends in turn on how precise its predictions are. Predicting that an experimental group will differ from a control group in a stated direction forbids only half of the possible outcomes; predicting a specific point value, or the ordering of several conditions, or the exact form of a dose-response curve, forbids almost all of them (Meehl, 1978). Because null-hypothesis significance testing rewards the weak directional prediction, the ritual of finding a significant difference can proceed for decades without ever subjecting a theory to a genuinely risky test. John Platt made the constructive counterpart of this argument in his account of strong inference: the fastest-moving sciences are those that frame competing hypotheses so that a single crucial experiment can exclude at least one of them, and then actually run it (Platt, 1964). The demonstration below lets the reader set how much of the outcome space a prediction rules out and see how the corroboration from a successful test rises as the prediction grows bolder.
Falsifiability: what a prediction forbids
A theory’s prediction carves a permitted band out of a 100-unit outcome space. The narrower the band, the more the theory forbids and the more a passed test corroborates it. Falsifiability is F = 1 − W/100. The observed result is fixed at 50; a prediction passes when 50 falls inside its band.
Illustrative, not measured: the arithmetic (p = W/100, F = 1 − p) is computed locally in your browser and nothing is stored. A confirmed narrow band earns more than a confirmed wide one, because passing it was less probable in advance — the point Meehl pressed against weak directional predictions.
Construct Validity and the Nomological Network
Psychology faces a difficulty the physical sciences largely escape: its theoretical terms name things that cannot be observed directly. Anxiety, working-memory capacity, extraversion, and intelligence are constructs, posited inner properties inferred from patterns of behaviour, and a theory that invokes them must somehow connect them to what can be measured. Lee Cronbach and Paul Meehl gave the classic account of how this connection licenses a construct's use: a construct earns its keep by being embedded in a nomological network, an interlocking system of lawful propositions relating the construct to other constructs and to observable measures (Cronbach & Meehl, 1955). To validate a measure of a construct is not to check it against a single criterion but to show that its pattern of relations to other measures matches what the network predicts: measures of the same construct should converge, measures of different constructs should diverge, and the construct should relate to its theoretical neighbours in the ways the theory says it must.
This reframing dissolved the idea that validity is a property a test either has or lacks, replacing it with the ongoing project of testing the theory in which the construct is embedded. Evidence that a supposed measure of anxiety correlates with physiological arousal, rises under threat, and predicts avoidance is simultaneously evidence for the validity of the measure and for the theory of anxiety, because the two cannot be tested separately (Cronbach & Meehl, 1955). The corollary is uncomfortable: when a measure behaves unexpectedly, the fault may lie in the measure, in the theory, or in the auxiliary assumptions linking them, and disentangling these is rarely straightforward. Modern critics argue that psychology's constructs are often too loosely specified for a genuine nomological network to be built at all, so that construct validation becomes a matter of accumulating suggestive correlations rather than testing a stated law (Fried, 2020). The second demonstration lets the reader populate a small nomological network with convergent and discriminant relations and read off whether the resulting pattern supports the construct.
Building a nomological network
Two measures (A and B) are meant to tap the same construct, so they should converge. A third measure taps a different construct, so our construct should diverge from it. The construct is supported only when the same-construct correlation is high and clearly exceeds the cross-construct one.
The pattern supports the construct: its indicators converge and stand apart from a different construct.
Illustrative, not measured: correlations are set by you and the verdict is computed locally in your browser, nothing stored. The thresholds (convergent r ≥ 0.50, separation ≥ 0.30) stand in for the judgement a real validation study makes; the logic — converge on the same, diverge from the different — is Cronbach and Meehl’s.
Paradigms and Research Programmes
Popper's picture of theories falling to single refutations faces a logical difficulty as well as a historical one. The Duhem-Quine thesis holds that a hypothesis is never tested in isolation but only in conjunction with a body of background assumptions, so a failed prediction shows that something in the whole bundle is false without saying what; the theory of interest can always be rescued by revising an auxiliary assumption instead (Quine, 1951). This underdetermination is why a refutation rarely falls cleanly on its intended target, and two philosophers supplied the correction by shifting attention from the isolated theory to the framework around it. Thomas Kuhn argued that mature science is organised around paradigms, shared frameworks of assumptions, exemplary problems, and accepted methods within which most research, the normal science of puzzle-solving, proceeds without questioning the framework itself (Kuhn, 2012). Anomalies that resist solution accumulate until the paradigm enters crisis and is eventually replaced in a scientific revolution, after which the discipline sees its subject matter through a new lens. Crucially, an anomaly does not by itself overthrow a paradigm; scientists rightly tolerate unexplained findings so long as the framework remains fruitful, because no framework explains everything and abandoning one prematurely would forfeit its successes.
Imre Lakatos reconciled Kuhn's history with Popper's rationalism by shifting the unit of appraisal from the single theory to the research programme (Lakatos, 1970). A programme has a hard core of central assumptions its adherents refuse to give up, surrounded by a protective belt of auxiliary hypotheses that absorb the impact of failed predictions and can be modified to fit new data. What distinguishes science from pseudoscience is not whether the belt is adjusted, since all programmes adjust it, but whether the adjustments are progressive or degenerating. A progressive programme's modifications predict novel facts that are subsequently confirmed, so the programme's empirical content grows; a degenerating programme's modifications only accommodate what has already been observed, adding nothing new and merely protecting the core from refutation. This gives a workable criterion for a discipline whose theories are rarely abandoned outright: judge a programme by whether it keeps predicting new things that turn out to be true. The third demonstration lets the reader run a programme through a series of anomalies, choosing at each step between a bold novel prediction and an ad-hoc patch, and watch the programme's content grow or wither in response.
Steering a research programme
Your programme has a fixed hard core. Each anomaly can be met with a bold novel prediction — which, when confirmed, grows the programme’s empirical content — or with an ad-hoc patch that only shields the core and predicts nothing new. Lakatos’s test is whether the content keeps growing.
Anomaly 1 of 5: A predicted effect fails to appear in a new population.
Illustrative, not measured: the anomaly sequence is fixed and the content score is computed locally in your browser, nothing stored. A bold prediction adds +6 and a patch −3 — stand-in weights for Lakatos’s real distinction between a programme that predicts novel facts and one that merely accommodates known ones.
Levels of Analysis
A further question about any psychological theory is not whether it is true but at what level it is pitched, because the same phenomenon can be explained in several complementary ways that do not compete. David Marr drew the canonical distinction for information-processing theories, separating three levels at which any system that computes can be understood (Marr, 2010). The computational level specifies what problem the system solves and why, the abstract goal and the logic of the mapping from input to output; the algorithmic level specifies the representations and procedures by which the system solves it; and the implementational level specifies how those representations and procedures are physically realised in neural tissue. A theory of reading, on this scheme, might state that the goal is to recover meaning from marks on a page (computational), that this is achieved by mapping letter strings onto stored word forms and thence onto meanings (algorithmic), and that the mapping is carried out by particular cortical circuits (implementational).
Marr's insight was that these levels are loosely coupled: a single computational theory can be realised by different algorithms, and a single algorithm by different hardware, so confusion follows when an explanation at one level is mistaken for a rival to an explanation at another. A dispute about whether a behaviour is produced by rule-following or by association is an algorithmic dispute; a dispute about which brain region implements it is an implementational one; and neither settles the computational question of what the behaviour is for. The framework remains a discipline against a recurring error in psychological theorising, the collapse of distinct questions into one, and it clarifies why neuroscientific findings, however detailed, do not automatically adjudicate between cognitive theories pitched at a higher level (McGuire, 1997). Good theory building, in Marr's spirit, begins by being explicit about which question is being answered before proposing a mechanism to answer it (van Rooij & Baggio, 2021).
Figure 1
Marr's Three Levels of Analysis for an Information-Processing System
Note. The three levels are loosely coupled: a single computational theory can be realised by many algorithms, and a single algorithm by different neural hardware, so an explanation at one level does not by itself settle a question at another (Marr, 2010).
Worked Example
The claim that a bolder prediction earns more from a successful test can be made exact, and doing so shows why Meehl thought the directional prediction so weak. Imagine a measured effect that, on everything known before the test, could plausibly fall anywhere in a range of 100 units; call this the outcome space. A theory's prediction carves out a permitted band within that space, and the width of the band determines how much the theory risks. The prior probability that a result would land in the permitted band by chance alone is the band's width divided by the total, p = W/100, and the theory's falsifiability, the fraction of outcomes it forbids, is F = 1 − p (Popper, 2002; Meehl, 1978).
Consider two theories tested against the same observed result of 50. Theory A makes a sharp point prediction, that the effect lies between 48 and 52, a permitted band of width 4; its prior probability of passing by chance is p = 4 ÷ 100 = 0.04, so its falsifiability is F = 1 − 0.04 = 0.96. Theory B makes only a loose prediction, that the effect lies somewhere between 20 and 80, a band of width 60; its prior probability of passing is p = 60 ÷ 100 = 0.60 and its falsifiability is F = 1 − 0.60 = 0.40. Both predictions are confirmed, because 50 falls inside both bands, and a significance test would register a 'success' for each. But the successes are not equal: Theory A survived a test it had a 96 per cent chance of failing, whereas Theory B survived a test it had only a 40 per cent chance of failing. The corroboration a theory draws from a passed test scales with how improbable that pass was in advance, so A's confirmation is worth more than twice B's, and a theory that merely predicts an effect in the expected direction over a permitted band approaching the whole space gains almost nothing when it passes (Meehl, 1978). Moving the boldness control in the first demonstration recomputes F by this same equation, and the contrast is exactly the one the numbers make here.
Discussion
Psychology entered the 2010s preoccupied with a replication crisis and left the decade increasingly convinced that beneath it lay a theory crisis. The argument, pressed from several directions at once, is that better methods cannot rescue a science whose theories are too vague to be decisively tested: preregistration and larger samples make findings more reliable, but a reliable estimate of an effect predicted by a theory that forbids almost nothing is still uninformative about that theory (Smaldino, 2019; Oberauer & Lewandowsky, 2019). Muthukrishna and Henrich put the point sharply: without a theoretical framework that constrains which hypotheses are worth testing, psychology accumulates isolated findings that neither cohere nor cumulate, and the absence of such frameworks, not the presence of questionable research practices, is the deeper problem (Muthukrishna & Henrich, 2019). The diagnoses converge even where the proposed cures differ.
Those cures define the field's current program of reform. One strand calls for formal theory: rendering verbal theories as mathematical or computational models forces the hidden assumptions into the open and reveals whether the theory can generate its claimed predictions at all (Guest & Martin, 2021; Robinaugh et al., 2021), a case that extends to formalising the methodology of reform itself (Devezer et al., 2021). A second strand offers explicit methodologies for building theories in the first place, treating theory construction as a skilled activity with its own steps rather than an act of inspiration (Borsboom et al., 2021). A third argues that the discipline should spend less of its effort testing hypotheses and more of it establishing the robust phenomena and well-specified constructs that make a hypothesis worth testing (Scheel et al., 2021; Eronen & Bringmann, 2021). What unites them is a shift of attention from the appraisal of finished theories, the concern of Popper and Lakatos, to the neglected craft of constructing theories good enough to appraise. The standards laid out in this article, falsifiability, construct validity, progressive content, and clarity about levels, remain the criteria; the contemporary recognition is that psychology has too rarely built theories capable of meeting them, and that repairing this is the precondition for the field's other repairs to matter.
Common Misconceptions
- A theory is just a guess or a hunch.
- In ordinary speech a theory means a speculation, but in science it means the opposite: a systematic explanation that has generated testable predictions and survived attempts to refute them. The colloquial sense invites the mistake of treating a well-corroborated theory as merely a theory, when corroboration through risky tests is exactly what gives it standing (Popper, 2002).
- A theory that explains everything is a strong theory.
- Universal explanatory reach is a weakness, not a strength. A theory compatible with every possible observation forbids none, and a theory that forbids nothing cannot be tested and explains nothing. The power of a theory lies in what it rules out (Popper, 2002; Meehl, 1978).
- A statistically significant result confirms the theory behind it.
- A significant difference in the predicted direction is a weak test, because a directional prediction forbids only half the outcome space and, with a large enough sample, some difference is nearly always found. Genuine confirmation requires a prediction specific enough that passing it was improbable in advance (Meehl, 1978; Scheel et al., 2021).
Glossary
- Auxiliary hypothesis.
- A subsidiary assumption, part of a research programme's protective belt, that links a theory's core to observations and can be modified to absorb a failed prediction.
- Construct validity.
- The degree to which a measure behaves as the theory of its construct requires, established by showing its relations to other measures match the nomological network's predictions.
- Construct.
- A posited inner property, such as anxiety or working-memory capacity, that cannot be observed directly and is inferred from patterns in behaviour and measurement.
- Corroboration.
- The status a theory earns by surviving a severe test it was likely to fail; Popper's alternative to proof, reflecting how improbable the successful outcome was in advance.
- Degenerating programme.
- A research programme whose modifications only accommodate known facts without predicting new ones, so its empirical content stagnates or shrinks.
- Demarcation.
- The problem of distinguishing scientific from non-scientific claims; Popper's proposed criterion is falsifiability.
- Duhem-Quine thesis.
- The claim that a hypothesis can be tested only together with auxiliary assumptions, so a failed prediction refutes the bundle as a whole without identifying which element is at fault; also called confirmation holism or underdetermination.
- Falsifiability.
- The property of a theory that it forbids some possible observations and could therefore in principle be refuted; the greater the fraction of outcomes forbidden, the more testable the theory.
- Hard core.
- The central assumptions of a research programme that its adherents treat as irrefutable, protected from refutation by the surrounding belt of auxiliary hypotheses.
- Levels of analysis.
- Marr's distinction among the computational, algorithmic, and implementational descriptions of an information-processing system, which explain the same phenomenon without competing.
- Nomological network.
- The interlocking system of lawful propositions relating a construct to other constructs and to observable measures, within which the construct acquires meaning and validity.
- Paradigm.
- Kuhn's term for the shared framework of assumptions, exemplary problems, and methods within which normal science proceeds until anomalies force a revolution.
- Progressive programme.
- A research programme whose modifications predict novel facts that are subsequently confirmed, so its empirical content grows; Lakatos's mark of good science.
- Research programme.
- Lakatos's unit of scientific appraisal, a sequence of theories sharing a hard core and a protective belt, judged progressive or degenerating over time rather than by a single test.
- Strong inference.
- Platt's prescription for rapid scientific progress: frame mutually exclusive hypotheses and design crucial experiments that can exclude at least one of them.
- Theory crisis.
- The contemporary argument that psychology's replication problems stem from a shortage of theories precise enough to be decisively tested, not from methodological failings alone.
Key Researchers
Karl R. Popper (1902-1994). Philosopher of science at the London School of Economics; his criterion of falsifiability reframed the demarcation of science and remains the standard against which a theory's testability is judged. Wikipedia - Wikidata
Thomas S. Kuhn (1922-1996). Historian and philosopher of science at the Massachusetts Institute of Technology; his account of paradigms and scientific revolutions transformed the understanding of how theoretical frameworks change. Wikipedia - Wikidata
Lee J. Cronbach (1916-2001). Professor of Education and Psychology at Stanford University; with Paul Meehl he formulated construct validity and the nomological network that ground the use of unobservable constructs. Wikipedia - Wikidata
Paul E. Meehl (1920-2003). Regents' Professor of Psychology at the University of Minnesota; co-author of construct validity, his critique of weak risky prediction in soft psychology anticipated the modern theory crisis by decades. Wikipedia - Wikidata
Denny Borsboom. Professor of Psychological Methods at the University of Amsterdam; he leads theory construction methodology and the network approach to psychological constructs. Faculty Page - ORCID - Google Scholar - Wikipedia
Klaus Oberauer. Professor of Cognitive Psychology at the University of Zurich; his agenda-setting analysis with Stephan Lewandowsky named the theory crisis and set out its cures. ORCID - Google Scholar - Wikidata
Eiko I. Fried. Associate Professor of Clinical Psychology at Leiden University; he argues that thin theory building limits progress in the factor and network literatures. Faculty Page - ORCID - Google Scholar - Wikidata
Iris van Rooij. Professor of Computational Cognitive Science at Radboud University; she argues for building high-verisimilitude explanatory theories before designing the test. Faculty Page - ORCID - Google Scholar
Frequently Asked Questions
What is a psychological theory?
A psychological theory is an organised system of concepts and propositions that explains a class of mental or behavioural phenomena by specifying the entities and mechanisms involved, in a form precise enough to yield testable predictions about new situations (Popper, 2002).
How does a theory differ from a hypothesis?
A hypothesis is a single testable proposition, whereas a theory is the larger explanatory system from which many hypotheses are derived; testing a hypothesis supports its parent theory only to the extent the test was risky (Meehl, 1978).
What makes a theory scientific?
On Popper's influential criterion, a theory is scientific to the degree that it is falsifiable, forbidding some possible observations so that it could in principle be refuted; a theory compatible with every outcome forbids nothing and explains nothing (Popper, 2002).
What is construct validity?
Construct validity is the degree to which a measure of an unobservable construct behaves as the theory of that construct requires, established by showing that its relations to other measures match the predictions of a nomological network (Cronbach & Meehl, 1955).
What is the difference between a progressive and a degenerating research programme?
In Lakatos's account, a progressive programme's adjustments predict novel facts that are later confirmed, so its empirical content grows, whereas a degenerating programme's adjustments only accommodate known facts without predicting anything new (Lakatos, 1970).
What are Marr's three levels of analysis?
Marr distinguished the computational level, which specifies what problem a system solves and why, the algorithmic level, which specifies the representations and procedures it uses, and the implementational level, which specifies how these are physically realised; the levels complement rather than compete (Marr, 2010).
What is the theory crisis in psychology?
The theory crisis is the contemporary argument that psychology's difficulty in replicating findings is downstream of a shortage of theories precise enough to be decisively tested, so that improved methods alone cannot secure cumulative progress (Oberauer & Lewandowsky, 2019; Muthukrishna & Henrich, 2019).
How can psychological theories be improved?
Proposed reforms include formalising verbal theories as computational models to expose hidden assumptions, adopting explicit methodologies for constructing theories, and investing more effort in establishing robust phenomena before testing hypotheses (Borsboom et al., 2021; Guest & Martin, 2021; Scheel et al., 2021).
References
Borsboom, D., van der Maas, H. L. J., Dalege, J., Kievit, R. A., & Haig, B. D. (2021). Theory construction methodology: A practical framework for building theories in psychology. Perspectives on Psychological Science, 16(4), 756-766. https://doi.org/10.1177/1745691620969647
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302. https://doi.org/10.1037/h0040957
Devezer, B., Navarro, D. J., Vandekerckhove, J., & Buzbas, E. O. (2021). The case for formal methodology in scientific reform. Royal Society Open Science, 8(3), 200805. https://doi.org/10.1098/rsos.200805
Eronen, M. I., & Bringmann, L. F. (2021). The theory crisis in psychology: How to move forward. Perspectives on Psychological Science, 16(4), 779-788. https://doi.org/10.1177/1745691620970586
Fried, E. I. (2020). Lack of theory building and testing impedes progress in the factor and network literature. Psychological Inquiry, 31(4), 271-288. https://doi.org/10.1080/1047840X.2020.1853461
Guest, O., & Martin, A. E. (2021). How computational modeling can force theory building in psychological science. Perspectives on Psychological Science, 16(4), 789-802. https://doi.org/10.1177/1745691620970585
Kuhn, T. S. (2012). The structure of scientific revolutions (4th ed.). University of Chicago Press. (Original work published 1962)
Lakatos, I. (1970). Falsification and the methodology of scientific research programmes. In I. Lakatos & A. Musgrave (Eds.), Criticism and the growth of knowledge (pp. 91-196). Cambridge University Press.
Marr, D. (2010). Vision: A computational investigation into the human representation and processing of visual information. MIT Press. (Original work published 1982)
McGuire, W. J. (1997). Creative hypothesis generating in psychology: Some useful heuristics. Annual Review of Psychology, 48, 1-30. https://doi.org/10.1146/annurev.psych.48.1.1
Meehl, P. E. (1978). Theoretical risks and tabular asterisks: Sir Karl, Sir Ronald, and the slow progress of soft psychology. Journal of Consulting and Clinical Psychology, 46(4), 806-834. https://doi.org/10.1037/0022-006X.46.4.806
Muthukrishna, M., & Henrich, J. (2019). A problem in theory. Nature Human Behaviour, 3(3), 221-229. https://doi.org/10.1038/s41562-018-0522-1
Oberauer, K., & Lewandowsky, S. (2019). Addressing the theory crisis in psychology. Psychonomic Bulletin & Review, 26(5), 1596-1618. https://doi.org/10.3758/s13423-019-01645-2
Platt, J. R. (1964). Strong inference. Science, 146(3642), 347-353. https://doi.org/10.1126/science.146.3642.347
Popper, K. R. (2002). The logic of scientific discovery. Routledge. (Original work published 1935)
Quine, W. V. (1951). Two dogmas of empiricism. The Philosophical Review, 60(1), 20-43. https://doi.org/10.2307/2181906
Robinaugh, D. J., Haslbeck, J. M. B., Ryan, O., Fried, E. I., & Waldorp, L. J. (2021). Invisible hands and fine calipers: A call to use formal theory as a toolkit for theory construction. Perspectives on Psychological Science, 16(4), 725-743. https://doi.org/10.1177/1745691620974697
Scheel, A. M., Tiokhin, L., Isager, P. M., & Lakens, D. (2021). Why hypothesis testers should spend less time testing hypotheses. Perspectives on Psychological Science, 16(4), 744-755. https://doi.org/10.1177/1745691620966795
Smaldino, P. E. (2019). Better methods can't make up for mediocre theory. Nature, 575(7781), 9. https://doi.org/10.1038/d41586-019-03350-5
van Rooij, I., & Baggio, G. (2021). Theory before the test: How to build high-verisimilitude explanatory theories in psychological science. Perspectives on Psychological Science, 16(4), 682-697. https://doi.org/10.1177/1745691620970604