Abstract
Psychology is a branch of the behavioral sciences: the scientific study of mind and behavior, spanning the mental and social processes that shape what organisms do. This article treats psychology as a discipline rather than a subject matter, tracing its emergence as an experimental science from the first laboratories of the late nineteenth century, the split between its experimental and correlational traditions, and the standards of measurement and inference on which its claims depend. It examines the statistical power that governs whether a real effect is detected, the researcher degrees of freedom that inflate false positives, and the sampling problems that limit how far a result generalizes—the three pressures whose interaction produced the reproducibility crisis and the reforms that followed. Three interactive demonstrations let the reader vary these parameters and watch power, false positives, and estimate bias respond.
Keywords: psychology, behavioral science, statistical power, reproducibility, research methods
Psychology is the science that takes mind and behavior as its object: how organisms perceive, learn, remember, reason, feel, develop, and act, and how those processes arise in the brain and play out in a social world. It is at once a natural science, continuous with biology and neuroscience, and a human science that must study meaning, culture, and individual difference, and this dual character runs through its whole history. Because its subject matter is the same set of faculties that cognitive psychology analyzes in detail, the parent discipline supplies the methods and the epistemic standards—operational definition, controlled experiment, quantified measurement, statistical inference—by which any claim about the mind is to be tested. Its defining challenge is not a shortage of phenomena but the difficulty of measuring them: the constructs psychology cares about are latent, variable, and entangled with the observer, so the discipline's progress has depended less on new observations than on ever more careful ways of deciding which observations to believe (Meehl, 1978).
- Psychology is the scientific study of mind and behavior, one of the behavioral sciences, and it is organized into many subfields defined by the process or population each studies.
- It emerged as an experimental discipline in the late nineteenth century and has swung between behaviorist and cognitive framings of its subject matter.
- Its two founding methodological traditions—the experimental study of general laws and the correlational study of individual differences—rest on different logics and have never fully merged.
- The credibility of a psychological finding depends on statistical power, restraint in analytic flexibility, and representative sampling; failures of all three drove the reproducibility crisis.
- The reform movement that followed—preregistration, larger samples, open data, and large-scale replication—has measurably changed how the field works.
What Psychology Is
Psychology is distinguished among the sciences less by a single method than by the peculiar position of its subject matter, which is both the instrument of inquiry and its object. The processes it studies—attention, memory, motivation, reasoning—are not directly observable and must be inferred from behavior, physiology, or report, so the discipline is built on the construction of measures that stand in for latent variables and on the statistical machinery that decides whether a measured difference reflects a real one. This gives psychology its characteristic breadth: it reaches down toward the neuron, where it meets neuroscience, and up toward the group and the culture, where it meets sociology and anthropology, and its subfields are carved out along that vertical span. It is also unusually reflexive, in that the reasoning biases, memory distortions, and motivational pressures it documents in its participants also operate on its investigators, which is why the field has had to build explicit safeguards—control conditions, blinding, preregistration—against its own capacity for self-deception (Simmons et al., 2011). The unifying commitment across every subfield is empirical: a claim about the mind, however plausible, is credited only insofar as it survives a fair test against evidence.
Types of Psychology
Psychology is a formal descriptor in the National Library of Medicine's Medical Subject Headings, filed beneath the Behavioral Sciences at tree position F04.096.628. Beneath the descriptor hangs the set of narrower headings listed in Table 1, the subfields into which the discipline is conventionally divided. Two cautions apply. The list is an indexing classification built to organize the biomedical literature, not a theory that carves psychology at its joints, and its members are neither mutually exclusive nor jointly exhaustive: cognitive and social psychology overlap wherever cognition is studied in a social setting, developmental psychology cuts across every content area, and applied subfields such as industrial and educational psychology draw their methods from the experimental core rather than standing apart from it. Only subtypes that are themselves live articles on this site are linked.
| Subtype | In brief |
|---|---|
| Adolescent Psychology | The mental life and behavior of the teenage years, when identity, autonomy, and many adult patterns first take shape. |
| Behavioral Economics | The study of how real psychological limits and biases shape economic choice, departing from the rational-agent ideal. |
| Child Psychology | The development of perception, cognition, emotion, and social behavior across childhood. |
| Clinical Psychology | The assessment and psychological treatment of mental disorder and distress. |
| Cognitive Psychology | The experimental study of internal mental processes—attention, memory, language, reasoning, and problem solving. |
| Cognitive Science | The interdisciplinary study of mind combining psychology, linguistics, computer science, neuroscience, and philosophy. |
| Comparative Psychology | The study of behavior across animal species, used to trace the evolution and generality of psychological processes. |
| Developmental Psychology | The study of how mind and behavior change across the whole lifespan, from infancy to old age. |
| Educational Psychology | The study of how people learn in instructional settings and how teaching and assessment can be improved. |
| Environmental Psychology | The study of the reciprocal relation between people and their physical surroundings, built and natural. |
| Ethnopsychology | The study of how culture and ethnicity shape mind, self, and behavior across human groups. |
| Experimental Psychology | The tradition that establishes general laws of mind and behavior through controlled laboratory manipulation. |
| Forensic Psychology | The application of psychology to the legal system, including testimony, competency, and risk assessment. |
| Industrial Psychology | The study of behavior in the workplace, covering selection, performance, motivation, and organizational life. |
| Medical Psychology | The application of psychological science to physical health, illness, and medical care. |
| Positive Psychology | The study of well-being, strengths, and the conditions under which people and communities flourish. |
| Social Psychology | The study of how the actual or imagined presence of others shapes thought, feeling, and behavior. |
| Sports Psychology | The study of psychological factors in athletic performance and the effects of exercise on mind. |
The Emergence of a Science
Psychology became an experimental science by convention in 1879, when Wilhelm Wundt opened the first laboratory dedicated to its study at Leipzig and set about measuring the timing and structure of elementary conscious processes. Within a generation William James had given the young field its first great synthesis, framing mental life as a continuous, functional stream serving the organism's adaptation, and the two emphases—Wundt's analytic, laboratory bent and James's functional, whole-person view—marked out tensions the discipline still carries. The early reliance on trained introspection soon drew a sharp reaction. In 1913 John B. Watson declared that a scientific psychology must abandon consciousness altogether and study only observable behavior and its lawful relation to stimuli, a program that dominated Anglo-American psychology for four decades and made conditioning and learning its central problems (Watson, 1913). The behaviorist ban on inner states was in turn overthrown by the cognitive revolution of the 1950s and 1960s, as work on information, memory, and language showed that internal representations could be studied rigorously through their behavioral signatures; George Miller's demonstration that immediate memory is limited to a small number of chunks was an early emblem of that shift, restoring mental structure to the center of the field (Miller, 1956). Figure 1 sets these turning points on a single timeline.
Figure 1
Turning Points in the Emergence of Scientific Psychology
Two Disciplines of Scientific Psychology
Psychology carries within it two research traditions that grew from different roots and reason in different ways, a division Lee Cronbach influentially named the two disciplines of scientific psychology (Cronbach, 1957). The experimental discipline manipulates a variable, holds others constant, and seeks general laws that hold for the average organism; individual differences are, for it, error variance to be minimized. The correlational discipline does the opposite: it treats the variation between people—in ability, personality, or development—as the very phenomenon of interest, and studies how naturally occurring differences covary. Each has a characteristic blindness. The experimentalist, chasing a main effect, can miss that a manipulation helps some people and harms others; the correlationalist, cataloguing differences, can never establish what causes them. Cronbach's call to unite the two around the interaction of person and situation named a synthesis the field has approached but never completed, and the split still structures how psychologists are trained, which journals they publish in, and which threats to validity they most fear. It also shapes the statistics each tradition leans on, and with them the ways each can go wrong—the theme of the sections that follow.
Measurement, Power, and Evidence
Whatever its tradition, a psychological study succeeds or fails on measurement and inference. Because its constructs are latent, the field depends on operational definitions and on the statistical logic that decides whether an observed difference is larger than chance would readily produce. The central quantity is statistical power: the probability that a study will detect an effect that is really there, given the size of that effect and the size of the sample. Power is not a technicality but a precondition for cumulative science, because an underpowered literature is one in which true effects are missed, published effects are overestimated, and even correct findings replicate erratically. Paul Meehl argued decades ago that psychology's reliance on null-hypothesis significance testing, combined with small samples and weak theories, made it far too easy to accumulate statistically significant results that did not correspond to durable facts (Meehl, 1978). The demonstration below makes power concrete, letting the reader set an effect's true size and the number of participants and read off the probability that a conventional test will detect it—showing how quickly power collapses when either the effect or the sample is small.
Power is the probability that a two-group study detects a real effect, for a two-sided test at α = 0.05. Adjust the true effect size and the sample.
Normal approximation: power = 1 − Φ(1.96 − d√(n/2)).
The Reproducibility Crisis
The abstract worries about power and inference became a concrete crisis in the 2010s, when systematic attempts to reproduce published findings failed at a startling rate. A demonstration that flexibility in data collection and analysis—optional stopping, dropping conditions, choosing among outcome measures after seeing the data—could push the false-positive rate far above the nominal five percent showed that ordinary, well-intentioned research practices were enough to manufacture significant results at will (Simmons et al., 2011). When the Open Science Collaboration then attempted to replicate one hundred studies from leading journals, fewer than half yielded a significant effect in the same direction, and the replicated effects were on average about half the original size (Open Science Collaboration, 2015). A survey of more than fifteen hundred scientists soon confirmed that the sense of a reproducibility problem was widespread across fields, with a majority reporting they had failed to reproduce another scientist's results and many their own (Baker, 2016). The diagnosis converged on the interaction of three forces: underpowered designs that could not reliably detect true effects, undisclosed analytic flexibility that inflated false positives, and publication incentives that rewarded novel, significant results over careful or negative ones. The demonstration below isolates the second force, letting the reader add independent analytic choices to a study and watch the probability of at least one spurious significant result climb well past the nominal rate.
Each independent analytic choice tested at α = 0.05 is another chance for a spurious “significant” result. Add choices and watch the real false-positive rate climb above the nominal 5%.
Independent-tests bound: rate = 1 − (1 − 0.05)k. Real flexibility is exploited jointly, so the empirical rate is higher still.
Generalizability and the WEIRD Problem
Even a well-powered, honestly analyzed finding is only as broad as the sample it came from, and here psychology has a structural bias. The overwhelming majority of its participants are drawn from Western, educated, industrialized, rich, and democratic societies—the acronym is WEIRD—and often from the still narrower pool of undergraduates at research universities, yet conclusions are routinely stated as though they held for human beings in general. A survey of the comparative literature found that these samples are frequently outliers rather than a representative baseline on dimensions from visual perception to fairness and moral reasoning, so that behavior measured in one unusual population is a poor guide to the species (Henrich et al., 2010). The problem is not only cultural but statistical and inferential: when the conditions, stimuli, and populations that a study actually sampled are treated as interchangeable with the vastly larger universe the verbal hypothesis refers to, the inference from sample to claim is far weaker than the reported statistics suggest, a mismatch that has been called the generalizability crisis (Yarkoni, 2022). The demonstration below makes the sampling side of this vivid, letting the reader draw a sample that is increasingly dominated by one unrepresentative subpopulation and watch the estimate drift away from the true value for the whole population.
A trait differs between WEIRD populations (mean 72) and the rest of the world (mean 41). WEIRD societies are about 12% of humanity, so the true global mean is 44.7. Skew the sample toward the WEIRD pool and watch the estimate drift.
Worked Example
The single most consequential number a psychologist can compute before running a study is its statistical power, and the calculation shows starkly why so much of the older literature failed to replicate. Consider a two-group experiment—a treatment and a control—designed to detect a medium-sized effect, conventionally a standardized mean difference of Cohen's d equal to 0.5, using twenty participants in each group, a sample size entirely typical of published psychology before the reforms. Power is the probability that such a study will return a statistically significant result when the effect is genuinely present. For a two-sided test at the usual significance level of 0.05, the critical value on the standard normal scale is 1.96. The study's ability to detect the effect is captured by a noncentrality term equal to d multiplied by the square root of the per-group sample size divided by two: that is 0.5 times the square root of twenty divided by two, or 0.5 times the square root of ten, which is 0.5 times 3.162, giving 1.581. Power is then the probability that a standard normal variable exceeds 1.96 minus 1.581, that is exceeds 0.379, which is approximately 0.35. So a typical two-group study of a medium effect had only about a thirty-five percent chance of detecting it—worse than a coin flip—and would miss a real effect roughly two times in three. That figure is close to the aggregate replication rate the Open Science Collaboration actually observed, and it explains the pattern without any appeal to fraud: a literature built from underpowered studies will be shot through with both missed real effects and overestimated published ones (Open Science Collaboration, 2015). The first demonstration computes exactly this power as the effect size and sample are varied.
Discussion
Psychology's history can be read as a sequence of corrections, each overshooting before the next pulls it back: introspection giving way to a behaviorism that denied the mind, behaviorism giving way to a cognitivism that risked forgetting the body and the group, and most recently a confident empirical literature giving way to a hard reckoning with its own methods. What is striking about the reproducibility crisis is that it was, in the main, self-inflicted and self-diagnosed—the same discipline that produced the unreliable literature produced the tools that exposed it and the reforms that are repairing it. That reflexivity is not incidental. Because psychology studies the very biases that threaten its own conduct, it is uniquely positioned to build those threats into its methodology rather than merely deplore them, and the current emphasis on power, transparency, and replication is best understood as the field applying its own findings to itself. The stakes for cognitive psychology in particular are high, because much of the evidence base for how mental processes work was collected under exactly the conditions—small samples, flexible analysis, homogeneous participants—that the crisis called into question, and the ongoing work of re-establishing which effects are real is a precondition for any cumulative theory. The open questions are correspondingly foundational. How can a science measure latent constructs with enough precision to support strong theory? How far do laboratory findings on unrepresentative samples generalize to human beings at large? And can the experimental and correlational traditions finally be integrated into an account that predicts not just the average person but the variation between people? What psychology offers is not a finished map of the mind but a maturing method for building one under genuine uncertainty about its own instruments.
Current Directions
The most active development in the discipline's methodology is the consolidation of the reforms provoked by the crisis into ordinary practice. A synthesis of the evidence distinguishes three targets that were once run together—replicability, whether an independent study finds the same effect; robustness, whether the same data analyzed differently yield the same conclusion; and reproducibility, whether the reported analysis can be rerun at all—and argues that each requires its own remedy, from preregistration and larger samples to open data and open materials (Nosek et al., 2022). There is now direct evidence that these remedies are changing the field rather than merely being discussed: a broad assessment finds measurable structural, procedural, and community change, including the rapid spread of preregistration and registered reports, the growth of large-scale collaborative projects that pool participants across many laboratories, and shifts in incentives and norms toward transparency (Korbmacher et al., 2023). Running alongside the procedural reforms is a more theoretical reappraisal of how psychological claims are stated in the first place, pressing the field to match the breadth of its verbal hypotheses to the far narrower range of conditions its studies actually sample, and to treat generalization as a claim requiring evidence rather than an assumption granted for free (Yarkoni, 2022). Together these lines are reshaping not only how psychologists gather data but what they take a confirmed finding to license.
Common Misconceptions
- Psychology is just common sense dressed up in jargon.
- Common sense is a source of hypotheses, not a substitute for testing them, and it routinely endorses mutually contradictory maxims. The discipline exists precisely because intuition is unreliable about the mind: the same reasoning and memory biases psychology documents in its participants operate on everyone's everyday judgments, which is why controlled measurement so often overturns the obvious (Meehl, 1978).
- A statistically significant result means the finding is true and important.
- Significance at the conventional level only bounds one kind of error under a set of assumptions that flexible analysis can quietly violate; ordinary undisclosed choices can push the false-positive rate far above five percent. A significant result from an underpowered, flexibly analyzed study is weak evidence, not a settled fact (Simmons et al., 2011).
- The reproducibility crisis shows that psychology is not a real science.
- The opposite inference is better supported. That the field systematically measured its own replication rate, identified the causes, and reformed its methods is a display of scientific self-correction, and surveys show reproducibility problems are shared across the natural sciences rather than peculiar to psychology (Baker, 2016).
Glossary
- Behavioral sciences.
- The group of disciplines that study the behavior of organisms, under which MeSH classifies psychology.
- Behaviorism.
- The program, launched by Watson, that restricts scientific psychology to observable behavior and its lawful relation to stimuli, setting inner states aside.
- Cognitive psychology.
- The subfield that studies internal mental processes—attention, memory, language, reasoning—through their behavioral signatures.
- Cognitive revolution.
- The mid-twentieth-century shift that displaced behaviorism by showing internal representations could be studied rigorously through behavior.
- Correlational discipline.
- Cronbach's term for the research tradition that studies naturally occurring individual differences and how they covary, rather than manipulating variables.
- Effect size.
- A standardized measure of the magnitude of a phenomenon, such as Cohen's d, independent of sample size.
- Experimental discipline.
- Cronbach's term for the tradition that manipulates variables under control to establish general laws for the average organism.
- Generalizability.
- The extent to which a finding obtained in one sample, setting, and set of stimuli holds across the broader population the claim refers to.
- Introspection.
- The early method of systematically observing and reporting one's own conscious experience, central to Wundt's laboratory and later rejected by behaviorism.
- Null-hypothesis significance testing.
- The dominant inferential procedure in which a result is judged against the probability it would occur if no effect existed.
- Operational definition.
- The specification of a latent construct in terms of the concrete, measurable operations used to index it.
- Preregistration.
- The practice of publicly specifying a study's hypotheses and analysis plan before collecting data, which constrains later analytic flexibility.
- Reproducibility.
- The ability to obtain the same result, whether by re-analyzing the original data or by independently repeating the study; the failure of which defined the crisis of the 2010s.
- Researcher degrees of freedom.
- The undisclosed analytic choices—when to stop collecting, which conditions to compare, which measures to report—that can inflate the false-positive rate.
- Statistical power.
- The probability that a study will detect an effect that is genuinely present, determined chiefly by the effect size and the sample size.
- WEIRD samples.
- Participants from Western, educated, industrialized, rich, and democratic societies, who dominate psychological research yet are often unrepresentative of humanity.
Key Researchers
William James (1842-1910). Philosopher and psychologist at Harvard University; his Principles of Psychology (1890) gave the young discipline its first comprehensive synthesis and framed mental life as a functional, adaptive stream of consciousness. Wikipedia
Elizabeth F. Loftus (b. 1944). Distinguished Professor at the University of California, Irvine; her experimental work on the misinformation effect and the malleability of memory reshaped both cognitive psychology and the law's understanding of eyewitness testimony. ORCID - Wikipedia
Brian A. Nosek (b. 1973). Professor at the University of Virginia and co-founder of the Center for Open Science; he led the large-scale replication effort and the open-science reforms that transformed psychology's methods. ORCID - Wikipedia
B. F. Skinner (1904-1990). Psychologist at Harvard University; the leading figure of radical behaviorism, he developed the analysis of operant conditioning and made the lawful control of behavior by its consequences a central problem of the field. Wikipedia
Simine Vazire (b. 1980). Professor at the University of Melbourne; a leader in research methods and metascience whose work on the credibility and self-correction of psychological science helped drive the reform movement. ORCID - Wikipedia
Wilhelm Wundt (1832-1920). Physiologist and psychologist at the University of Leipzig; by founding the first laboratory dedicated to experimental psychology in 1879 he is conventionally credited with establishing the discipline as an independent science. Wikipedia
Frequently Asked Questions
What is psychology?
Psychology is the scientific study of mind and behavior, encompassing perception, cognition, emotion, development, and social interaction, and it is one of the behavioral sciences. It studies latent mental processes through their measurable effects on behavior, physiology, and report, using controlled observation and statistical inference (Meehl, 1978).
What are the main branches of psychology?
The discipline divides into subfields defined by the process or population each studies, including cognitive, developmental, social, clinical, educational, and industrial psychology, among others. These subfields overlap heavily and share a common experimental and statistical core rather than standing as separate sciences (Cronbach, 1957).
How is psychology different from psychiatry?
Psychiatry is a branch of medicine whose practitioners are physicians who can prescribe drugs, whereas psychology is the broader science of mind and behavior, whose clinical branch delivers assessment and psychotherapy. Their subject matter overlaps but their training, methods, and scope differ (Meehl, 1978).
When did psychology become a science?
It is conventionally dated to 1879, when Wilhelm Wundt opened the first laboratory devoted to experimental psychology at Leipzig and began measuring elementary mental processes. The field has since moved through behaviorist and cognitive framings of its subject matter (Watson, 1913).
What is statistical power and why does it matter?
Statistical power is the probability that a study detects an effect that is really present, set chiefly by the effect size and the sample size. Low power means true effects are missed and published ones are exaggerated, which is a central reason many older findings failed to replicate (Open Science Collaboration, 2015).
What is the reproducibility crisis?
It is the finding, prominent from the 2010s, that a large fraction of published results could not be reproduced when studies were repeated or reanalyzed. Systematic replication of one hundred studies reproduced fewer than half, with replicated effects about half the original size (Open Science Collaboration, 2015).
Why are WEIRD samples a problem?
Most participants come from Western, educated, industrialized, rich, and democratic societies, which are often unrepresentative of humanity on the very dimensions being studied. Conclusions stated as universal may hold only for this narrow population, weakening the inference from sample to species (Henrich et al., 2010).
How is psychology responding to its methodological problems?
Through reforms that include preregistering hypotheses and analysis plans, running larger and better-powered studies, sharing data and materials openly, and mounting large collaborative replications. Assessments find these practices are spreading and measurably changing the field's structure and norms (Korbmacher et al., 2023).
References
Baker, M. (2016). 1,500 scientists lift the lid on reproducibility. Nature, 533(7604), 452-454. https://doi.org/10.1038/533452a
Cronbach, L. J. (1957). The two disciplines of scientific psychology. American Psychologist, 12(11), 671-684. https://doi.org/10.1037/h0043943
Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2-3), 61-83. https://doi.org/10.1017/S0140525X0999152X
Korbmacher, M., Azevedo, F., Pennington, C. R., Hartmann, H., Pownall, M., Schmidt, K., Elsherif, M., Breznau, N., Robertson, O., Kalandadze, T., Yu, S., Baker, B. J., O'Mahony, A., Olsnes, J. O.-S., Shaw, J. J., Gjoneska, B., Yamada, Y., Rosetti, F., Karhulahti, V.-M., ... Evans, T. (2023). The replication crisis has led to positive structural, procedural, and community changes. Communications Psychology, 1(1), 3. https://doi.org/10.1038/s44271-023-00003-2
Meehl, P. E. (1978). Theoretical risks and tabular asterisks: Sir Karl, Sir Ronald, and the slow progress of soft psychology. Journal of Consulting and Clinical Psychology, 46(4), 806-834. https://doi.org/10.1037/0022-006X.46.4.806
Miller, G. A. (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review, 63(2), 81-97. https://doi.org/10.1037/h0043158
Nosek, B. A., Hardwicke, T. E., Moshontz, H., Allard, A., Corker, K. S., Dreber, A., Fidler, F., Hilgard, J., Kline Struhl, M., Nuijten, M. B., Rohrer, J. M., Romero, F., Scheel, A. M., Scherer, L. D., Schönbrodt, F. D., & Vazire, S. (2022). Replicability, robustness, and reproducibility in psychological science. Annual Review of Psychology, 73, 719-748. https://doi.org/10.1146/annurev-psych-020821-114157
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366. https://doi.org/10.1177/0956797611417632
Watson, J. B. (1913). Psychology as the behaviorist views it. Psychological Review, 20(2), 158-177. https://doi.org/10.1037/h0074428
Yarkoni, T. (2022). The generalizability crisis. Behavioral and Brain Sciences, 45, e1. https://doi.org/10.1017/S0140525X20001685