Abstract

The behavioral sciences are the family of empirical disciplines that study the behavior of organisms and the mental processes behind it through systematic observation and experiment. They span psychology, ethology, behavioral economics, behavioral genetics, and parts of anthropology and sociology, united less by a shared subject than by a shared commitment to explaining behavior from evidence rather than intuition. This article traces the field from the behaviorist program that made observable action the unit of study, through the cognitive and biological expansions that followed, to the methodological machinery of construct validity, levels of analysis, and representative sampling on which its claims depend. It closes with the replication crisis that forced the behavioral sciences to confront how they generate and test knowledge. Three interactive demonstrations let the reader manipulate a validity matrix, a biased sample, and a literature's true-finding rate.

Keywords: behavioral sciences, behaviorism, replication crisis, construct validity, generalizability

The behavioral sciences are the disciplines that seek lawful, evidence-based accounts of what organisms do and why. The label gathers a wide range of fields, from experimental psychology and ethology to behavioral economics and behavioral genetics, that differ in their preferred organisms, timescales, and methods but agree on a single methodological posture: behavior is a natural phenomenon to be measured, modeled, and explained, not merely interpreted. The term rose to prominence in mid-twentieth-century American science as a deliberately neutral banner under which psychologists, sociologists, anthropologists, and biologists could pursue a common program, and it still marks that program's ambition to be as rigorous about behavior as the natural sciences are about matter. What distinguishes a behavioral science from a humanistic study of the same conduct is the insistence that claims answer to systematic data and survive attempts to falsify them, a discipline whose founding statement in psychology was John Watson's demand that the field restrict itself to the prediction and control of observable behavior (Watson, 1913).

Key Takeaways
  • The behavioral sciences are unified by method rather than subject: a commitment to explaining behavior through systematic measurement and experiment rather than intuition or authority.
  • The field's modern arc runs from behaviorism, which made observable action the unit of analysis, through the cognitive revolution and the biological turn, each widening what counts as legitimate evidence about behavior.
  • Because its constructs are not directly observed, the field depends on construct validity: the disciplined demonstration that a measure converges with other measures of the same thing and diverges from measures of different things.
  • Its claims generalize only as far as its samples, and the reliance on narrow, industrialized populations has made generalizability a central methodological concern.
  • The replication crisis exposed how low statistical power, publication bias, and weak theory testing can fill a literature with findings that do not hold up, prompting a broad reform of research practice.

What the Behavioral Sciences Are

The behavioral sciences are best defined by a method and an attitude rather than by a fixed list of topics. The attitude is naturalism: the conviction that behavior, including the covert behavior of thinking and feeling, is part of the natural order and therefore open to the same empirical scrutiny as any other phenomenon. The method is the disciplined cycle of observation, hypothesis, controlled test, and revision, applied to organisms ranging from the nematode to the human. On this view the boundary that matters is not between subject matters but between a study that answers to data and one that does not: a laboratory experiment on memory, an ethological field study of birdsong, an economic experiment on bargaining, and a twin study of personality are all behavioral science because each stakes its claims on evidence that could in principle overturn them. The field acquired its name and much of its self-conception in the decades after the Second World War, when the neutral term behavioral sciences was adopted to unite psychology with the empirical wings of sociology, anthropology, and biology under a shared scientific standard. That unity has always been aspirational rather than complete, and the field remains a federation of disciplines with distinct traditions; but the aspiration is real, and it explains why methodological questions, how to measure, how to sample, how to test a theory, recur across every member of the family. It also has an applied edge, most visible in medicine, where the recognition that illness has behavioral and social determinants as well as biological ones reshaped clinical thinking; Engel's biopsychosocial model made this multi-level view explicit, holding that health and disease arise from interacting biological, psychological, and social factors rather than from biology alone (Engel, 1977).

Types of Behavioral Sciences

Behavioral Sciences is also a formal descriptor in the National Library of Medicine's Medical Subject Headings, which files it beneath Behavioral Disciplines and Activities at tree position F04.096. Beneath the descriptor MeSH hangs the set of narrower headings listed in Table 1. Two cautions apply. The list is an indexing classification built to organize the biomedical literature, not a theory that carves the sciences at their joints, and its members are neither mutually exclusive nor jointly exhaustive: psycholinguistics draws on psychology and the social sciences at once, and behavioral genetics straddles this branch and genetics proper. Only subtypes that are themselves live articles on this site are linked, and at present none of these descriptors has its own page.

Table 1. Direct subtypes of Behavioral Sciences in the MeSH classification (tree F04.096).
Subtype In brief
Behavioral MedicineThe integration of behavioral and biomedical knowledge in the prevention, diagnosis, and treatment of disease.
Behavioral ResearchThe systematic empirical investigation of behavior, encompassing the field's shared methods of measurement and inference.
EthologyThe biological study of animal behavior in its natural context, emphasizing function and evolutionary origin.
Genetics, BehavioralThe study of how genetic variation contributes to differences in behavior, chiefly through twin and adoption designs.
ParapsychologyThe investigation of putative anomalous phenomena such as telepathy, retained in the index though outside mainstream science.
PsychiatryThe medical specialty concerned with the diagnosis, treatment, and prevention of mental disorders.
PsycholinguisticsThe study of the psychological processes by which language is acquired, produced, and understood.
PsychologyThe science of mind and behavior, the central and largest of the behavioral sciences.
PsychopathologyThe study of the nature, development, and manifestations of mental disorder.
PsychopharmacologyThe study of how drugs affect mood, cognition, and behavior.
PsychophysicsThe quantitative study of the relation between physical stimuli and the sensations and perceptions they evoke.
PsychophysiologyThe inference of psychological states from the physiological signals of the intact, behaving organism.
SexologyThe scientific study of human sexuality, its behavior, function, and variation.
Social SciencesThe disciplines studying human society and social relationships, including sociology, anthropology, and economics.
SociobiologyThe study of the biological, especially evolutionary, basis of social behavior in animals and humans.

From Behaviorism to Cognition

The modern behavioral sciences begin with a methodological revolt. In 1913 John Watson declared that psychology should abandon introspection and the study of consciousness altogether and become a purely objective, experimental branch of natural science whose goal was the prediction and control of behavior (Watson, 1913). The move was radical and clarifying: by restricting the field to publicly observable stimuli and responses, behaviorism gave psychology a subject matter it could measure and a standard of evidence it could enforce, and for four decades it set the agenda for American experimental psychology. Its most systematic development came from B. F. Skinner, whose analysis of operant conditioning showed how behavior is shaped by its consequences and who later cast that principle as one instance of a general causal mode, selection by consequences, that also governs natural selection and cultural evolution (Skinner, 1981). The power of the behaviorist program lay in its refusal to invoke unobserved inner causes; its limitation lay in the same refusal, which left it unable to account for language, planning, and the structured knowledge that observably guides behavior. The cognitive revolution of the 1950s and 1960s restored the mind to the behavioral sciences by treating internal representations and computations as legitimate, if unobserved, theoretical entities inferred from behavior in the same way physics infers unobserved particles from their traces (Miller, 2003). The result was not a rejection of measurement but its extension: the mental process became a construct to be operationalized and tested, and the behavioral sciences kept behaviorism's rigor while regaining the inner machinery it had exiled.

Levels of Analysis

A recurring source of confusion in the behavioral sciences is that a single behavior can be explained correctly in several different ways at once, and the explanations do not compete. The clearest statement of this comes from ethology. Niko Tinbergen argued that any behavior admits four distinct and complementary questions: its mechanism (what physiological and psychological machinery produces it), its development (how it arises over the life of the individual), its function (what survival or reproductive advantage it confers), and its evolution (how it arose over phylogenetic time) (Tinbergen, 1963). A full account answers all four, and a dispute that appears to be about which explanation is right is often merely about which question is being asked. The framework has proved durable enough that later biologists have refined rather than replaced it, sharpening the distinction between the proximate causes that Tinbergen's first two questions address and the ultimate, evolutionary causes of the last two (Bateson & Laland, 2013). Psychology has its own version of the levels problem, expressed by Lee Cronbach as the split between two disciplines of scientific psychology: an experimental tradition that manipulates variables to find general laws and a correlational tradition that studies the naturally occurring differences among individuals (Cronbach, 1957). Cronbach's plea was that the two traditions, long estranged, be integrated, because a complete science of behavior needs both the laws that hold across people and the dimensions along which people reliably differ. The lesson common to Tinbergen and Cronbach is that behavioral explanation is layered, and that clarity about which level one is working at prevents a great deal of spurious disagreement.

Figure 1

Tinbergen's Four Questions as a Two-by-Two of Explanatory Levels

Tinbergen's four questions arranged in a two-by-two grid A two-by-two matrix. The columns divide proximate causes, meaning how a behavior works, from ultimate causes, meaning why it exists. The rows divide a static snapshot of current form from a dynamic account of change over time. The four cells are: mechanism, the proximate and static cell; development or ontogeny, the proximate and dynamic cell; function or adaptation, the ultimate and static cell; and evolution or phylogeny, the ultimate and dynamic cell. Together the four are complementary rather than competing explanations of the same behavior. Proximate (how it works) Ultimate (why it exists) Static (current form) Dynamic (change over time) Mechanism physiological and psychological machinery Function (adaptation) survival or reproductive advantage Development (ontogeny) how it arises across the lifespan Evolution (phylogeny) how it arose over phylogenetic time
Note. The columns separate proximate causes (how a behavior works) from ultimate causes (why it exists); the rows separate a snapshot of current form from an account of change over time. A complete explanation answers all four cells, and an apparent dispute over which answer is correct is often only a disagreement about which cell is in question. Original schematic after Tinbergen (1963), with the proximate-ultimate refinement of Bateson and Laland (2013).

Explanatory Programs Across the Field

Within this layered structure, several large research programs have organized the behavioral sciences of the past half-century. Behavioral economics grew from the discovery that human choice departs systematically from the rational-agent model economics had assumed, and it now forms a bridge between psychology and economics built on documented, replicable deviations from that ideal (Thaler, 2016). Its account of how people actually decide has itself divided into camps, with one tradition cataloguing the biases that heuristics produce and another arguing that simple heuristics are ecologically rational, well matched to the structure of real environments and often outperforming more complex strategies (Gigerenzer & Gaissmaier, 2011). A second program, behavioral genetics, uses twin and adoption designs to partition the variation in behavior into genetic and environmental components; its accumulated findings, that essentially all behavioral traits are substantially heritable and that much of the environmental variance is not shared between siblings raised together, are among the most robust in the field (Plomin, 2016). A third, the social psychology of influence, maps the principles by which people's behavior is shaped by others, from the reciprocity and consistency norms that drive compliance to the informational and normative pressures that produce conformity (Cialdini & Goldstein, 2004). These programs differ in organism, method, and timescale, but each exemplifies the same commitment: to replace intuition about behavior with a structured body of tested, quantitative claims.

Measurement and Construct Validity

The behavioral sciences face a measurement problem the physical sciences largely escape: their central variables, anxiety, intelligence, attitude, working memory, are not observed directly but inferred from behavior, and the inference can go wrong. A questionnaire labeled a measure of anxiety might instead be measuring the willingness to admit distress, or the tendency to endorse any statement, or simply the format of the instrument. The behavioral sciences answer this with the concept of construct validity: the requirement that a measure be shown to track the theoretical construct it claims to and not something else. Donald Campbell and Donald Fiske gave the idea an operational form in the multitrait-multimethod matrix, which assesses a measure by correlating several traits each assessed by several methods (Campbell & Fiske, 1959). Their logic has two prongs. Convergent validity requires that different methods of measuring the same trait agree: two anxiety measures, one self-report and one behavioral, should correlate substantially. Discriminant validity requires that measures of different traits diverge even when they share a method: an anxiety measure and a depression measure taken by the same questionnaire should correlate less than either does with an alternative measure of itself. When the correlations driven by shared method rival those driven by shared trait, the instrument is measuring the method, not the construct, and any theory built on it inherits the defect. This machinery is not a technicality; it is the behavioral sciences' primary defense against the ever-present danger of naming a construct and then measuring something else. The demonstration below builds a multitrait-multimethod matrix so the reader can watch construct validity survive or collapse as method variance grows.

Convergent and discriminant validity: the MTMM matrix

Two traits (A, B) are each measured by two methods (1, 2). The gold cells are convergent-validity coefficients (same trait, different method); they should exceed both the different-trait cells.

A1A2B1B2A1A2B1B20.850.550.350.100.550.850.100.350.350.100.850.550.100.350.550.85
Convergent validity is substantial (> 0.30).
Validity exceeds the heterotrait-heteromethod values.
Validity exceeds the method-variance (heterotrait-monomethod) values.
Campbell-Fiske criteria met: the construct shows convergent and discriminant validity.

Note. When method variance (gray) rivals or exceeds convergent validity (gold), the measures are tracking the method more than the trait. Original schematic after Campbell and Fiske (1959).

The Problem of Generalizability

A behavioral finding is only as broad as the sample it came from, and for most of the field's history the samples have been remarkably narrow. Joseph Henrich and colleagues documented that the overwhelming majority of published behavioral research draws its participants from societies that are Western, educated, industrialized, rich, and democratic, and that on many measures, visual perception, fairness, cooperation, moral reasoning, spatial cognition, these populations are not a representative sample of humanity but a frequent outlier (Henrich et al., 2010). The consequence is a systematic overreach: a claim established on undergraduates at a research university is announced as a fact about human nature, when it may be a fact about that unusual population. The demonstration below makes the arithmetic of this bias visible, showing how an estimate drifts from the true cross-cultural average as the sample is drawn more heavily from a single outlier population. The problem is not confined to cross-cultural sampling. Tal Yarkoni has argued that it is a special case of a deeper generalizability crisis, in which researchers routinely generalize far beyond the specific stimuli, tasks, and conditions their studies actually sampled, treating a fixed set of experimental materials as if it were a random draw from all possible ones (Yarkoni, 2022). On this analysis the verbal breadth of behavioral theories vastly outruns the narrowness of the designs used to test them, and closing the gap requires either humbler claims or genuinely representative sampling of everything a claim ranges over, from people to items to situations.

Sampling from WEIRD populations biases the estimate

Each bar is a society's value on a behavioral measure. As the sample is drawn ever more heavily from the industrialized outlier, the estimated “human” average drifts away from the true cross-cultural mean.

WEIRDHadzaTsimanAuMachigSanguLamaleAchuarShonatrue mean 34.2estimate 45.7
Sampling bias: +11.5 points away from the cross-cultural mean.

Note. Values are illustrative, not empirical. The pattern reproduces the concern of Henrich, Heine, and Norenzayan (2010) that a science built on one unrepresentative population mistakes its outlier for the species.

The Replication Crisis and Reform

Beginning around 2011 the behavioral sciences confronted evidence that a substantial fraction of their published findings could not be reproduced. The pivotal demonstration came from the Open Science Collaboration, which ran high-powered replications of one hundred studies from leading psychology journals and found that, while ninety-seven of the originals had reported statistically significant effects, only about thirty-six of the replications did, with effect sizes roughly half the original magnitude on average (Open Science Collaboration, 2015). Large multi-laboratory efforts confirmed the pattern for specific celebrated effects, as when a preregistered replication across more than twenty laboratories failed to find the ego-depletion effect at anything like its published size (Hagger et al., 2016), and a systematic replication of social-science experiments in the most prestigious journals reproduced only about sixty percent of them, again at diminished magnitude (Camerer et al., 2018). The causes were structural rather than fraudulent. Chronically low statistical power meant that even real effects were often missed and that significant results were disproportionately inflated; the many undisclosed choices open to an analyst, the researcher degrees of freedom of which stopping rules, outlier exclusions, and the selection among measures are examples, let researchers find significance where none existed, a practice later named p-hacking after the demonstration that such flexibility can push the false-positive rate above sixty percent while every individual decision looks defensible (Simmons et al., 2011); and publication bias ensured that only positive results reached print, so the literature became a biased sample of the studies actually run. Beneath these lay a diagnosis Paul Meehl had delivered decades earlier: that soft psychology tests its theories so weakly, against a null hypothesis almost always false to begin with, that a significant result provides little evidence for the theory that predicted it (Meehl, 1978). The reform movement that followed reworked the field's methods: preregistration to separate prediction from postdiction, larger samples, open data and materials, and a renewed insistence that theory make risky, specific predictions (Munafò et al., 2017). Some have argued that the deepest fix is theoretical, that the behavioral sciences replicate poorly in part because they formalize their theories poorly, leaving predictions too vague to be sharply tested (Muthukrishna & Henrich, 2019). The demonstration below shows why weak power and a low base rate of true effects can fill a literature with false positives even when every individual study plays by the rules.

How many significant findings are true?

Each square is one of 1,000 tested hypotheses. Green squares are true effects correctly detected; red squares are false alarms. The positive predictive value is the share of significant results (green + red) that are green.

True positives: 100False positives: 40Positive predictive value: 71.4%

Note. With low power and a modest base rate of true effects, a large fraction of published “discoveries” are false, the arithmetic behind the reproducibility findings of the Open Science Collaboration (2015).

Worked Example

The reproducibility problem is often blamed on misconduct, but simple arithmetic shows how a field of honest researchers can still produce a literature riddled with false positives. Suppose that among the hypotheses a field tests, only one in five corresponds to a real effect, so that of 1,000 tested hypotheses, 200 are true and 800 are false. Suppose too that the typical study has 50 percent statistical power, a realistic figure for the behavioral sciences, and uses the conventional 5 percent significance threshold. Among the 200 true hypotheses, a study with 50 percent power detects half, yielding 200 times 0.50, or 100, true positives. Among the 800 false hypotheses, the 5 percent false-positive rate produces 800 times 0.05, or 40, false positives. The literature therefore contains 100 plus 40, or 140, statistically significant findings, of which only 100 are real. The positive predictive value, the probability that a significant result reflects a true effect, is 100 divided by 140, which is about 0.714, or 71.4 percent. Put the other way, nearly three in ten of the published discoveries in this field are false, and that is before any questionable research practice or publication bias, both of which push the fraction higher. Raising statistical power to 80 percent would lift the true positives to 160 and the predictive value to 160 divided by 200, or 80 percent; lowering the base rate of true hypotheses, as happens when a field chases surprising, a priori unlikely effects, would drive it down. The single lesson is that the trustworthiness of a literature depends not only on the honesty of its researchers but on the power of their studies and the prior plausibility of their hypotheses (Open Science Collaboration, 2015).

Discussion

The behavioral sciences occupy an unusual position among the sciences: their subject matter is the most familiar thing in the world, our own conduct, and yet the disciplined study of it is barely more than a century old and still contested in its methods. What the field has achieved is real and cumulative. It made behavior measurable, first by stripping away the unobservable and later by learning to infer inner constructs rigorously from outward signs; it produced durable bodies of quantitative fact, from the heritability of traits to the principles of social influence; and it supplied a set of conceptual tools, levels of analysis, construct validity, statistical power, that discipline explanation itself. The crises of the past fifteen years are best read not as evidence that the enterprise has failed but as evidence that it is holding itself to account: a field that could measure its own replication rate, publish the disappointing number, and reform its practices in response is behaving as a science should. The open questions are correspondingly methodological as much as substantive. How far can findings from narrow samples be generalized, and what would genuinely representative sampling of people, stimuli, and situations require (Henrich et al., 2010; Yarkoni, 2022)? Can the behavioral sciences build theories formal enough to make the risky predictions that strong tests demand (Muthukrishna & Henrich, 2019)? And how should a federation of disciplines with different organisms and timescales integrate its layered explanations into a coherent account of behavior? What is settled is the field's founding wager, now so ordinary its boldness is forgotten: that behavior, in all its variety, is a natural phenomenon, and that patient, self-correcting measurement is the way to understand it.

Current Directions

The most consequential movement in the contemporary behavioral sciences is the metascientific reform touched off by the replication crisis, which has hardened from critique into infrastructure. Preregistration and registered reports, in which a study's hypotheses and analysis plan are reviewed and accepted before the data exist, have moved from novelty to institution, and large-scale collaborative replication has become a standing feature of the field rather than an occasional audit; reviews now treat replicability, robustness, and reproducibility as distinct, separately measurable properties of a result rather than a single vague virtue (Nosek et al., 2022). A second current is the push toward stronger theory. The recognition that vague verbal theories cannot be sharply tested has spurred interest in formal and computational models that commit to precise predictions, and in the broader argument that the behavioral sciences need better theoretical scaffolding, drawn where possible from the cumulative frameworks of evolutionary and cultural science, to organize their findings and constrain their claims (Muthukrishna & Henrich, 2019). A third runs toward integration with economics and policy: behavioral economics has matured from a catalogue of anomalies into an applied discipline that designs choice environments and evaluates them with the same experimental rigor the field demands of its basic research, even as it confronts its own replication questions (Thaler, 2016). Together these directions describe a field turning its analytic tools on itself, and treating the reliability of its own knowledge as an object of study.

Common Misconceptions

The behavioral sciences are not real sciences because behavior cannot be measured objectively.
Behavior is measured objectively as a matter of routine, from reaction times and error rates to physiological signals and choices with real stakes. The genuine difficulty is not measurement but inference from a measure to an unobserved construct, and the field has built explicit machinery, construct validity and the multitrait-multimethod logic, to discipline exactly that step (Campbell & Fiske, 1959).
The replication crisis shows that behavioral research is mostly fraud.
Fabrication is rare; the replication problem arises overwhelmingly from honest practice under bad incentives, chiefly low statistical power, analytic flexibility, and publication bias. The arithmetic of predictive value shows that a field of scrupulous researchers can still produce many false positives when studies are underpowered and true effects are uncommon (Open Science Collaboration, 2015).
A finding from a psychology experiment is a fact about human nature.
A finding generalizes only as far as its sample, and most behavioral research has sampled a narrow, industrialized slice of humanity that is on many measures a global outlier. A result established on that population is a fact about it until representative sampling shows otherwise (Henrich et al., 2010).

Glossary

Behavioral economics.
The field bridging psychology and economics that studies the systematic ways human choice departs from the idealized rational agent.
Behavioral genetics.
The study of the genetic and environmental contributions to variation in behavior, chiefly through twin and adoption designs.
Behaviorism.
The program, founded by Watson, that restricts psychology to the objective study of observable stimuli and responses and the prediction and control of behavior.
Biopsychosocial model.
Engel's framework holding that health and illness arise from interacting biological, psychological, and social factors rather than biology alone.
Construct validity.
The degree to which a measure actually reflects the theoretical construct it is intended to assess rather than something else.
Convergent validity.
The agreement among different methods of measuring the same construct, one of the two demands of construct validity.
Discriminant validity.
The divergence of measures of different constructs even when they share a method, the complement of convergent validity.
Ethology.
The biological study of animal behavior in its natural setting, emphasizing its function and evolutionary origin.
Generalizability.
The extent to which a finding holds beyond the specific participants, stimuli, and conditions of the study that produced it.
Heuristic.
A simple decision rule that ignores part of the information, sometimes yielding bias and sometimes ecologically rational efficiency.
Levels of analysis.
The complementary questions, such as Tinbergen's four, that a behavior can be asked, spanning mechanism, development, function, and evolution.
Multitrait-multimethod matrix.
Campbell and Fiske's table of correlations among several traits measured by several methods, used to assess convergent and discriminant validity.
Operant conditioning.
Skinner's process by which behavior is strengthened or weakened by its consequences, the core mechanism of radical behaviorism.
P-hacking.
Exploiting the many undisclosed analytic choices (the researcher degrees of freedom) until a result crosses the significance threshold, inflating the false-positive rate while each individual decision looks reasonable.
Positive predictive value.
The probability that a statistically significant finding reflects a true effect, determined by power, the significance level, and the base rate of true hypotheses.
Preregistration.
The practice of specifying hypotheses and analysis plans before collecting data, separating confirmatory prediction from exploratory postdiction.
Publication bias.
The tendency for statistically significant results to be published while null results are not, making the literature an unrepresentative sample of the studies actually conducted.
Statistical power.
The probability that a study will detect a true effect of a given size; chronically low power is a principal driver of the replication crisis.
WEIRD samples.
Participants from Western, educated, industrialized, rich, and democratic societies, who dominate behavioral research yet are a global outlier on many measures.

Key Researchers

Robert B. Cialdini. Social psychologist and Regents' Professor Emeritus at Arizona State University; his synthesis of the principles of influence organized the study of compliance and conformity. Faculty Page - Google Scholar - Wikipedia

Lee J. Cronbach (1916-2001). Psychometrician at Stanford University; his analysis of the two disciplines of scientific psychology and his work on validity and reliability shaped how the field measures its constructs. Wikipedia - Wikidata

George L. Engel (1913-1999). Psychiatrist and internist at the University of Rochester; his biopsychosocial model brought the behavioral and social determinants of illness into mainstream medicine. Wikipedia - Wikidata

Gerd Gigerenzer. Psychologist at the Max Planck Institute for Human Development; his account of ecological rationality argues that simple heuristics are well adapted to real environments and often outperform complex strategies. Faculty Page - ORCID - Google Scholar - Wikipedia

Joseph Henrich. Evolutionary anthropologist at Harvard University; his documentation of the WEIRD-sample problem reframed the question of how far behavioral findings generalize across human populations. Faculty Page - ORCID - Google Scholar - Wikipedia

Daniel Kahneman (1934-2024). Psychologist at Princeton University and Nobel laureate in economics; with Amos Tversky he founded the study of judgment under uncertainty that became behavioral economics. Wikipedia - Wikidata

Paul E. Meehl (1920-2003). Clinical psychologist at the University of Minnesota; his critique of weak theory testing in soft psychology anticipated the diagnosis behind the replication crisis by decades. Wikipedia - Wikidata

Brian A. Nosek. Social psychologist at the University of Virginia and co-founder of the Center for Open Science; he led the large-scale replication and reform efforts that reshaped research practice across the behavioral sciences. Faculty Page - ORCID - Google Scholar - Wikipedia

Robert Plomin. Behavioral geneticist at King's College London; his twin and adoption research established the pervasive heritability of behavioral traits and the importance of the nonshared environment. Faculty Page - ORCID - Google Scholar - Wikipedia

B. F. Skinner (1904-1990). Psychologist at Harvard University; his analysis of operant conditioning and the principle of selection by consequences was the most systematic development of the behaviorist program. Wikipedia - Wikidata

Richard H. Thaler. Economist at the University of Chicago Booth School of Business and Nobel laureate; a founder of behavioral economics who charted its growth from anomaly to applied policy discipline. Faculty Page - Google Scholar - Wikipedia - Wikidata

Nikolaas Tinbergen (1907-1988). Ethologist at the University of Oxford and Nobel laureate; his four questions gave the behavioral sciences their canonical statement of complementary levels of explanation. Wikipedia - Wikidata

John B. Watson (1878-1958). Psychologist at Johns Hopkins University whose 1913 behaviorist manifesto redefined psychology as the objective, experimental study of behavior and set the field's methodological course for decades. Wikipedia - Wikidata

Frequently Asked Questions

What are the behavioral sciences?
The behavioral sciences are the family of empirical disciplines that study the behavior of organisms and the mental processes behind it, including psychology, ethology, behavioral economics, behavioral genetics, and the empirical parts of anthropology and sociology. They are united not by a single subject but by a shared method: explaining behavior through systematic observation and experiment rather than intuition (Watson, 1913).

How do the behavioral sciences differ from the social sciences?
The two overlap heavily and the terms are often used interchangeably, but the behavioral sciences emphasize the behavior of the individual organism and its immediate causes, drawing on biology as well as social context, whereas the social sciences emphasize collective structures such as institutions, economies, and cultures. In the MeSH classification the social sciences are in fact filed as a subtype of the behavioral sciences.

What is construct validity and why does it matter?
Construct validity is the degree to which a measure reflects the theoretical construct it claims to assess rather than something else, such as the response format or a willingness to admit distress. It matters because behavioral variables are inferred, not observed directly, so a theory built on an invalid measure may be tracking an artifact (Campbell & Fiske, 1959).

What was the cognitive revolution?
The cognitive revolution was the mid-twentieth-century shift that restored the mind to psychology after behaviorism had exiled it, treating internal representations and computations as legitimate theoretical entities inferred from behavior. It kept behaviorism's insistence on measurement while regaining the ability to explain language, memory, and reasoning (Miller, 2003).

Why do so many psychology studies use Western university students?
Undergraduates are convenient and cheap participants, so the field came to rely on them, but they belong to a narrow slice of humanity, Western, educated, industrialized, rich, and democratic, that is a global outlier on many behavioral measures. Findings from this population may not generalize to humanity as a whole (Henrich et al., 2010).

What caused the replication crisis?
The replication crisis arose mainly from structural problems in honest research: chronically low statistical power, flexibility in data analysis, and publication bias that suppressed null results. Together these filled the literature with findings, including many false positives and inflated effect sizes, that failed to reproduce in high-powered replications (Open Science Collaboration, 2015).

Does the replication crisis mean psychology is not a science?
No; the opposite reading is more apt. A field that can measure its own replication rate, publish the disappointing figure, and reform its methods in response is exhibiting the self-correction that defines science. The reforms of preregistration, larger samples, and open data are that correction underway (Munafò et al., 2017).

What is behavioral economics?
Behavioral economics is the discipline that studies how human choice systematically departs from the rational-agent model economics traditionally assumed. Built on documented, replicable deviations from that ideal, it has grown from a catalogue of anomalies into an applied field that designs and evaluates choice environments (Thaler, 2016).

References

Bateson, P., & Laland, K. N. (2013). Tinbergen's four questions: An appreciation and an update. Trends in Ecology & Evolution, 28(12), 712-718. https://doi.org/10.1016/j.tree.2013.09.013

Camerer, C. F., Dreber, A., Holzmeister, F., Ho, T.-H., Huber, J., Johannesson, M., et al. (2018). Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015. Nature Human Behaviour, 2(9), 637-644. https://doi.org/10.1038/s41562-018-0399-z

Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81-105. https://doi.org/10.1037/h0046016

Cialdini, R. B., & Goldstein, N. J. (2004). Social influence: Compliance and conformity. Annual Review of Psychology, 55, 591-621. https://doi.org/10.1146/annurev.psych.55.090902.142015

Cronbach, L. J. (1957). The two disciplines of scientific psychology. American Psychologist, 12(11), 671-684. https://doi.org/10.1037/h0043943

Engel, G. L. (1977). The need for a new medical model: A challenge for biomedicine. Science, 196(4286), 129-136. https://doi.org/10.1126/science.847460

Gigerenzer, G., & Gaissmaier, W. (2011). Heuristic decision making. Annual Review of Psychology, 62, 451-482. https://doi.org/10.1146/annurev-psych-120709-145346

Hagger, M. S., Chatzisarantis, N. L. D., Alberts, H., Anggono, C. O., Batailler, C., Birt, A. R., et al. (2016). A multilab preregistered replication of the ego-depletion effect. Perspectives on Psychological Science, 11(4), 546-573. https://doi.org/10.1177/1745691616652873

Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2-3), 61-83. https://doi.org/10.1017/S0140525X0999152X

Meehl, P. E. (1978). Theoretical risks and tabular asterisks: Sir Karl, Sir Ronald, and the slow progress of soft psychology. Journal of Consulting and Clinical Psychology, 46(4), 806-834. https://doi.org/10.1037/0022-006X.46.4.806

Miller, G. A. (2003). The cognitive revolution: A historical perspective. Trends in Cognitive Sciences, 7(3), 141-144. https://doi.org/10.1016/S1364-6613(03)00029-9

Munafo, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers, C. D., Percie du Sert, N., et al. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1(1), 0021. https://doi.org/10.1038/s41562-016-0021

Muthukrishna, M., & Henrich, J. (2019). A problem in theory. Nature Human Behaviour, 3(3), 221-229. https://doi.org/10.1038/s41562-018-0522-1

Nosek, B. A., Hardwicke, T. E., Moshontz, H., Allard, A., Corker, K. S., Dreber, A., et al. (2022). Replicability, robustness, and reproducibility in psychological science. Annual Review of Psychology, 73, 719-748. https://doi.org/10.1146/annurev-psych-020821-114157

Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716

Plomin, R., DeFries, J. C., Knopik, V. S., & Neiderhiser, J. M. (2016). Top 10 replicated findings from behavioral genetics. Perspectives on Psychological Science, 11(1), 3-23. https://doi.org/10.1177/1745691615617439

Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366. https://doi.org/10.1177/0956797611417632

Skinner, B. F. (1981). Selection by consequences. Science, 213(4507), 501-504. https://doi.org/10.1126/science.7244649

Thaler, R. H. (2016). Behavioral economics: Past, present, and future. American Economic Review, 106(7), 1577-1600. https://doi.org/10.1257/aer.106.7.1577

Tinbergen, N. (1963). On aims and methods of ethology. Zeitschrift fur Tierpsychologie, 20(4), 410-433. https://doi.org/10.1111/j.1439-0310.1963.tb01161.x

Watson, J. B. (1913). Psychology as the behaviorist views it. Psychological Review, 20(2), 158-177. https://doi.org/10.1037/h0074428

Yarkoni, T. (2022). The generalizability crisis. Behavioral and Brain Sciences, 45, e1. https://doi.org/10.1017/S0140525X20001685