Abstract
Behavioral research is the branch of the behavioral sciences concerned with how behavior is studied: the designs, measurements, and inferential rules by which claims about what organisms do are established and tested. It spans the controlled experiment, the quasi-experiment, the correlational survey, and the single-subject reversal design, and it rests on a measurement theory that turns unobservable constructs into quantities and on a statistics that separates signal from noise. This article surveys those designs, the logic of statistical power, and the ethical constraints that govern research on living participants, before turning to the replication crisis that forced the field to examine how it generates knowledge. Three interactive demonstrations let the reader balance a confounder by randomization, trace a statistical power curve, and read the experimental control of a single-subject design.
Keywords: behavioral research, research design, statistical power, replication crisis, research ethics
Behavioral research is the systematic empirical investigation of behavior, the working machinery beneath every claim the behavioral sciences make. Where a substantive theory asserts that a treatment reduces anxiety or that a cue speeds a response, behavioral research is the discipline that decides whether such an assertion has been established: how the study was designed, how the variables were measured, how many participants were needed, what could have confounded the result, and how likely the finding is to hold on repetition. Its subject is therefore method rather than any particular phenomenon, and its history is a steady accumulation of tools for wringing trustworthy inference from the noisy, variable conduct of living things. The field inherits its founding standard from John Watson's insistence that psychology restrict itself to the prediction and control of observable behavior, a demand that made the response a measurable unit and the experiment its natural test (Watson, 1913).
- Behavioral research is defined by method, not topic: it is the body of designs, measures, and inferential rules that turn observations of behavior into defensible claims.
- Its designs trade off internal validity, the confidence that the manipulation caused the effect, against external validity, the confidence that the effect generalizes beyond the study.
- Random assignment is the single most powerful device for internal validity, because it balances every confounder, known and unknown, across conditions in expectation.
- Statistical power, the probability of detecting a real effect, depends on effect size, sample size, and the significance threshold; chronically underpowered studies both miss true effects and exaggerate the ones they find.
- The replication crisis exposed how low power, analytic flexibility, and publication bias can fill a literature with unreliable findings, prompting reforms such as preregistration, registered reports, and large-scale replication.
What Behavioral Research Is
Behavioral research is the enterprise of studying behavior scientifically, which means holding every claim answerable to systematically gathered evidence that could in principle overturn it. It is not a single technique but a graded family of them, ordered roughly by how much control the researcher exerts over the situation. At one extreme sits the true experiment, in which the investigator manipulates an independent variable and randomly assigns participants to its levels, so that any resulting difference can be attributed to the manipulation rather than to preexisting differences among participants. At the other sits pure observation, in which behavior is recorded as it naturally unfolds, sacrificing causal control for fidelity to the real conditions under which behavior occurs, a tradition whose canonical statement is Niko Tinbergen's program for the objective observation of animals in their natural context (Tinbergen, 1963). Between these poles lie the quasi-experiment, which compares groups that were not randomly formed, and the correlational study, which measures variables without manipulating any of them. What unifies the family is a shared preoccupation with validity: the disciplined worry that an observed pattern might be produced by something other than the process the researcher means to study, and a corresponding toolkit, from control conditions to statistical adjustment, built to rule those alternatives out.
Types of Behavioral Research
Behavioral Research is also a formal descriptor in the National Library of Medicine's Medical Subject Headings, which files it beneath Behavioral Sciences at tree position F04.096.144 and cross-lists it under Research at H01.770.644.108. Beneath the descriptor MeSH hangs the single narrower heading listed in Table 1. Two cautions apply. The list is an indexing classification built to organize the biomedical literature, not a theory that carves the field at its joints, and its one current member overlaps heavily with the parent term and with neighboring branches of the tree. Only subtypes that are themselves live articles on this site are linked, and at present this descriptor has none.
| Subtype | In brief |
|---|---|
| Biobehavioral Sciences | The interdisciplinary study of the reciprocal relations between biological processes and behavior, joining the methods of behavioral research to those of the neural and physiological sciences. |
Designs and the Logic of Validity
The central problem of behavioral research is confounding: whenever two groups differ in the behavior of interest, they may also differ in some other respect that is the true cause, and a naive comparison cannot tell the two apart. The designs of the field are ranked by how effectively they defuse this problem. Donald Campbell's analysis of experimental and quasi-experimental designs gave the discipline its enduring vocabulary, distinguishing threats to internal validity, the many ways a study can create a spurious effect, from threats to external validity, the many ways a real effect can fail to generalize (Campbell & Stanley, 1963). The true experiment answers the internal-validity threats at a stroke through random assignment: by allocating participants to conditions by chance, it makes the groups equivalent in expectation on every variable, including those the researcher never thought to measure, so that a post-treatment difference has no plausible source but the treatment. The quasi-experiment, used when random assignment is impossible or unethical, must instead rule out confounders one at a time, through matching, statistical control, or clever comparison conditions, and its conclusions are correspondingly more tentative. The correlational study, which only measures, can establish that two variables travel together but cannot by itself say which causes which, or whether a third variable causes both. This ordering is not a ranking of worth, because the design best suited to a question depends on what can be manipulated and what must be observed; it is a ranking of the inferential price each design pays, and a competent researcher chooses with that price in view.
Figure 1
The Internal-External Validity Tradeoff Across Behavioral Research Designs
Random assignment deserves its central place because of a property no other device shares: it controls for confounders the researcher has not even imagined. Matching and statistical adjustment can only balance variables that have been identified and measured, but chance allocation balances everything in expectation, which is why the randomized experiment remains the field's strongest warrant for a causal claim. The demonstration below makes this visible by letting the reader assign a sample to two conditions with and without randomization and watch an unmeasured confounder come into or out of balance as the sample grows.
Every participant carries a hidden confounder (think of it as age). The two bars are its average in the treatment and control groups. Under random assignment the groups converge as the sample grows; under a biased rule they stay apart no matter how large the study.
Note. The confounder values are fixed and illustrative. Random assignment does not force exact equality; it makes the expected difference zero, so imbalance shrinks with sample size instead of persisting as it does under a systematic rule.
Measurement, Effect Size, and Statistical Power
A design is only half of an inference; the other half is measurement and the statistics that interpret it. Because behavioral variables are noisy, a single observation carries little information, and the field relies on aggregation and on the machinery of statistical testing to decide whether an apparent effect exceeds what chance alone would produce. The pivotal quantity is the effect size, the magnitude of a difference or association expressed in standardized units, which Jacob Cohen urged researchers to report and reason about in place of a bare verdict of significance (Cohen, 1992). From the effect size, the sample size, and the significance threshold follows statistical power: the probability that a study will detect a real effect of a given size. Power is the most neglected parameter in behavioral research and the most consequential, because a study with low power is a poor instrument in two distinct ways. It frequently misses true effects, producing false negatives that clutter the literature with failures to replicate; and, more insidiously, when an underpowered study does reach significance, the effect it reports is necessarily inflated, since only the largest, luckiest sample estimates clear the threshold. Katherine Button and colleagues demonstrated that across whole fields the median power to detect plausible effects can fall well below fifty percent, so that a large fraction of published positive findings are either false or exaggerated (Button et al., 2013). The remedy is not subtle. Powering a study adequately, usually by enlarging the sample, is the single most reliable way to make its results trustworthy, and the worked example below shows exactly how many participants a conventional target demands. The demonstration that follows lets the reader trace the power curve directly, watching detection probability climb as sample size and effect size vary.
Power is the probability of detecting a real effect of the chosen size. The curve climbs with sample size; the marker is the study you have specified. The dashed line is the 80 percent convention.
Note. Computed from the two-sample normal approximation, power = Φ(d√(n/2) − zα/2). A medium effect (d = 0.5) reaches 80 percent power near 63 to 64 participants per group, matching the worked example.
Single-Subject Designs and Experimental Control
Not all rigorous behavioral research compares groups. A distinct and powerful tradition, developed within the experimental analysis of behavior and formalized in applied behavior analysis, establishes causal control within a single organism by manipulating conditions over time rather than across participants. Donald Baer, Montrose Wolf, and Todd Risley set out the standards of this approach, in which behavior is measured repeatedly across alternating phases, a baseline followed by an intervention followed by a return to baseline, so that the behavior can be seen to rise and fall with the manipulation (Baer et al., 1968). The logic is that of replication within the individual: if the target behavior changes when and only when the intervention is introduced and withdrawn, the intervention, not some coincident event, is the credible cause. This reversal or ABAB design substitutes the participant's own baseline for a separate control group, which makes it invaluable where large samples are impossible, as in the study of rare conditions or intensive clinical treatment, and where the effect of interest is large and reversible. Its limits mirror its strengths: a treatment whose effect persists after withdrawal cannot be demonstrated by reversal, and generalization from one organism to a population requires the accumulation of many such demonstrations. The demonstration below renders an idealized ABAB record so the reader can read experimental control at the level of the single case, seeing how the pattern across phases licenses the causal claim.
One participant is measured repeatedly across a baseline (A), an intervention (B), a return to baseline (A), and a reinstatement (B). If the behavior tracks the phases, the intervention is the credible cause.
Note. An idealized record with fixed session wobble. Causal force in a single-subject design comes from the replicated reversal: the effect appears in both B phases and disappears in both A phases, after Baer, Wolf, and Risley (1968).
Research Ethics and the Constraints on Method
Behavioral research studies living participants, and the freedom to manipulate their circumstances is bounded by an ethical framework that the field learned, in part, from its own failures. Stanley Milgram's obedience studies, in which participants were led to believe they were delivering painful shocks to another person, produced findings of lasting importance about the situational power of authority while also subjecting participants to severe distress under deception (Milgram, 1963). Studies of this kind, alongside notorious abuses in biomedical research, drove the codification of principles that now govern work with human participants. The Belmont Report distilled them into three: respect for persons, expressed through voluntary informed consent; beneficence, the obligation to minimize harm and maximize benefit; and justice in the distribution of research burdens (National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research, 1979). These principles are enforced through institutional review, and they impose real methodological constraints, since the strongest design for a question is often the one ethics forbids. The resulting tension is intrinsic to the field rather than incidental: a discipline that could do anything to its participants could answer more questions cleanly, and the ethical limits on manipulation are precisely what force behavioral researchers toward quasi-experimental and observational designs, and toward the ingenuity that makes weaker designs informative. Research on nonhuman animals carries its own parallel framework of welfare constraints, and the observational ethology that Tinbergen championed is in part a response to the recognition that much cannot, and need not, be studied through intervention.
The Replication Crisis and Methodological Reform
Beginning around 2011, behavioral research turned its instruments on itself and found the results disquieting. John Ioannidis had argued from first principles that under common conditions of low power, analytic flexibility, and low prior odds, most published research findings in a field will be false (Ioannidis, 2005), and a concrete demonstration soon followed: Joseph Simmons and colleagues showed that the undisclosed choices open to an analyst, which measures to report, when to stop collecting data, which observations to exclude, can be exploited until a result crosses the significance threshold, driving the false-positive rate far above its nominal level while every individual decision looks defensible (Simmons et al., 2011). The empirical scale of the problem became visible when the Open Science Collaboration ran high-powered replications of one hundred psychology studies and reproduced only about a third of the originally significant effects, at roughly half their published magnitude (Open Science Collaboration, 2015), and when a coordinated replication of many effects across dozens of laboratories confirmed that some celebrated findings were robust while others essentially vanished under scrutiny (Klein et al., 2018). A parallel effort in the social sciences reproduced only about sixty percent of experiments drawn from the most prestigious journals, again at diminished magnitude (Camerer et al., 2018). The causes were understood to be structural rather than fraudulent, rooted in incentives that reward novel positive results over reliable ones, an arrangement that Paul Smaldino and Richard McElreath modeled as a kind of natural selection for bad science, in which methods that produce publishable but unreliable findings outcompete more rigorous ones (Smaldino & McElreath, 2016). The reform that followed reworked the field's procedures rather than merely exhorting its members: a manifesto for reproducible science laid out concrete changes to methods, reporting, reproducibility, and incentives (Munafò et al., 2017); preregistration was advanced as a device to separate confirmatory prediction from exploratory analysis by fixing hypotheses and analysis plans before the data exist (Nosek et al., 2018); and the registered report, in which a study is reviewed and accepted on the strength of its question and design before its results are known, was introduced to break the link between a finding's direction and its publishability (Chambers, 2013). Commentators came to describe the upheaval not as a collapse but as a renaissance, a period in which the field rebuilt its methods on firmer ground (Nelson et al., 2018).
Worked Example
The abstract injunction to power a study adequately becomes concrete the moment a researcher asks how many participants a planned study requires. Consider the most common design in behavioral research: a comparison of two independent groups, a treatment and a control, on a continuous outcome, tested with a two-tailed t test at the conventional significance level of 0.05. Suppose the researcher judges, from theory and prior work, that an effect worth detecting is of medium size, which in standardized units is a difference of 0.5 standard deviations, that is, Cohen's d equal to 0.5. The question is the sample size per group needed to achieve 80 percent power, the widely adopted convention. The normal approximation to the required per-group sample size is n equals 2 times the quantity (z-alpha plus z-beta) squared, divided by d squared, where z-alpha is the standard normal value cutting off the two-tailed significance region and z-beta the value corresponding to the desired power. For a two-tailed test at 0.05, z-alpha is 1.96; for 80 percent power, z-beta is 0.84. Substituting, the sum 1.96 plus 0.84 equals 2.80, and its square is 7.84. Multiplying by 2 gives 15.68, and dividing by d squared, which is 0.25, yields 62.7, so the study needs 63 participants per group, or 64 by Cohen's slightly more conservative tabled value (Cohen, 1992). The lesson is bracing: detecting a medium effect with adequate power takes roughly 128 participants in total, far more than the underpowered samples that long typified the field, and detecting a small effect, d equal to 0.2, would require about 393 per group. Halving the demanded sample by relaxing power to a coin-flip 50 percent is exactly the false economy that fills a literature with irreproducible results, because the studies that survive to publication are then a biased sample of overestimates (Button et al., 2013).
Discussion
Behavioral research occupies a peculiar position: it is the part of the behavioral sciences that studies not behavior itself but the act of studying behavior, and its progress is measured less in discoveries than in the growing trustworthiness of discovery. What the field has built is a layered defense against error. Design controls for confounding, most decisively through random assignment; measurement theory disciplines the leap from an observable indicator to an unobservable construct; power analysis ensures that a study is a sensitive enough instrument to detect what it seeks; and ethical review constrains the whole enterprise to what may permissibly be done to participants. The crises of the past fifteen years are best read not as a refutation of this apparatus but as its most rigorous application, the field turning its own tools of measurement and inference on its published record and reforming its practice in light of what it found. The open questions are correspondingly methodological. How much of the reform agenda, preregistration, registered reports, larger samples, open data, will harden into permanent infrastructure rather than fading as attention moves on? Can the incentives that reward novelty over reliability be restructured at the level of hiring, funding, and publication, where Smaldino and McElreath locate the true source of unreliable science (Smaldino & McElreath, 2016)? And how should the field balance the internal validity of tightly controlled designs against the external validity that its ambition to explain real behavior demands? What is settled is that behavioral research has become self-conscious about its own methods to a degree few sciences match, and that this reflexivity, uncomfortable as it has been, is the mechanism by which the field corrects itself.
Current Directions
The most active front in contemporary behavioral research is the consolidation of the reform movement from critique into standing infrastructure. Reviews now treat replicability, robustness, and reproducibility as distinct, separately measurable properties of a result rather than a single vague virtue, and preregistration, registered reports, and large collaborative replication have moved from novelty to routine (Nosek et al., 2022). A second current runs through the statistics of the field, where the reflexive reliance on null-hypothesis significance testing is under sustained pressure from proposals to report estimation and uncertainty directly, to justify sample sizes through formal design analysis, and to adopt Bayesian alternatives that quantify evidence rather than merely rejecting a null. A third concerns the machinery of research itself: preregistration platforms, shared data repositories, and analytic pipelines that record every decision have begun to make the full provenance of a finding inspectable, turning the researcher degrees of freedom that once enabled p-hacking into a documented and reviewable record. Running beneath all three is a growing interest in metascience, the empirical study of research practice as itself an object of behavioral research, which measures how scientists actually design, analyze, and report their work and tests interventions meant to improve it. The field, in short, has extended its methods to its own conduct, and treats the reliability of its knowledge as a question to be investigated rather than assumed.
Common Misconceptions
- Correlation and causation are simply unrelated, so a correlational study tells us nothing.
- A correlation is genuine evidence; what it cannot do by itself is fix the direction of causation or exclude a common cause. Correlational designs remain indispensable where manipulation is impossible or unethical, and modern methods can strengthen causal inference from them, but the inference is always more fragile than that from a randomized experiment (Campbell & Stanley, 1963).
- A statistically significant result is an important or a large one.
- Significance reports only that an effect is unlikely under the null hypothesis, not that it is large or consequential; with a big enough sample a trivial effect is significant, and with a small sample an important one is not. This is why the field increasingly asks for effect sizes and confidence intervals alongside the verdict of a test (Cohen, 1992).
- A failure to replicate proves the original researchers committed fraud.
- Non-replication arises overwhelmingly from honest practice under bad incentives, chiefly low power, analytic flexibility, and publication bias, not from misconduct. A field of scrupulous researchers can still produce many unreliable findings when its studies are underpowered and only positive results reach print (Open Science Collaboration, 2015).
Glossary
- Confounding variable.
- A variable that influences both the presumed cause and the presumed effect, creating a spurious association that a design must rule out.
- Correlational study.
- A design that measures variables without manipulating any of them, able to establish association but not, by itself, causal direction.
- Effect size.
- A standardized measure of the magnitude of a difference or association, such as Cohen's d, independent of sample size.
- External validity.
- The extent to which a finding generalizes beyond the specific participants, stimuli, and setting of the study that produced it.
- Informed consent.
- The voluntary agreement to participate given after disclosure of a study's risks and purposes, the operational form of respect for persons.
- Internal validity.
- The confidence that the manipulation, rather than a confound, produced the observed effect within the study.
- P-hacking.
- Exploiting undisclosed analytic choices until a result crosses the significance threshold, inflating the false-positive rate while each decision looks reasonable.
- Preregistration.
- The practice of specifying hypotheses and an analysis plan before collecting data, separating confirmatory prediction from exploratory postdiction.
- Quasi-experiment.
- A design that compares conditions without random assignment, requiring confounders to be ruled out individually rather than by chance allocation.
- Random assignment.
- The allocation of participants to conditions by chance, which balances all confounders, known and unknown, across groups in expectation.
- Registered report.
- A publication format in which a study is reviewed and accepted on its question and design before results are known, breaking the link between outcome and publishability.
- Replication.
- The repetition of a study to test whether its finding holds, the central check on the reliability of a behavioral claim.
- Reversal (ABAB) design.
- A single-subject design that alternates baseline and intervention phases so a behavior can be seen to change with, and only with, the manipulation.
- Single-subject design.
- A family of designs that establish experimental control within one organism through repeated measurement across manipulated phases.
- Statistical power.
- The probability that a study will detect a true effect of a given size; it rises with effect size, sample size, and the significance level.
- True experiment.
- A design in which the researcher manipulates an independent variable and randomly assigns participants to its levels, the strongest warrant for a causal claim.
Key Researchers
Donald T. Campbell (1916-1996). Social scientist at Northwestern University; his codification of experimental and quasi-experimental designs and of the threats to internal and external validity gave behavioral research its enduring methodological vocabulary. Wikipedia - Wikidata
Christopher D. Chambers. Cognitive neuroscientist at Cardiff University; the originator of the registered report format, which reviews and accepts a study on its design before its results are known. Faculty Page - ORCID - Google Scholar - Wikipedia
Jacob Cohen (1923-1998). Psychologist and statistician at New York University; his advocacy of effect sizes, power analysis, and the primacy of estimation over bare significance reshaped how behavioral researchers plan and interpret studies. Wikipedia - Wikidata
John P. A. Ioannidis. Physician and meta-researcher at Stanford University; his argument that most published research findings may be false crystallized the case for reforming how behavioral and biomedical research is conducted. Faculty Page - ORCID - Google Scholar - Wikipedia
Brian A. Nosek. Social psychologist at the University of Virginia and co-founder of the Center for Open Science; he led the large-scale replication and preregistration efforts that reshaped research practice across the behavioral sciences. Faculty Page - ORCID - Google Scholar - Wikipedia
Uri Simonsohn. Behavioral scientist at Esade Business School; his demonstration of how researcher degrees of freedom inflate false positives named and diagnosed p-hacking, a turning point in the field's self-examination. Faculty Page - ORCID - Google Scholar - Wikipedia
Nikolaas Tinbergen (1907-1988). Ethologist at the University of Oxford and Nobel laureate; his program for the objective observation of animals in their natural context established the rigorous end of the observational tradition in behavioral research. Wikipedia - Wikidata
John B. Watson (1878-1958). Psychologist at Johns Hopkins University whose 1913 behaviorist manifesto redefined psychology as the objective, experimental study of behavior and set the methodological course that behavioral research still follows. Wikipedia - Wikidata
Frequently Asked Questions
What is behavioral research?
Behavioral research is the systematic empirical investigation of behavior: the designs, measurements, and inferential rules by which claims about what organisms do are established and tested. Its subject is method rather than any single phenomenon, and it supplies the machinery, from controlled experiments to statistical power analysis, on which the substantive claims of the behavioral sciences depend (Watson, 1913).
What is the difference between internal and external validity?
Internal validity is the confidence that the manipulation, and not some confound, produced the effect within a study; external validity is the confidence that the effect generalizes beyond the study to other people, stimuli, and settings. The two often trade off, because the tightly controlled designs that secure internal validity tend to sample narrow, artificial conditions (Campbell & Stanley, 1963).
Why is random assignment so important?
Random assignment allocates participants to conditions by chance, which makes the groups equivalent in expectation on every variable, including confounders no one thought to measure. Matching and statistical adjustment can only balance identified variables, so randomization remains the strongest single warrant for a causal claim (Campbell & Stanley, 1963).
What is statistical power and why does it matter?
Statistical power is the probability that a study will detect a real effect of a given size, and it rises with effect size, sample size, and the significance level. Low power both misses true effects and exaggerates the ones it detects, because only inflated estimates clear the threshold, which is why underpowered research is a principal source of irreproducible findings (Button et al., 2013).
How large a sample does a study need?
It depends on the effect size sought, the desired power, and the significance level. For the common case of detecting a medium effect, Cohen's d of 0.5, with 80 percent power at a two-tailed 0.05 level, a two-group study needs roughly 63 to 64 participants per group, about 128 in total; smaller effects demand far larger samples (Cohen, 1992).
What is a single-subject design?
A single-subject design establishes experimental control within one organism by measuring behavior repeatedly across alternating phases, such as a baseline, an intervention, and a return to baseline. If the behavior changes when and only when the intervention is introduced and withdrawn, the intervention is the credible cause, which makes the approach valuable where large samples are impossible (Baer et al., 1968).
What caused the replication crisis?
The crisis arose mainly from structural features of honest research: chronically low statistical power, flexibility in data analysis that permits p-hacking, and publication bias that suppresses null results. Together these filled the literature with false positives and inflated effects that failed to reproduce in high-powered replications (Open Science Collaboration, 2015; Simmons et al., 2011).
What is a registered report?
A registered report is a publication format in which a study's question and design are peer reviewed and provisionally accepted before the data are collected, so that publication depends on the quality of the plan rather than the direction of the results. It is one of the central reforms aimed at curbing publication bias and analytic flexibility (Chambers, 2013).
References
Baer, D. M., Wolf, M. M., & Risley, T. R. (1968). Some current dimensions of applied behavior analysis. Journal of Applied Behavior Analysis, 1(1), 91-97. https://doi.org/10.1901/jaba.1968.1-91
Button, K. S., Ioannidis, J. P. A., Mokrysz, C., Nosek, B. A., Flint, J., Robinson, E. S. J., & Munafo, M. R. (2013). Power failure: Why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 14(5), 365-376. https://doi.org/10.1038/nrn3475
Camerer, C. F., Dreber, A., Holzmeister, F., Ho, T.-H., Huber, J., Johannesson, M., et al. (2018). Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015. Nature Human Behaviour, 2(9), 637-644. https://doi.org/10.1038/s41562-018-0399-z
Campbell, D. T., & Stanley, J. C. (1963). Experimental and quasi-experimental designs for research. In N. L. Gage (Ed.), Handbook of research on teaching (pp. 171-246). Rand McNally.
Chambers, C. D. (2013). Registered reports: A new publishing initiative at Cortex. Cortex, 49(3), 609-610. https://doi.org/10.1016/j.cortex.2012.12.016
Cohen, J. (1992). A power primer. Psychological Bulletin, 112(1), 155-159. https://doi.org/10.1037/0033-2909.112.1.155
Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124. https://doi.org/10.1371/journal.pmed.0020124
Klein, R. A., Vianello, M., Hasselman, F., Adams, B. G., Adams, R. B., Alper, S., et al. (2018). Many Labs 2: Investigating variation in replicability across samples and settings. Advances in Methods and Practices in Psychological Science, 1(4), 443-490. https://doi.org/10.1177/2515245918810225
Milgram, S. (1963). Behavioral study of obedience. Journal of Abnormal and Social Psychology, 67(4), 371-378. https://doi.org/10.1037/h0040525
Munafo, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers, C. D., Percie du Sert, N., et al. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1(1), 0021. https://doi.org/10.1038/s41562-016-0021
National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research. (1979). The Belmont Report: Ethical principles and guidelines for the protection of human subjects of research. U.S. Department of Health, Education, and Welfare (DHEW Publication No. (OS) 78-0012).
Nelson, L. D., Simmons, J. P., & Simonsohn, U. (2018). Psychology's renaissance. Annual Review of Psychology, 69, 511-534. https://doi.org/10.1146/annurev-psych-122216-011836
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600-2606. https://doi.org/10.1073/pnas.1708274114
Nosek, B. A., Hardwicke, T. E., Moshontz, H., Allard, A., Corker, K. S., Dreber, A., et al. (2022). Replicability, robustness, and reproducibility in psychological science. Annual Review of Psychology, 73, 719-748. https://doi.org/10.1146/annurev-psych-020821-114157
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366. https://doi.org/10.1177/0956797611417632
Smaldino, P. E., & McElreath, R. (2016). The natural selection of bad science. Royal Society Open Science, 3(9), 160384. https://doi.org/10.1098/rsos.160384
Tinbergen, N. (1963). On aims and methods of ethology. Zeitschrift fur Tierpsychologie, 20(4), 410-433. https://doi.org/10.1111/j.1439-0310.1963.tb01161.x
Watson, J. B. (1913). Psychology as the behaviorist views it. Psychological Review, 20(2), 158-177. https://doi.org/10.1037/h0074428