Abstract

Behavior observation techniques are a family of psychological techniques for systematically watching, coding, and recording behavior as it occurs, whether in natural settings or in the laboratory. They convert a continuous stream of action into quantifiable data through three linked decisions: what to record, when to sample it, and how to check that two observers agree. The methods range from the ethologist's field ethogram to interval time-sampling, from a chance-corrected reliability coefficient to automated video coding of facial movement. Because an observer is a measuring instrument, the central methodological problems are reliability, reactivity, and bias rather than sensitivity or resolution. Rigorous observation therefore rests on explicit operational definitions, disciplined sampling rules, and quantified interobserver agreement, and it remains the primary route to measuring behavior that cannot be captured by self-report.

Keywords: behavior observation, coding scheme, interobserver reliability, sampling methods, observer bias

Behavior observation techniques are procedures for measuring behavior by watching it directly and recording it against a defined scheme, rather than by inferring it from a questionnaire, a physiological trace, or a task score. Their intellectual roots lie in ethology, where the disciplined description of naturally occurring behavior was made a science in its own right: Tinbergen's programmatic statement of the aims and methods of the field established that observation, properly constrained, answers questions about causation, development, function, and evolution (Tinbergen, 1963). The same logic was imported into developmental, clinical, and social psychology, where much of what matters — a child's aggression, a couple's conflict, a patient's compliance — is more validly seen than reported. The defining move is to treat the human observer as an instrument, and to build the safeguards that any instrument requires.

Key Takeaways
  • Behavior observation turns a continuous stream of action into data through explicit choices about what, when, and how to record.
  • Sampling rules — continuous, momentary, one-zero, focal, and scan — each answer different questions and carry different biases.
  • Because the observer is the instrument, interobserver reliability must be quantified with a chance-corrected coefficient, not raw percent agreement.
  • Reactivity, observer drift, and expectation bias degrade data unless designs and covert reliability checks control them.
  • Automated video and facial-action coding now extend observation to a scale and consistency manual coders cannot match.

Sampling Methods

Behavior is continuous, but recording cannot be; the first methodological decision is how to sample the stream. Altmann's analysis of sampling methods remains the canonical treatment, distinguishing the designs by what unit is watched and how time is partitioned (Altmann, 1974). Ad libitum sampling records whatever is conspicuous and is useful only for generating hypotheses, because salient events are overrepresented. Focal-animal sampling fixes attention on one individual for a set period and records everything it does, giving unbiased rate and duration data at the cost of coverage. Scan sampling sweeps a whole group at fixed instants and records each member's current state, trading detail for breadth.

Orthogonal to who is watched is how time is treated. Continuous recording logs every onset and offset, yielding true frequencies and durations but demanding the most from the coder. Interval recording divides the session into short blocks and asks only whether a behavior occurred. Under momentary time sampling the coder records the state at the instant each interval begins; under one-zero sampling the coder marks an interval if the behavior occurred at any point within it. The distinction is not pedantic: one-zero sampling systematically overestimates the proportion of time a behavior occupies, and the overestimation grows with interval width, whereas momentary sampling is approximately unbiased for that quantity (Altmann, 1974). Choosing a rule is therefore choosing which quantity can be estimated without bias.

Sampling the Same Behavior Three Ways

A fixed 60-second session contains five bouts of the target behavior (gold). Continuous recording sees all of it; interval methods sample it. Adjust the interval width to watch one-zero sampling inflate the estimate.

Behavior streamGreen dots = momentary samples at each interval start (filled = in a bout)True 52%
True proportion of time on the behavior: 51.7%. Momentary time sampling: 50.0% (close: unbiased in expectation, though noisy with few intervals). One-zero sampling: 75.0% (overestimate, and it grows as the interval widens).

Illustrative of Altmann (1974): one-zero sampling counts an interval whenever the behavior touches it, so wider intervals sweep in more and overstate duration. Values computed locally, not stored.

Coding Schemes and Operational Definitions

A sampling rule is worthless without a coding scheme: the finite set of mutually exclusive, exhaustively defined categories into which behavior is sorted. The ethological ancestor is the ethogram, a catalogue of a species' behavioral repertoire in which each act is described by its form precisely enough that another observer would identify it the same way (Tinbergen, 1963). The modern equivalent in human research is a codebook of operational definitions, each specifying the observable features that qualify an instance and, crucially, the boundary cases that do not. Developing such a scheme is iterative — categories are drafted, tested against pilot recordings, split when they prove heterogeneous, and merged when coders cannot tell them apart — and the process is itself a reportable part of method (Chorney et al., 2015).

Codes divide into two kinds whose statistics differ. Events are discrete, near-instantaneous acts counted as frequencies; states are behaviors with appreciable duration, for which onset and offset define a bout. The distinction governs which measures are meaningful and which sequential analyses apply, because the dependency structure of a stream of timed states differs from that of a series of point events (Bakeman & Gottman, 1997). Where the research question concerns not merely how often behaviors occur but how they follow one another — whether a mother's bid reliably precedes an infant's response — the coded stream becomes the input to sequential analysis, which tests temporal contingencies with contingency tables and lag statistics (Bakeman & Gottman, 1997).

Interobserver Reliability

If two trained observers watching the same behavior produce different records, the scheme is not measuring anything stable. Interobserver reliability quantifies their agreement, and the central lesson is that raw percent agreement is misleading, because two coders will agree by chance at a rate set by the base rates of the categories. Cohen's kappa corrects for that chance agreement, expressing the excess of observed over expected agreement as a fraction of the maximum possible excess (Cohen, 1960). A kappa of zero means agreement no better than chance; one means perfect agreement. Fleiss generalised the coefficient to any number of raters, so that a coding team rather than a pair can be evaluated (Fleiss, 1971).

Kappa is a number, not a verdict, and its interpretation needs a benchmark. Landis and Koch proposed the widely used bands — slight, fair, moderate, substantial, and almost perfect agreement — that give a kappa value a qualitative reading (Landis & Koch, 1977). The choice of reliability statistic is itself consequential: percent agreement, kappa, and the several intraclass correlations answer subtly different questions, and reporting the wrong one can flatter or unfairly penalise a coding scheme, so the estimate must be matched to the data type and the claim (Hartmann, 1977). None of this is a one-time hurdle; reliability is a property of a particular team coding particular material and must be re-established rather than assumed (Kazdin, 1977). High agreement is also necessary but not sufficient: two observers can reliably code a category that fails to capture the construct it names, so the validity of a coding scheme — whether its codes mean what the research question requires — is a separate question from its reliability (Girard & Cohn, 2016).

Cohen's Kappa: Agreement Beyond Chance

Two observers each code every interval as on-task or off-task. Enter how many intervals fall in each cell of their agreement table. The defaults are the article's worked example.

B: on-taskB: off-task
A: on-task4515
A: off-task1030
N = 100 intervals. Observed agreement po = 0.750 (75%). Chance agreement pe = 0.510. Cohen’s kappa = 0.490 moderate agreement.

Notice how raw percent agreement can stay high while kappa falls, because common categories make chance agreement large (Cohen, 1960; bands after Landis & Koch, 1977). Computed locally, not stored.

Reactivity, Bias, and Observer Drift

Treating the observer as an instrument makes its failure modes explicit, and they are not the sensitivity limits of a physical sensor but distinctively human distortions. Reactivity is the change in behavior produced by being watched: participants who know they are observed may suppress or amplify the very behavior of interest, so unobtrusive methods and habituation periods are used to let behavior return to baseline. Observer bias is systematic error introduced by the coder's expectations. When observers are led to expect improvement, their records drift toward it even when the behavior does not change, an expectation effect demonstrated directly in observational evaluation of therapy (Kent et al., 1974).

A subtler problem is observer drift: the gradual, shared redefinition of a category by coders who calibrate against one another rather than the codebook. Reid showed that observers whose reliability was assessed openly maintained high agreement, but that agreement fell sharply when they believed they were no longer being checked — reliability measured under overt conditions overstated the reliability of the data actually collected (Reid, 1970). The remedy is designed into the protocol: covert or unpredictable reliability checks, periodic recalibration against the original definitions, and reporting reliability on a random sample of the real sessions rather than on a rehearsal. These artifacts, and the safeguards against them, are as much a part of observational method as the coding scheme itself (Kazdin, 1977).

Observer Drift and the Covert Reliability Check

Two coders start well calibrated (kappa 0.85). Left unchecked, their agreement with the original codebook decays as they redefine categories against each other. Turn on periodic reliability checks to recalibrate.

substantial (0.6)moderate (0.4)0.80.4Session
Final agreement after 20 sessions: kappa 0.28 — unchecked drift has eroded it below the substantial threshold.

Illustrative linear-decay model of the effect Reid (1970) documented: reliability assessed only overtly overstates the data actually collected. Values computed locally, not stored.

Automated Behavior Measurement

The manual observer's limits — fatigue, drift, and the sheer cost of coding — motivated a century-long effort to make behavioral measurement objective and mechanical, nowhere more fully than for the face. Ekman and Friesen's Facial Action Coding System decomposed any facial movement into a fixed inventory of anatomically defined action units, turning expression from an impression into a codeable variable and giving human coders a common language (Ekman, 1993). Manual action-unit coding is nonetheless slow, and its reliability still had to be established coder by coder, which is precisely the burden that automation promised to lift.

Computer-vision systems now detect facial landmarks, head pose, gaze, and action-unit intensity from ordinary video, and open toolkits have made such coding routine rather than specialist (Baltrusaitis et al., 2018). A body of work on the automatic analysis of facial actions documents both the progress and the residual gap: machine coding is fast, tireless, and perfectly consistent with itself, but its accuracy still depends on lighting, pose, and the representativeness of the training data, and it inherits any bias in the human-coded material it learns from (Martinez et al., 2019). Automation therefore relocates the classic reliability question rather than dissolving it, and the field now asks how well an algorithm agrees with expert human coders and how measurement should be reported when one instrument is a model (Girard & Cohn, 2016).

Figure 1

The Observation Pipeline: From Behavior Stream to Verified Data

Four-stage pipeline of behavior observation A left-to-right flow: a continuous behavior stream is sampled by a rule, sorted by a coding scheme into categories, and checked by a second observer for interobserver reliability before becoming verified data. Behavior stream continuous action Sampling rule what, when Coding scheme defined categories Reliability second coder, kappa verified data Threats: reactivity, observer bias, drift
Note. Each stage adds a decision that shapes the data; the dashed line marks the threats that operate throughout and that reliability checks are designed to catch. Original schematic.
Table 1. Common sampling rules and the quantity each estimates without bias.
Sampling rule What is recorded Estimates well Main limitation
Continuous recordingEvery onset and offsetFrequency and durationHeaviest coder load
Momentary time samplingState at each interval's instantProportion of time (unbiased)Misses brief events
One-zero samplingWhether it occurred in intervalPresence, roughlyOverestimates time; grows with interval
Focal-animal samplingAll acts of one individualIndividual rates, unbiasedLimited group coverage
Scan samplingWhole group at fixed instantsGroup state proportionsLittle sequential detail

Worked Example

Two observers independently code 100 one-minute intervals of a classroom session, marking each interval as on-task or off-task. Their records cross-tabulate as follows: both call it on-task in 45 intervals, both off-task in 30; observer A codes on-task while B codes off-task in 15 intervals, and A off-task while B on-task in 10. Raw agreement is the sum of the diagonal, (45 + 30) / 100 = 0.75, an apparently strong 75 percent.

Cohen's kappa asks how much of that agreement beats chance. Observer A coded on-task in 45 + 15 = 60 intervals (0.60) and off-task in 40 (0.40); observer B coded on-task in 45 + 10 = 55 (0.55) and off-task in 45 (0.45). The chance agreement is the sum over categories of the product of the marginals: (0.60 x 0.55) + (0.40 x 0.45) = 0.33 + 0.18 = 0.51. Kappa is then the observed agreement in excess of chance, scaled by the room left above chance: (0.75 - 0.51) / (1 - 0.51) = 0.24 / 0.49 = 0.49 (Cohen, 1960). By the Landis and Koch bands a kappa of 0.49 is only moderate agreement (Landis & Koch, 1977), so a headline figure of 75 percent agreement conceals a coding scheme that is merely adequate. The Kappa Calculator demonstration reproduces this table and recomputes the coefficient as the cell counts change.

Discussion

The through-line of behavior observation is that the method's rigor lives in its safeguards, not in the act of watching. Every stage of the pipeline in Figure 1 embeds a decision that can bias the result — a sampling rule that estimates the wrong quantity, a category so loose that coders diverge, a reliability figure inflated by chance or by overt assessment — and the discipline of the field is the accumulated set of corrections for each. This is what separates systematic observation from mere description, and it is why an observational study reports its coding scheme, its sampling rule, and its interobserver reliability as a matter of course.

Observation retains a standing that neither self-report nor instrumentation can displace, because some constructs are defined by behavior and are distorted the moment they are queried. A great deal of applied work — in operant conditioning and behavior analysis, in developmental and clinical assessment, in the behavioral sciences broadly — rests on it. At the same time the boundary between human and machine observation is dissolving: as automated systems take over the coding of movement and expression, the field's oldest question, whether two observers agree, reappears as the question of whether an algorithm agrees with an expert. The safeguards change form but not purpose.

Current Directions

The active frontier is the migration of coding from human judgment to machine measurement, and the methodological questions that migration raises. Open, validated toolkits now recover facial landmarks, head pose, gaze direction, and action-unit intensity from unconstrained video, bringing automated behavioral coding within reach of laboratories without computer-vision expertise (Baltrusaitis et al., 2018). The consolidating survey literature makes the trade-offs explicit: automated facial-action analysis has advanced rapidly on benchmark datasets, yet performance degrades under naturalistic lighting and pose, and models trained on non-representative samples carry that bias forward into every session they code (Martinez et al., 2019). The result is a renewed emphasis on measurement thinking — treating an automated coder as an instrument whose agreement with human experts must be quantified, whose error must be characterised, and whose outputs must be reported with the same candor once reserved for kappa between two graduate students (Girard & Cohn, 2016). The open problem is not whether machines can code behavior but how to certify that their codes mean what the construct requires.

Common Misconceptions

High percent agreement between observers means the data are reliable.
Percent agreement ignores the agreement two coders reach by chance, which is large when one category is common. Cohen's kappa corrects for it, and a scheme with 90 percent raw agreement can have a mediocre kappa when base rates are skewed (Cohen, 1960).
Establishing reliability once at the start certifies the whole dataset.
Agreement drifts as coders unconsciously redefine categories, and it drops when observers believe they are no longer being checked, so reliability assessed only overtly overstates the reliability of the data actually collected (Reid, 1970).
A trained observer records behavior objectively, free of expectation.
Observers led to expect a change record it even when none occurs; expectation biases observational coding measurably, which is why coders are kept blind to condition and hypothesis wherever possible (Kent et al., 1974).

Glossary

Ad libitum sampling.
Recording whatever behavior is conspicuous, with no fixed rule; useful for generating hypotheses but biased toward salient events.
Coding scheme.
A finite set of mutually exclusive, exhaustively defined behavior categories into which observed action is sorted.
Cohen's kappa.
A chance-corrected index of agreement between two observers coding the same categorical behavior.
Continuous recording.
Logging every onset and offset of the target behavior, yielding true frequencies and durations.
Ethogram.
A catalogue of a species' behavioral repertoire, each act described by its form precisely enough to be recognised by another observer.
Event.
A discrete, near-instantaneous act counted as a frequency, as distinct from a durational state.
Facial Action Coding System.
An anatomically based scheme that decomposes any facial movement into component action units.
Focal-animal sampling.
Watching one individual for a set period and recording all of its behavior, giving unbiased individual rates.
Interobserver reliability.
The degree to which two independent observers produce the same record of the same behavior.
Momentary time sampling.
Recording the behavioral state at the instant each interval begins; approximately unbiased for proportion of time.
Observer drift.
The gradual shared redefinition of a category by coders who calibrate against one another rather than the codebook.
One-zero sampling.
Marking an interval if the behavior occurred at any point within it; systematically overestimates proportion of time.
Operational definition.
A statement of the observable features that qualify an instance of a behavior, including its boundary cases.
Reactivity.
Change in the behavior of interest produced by the awareness of being observed.
Scan sampling.
Sweeping a whole group at fixed instants and recording each member's current state.
Sequential analysis.
Statistical testing of temporal contingencies in a coded stream, asking whether one behavior reliably follows another.
State.
A behavior of appreciable duration, defined by an onset and an offset that bound a bout.

Key Researchers

Jeanne Altmann. Eugene Higgins Professor of Ecology and Evolutionary Biology, Emerita, at Princeton University; her 1974 analysis of sampling methods gave observational research its methodological grammar, distinguishing focal, scan, ad libitum, and one-zero designs and showing which quantity each can estimate without bias. Faculty Page - Wikipedia

Roger Bakeman. Professor Emeritus of Psychology at Georgia State University; with John Gottman he formalised the sequential analysis of observational data, turning coded interaction streams into testable temporal structure through contingency tables and lag analysis. Faculty Page - ORCID

Jacob Cohen (1923-1998). Psychologist and statistician at New York University; his 1960 coefficient of agreement, Cohen's kappa, gave observational reliability a chance-corrected index that remains the standard measure of inter-observer agreement for categorical codes. Wikipedia - Wikidata

Jeffrey F. Cohn. Professor Emeritus of Psychology at the University of Pittsburgh and a member of the Carnegie Mellon Robotics Institute; he co-developed the automated coding of facial actions from video, building the datasets and computer-vision methods that let facial behavior be measured objectively at scale. Faculty Page - ORCID

Paul Ekman (1934-2025). Professor Emeritus at the University of California, San Francisco; with Wallace Friesen he built the Facial Action Coding System, an anatomically grounded scheme that made facial behavior an observable, codeable variable. Wikipedia - Wikidata

Nikolaas Tinbergen (1907-1988). Dutch-British ethologist and 1973 Nobel laureate; his framing of ethology's questions and its observational program established the disciplined, minimally intrusive field observation from which systematic behavior-observation methods descend. Wikipedia - Wikidata

Frequently Asked Questions

What are behavior observation techniques?
They are systematic procedures for watching, coding, and recording behavior as it occurs, in natural or laboratory settings, converting a continuous stream of action into quantifiable data through explicit rules about what and when to record (Altmann, 1974).

Why is percent agreement not enough to show observers are reliable?
Because two coders agree by chance at a rate set by how common each category is, so raw percent agreement can look high while true agreement is weak; Cohen's kappa corrects for that chance agreement (Cohen, 1960).

What is the difference between momentary and one-zero time sampling?
Momentary sampling records the state at the instant each interval begins and is approximately unbiased for proportion of time, whereas one-zero sampling marks any interval in which the behavior occurred and systematically overestimates that proportion (Altmann, 1974).

What is observer drift?
It is the gradual, shared redefinition of a behavior category by coders who calibrate against each other rather than the codebook; agreement can stay high among the drifting coders while diverging from the original definitions (Reid, 1970).

Can an observer's expectations bias the data?
Yes; observers led to expect a change tend to record it even when the behavior does not change, so coders are kept blind to condition and hypothesis wherever possible (Kent et al., 1974).

What is an ethogram?
It is a catalogue of a species' behavioral repertoire in which each act is described by its form precisely enough that another observer would identify it the same way, the ethological ancestor of the modern coding scheme (Tinbergen, 1963).

How is agreement measured when more than two observers code the data?
Fleiss generalised Cohen's kappa to any number of raters, allowing a whole coding team rather than a single pair to be evaluated for chance-corrected agreement (Fleiss, 1971).

Can software replace human observers?
Automated systems now code facial landmarks and action units from video quickly and consistently, but their accuracy depends on recording conditions and training data, so their agreement with expert human coders must still be quantified (Martinez et al., 2019).

References

Altmann, J. (1974). Observational study of behavior: Sampling methods. Behaviour, 49(3-4), 227-267. https://doi.org/10.1163/156853974X00534

Bakeman, R., & Gottman, J. M. (1997). Observing interaction: An introduction to sequential analysis (2nd ed.). Cambridge University Press. https://doi.org/10.1017/CBO9780511527685

Baltrusaitis, T., Zadeh, A., Lim, Y. C., & Morency, L.-P. (2018). OpenFace 2.0: Facial behavior analysis toolkit. In 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018) (pp. 59-66). IEEE. https://doi.org/10.1109/FG.2018.00019

Chorney, J. M., McMurtry, C. M., Chambers, C. T., & Bakeman, R. (2015). Developing and modifying behavioral coding schemes in pediatric psychology: A practical guide. Journal of Pediatric Psychology, 40(1), 154-164. https://doi.org/10.1093/jpepsy/jsu099

Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37-46. https://doi.org/10.1177/001316446002000104

Ekman, P. (1993). Facial expression and emotion. American Psychologist, 48(4), 384-392. https://doi.org/10.1037/0003-066X.48.4.384

Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5), 378-382. https://doi.org/10.1037/h0031619

Girard, J. M., & Cohn, J. F. (2016). A primer on observational measurement. Assessment, 23(4), 404-413. https://doi.org/10.1177/1073191116635807

Hartmann, D. P. (1977). Considerations in the choice of interobserver reliability estimates. Journal of Applied Behavior Analysis, 10(1), 103-116. https://doi.org/10.1901/jaba.1977.10-103

Kazdin, A. E. (1977). Artifact, bias, and complexity of assessment: The ABCs of reliability. Journal of Applied Behavior Analysis, 10(1), 141-150. https://doi.org/10.1901/jaba.1977.10-141

Kent, R. N., O'Leary, K. D., Diament, C., & Dietz, A. (1974). Expectation biases in observational evaluation of therapeutic change. Journal of Consulting and Clinical Psychology, 42(6), 774-780. https://doi.org/10.1037/h0037516

Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159-174. https://doi.org/10.2307/2529310

Martinez, B., Valstar, M. F., Jiang, B., & Pantic, M. (2019). Automatic analysis of facial actions: A survey. IEEE Transactions on Affective Computing, 10(3), 325-347. https://doi.org/10.1109/TAFFC.2017.2731763

Reid, J. B. (1970). Reliability assessment of observation data: A possible methodological problem. Child Development, 41(4), 1143-1150. https://doi.org/10.2307/1127341

Tinbergen, N. (1963). On aims and methods of ethology. Zeitschrift fur Tierpsychologie, 20(4), 410-433. https://doi.org/10.1111/j.1439-0310.1963.tb01161.x