Abstract
The semantic differential, which the Medical Subject Headings classify under psycholinguistics, is a rating technique that measures the connotative meaning of a concept by having respondents place it on a series of seven-point scales anchored by opposing adjectives. Charles Osgood and his colleagues introduced it in 1957, and factor analysis of many such scales repeatedly recovered three dimensions of affective meaning: evaluation, potency, and activity. These three axes locate any concept as a point in a shared semantic space, so the distance between two concepts becomes a measurable index of how differently they feel. The method travelled from cross-cultural studies of meaning to brand research, emotion measurement, and the affect control theory that models social interaction. Debate continues over the technique's dimensionality, its concept-scale interaction, and how far the three factors are genuinely universal.
Keywords: semantic differential, connotative meaning, evaluation-potency-activity, affective meaning
What the Semantic Differential Is
The semantic differential is a method for measuring the connotative meaning of a concept — the affective colouring a word carries, what it feels like — as distinct from its denotative meaning, the dictionary reference or factual belief a person holds about it. Charles Osgood built the technique to make that felt quality quantifiable: instead of asking what a respondent knows about a concept, it asks how the concept sits on a set of feeling-toned dimensions (#ref-osgood-1952). The word father, tornado, or democracy each carries a connotation that two people can share even when their factual knowledge differs, and it is that shared affective signal the instrument is designed to capture.
The mechanism is a bipolar adjective scale: the concept is printed at the top, and beneath it runs a stack of rows, each stretched between two antonymous adjectives — good–bad, weak–strong, slow–fast — with seven positions in between. The respondent marks one position on every row, closer to one pole, closer to the other, or neutral at the centre, and the marks are scored from −3 through 0 to +3. A completed form is thus a profile of numbers, one per scale, and averaging the scales that belong together yields a small set of factor scores that summarise where the concept stands (#ref-osgood-1957).
Figure 1
A Semantic Differential Rating Form With a Marked Profile
Historical Origins
The technique grew out of Osgood's programme on the psychology of meaning at the University of Illinois. In a 1952 paper he laid out the problem — meaning had long been treated as private and unmeasurable — and proposed that the connotative aspect could be indexed by the way people apply descriptive adjectives to concepts (#ref-osgood-1952). The full method arrived with the 1957 monograph The Measurement of Meaning, written with George Suci and Percy Tannenbaum, which set out the scaling procedure, the factor-analytic evidence, and the applications that would define the field for a generation (#ref-osgood-1957).
The book's decisive contribution was empirical rather than merely procedural. Osgood and his co-authors gathered ratings of many concepts on many bipolar scales and submitted the resulting correlations to factor analysis, a statistical method that finds the small number of underlying dimensions accounting for how the scales covary. Across samples and concepts the same three factors kept re-emerging, a robustness that turned a rating gadget into a claim about the structure of affective meaning itself, and the sourcebook of studies that followed consolidated the approach into a standard research tool (#ref-snider-osgood-1969).
The E-P-A Structure
The three recurring factors are evaluation, potency, and activity — the E-P-A structure that is the technique's signature finding. Evaluation, the largest and most stable factor, is the good–bad dimension: it captures whether a concept is liked or disliked, and scales such as good–bad, pleasant–unpleasant, and kind–cruel load on it heavily. Potency is the strong–weak dimension of power and size, marked by scales like strong–weak, large–small, and heavy–light. Activity is the fast–slow dimension of movement and energy, carried by active–passive, fast–slow, and sharp–dull (#ref-osgood-1957). Osgood later set out explicitly why these particular three dimensions recur, grounding them in his mediational account of meaning: the factors index the dominant components of the affective reaction a concept evokes, which is why evaluation — the good–bad response — is consistently the largest and most stable of the three (#ref-osgood-1969).
Because the three factors are close to statistically independent, they behave like the axes of a space: any concept, once rated, is a point whose coordinates are its evaluation, potency, and activity scores. This is the semantic space, and it is what lets the method treat meaning geometrically — two concepts that draw similar profiles sit near each other, and concepts that feel opposite sit far apart. The demonstration below builds a profile from a handful of bipolar scales and reports the evaluation, potency, and activity scores it implies, making the move from individual marks to factor coordinates concrete.
Rate a concept, read its meaning profile
The semantic differential measures a concept’s connotative meaning by having raters mark it on bipolar adjective scales. Move each slider to rate an object — say my new phone — and watch the six ratings collapse onto Osgood’s three factors: Evaluation, Potency, and Activity.
Six scales became three numbers. That collapse is the technique’s whole point: however many adjective pairs a study uses, their shared variance falls almost entirely onto Evaluation, Potency, and Activity.
Constructing and Scoring the Scale
Building a semantic differential means choosing bipolar scales whose factor loadings are known — the loading being the correlation between a scale and a factor, a number that says how purely that scale measures evaluation, potency, or activity. A well-designed instrument samples all three factors with scales that load cleanly on one and weakly on the others, so that averaging the evaluation scales gives an evaluation score uncontaminated by potency or activity. Scales are usually scored from −3 to +3, reverse-scoring any row printed with its positive pole on the left so that higher always means more of the factor (#ref-osgood-1957).
The complication Osgood himself flagged is concept-scale interaction: a scale does not always load the same way for every concept. The hot–cold scale is evaluative for soup but almost purely denotative for stove, and the adjective hard means one thing applied to rock and another applied to exam. This context sensitivity means factor loadings established on one set of concepts can shift when the instrument is carried to a new domain, and it is the central methodological caution in the technique's use (#ref-heise-1969). The demonstration below lets a reader select a scale and see its evaluation, potency, and activity loadings side by side, illustrating why scale selection is the craft at the heart of the method.
Why a dozen scales become three numbers
Factor analysis of many bipolar scales repeatedly finds the same three underlying dimensions. Each scale mostly “loads” on one of them. Step through the scales and watch which factor each one measures — and notice the mixed cases that make scale choice matter.
This scale measures mainly Evaluation. Clean scales load on a single factor; a mixed scale like strong – weak, which carries both Potency and Activity, is why the meaning of a scale can shift with the concept being rated.
Cross-Cultural Universality
Osgood's boldest claim was that the E-P-A structure is not an artefact of English but a property of human affective meaning generally. To test it he mounted a vast cross-cultural programme, administering translated instruments across dozens of language-culture communities and asking whether the same three factors would emerge everywhere (#ref-osgood-1964). The comparative results were summarised in the 1975 book Cross-Cultural Universals of Affective Meaning, which reported that evaluation, potency, and activity recur across the sampled cultures with remarkable consistency, whatever the surface differences in which concepts are liked or feared (#ref-osgood-1975).
The finding is a qualified universal. What appears to be shared is the framework — the three dimensions along which affective meaning is organised — not the placement of particular concepts, which varies with culture as one would expect. A snake, a mother, or a colour can occupy very different points in different societies while the axes of the space stay the same, and it is that separation of a common structure from culturally variable content that gives the claim its force and its limits.
Applications and Descendants
The semantic differential seeded a lineage of affective-meaning models that reaches into contemporary emotion science. Mehrabian and Russell recast Osgood's three factors as the PAD model — pleasure, arousal, and dominance — and used it to describe emotional responses to environments, mapping evaluation onto pleasure, activity onto arousal, and potency onto dominance (#ref-mehrabian-1974). Russell then compressed the affective plane into the circumplex model of affect, arranging emotion words in a circle spanned by valence and arousal, a structure that remains a workhorse of emotion research (#ref-russell-1980). Bradley and Lang built the Self-Assessment Manikin, a pictorial semantic differential that measures pleasure, arousal, and dominance with cartoon figures rather than words, extending the method to respondents and settings where verbal scales fail (#ref-bradley-1994).
The most systematic descendant is affect control theory, which treats E-P-A ratings of identities, behaviours, and settings as the raw material of a generative model of social interaction, worked out at book length in Heise's Surveying Cultures (#ref-heise-2010). Because every concept is a point in the space, the generalized distance between two points — the straight-line distance computed from their evaluation, potency, and activity differences — becomes a usable quantity: it indexes how affectively far apart two identities or events are, and drives the theory's predictions about what people do to restore disturbed meanings. The demonstration places several concepts in the space and computes that distance for any selected pair.
Meaning as a point in semantic space
Each concept’s three factor scores place it at a point in a three-dimensional space. The plane below shows Evaluation against Activity, with Potency drawn as marker size. Pick two concepts to see the generalised distance D — the straight-line gap between their meanings — computed across all three dimensions.
| E | P | A | |
|---|---|---|---|
| HERO | 2.6 | 2.1 | 1.4 |
| TYRANT | -2.7 | 2.3 | 0.8 |
Distance D = 5.34
D is the square root of the summed squared differences on E, P, and A. Concepts that feel alike sit close together; opposites like a hero and a tyrant — near-identical in potency and activity but reversed on evaluation — sit far apart.
Measurement Properties and Cautions
As a measurement instrument the semantic differential is prized for being quick, flexible, and comparatively robust: a handful of bipolar scales yields interval-level data on any concept a researcher can name, and the E-P-A framework gives those numbers a ready interpretation. Heise's early review catalogued the methodological choices that determine whether the numbers mean anything — how many scale steps to use, how to handle the neutral midpoint, how concept-scale interaction threatens the assumption that a scale measures the same factor throughout a study (#ref-heise-1969). These are not fatal flaws but design decisions, and the technique's reliability depends on making them deliberately.
Modern methodological work has pressed the point that the instrument is often used carelessly. Verhagen and colleagues, reviewing its use in information-systems research, found scales assembled without checking that their factor structure held in the new domain, and offered an integrative framework for constructing and validating a semantic differential properly rather than borrowing adjectives on faith (#ref-verhagen-2015). The reference literature now treats semantic differential scaling as a distinct measurement approach with its own validity requirements, catalogued alongside other scaling methods in research-methods encyclopedias (#ref-rosenberg-2018).
| Factor | What it captures | Representative scales |
|---|---|---|
| Evaluation (E) | Goodness: liking, favourability, worth | good–bad, pleasant–unpleasant, kind–cruel |
| Potency (P) | Power: strength, size, weight | strong–weak, large–small, heavy–light |
| Activity (A) | Energy: movement, speed, arousal | active–passive, fast–slow, sharp–dull |
Worked Example
Because every concept is a point in E-P-A space, the affective difference between two concepts can be computed as an ordinary straight-line distance. Take two identities rated on the three factors: HERO at evaluation 2.6, potency 2.1, activity 1.4, and TYRANT at evaluation −2.7, potency 2.3, activity 0.8. The question is how far apart they feel.
Subtract coordinate by coordinate. The evaluation difference is 2.6 − (−2.7) = 5.3, the potency difference is 2.1 − 2.3 = −0.2, and the activity difference is 1.4 − 0.8 = 0.6. Almost all the separation is on evaluation, which is exactly right: a hero and a tyrant can be equally potent and similarly active, and it is their goodness that opposes them.
Now square each difference, add, and take the root — the generalized distance formula. The squares are 5.3² = 28.09, (−0.2)² = 0.04, and 0.6² = 0.36, summing to 28.49. The square root of 28.49 is 5.34. That single number, D = 5.34, is the affective distance between the two identities, dominated by the evaluative gap and barely moved by their near-equal potency and activity — the same quantity affect control theory uses to gauge how badly one identity disturbs the meaning of another. The interactive EPA-space demonstration above computes this distance for any pair of its concepts, and reproduces D = 5.34 for HERO and TYRANT.
Discussion
The semantic differential's staying power rests on a rare combination: a simple procedure that produces numbers, backed by an empirical regularity — the recurrence of evaluation, potency, and activity — that gives those numbers a stable meaning. Where many attitude-measurement schemes are ad hoc, Osgood's method carries a theory of what it measures, and the E-P-A structure has proved robust enough to survive translation across languages and reincarnation as the PAD model, the circumplex, and affect control theory. That lineage is the strongest evidence that the technique captured something real about how affective meaning is organised.
The cautions are equally durable. Evaluation dominates so heavily that potency and activity are sometimes hard to recover cleanly, concept-scale interaction means a validated instrument can mislead when carried to a new domain, and the whole apparatus measures connotation, not denotation — it reports how communism or chemotherapy feels to a sample, not what the respondents believe to be true of it. Read with those limits in mind, the semantic differential remains one of psychology's most transportable measurement ideas, precisely because it measures the affective residue that ordinary questionnaires miss.
Current Directions
The liveliest current work turns Osgood's static space into a dynamic, computational model of social life. Affect control theory has been reformulated in explicitly probabilistic terms: Schröder, Hoey, and Rogers built a Bayesian affect control theory that represents identities and their uncertainty as probability distributions in E-P-A space and updates them as an interaction unfolds (#ref-schroder-2016). Hoey and colleagues embedded the same machinery in a partially observable Markov decision process, letting an artificial agent reason about affective meaning while it acts, which carries Osgood's three factors into affective computing and human-robot interaction (#ref-hoey-2016).
A second front is scale and automation. Where Osgood rated concepts by hand a few at a time, natural-language-processing researchers now derive evaluation-like, potency-like, and activity-like values for tens of thousands of words at once: Mohammad's crowd-sourced lexicon supplies reliable valence, arousal, and dominance ratings for 20,000 English words, a direct computational heir to the semantic differential's three dimensions (#ref-mohammad-2018). The open questions remain the old ones sharpened by scale — how many dimensions truly underlie affective meaning, how far the three factors generalise across languages and modalities, and how to keep concept-scale interaction from corrupting automatically harvested ratings.
Common Misconceptions
- The semantic differential is just a Likert scale with adjectives.
- It is a different instrument with a different target. A Likert scale measures agreement with a statement; the semantic differential rates a concept between two antonymous adjectives to recover its connotative meaning along the evaluation, potency, and activity factors, which agreement scales do not isolate (#ref-rosenberg-2018).
- A scale measures the same factor no matter what concept it rates.
- Concept-scale interaction says otherwise. The hot–cold scale is evaluative for soup but denotative for stove, so factor loadings established on one set of concepts can shift on another, which is why scales must be revalidated in each new domain (#ref-heise-1969).
- The three factors mean every culture likes and fears the same things.
- The universal is the framework, not the content. Evaluation, potency, and activity recur across cultures, but where a given concept falls within that space varies from society to society exactly as expected (#ref-osgood-1975).
Glossary
- Activity (A).
- The third E-P-A factor, the fast–slow dimension of movement and energy, carried by scales such as active–passive and fast–slow.
- Affect control theory.
- A generative theory of social interaction that treats E-P-A ratings of identities, behaviours, and settings as data and predicts action from the drive to maintain affective meanings.
- Bipolar adjective scale.
- A rating row stretched between two antonymous adjectives with (usually seven) positions between them; the basic unit of a semantic differential form.
- Circumplex model of affect.
- Russell's arrangement of emotion words in a circle spanned by valence and arousal, a two-dimensional descendant of Osgood's affective space.
- Concept-scale interaction.
- The fact that a bipolar scale can load on different factors for different concepts, so a scale's meaning is not fixed independently of what it rates.
- Connotative meaning.
- The affective, feeling-toned sense a concept carries, as opposed to its dictionary reference; the property the semantic differential is built to measure.
- Denotative meaning.
- The referential, factual sense of a concept — what it points to and what is literally true of it — distinct from its connotation.
- Evaluation (E).
- The first and largest E-P-A factor, the good–bad dimension of liking and worth; the most stable factor the technique recovers.
- Factor analysis.
- A statistical method that finds the small number of underlying dimensions accounting for how many measured variables covary; the tool that revealed the E-P-A structure.
- Factor loading.
- The correlation between a scale and a factor, expressing how purely that scale measures evaluation, potency, or activity.
- Generalized distance (D).
- The straight-line distance between two concepts in E-P-A space, computed as the square root of the summed squared differences on the three factors.
- PAD model.
- Mehrabian and Russell's pleasure–arousal–dominance framework, a relabelling of evaluation, activity, and potency for describing emotional responses to environments.
- Potency (P).
- The second E-P-A factor, the strong–weak dimension of power, size, and weight, marked by scales such as strong–weak and large–small.
- Self-Assessment Manikin.
- Bradley and Lang's pictorial semantic differential that measures pleasure, arousal, and dominance with cartoon figures rather than verbal scales.
- Semantic space.
- The geometric space whose axes are the E-P-A factors, in which each rated concept is a point and affective similarity is nearness.
Key Researchers
David R. Heise (1937–2021). Rudy Professor of Sociology Emeritus at Indiana University Bloomington; advanced semantic-differential measurement and founded affect control theory, formalising E-P-A ratings into a generative model of social interaction. Wikipedia - Wikidata
Charles E. Osgood (1916–1991). Psychologist at the University of Illinois Urbana-Champaign; invented the semantic differential and led the cross-cultural programme, lead author of The Measurement of Meaning (1957). Wikipedia - Wikidata
Dawn T. Robinson (living). Professor of Sociology at the University of Georgia; a contemporary affect-control-theory and emotion researcher extending E-P-A measurement of affective meaning. ORCID - Google Scholar
Lynn Smith-Lovin (living). Robert L. Wilson Distinguished Professor Emeritus of Sociology at Duke University; a leading contemporary figure in affect control theory and the E-P-A measurement of affective meaning. ORCID - Google Scholar
George J. Suci (1925–1998). Professor emeritus of human development at Cornell University; co-author of The Measurement of Meaning (1957) and of the factor-analytic work that established the E-P-A structure. Cornell Memorial
Percy H. Tannenbaum (1927–2009). Communications researcher at the University of California, Berkeley; co-author of The Measurement of Meaning (1957) who later applied the method in mass-communication research. Wissenschaftskolleg zu Berlin - CASBS
Frequently Asked Questions
What is the semantic differential in simple terms?
It is a rating method that measures how a concept feels rather than what it means literally. Respondents mark a concept on a set of seven-point scales anchored by opposite adjectives, such as good to bad or weak to strong, and the marks are scored to show where the concept sits on a few underlying dimensions of feeling (Osgood et al., 1957).
Who invented it?
Charles Osgood devised the technique at the University of Illinois, setting out the idea in a 1952 paper and the full method in the 1957 book The Measurement of Meaning, written with George Suci and Percy Tannenbaum (Osgood, 1952).
What are evaluation, potency, and activity?
They are the three dimensions that factor analysis repeatedly finds beneath semantic differential ratings. Evaluation is the good to bad dimension of liking, potency is the strong to weak dimension of power, and activity is the fast to slow dimension of energy (Osgood et al., 1957).
How is the semantic differential different from a Likert scale?
A Likert scale measures how strongly someone agrees with a statement. The semantic differential rates a concept between two opposing adjectives to recover its connotative meaning along the evaluation, potency, and activity factors, which agreement scales are not designed to isolate (Rosenberg & Navarro, 2018).
Are the three dimensions the same across cultures?
The framework appears to be. Osgood's cross-cultural programme found evaluation, potency, and activity recurring across many language communities, although where any particular concept falls within that space still varies from culture to culture (Osgood et al., 1975).
What is concept-scale interaction?
It is the finding that a scale can measure different factors for different concepts. The hot to cold scale is evaluative for soup but merely descriptive for stove, so a scale validated on one set of concepts may not behave the same way on another (Heise, 1969).
What modern methods came out of the semantic differential?
Its three dimensions were relabelled as the pleasure, arousal, and dominance model, compressed into the circumplex model of affect, and turned into the Self-Assessment Manikin for measuring emotion with pictures (Bradley & Lang, 1994).
Is the semantic differential still used in research today?
Yes. It underpins affect control theory and its Bayesian and computational extensions, and its dimensions survive in large word-rating lexicons that supply valence, arousal, and dominance values for tens of thousands of words (Mohammad, 2018).
References
Bradley, M. M., & Lang, P. J. (1994). Measuring emotion: The Self-Assessment Manikin and the semantic differential. Journal of Behavior Therapy and Experimental Psychiatry, 25(1), 49-59. https://doi.org/10.1016/0005-7916(94)90063-9
Heise, D. R. (1969). Some methodological issues in semantic differential research. Psychological Bulletin, 72(6), 406-422. https://doi.org/10.1037/h0028448
Heise, D. R. (2010). Surveying cultures: Discovering shared conceptions and sentiments. Wiley. https://doi.org/10.1002/9780470575789
Hoey, J., Schröder, T., & Alhothali, A. (2016). Affect control processes: Intelligent affective interaction using a partially observable Markov decision process. Artificial Intelligence, 230, 134-172. https://doi.org/10.1016/j.artint.2015.09.004
Mehrabian, A., & Russell, J. A. (1974). An approach to environmental psychology. MIT Press.
Mohammad, S. M. (2018). Obtaining reliable human ratings of valence, arousal, and dominance for 20,000 English words. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 174-184). Association for Computational Linguistics. https://doi.org/10.18653/v1/P18-1017
Osgood, C. E. (1952). The nature and measurement of meaning. Psychological Bulletin, 49(3), 197-237. https://doi.org/10.1037/h0055737
Osgood, C. E., Suci, G. J., & Tannenbaum, P. H. (1957). The measurement of meaning. University of Illinois Press.
Osgood, C. E. (1964). Semantic differential technique in the comparative study of cultures. American Anthropologist, 66(3), 171-200. https://doi.org/10.1525/aa.1964.66.3.02a00880
Osgood, C. E. (1969). On the whys and wherefores of E, P, and A. Journal of Personality and Social Psychology, 12(3), 194-199. https://doi.org/10.1037/h0027715
Osgood, C. E., May, W. H., & Miron, M. S. (1975). Cross-cultural universals of affective meaning. University of Illinois Press.
Rosenberg, B. D., & Navarro, M. A. (2018). Semantic differential scaling. In B. B. Frey (Ed.), The SAGE encyclopedia of educational research, measurement, and evaluation. SAGE Publications. https://doi.org/10.4135/9781506326139.n624
Russell, J. A. (1980). A circumplex model of affect. Journal of Personality and Social Psychology, 39(6), 1161-1178. https://doi.org/10.1037/h0077714
Schröder, T., Hoey, J., & Rogers, K. B. (2016). Modeling dynamic identities and uncertainty in social interactions: Bayesian affect control theory. American Sociological Review, 81(4), 828-855. https://doi.org/10.1177/0003122416650963
Snider, J. G., & Osgood, C. E. (Eds.). (1969). Semantic differential technique: A sourcebook. Aldine Publishing Company.
Verhagen, T., van den Hooff, B., & Meents, S. (2015). Toward a better use of the semantic differential in IS research: An integrative framework of suggested action. Journal of the Association for Information Systems, 16(2), 108-143. https://doi.org/10.17705/1jais.00388