Abstract
Pattern recognition is the process by which the visual system assigns a variable sensory input to a known category, solving the problem that one object projects endlessly different images onto the eye. This article traces the major theories: rigid template matching and its failure on natural variability; the abstraction of a central prototype from variable instances; feature analysis, from Selfridge's Pandemonium and Hubel and Wiesel's cortical detectors to Treisman's feature-integration theory and the binding problem; and structural description, from Marr's three-dimensional models to Biederman's recognition-by-components. It sets out how word context reshapes letter recognition in the interactive-activation model, the dispute over whether stored representations are viewpoint-invariant or view-based, and the hierarchy of the ventral visual stream that untangles object identity across images. Three demonstrations model visual search, recognition-by-components, and the word-superiority effect.
Keywords: pattern recognition, feature integration, recognition-by-components, prototype, ventral stream
A capital letter A can be printed, scrawled, italicised, lit from any angle, and thrown onto the retina at any size or position, yet a reader recognises it in an instant. The same is true of a face, a chair, or a dog: each names a category whose members, and whose retinal projections, vary without limit. Pattern recognition is the process that bridges that gap, matching a variable and never-exactly-repeated sensory input to a stored category so that perception can say what a thing is (Palmeri & Gauthier, 2004). It is the culminating step of visual perception: after the scene has been segmented into surfaces and objects, recognition assigns each to a class, connecting sensation to memory, language, and action. The central obstacle is invariance — how the system treats endlessly different images as instances of one thing — and the history of the field is a sequence of increasingly capable answers to it.
- Pattern recognition is the assignment of a variable sensory input to a category; its central difficulty is invariance — recognising the same object across changes in size, position, lighting, and viewpoint.
- Template theories match input against stored copies and fail on natural variability; prototype theories store an abstracted central tendency, which people recognise even when they have never seen that exact instance.
- Feature theories decompose a pattern into elementary parts detected in parallel; feature-integration theory holds that focal attention is required to bind separate features into an object, explaining conjunction search and illusory conjunctions.
- Structural-description theories represent objects as arrangements of viewpoint-invariant volumetric parts — Marr's three-dimensional models and Biederman's geons — while view-based accounts hold that recognition is tied to stored views.
- Recognition is not purely bottom-up: word context raises the recognisability of a letter, modelled by interactive activation, and in the brain a hierarchy in the ventral stream progressively untangles object identity.
What Pattern Recognition Is
Pattern recognition is the assignment of a perceptual input to a category on the basis of stored knowledge. The problem it solves is not detection — deciding that something is present — but classification: deciding what the something is, given that the very same object never casts the same image twice. Size, position, orientation, illumination, and viewpoint all vary, as does the object itself when the category is natural rather than a fixed symbol, so the recognition system must map a vast and continuous space of images onto a discrete and stable set of categories. This mapping is what any theory of pattern recognition must deliver, and the differences among theories are differences in what is stored and how the input is compared to it (Palmeri & Gauthier, 2004).
Recognition does not operate on the raw image but on an already-organised one. Before a pattern can be matched, the visual field must be parsed into figures and grounds, and the parts that belong to one object grouped together, a stage the Gestalt psychologists described in their laws of perceptual organisation: elements that are near, similar, or move together are seen as a unit (Wertheimer, 1923). This grouping is the input to recognition, and its failures — a camouflaged animal, an ambiguous figure — are failures of recognition as much as of organisation. The theories that follow divide into four families by what they take the stored category to be: a whole template, an abstracted prototype, a set of features, or a structural description of parts and their relations. Table 1 sets the four families side by side.
| Approach | What is stored | How a new instance is recognised | Illustrative source |
|---|---|---|---|
| Template matching | A whole-pattern copy of each known form | Normalise the input, then find the best-overlapping template | Constrained inputs (barcodes, standardised numerals) |
| Prototype | The abstracted central tendency of the category | By its similarity to the stored prototype | Posner & Keele (1968); Rosch (1975) |
| Feature analysis | A set of elementary features and their detectors | Detect features in parallel, then bind them with attention | Selfridge (1959); Treisman & Gelade (1980) |
| Structural description | Volumetric parts (geons) and their spatial relations | Match the recovered parts and relations to a stored description | Marr & Nishihara (1978); Biederman (1987) |
Template and Prototype Theories
The simplest theory is template matching: the system stores a copy, a template, of each known pattern and recognises an input by finding the template it overlaps best. Template matching works for constrained inputs — the stylised numerals on a cheque, a barcode — but it fails on natural vision, because a single template cannot absorb the variation in size, position, orientation, and form that a real category shows. Adding a normalisation stage that rescales and reorients the input before matching helps, but the number of templates needed to cover every distortion of every object explodes, and the theory has no principled account of how a never-seen instance is recognised at all.
Prototype theory answers that failure by storing not every instance but an abstracted central tendency. Posner and Keele trained people to sort dot patterns that were random distortions of an unseen prototype into categories; afterwards, people classified the prototype itself — which they had never actually seen — as accurately as the training items, and more accurately than new distortions, showing that what had been learned was the abstracted central form rather than the specific instances (Posner & Keele, 1968). Rosch extended the idea from artificial dot patterns to natural categories, showing that everyday categories such as bird or furniture are organised around prototypical members: a robin is judged a better example of bird than a penguin, is recognised faster, and comes to mind first, so category structure is graded rather than all-or-none (Rosch, 1975). Prototype abstraction gives recognition its tolerance for novelty, being able to classify a member it has never seen before.
Storing an abstracted prototype is not the only alternative to a rigid template. Exemplar theories hold that a category is represented by its remembered individual instances, and that a new item is classified by its summed similarity to those stored exemplars rather than by comparison to a single averaged form (Medin & Schaffer, 1978). The prototype-endorsement result does not by itself decide between the two, because a stored collection of exemplars also classifies the never-seen central tendency well, and the modern categorisation literature keeps both, with exemplar models often fitting the finer grain of human classification more closely. What the accounts share, and what breaks decisively with template matching, is that the stored category is abstracted from experience rather than held as a fixed copy. Neither, though, says much about how a pattern is decomposed in the first place, which is the problem feature theories take up.
Feature Analysis and Feature Integration
Feature theories hold that recognition begins by decomposing a pattern into elementary parts — line segments, curves, angles, colours — that are detected in parallel across the field, with the identity of the whole computed from which features are present. The founding model is Selfridge's Pandemonium, a hierarchy in which “demons” at a feature level each shout in proportion to how strongly their feature is present, cognitive demons above them listen for the feature combinations that define each letter, and a decision demon picks the loudest; the metaphor captured parallel, weighted, bottom-up feature extraction feeding a recognition decision (Selfridge, 1959). The proposal gained biological force when Hubel and Wiesel found neurons in the cat's primary visual cortex that fire selectively to edges and bars at particular orientations and positions — literal feature detectors at the first cortical stage of vision — establishing that the brain does begin by analysing the image into oriented features (Hubel & Wiesel, 1962).
Detecting features is not the same as recognising an object, because an object is a particular conjunction of features, and the features must be bound to the right locations and to each other. Treisman and Gelade's feature-integration theory drew the line precisely here: individual features are registered pre-attentively and in parallel across the whole field, so a target defined by a unique feature “pops out” and is found equally fast whatever the number of distractors, but a target defined by a conjunction of features requires focal attention to be directed to each location in turn to bind its features, so conjunction search is slow and its time rises with the number of items (Treisman & Gelade, 1980). When attention is overloaded or diverted, features from different objects can be miscombined into illusory conjunctions — a red X and a green O reported as a green X — direct evidence that binding is a distinct, attention-demanding step. The first demonstration implements this contrast, letting the reader run a feature search and a conjunction search over increasing set sizes and compare the resulting response times.
Search the Display, Watch the Slope
Feature Search versus Conjunction Search
Choose a search type and whether the target is present, then vary the set size. In feature search the target is the one red item; in conjunction search it is the red square among red circles and blue squares. The modelled response time is 400 ms plus a per-item slope times the set size.
The strict split between a parallel feature stage and a purely serial conjunction search proved too clean. Wolfe's Guided Search revised the account: the parallel stage does not merely flag single features but builds a coarse activation map that ranks locations by how well they match the target on each feature dimension, and attention is then directed to the highest-ranked locations first rather than at random (Wolfe, 1994). Conjunction search is therefore neither strictly parallel nor blindly serial but guided: top-down knowledge of the target's features steers a limited-capacity stage, which is why real conjunction searches are usually faster and their set-size slopes shallower than a pure item-by-item scan predicts. Guided search keeps feature-integration theory's core division between registering features and binding them, while making the binding stage efficient rather than exhaustive.
Structural Descriptions: Marr and Recognition-by-Components
Features and prototypes both treat the stored pattern as essentially two-dimensional, a picture or a list. Structural-description theories instead represent an object as an arrangement of three-dimensional parts and the spatial relations among them, aiming for a description that stays constant as the viewpoint changes. Marr set this work in a wider frame, arguing that any information-processing system must be understood at three separate levels: a computational theory of what is computed and why, an algorithmic level of the representations and processes that carry it out, and an implementational level of the physical mechanism that runs them (Marr, 1982). Recognition, on this view, is one computational problem — turning a variable image into a stable identity — that can be attacked at the middle level with representations and at the lowest with neurons. His account of shape with Nishihara lives at that middle level: the visual system builds an object-centred, three-dimensional model in which shape is described as a hierarchy of generalised cones — volumetric primitives — arranged on axes defined by the object itself rather than by the viewer, so that the same description is recovered from any vantage point (Marr & Nishihara, 1978). Because the representation is anchored to the object's own axes, it is in principle viewpoint-invariant, and it decomposes a complex shape into a manageable set of parts.
Biederman turned this program into a concrete recognition theory with recognition-by-components. He proposed a small alphabet of roughly three dozen volumetric primitives, called geons — a brick, a cylinder, a cone, a wedge, and so on — distinguished by non-accidental properties of the image, features such as whether an edge is straight or curved and whether contours are parallel, that are preserved across most viewpoints and so can be read directly off the two-dimensional image (Biederman, 1987). On this account an object is recognised from a handful of geons and their relations — a cup is a cylinder with a curved handle at its side — which explains why line drawings are recognised almost as fast as photographs, why recognition survives heavy occlusion as long as a few geons and their junctions remain visible, and why degrading the image at the geon-defining vertices is far more damaging than degrading it elsewhere. The second demonstration builds an object from geons and lets the reader remove or occlude components to see how few, in the right relations, still support recognition.
Remove Parts, Test Recognition
Recognition-by-Components
Toggle each geon to occlude it. The lamp is recognised as long as enough diagnostic parts remain in their relations, even with several removed — the robustness to occlusion that recognition-by-components predicts.
Context and the Interactive-Activation Model
The theories so far run bottom-up, from image to category, but recognition is also shaped from the top down by context. The clearest case is the word-superiority effect: Reicher found that a letter is identified more accurately when it appears in a real word than when it appears alone or in a random string, even though the word carries more to process, so that a D is reported more reliably in WORD than in isolation or in ORWD (Reicher, 1969). A higher level of structure was helping to recognise its own parts, which no purely bottom-up feature model predicts. McClelland and Rumelhart modelled the effect with the interactive-activation model, a network of three levels — features, letters, and words — in which detecting features excites the letters they belong to and inhibits the rest, active letters excite the words consistent with them, and, crucially, active words send excitation back down to their constituent letters (McClelland & Rumelhart, 1981). A letter embedded in a word therefore receives top-down support from the word node it helps activate, so it is recognised better than the same letter alone, and the model reproduced the word-superiority data quantitatively.
Context operates at the level of whole configurations as well. Navon showed that when a large letter is made out of small letters, observers identify the large, global letter faster than the small, local ones, and the global form interferes with naming the local ones more than the reverse — the “forest before the trees,” a global precedence in which the overall structure is available before its parts (Navon, 1977). Together the two findings establish that recognition is an interaction of bottom-up feature evidence and top-down structural context rather than a one-way flow. The third demonstration runs a reduced interactive-activation network, presenting a target letter in a word, a nonword, and alone, and showing how top-down feedback raises the activation of the letter in the word.
Feed Context Back to the Letter
The Word-Superiority Effect
The target letter D receives a fixed bottom-up feature evidence of 0.40 in every context. Only inside a real word does an active word node feed activation back down to it. Adjust the strength of that top-down feedback and watch the word context pull ahead of the nonword and the single letter.
Viewpoint and the Ventral Stream
Whether the stored representation is truly viewpoint-invariant, as Marr and Biederman held, or tied to particular views, is the field's longest-running dispute. Tarr and Pinker gave the case for view-dependence: when people learn novel shapes at one orientation and are then tested at others, recognition time rises roughly linearly with the angular distance from the trained view, as though the input were mentally rotated to a stored view before matching — a cost that a genuinely viewpoint-invariant representation should not incur (Tarr & Pinker, 1989). The current consensus is mixed: recognition can be nearly viewpoint-invariant for distinguishing broad categories from a few geons, but becomes viewpoint-dependent for fine discriminations among similar objects, so the two kinds of representation coexist and dominate under different task demands.
Physiology has grounded the debate in the ventral visual stream. Beyond primary visual cortex the cortical visual system divides into two broad pathways: a dorsal stream running towards the parietal lobe that codes where things are and guides visually directed action, and a ventral stream running forward through the occipitotemporal lobe that codes what things are — the where and how of vision against its what (Goodale & Milner, 1992). Recognition is the work of the ventral stream. Building on Hubel and Wiesel's oriented detectors, hierarchical models propose that recognition is computed by alternating stages: template-like tuning that makes cells selective for particular feature combinations, interleaved with pooling operations that make the response tolerant to the feature's exact position and size, so that selectivity and invariance are built up together over successive layers (Riesenhuber & Poggio, 1999). DiCarlo and colleagues framed the stream's job geometrically as untangling: the images of a given object trace out a tangled manifold in the space of retinal activity that overlaps the manifolds of other objects, and each stage of the ventral stream re-represents the input so that, by the highest stages, a simple linear readout can separate one object's manifold from another's (DiCarlo et al., 2012). These hierarchical ideas found their most powerful expression in deep convolutional neural networks: a network trained only to classify objects, through many alternating layers of convolution and pooling, comes to predict the responses of individual neurons along the ventral stream better than any model hand-designed for the purpose, so that a single class of architecture now serves both as a working recognition system and as the leading quantitative model of the biological one (Yamins & DiCarlo, 2016). Human imaging maps this hierarchy onto a mosaic of areas in occipitotemporal cortex, some tuned to general objects and others to specific classes such as faces, places, and body parts (Grill-Spector & Malach, 2004). Whether the class-specific regions reflect innate modules or acquired perceptual expertise — the same debate that surrounds face perception — remains open, and both readings are argued from the same data (Palmeri & Gauthier, 2004). Figure 1 traces the hierarchy from the variable retinal image to a stable category.
Figure 1
The Ventral-Stream Recognition Hierarchy
Worked Example
Feature-integration theory makes a quantitative prediction that the first demonstration reproduces, and it can be worked by hand. Model the time to search a display as a straight line, response time equals an intercept plus a per-item slope multiplied by the number of items: RT = a + b·N. The intercept a collects everything independent of set size — perceiving the display, deciding, and responding — and is taken here as 400 ms. The slope b is the theory's diagnostic quantity. For a feature search, where the target pops out in parallel, the slope is near zero; take it as 5 ms per item. For a conjunction search, where attention must visit items serially, the slope is substantial and positive, and because a self-terminating serial search inspects, on average, about half the items before finding a present target but all of them to be sure a target is absent, the target-absent slope is about twice the target-present slope; take them as 30 and 60 ms per item.
Consider a display of 12 items. In feature search the model gives 400 + 5 × 12 = 400 + 60 = 460 ms, essentially the same as it would give for two items or twenty, the flat function that is the signature of parallel, pre-attentive detection. In conjunction search with the target present it gives 400 + 30 × 12 = 400 + 360 = 760 ms, and with the target absent 400 + 60 × 12 = 400 + 720 = 1120 ms. The two conjunction slopes stand in a ratio of 60 to 30, that is 2 to 1, the pattern expected if attention is deployed serially and the search stops as soon as the target is found. The contrast between a flat feature-search function and a steep, roughly two-to-one conjunction-search function is the empirical core of feature-integration theory: the first needs no attention, the second is limited by it. The strictly serial, self-terminating model used here is an idealisation that makes the arithmetic transparent; guided search predicts that real conjunction slopes are shallower, because feature guidance lets attention skip locations that cannot contain the target (Wolfe, 1994).
Discussion
The theories of pattern recognition are less a sequence of refutations than a layering of partial answers to the invariance problem. Template matching fails on natural variability but survives wherever the input is constrained; prototype abstraction explains tolerance for novelty and the graded structure of categories but not how a pattern is decomposed; feature analysis supplies the decomposition and has a clear neural correlate in oriented detectors, but leaves the binding of features to attention; structural description delivers parts and relations that can be viewpoint-invariant, at the cost of a harder recovery problem from the image (Biederman, 1987; Marr & Nishihara, 1978). Each captures a real aspect of a system that plainly uses several strategies at once, and the modern view treats them as complementary rather than mutually exclusive.
Two tensions continue to organise the field. The first is viewpoint: the evidence supports neither pure object-centred invariance nor pure view dependence but a task-graded mixture, invariant enough to place an object in a broad class from a glance and view-dependent enough to make fine discriminations costly across rotation (Tarr & Pinker, 1989). The second is the relation between psychological models and cortical mechanism — in Marr's terms, between the algorithmic and implementational levels (Marr, 1982). The hierarchical, interactive architectures that psychologists inferred from behaviour — parallel feature extraction, attentional binding, top-down context, successive stages of increasing invariance — turn out to describe the ventral stream well, and the untangling account, now instantiated in deep networks that predict neural responses directly, gives a single geometric statement of what that hierarchy computes (McClelland & Rumelhart, 1981; DiCarlo et al., 2012; Yamins & DiCarlo, 2016). What began as a set of competing psychological theories has become, without discarding them, a layered account in which decomposition, abstraction, binding, structural description, and hierarchical invariance are stages of one process that turns a variable image into a stable name.
Common Misconceptions
- Recognition works by matching the input against a stored picture of the object.
- That is template matching, and it fails on natural vision: a single stored picture cannot absorb the variation in size, position, orientation, and form a real category shows, and covering every distortion with its own template is combinatorially hopeless. The system instead abstracts prototypes, analyses features, and builds structural descriptions (Posner & Keele, 1968; Biederman, 1987).
- Once the visual system has detected an object's features, it has recognised the object.
- Detecting features is not enough, because an object is a particular conjunction of features bound to the right locations. Feature-integration theory shows that binding requires focal attention: without it, features from different objects can be miscombined into illusory conjunctions (Treisman & Gelade, 1980).
- Objects are stored in one viewpoint-invariant form, so orientation does not matter.
- Recognition time rises with the angular distance of a shape from a trained view, a cost a fully viewpoint-invariant representation should not incur. Representations appear partly view-based, invariant enough for broad categorisation but view-dependent for fine discrimination (Tarr & Pinker, 1989).
Glossary
- Binding problem.
- The problem of combining the separately registered features of an object — its colour, shape, and location — into a single unified percept; feature-integration theory holds that focal attention solves it.
- Complex cell.
- A neuron in primary visual cortex, described by Hubel and Wiesel, that responds to an edge or bar of a preferred orientation across a range of positions, giving orientation selectivity with some positional tolerance.
- Deep convolutional neural network.
- A layered image-classification network whose successive convolution-and-pooling stages parallel the ventral stream; trained only to recognise objects, it predicts ventral-stream neural responses better than models hand-designed for the purpose.
- Dorsal stream.
- The occipitoparietal visual pathway that codes the spatial location of objects and guides visually directed action — the where/how stream — distinguished from the ventral what stream that supports recognition.
- Exemplar theory.
- The account on which a category is represented by its stored individual instances rather than an abstracted prototype, a new item being classified by its summed similarity to those remembered exemplars.
- Feature-integration theory.
- Treisman and Gelade's theory that simple features are registered pre-attentively in parallel, but that binding features into an object requires focal attention directed to its location.
- Geon.
- One of the roughly three dozen viewpoint-stable volumetric primitives — brick, cylinder, cone, wedge, and the like — from which recognition-by-components composes objects.
- Global precedence.
- Navon's finding that the overall, global level of a configuration is processed before its local parts and interferes with them asymmetrically — the “forest before the trees.”
- Guided search.
- Wolfe's revision of feature-integration theory in which a parallel stage builds an activation map ranking locations by target similarity, so attention is guided to the most promising items rather than deployed at random, making conjunction search more efficient than a pure serial scan.
- Illusory conjunction.
- A misperception in which features belonging to different objects are wrongly combined — a red X and a green O reported as a green X — occurring when attention cannot bind features correctly.
- Interactive activation.
- McClelland and Rumelhart's model in which feature, letter, and word levels excite and inhibit one another, with words feeding activation back down to their constituent letters.
- Levels of analysis.
- Marr's proposal that an information-processing system be understood at three levels: the computational (what is computed and why), the algorithmic (the representations and processes), and the implementational (the physical mechanism).
- Non-accidental property.
- An image feature, such as a straight versus curved edge or parallel contours, that is preserved across most viewpoints and so can be read reliably off the two-dimensional image to identify a geon.
- Pandemonium.
- Selfridge's feature-analysis model in which a hierarchy of “demons” detect features and feature combinations in parallel, with a decision demon selecting the loudest to recognise a pattern.
- Prototype.
- The abstracted central tendency of a category, recognised even when never seen as an instance; category membership is graded by similarity to the prototype rather than all-or-none.
- Recognition-by-components.
- Biederman's theory that objects are recognised as arrangements of a small alphabet of geons defined by non-accidental properties, yielding robustness to occlusion and viewpoint.
- Template matching.
- The theory that recognition compares an input to stored whole-pattern copies (templates); effective for constrained inputs but defeated by natural variation in size, orientation, and form.
- Ventral stream.
- The occipitotemporal visual pathway, running from primary visual cortex forward, that supports object recognition by building selectivity and invariance across successive stages.
- Viewpoint invariance.
- The property of a representation that stays constant as the vantage point changes; debated against view-based accounts on which recognition is tied to stored views and costs rise with rotation.
- Visual search.
- A task in which an observer looks for a target among distractors; the slope of response time against set size distinguishes parallel feature search from serial conjunction search.
- Word-superiority effect.
- Reicher's finding that a letter is identified more accurately within a real word than alone or in a nonword, evidence that word-level context feeds back to letter recognition.
Key Researchers
Irving Biederman (1939-2022). Harold W. Dornsife Professor of Neuroscience at the University of Southern California; he proposed recognition-by-components, the account of object recognition as the parsing of a shape into a small alphabet of viewpoint-stable geons defined by non-accidental properties. Google Scholar - Wikipedia - Wikidata
James J. DiCarlo. Professor in the Department of Brain and Cognitive Sciences at the Massachusetts Institute of Technology; he framed object recognition as the untangling of object identity along the ventral visual stream, linking primate physiology to hierarchical computational models. Faculty Page - Google Scholar - Wikipedia - ORCID
David Marr (1945-1980). Vision scientist at the Massachusetts Institute of Technology; with Nishihara he set out the object-centred, three-dimensional model representation of shape and the three levels of analysis that framed recognition as a computational problem. Wikipedia - Wikidata
James L. McClelland. Lucie Stern Professor in the Social Sciences at Stanford University; with Rumelhart he built the interactive-activation model of letter perception, showing how top-down word context and bottom-up feature evidence combine to recognise letters. Faculty Page - Google Scholar - Wikipedia - ORCID
Michael I. Posner. Professor Emeritus of Psychology at the University of Oregon; with Keele he provided the classic prototype-abstraction evidence — learning a category from distortions, then recognising the never-seen prototype — a foundation of prototype theory. Faculty Page - Google Scholar - Wikipedia
Anne Treisman (1935-2018). Professor of Psychology at Princeton University; with Gelade she proposed feature-integration theory, on which separable features are registered in parallel but must be bound by focal attention, explaining illusory conjunctions and serial conjunction search. Wikipedia - Wikidata
Frequently Asked Questions
What is pattern recognition in cognitive psychology?
It is the process by which the visual system assigns a variable sensory input to a known category, deciding what a thing is, despite endless variation in the images a single object can cast on the eye (Palmeri & Gauthier, 2004).
Why is pattern recognition considered a hard problem?
Because of invariance: the same object projects a different image every time it changes size, position, orientation, lighting, or viewpoint, so the system must map a vast, continuous space of images onto a small, stable set of categories (Marr & Nishihara, 1978).
What is wrong with template-matching theories?
A stored whole-pattern template cannot absorb natural variation, and covering every distortion of every object with its own template is combinatorially hopeless; template matching survives only for constrained inputs such as barcodes or standardised numerals.
What is a prototype?
A prototype is the abstracted central tendency of a category. Posner and Keele showed people recognise a never-seen prototype as well as the training items, and Rosch showed natural categories are graded around their most typical members (Posner & Keele, 1968; Rosch, 1975).
What is feature-integration theory?
Treisman and Gelade's theory that simple features are detected in parallel across the visual field, but that binding them into an object requires focal attention; without it, features can be miscombined into illusory conjunctions (Treisman & Gelade, 1980).
What are geons?
Geons are the roughly three dozen viewpoint-stable volumetric primitives, such as cylinders, bricks, cones, and wedges, from which Biederman's recognition-by-components composes objects, identified from non-accidental image properties (Biederman, 1987).
What is the word-superiority effect?
A letter is identified more accurately inside a real word than alone or in a nonword; McClelland and Rumelhart modelled it with interactive activation, in which word-level context feeds activation back down to constituent letters (Reicher, 1969; McClelland & Rumelhart, 1981).
How does the brain recognise objects?
Through the ventral visual stream, a hierarchy from primary visual cortex forward that builds selectivity and invariance across stages, progressively untangling object identity so that a simple readout can separate one object from another (DiCarlo et al., 2012; Grill-Spector & Malach, 2004).
References
Biederman, I. (1987). Recognition-by-components: A theory of human image understanding. Psychological Review, 94(2), 115-147. https://doi.org/10.1037/0033-295X.94.2.115
DiCarlo, J. J., Zoccolan, D., & Rust, N. C. (2012). How does the brain solve visual object recognition? Neuron, 73(3), 415-434. https://doi.org/10.1016/j.neuron.2012.01.010
Goodale, M. A., & Milner, A. D. (1992). Separate visual pathways for perception and action. Trends in Neurosciences, 15(1), 20-25. https://doi.org/10.1016/0166-2236(92)90344-8
Grill-Spector, K., & Malach, R. (2004). The human visual cortex. Annual Review of Neuroscience, 27, 649-677. https://doi.org/10.1146/annurev.neuro.27.070203.144220
Hubel, D. H., & Wiesel, T. N. (1962). Receptive fields, binocular interaction and functional architecture in the cat's visual cortex. The Journal of Physiology, 160(1), 106-154. https://doi.org/10.1113/jphysiol.1962.sp006837
Marr, D. (1982). Vision: A computational investigation into the human representation and processing of visual information. W. H. Freeman.
Marr, D., & Nishihara, H. K. (1978). Representation and recognition of the spatial organization of three-dimensional shapes. Proceedings of the Royal Society of London. Series B, Biological Sciences, 200(1140), 269-294. https://doi.org/10.1098/rspb.1978.0020
McClelland, J. L., & Rumelhart, D. E. (1981). An interactive activation model of context effects in letter perception: I. An account of basic findings. Psychological Review, 88(5), 375-407. https://doi.org/10.1037/0033-295X.88.5.375
Medin, D. L., & Schaffer, M. M. (1978). Context theory of classification learning. Psychological Review, 85(3), 207-238. https://doi.org/10.1037/0033-295X.85.3.207
Navon, D. (1977). Forest before trees: The precedence of global features in visual perception. Cognitive Psychology, 9(3), 353-383. https://doi.org/10.1016/0010-0285(77)90012-3
Palmeri, T. J., & Gauthier, I. (2004). Visual object understanding. Nature Reviews Neuroscience, 5(4), 291-303. https://doi.org/10.1038/nrn1364
Posner, M. I., & Keele, S. W. (1968). On the genesis of abstract ideas. Journal of Experimental Psychology, 77(3, Pt.1), 353-363. https://doi.org/10.1037/h0025953
Reicher, G. M. (1969). Perceptual recognition as a function of meaningfulness of stimulus material. Journal of Experimental Psychology, 81(2), 275-280. https://doi.org/10.1037/h0027768
Riesenhuber, M., & Poggio, T. (1999). Hierarchical models of object recognition in cortex. Nature Neuroscience, 2(11), 1019-1025. https://doi.org/10.1038/14819
Rosch, E. (1975). Cognitive representations of semantic categories. Journal of Experimental Psychology: General, 104(3), 192-233. https://doi.org/10.1037/0096-3445.104.3.192
Selfridge, O. G. (1959). Pandemonium: A paradigm for learning. In Symposium on the mechanisation of thought processes (Vol. 1, pp. 511-526). Her Majesty's Stationery Office.
Tarr, M. J., & Pinker, S. (1989). Mental rotation and orientation-dependence in shape recognition. Cognitive Psychology, 21(2), 233-282. https://doi.org/10.1016/0010-0285(89)90009-1
Treisman, A. M., & Gelade, G. (1980). A feature-integration theory of attention. Cognitive Psychology, 12(1), 97-136. https://doi.org/10.1016/0010-0285(80)90005-5
Wertheimer, M. (1923). Untersuchungen zur Lehre von der Gestalt. II. Psychologische Forschung, 4(1), 301-350. https://doi.org/10.1007/BF00410640
Wolfe, J. M. (1994). Guided Search 2.0: A revised model of visual search. Psychonomic Bulletin & Review, 1(2), 202-238. https://doi.org/10.3758/BF03200774
Yamins, D. L. K., & DiCarlo, J. J. (2016). Using goal-driven deep learning models to understand sensory cortex. Nature Neuroscience, 19(3), 356-365. https://doi.org/10.1038/nn.4244