Abstract

Visual perception is the process by which the brain constructs a stable, meaningful representation of the world from the ambiguous, ever-changing pattern of light that reaches the eyes. Because any retinal image is consistent with countless arrangements of surfaces, sizes, and distances, perception is an ill-posed inverse problem that the visual system solves by combining incoming evidence with prior knowledge of how scenes are structured. The article follows this construction from the receptive fields of early visual cortex, through Gestalt perceptual organization and the division of labor between the ventral and dorsal streams, to theories of object recognition and the Bayesian and predictive-coding accounts that cast perception as inference. Three interactive demonstrations model orientation tuning in a cortical cell, perceptual grouping by proximity, and the reliability-weighted fusion of depth cues.

Keywords: visual perception, receptive field, object recognition, Bayesian inference, perceptual organization

Visual perception is the construction of a description of the world from the light captured by the eyes, and its central puzzle is that the description is far richer and more certain than the image it is built from. A single two-dimensional pattern on the retina could have been produced by an unlimited number of three-dimensional scenes, yet observers see one scene, confidently and almost instantly. The visual system resolves this ambiguity by treating the image as evidence to be interpreted rather than a picture to be copied, extracting simple features early and in parallel, organizing them into surfaces and objects, and settling on the interpretation most consistent with how the world is usually arranged. This article traces that process from the earliest cortical feature maps, whose oriented detectors also drive the visual search that attention research measures, up to object recognition and the inferential theories that unify the whole.

Key Takeaways
  • Visual perception is an inverse inference: the retinal image underdetermines the scene, so the brain must reconstruct three-dimensional structure from ambiguous two-dimensional evidence.
  • Early visual cortex analyzes the image into local features through oriented receptive fields, discovered by Hubel and Wiesel, arranged in a hierarchy of increasing receptive-field size and complexity.
  • Perception organizes features into wholes by Gestalt principles such as proximity and similarity, and divides the labor between a ventral stream for recognizing objects and a dorsal stream for guiding action.
  • Object recognition has been modeled as matching part-based structural descriptions, as feedforward hierarchical processing, and as generative inference, with converging evidence for a largely feedforward ventral hierarchy.
  • The Bayesian and predictive-coding frameworks recast perception as the combination of sensory evidence with prior knowledge, explaining constancies, illusions, and the reliability-weighted fusion of cues.

What Visual Perception Is

Visual perception is the set of processes that turn the light imaged on the retina into an experienced world of surfaces, objects, and events. It is not the passive reception of a picture but an active construction: the optic array is sampled by roughly a hundred million photoreceptors, compressed into the signals of about a million ganglion-cell axons, and then elaborated through a cascade of cortical areas that progressively recode the image into representations of edges, regions, shapes, and finally recognizable things. What reaches awareness is the end product of this cascade, a description in which the accidental properties of a particular image, such as its exact retinal size or its momentary shading, have largely been discounted in favor of the stable properties of the world.

The hallmark of perception is therefore its constancy. A door is seen as rectangular though it projects a trapezoid, a piece of coal in sunlight is seen as black though it reflects more light than white paper in shadow, and a friend at the far end of a room is seen at true size though the retinal image has halved. These constancies show that perception delivers the distal object rather than the proximal image, and they are the everyday evidence that the visual system infers the causes of its input rather than transcribing it (Kersten et al., 2004). Explaining how the brain achieves this inference, quickly and largely without error, is the organizing problem of the field.

The Inverse Problem of Vision

The reason perception must be constructive is that image formation destroys information. Projecting a three-dimensional scene onto a two-dimensional surface collapses depth, so an infinite family of scenes, differing in the sizes and distances of their parts, produces exactly the same image; recovering the scene from the image is the inverse of this projection and is mathematically ill-posed, admitting no unique solution from the data alone (Marr & Nishihara, 1978). The same ambiguity recurs at every level: a given contour could be an object boundary or a shadow edge, a given gray could be a dark surface brightly lit or a light surface dimly lit, and a given retinal motion could arise from object movement or from the observer's own.

The visual system makes the problem tractable by adding assumptions. If it presumes that light usually comes from above, that surfaces are usually smooth, that objects are usually rigid, and that viewpoints are usually generic rather than accidental, then most images have a single interpretation that satisfies both the evidence and the assumptions, and that interpretation is what is seen (Yuille & Kersten, 2006). This is why the study of perception leans so heavily on illusions: they are the cases where the built-in assumptions are wrong for the particular stimulus, so the percept departs from reality in a way that reveals the assumption. Perception is veridical in the ordinary environment precisely because the assumptions are true of that environment.

The Visual Pathway and Its Hierarchy

The anatomical substrate of this construction is a hierarchy. Signals pass from the retina through the lateral geniculate nucleus of the thalamus to primary visual cortex, area V1, and from there through a mesh of extrastriate areas including V2, V4, and the middle temporal area, up to inferior temporal cortex. A systematic analysis of the macaque cortex identified more than thirty visual areas linked by over three hundred pathways, arranged with remarkable consistency into an ascending hierarchy defined by the laminar origin and termination of the connections between areas (Felleman & Van Essen, 1991). The hierarchy is not a simple ladder: most connections are reciprocal, so that every feedforward projection carrying signals up the hierarchy is matched by a feedback projection carrying signals down.

Two properties change systematically as one ascends. Receptive fields grow larger, so that a V1 neuron responds only to a tiny patch of the field while an inferior-temporal neuron responds across much of it, and the features that drive a neuron grow more complex, from oriented edges early to object parts and whole objects late (Grill-Spector & Malach, 2004). This progression from small, simple, and local to large, complex, and global is the physiological expression of building objects out of features. The abundance of feedback connections, however, warns against reading the hierarchy as a one-way assembly line, and it is the anatomical basis for the top-down influences that later sections take up.

Receptive Fields and Feature Detection

The first cortical stage of perception was opened up by the discovery of what individual visual neurons respond to. Recording from the cat's striate cortex, Hubel and Wiesel found that most cells are indifferent to diffuse light but fire briskly to a bar or edge at a particular orientation, and that a small change in orientation sharply reduces the response; they distinguished simple cells, whose elongated on and off subregions predict their preference from the layout of the receptive field, from complex cells, which signal an oriented feature anywhere within a larger region (Hubel & Wiesel, 1962). Orientation-selective neurons of this kind are the elementary feature detectors of vision, decomposing the image into local measurements of edge orientation, and they are organized into orderly columns so that nearby patches of cortex prefer nearby orientations.

The importance of the finding is that it made feature extraction physiological and measurable. An oriented receptive field is, in effect, a filter that reports how much edge energy of a given orientation is present at a given location, and a bank of such filters spanning orientations and scales provides a rich local description from which higher stages can build. The orientation tuning of a single cell is the atom of that description: its firing rate falls off as the stimulus orientation departs from the cell's preference, tracing a tuning curve whose width sets how finely orientation is coded. The demonstration below realizes this directly, letting the reader rotate a stimulus bar and read off the modeled firing rate of an orientation-tuned cell as the response peaks at the preferred orientation and decays smoothly away from it.

The Feature Detector

Orientation Tuning of a Cortical Cell

A single neuron in primary visual cortex fires most strongly to an edge at its preferred orientation, here vertical at 0 degrees, and falls silent as the edge is rotated away. Rotate the stimulus bar and read off the modeled firing rate: the response peaks at the preferred orientation and decays smoothly, tracing the tuning curve that defines how finely the cell codes orientation.

Firing rate16.2 sp/s
02550Firing rate (sp/s)Stimulus orientation (deg)0306090120150180
Stimulus orientation30°
The stimulus is 30° from the cell's preferred orientation, so the modeled rate is 16.2 spikes per second, or 32% of the peak. Off the preferred orientation the response is graded, which is how a population of such cells codes orientation precisely.
An illustrative implementation of orientation tuning in a V1 simple cell (form after Hubel & Wiesel, 1962), with representative constants (peak 50 spikes/s, tuning width 20 degrees). The modeled firing rate is a Gaussian function of the stimulus orientation relative to the cell's preferred orientation. Values are computed locally, not stored.

Perceptual Organization

Local features do not add up to a scene until they are grouped, and the rules by which the visual system decides what belongs with what were first catalogued by the Gestalt psychologists. Elements that are near one another, similar to one another, moving together, or arranged along a smooth continuation tend to be perceived as a single unit, and a bounded region tends to be seen as a figure lying in front of a formless ground (Wagemans et al., 2012). These grouping principles are not decorative; they are the visual system's solution to the problem of segmentation, the partitioning of the image into the surfaces and objects that further processing will treat as things. Because the image itself does not label its parts, grouping must be imposed, and the Gestalt principles express the regularities the system relies on to impose it.

Modern work has recast grouping as another instance of inference under ambiguity, in which the principles reflect the statistics of natural scenes: edges really do tend to continue smoothly, and features of one object really are more similar to each other than to features of the background, so the grouping the principles prescribe is usually the correct one (Wagemans et al., 2012). Figure-ground assignment in particular is a decision about which region owns the contour between two regions, and it can flip when the cues are balanced, as in the classic vase-faces figure. The demonstration below isolates the principle of proximity: the reader varies the horizontal and vertical spacing of a dot lattice and watches it reorganize into rows or columns as the nearer neighbors capture the grouping.

Organizing The Image

Perceptual Grouping by Proximity

The dots are identical, so nothing about them individually says how they belong together. Yet the lattice is seen as rows when the horizontal neighbours are closer and as columns when the vertical neighbours are closer. Adjust the two spacings and watch the same dots reorganize, a demonstration that grouping is imposed by the visual system rather than given by the elements.

Horizontal spacing22 px
Vertical spacing44 px
The vertical-to-horizontal spacing ratio is 2.00, so the lattice groups into rows. Because the horizontal neighbours are closer, proximity binds each row into a unit.
An illustrative implementation of grouping by proximity, one of the Gestalt principles (form after Wagemans et al., 2012). A regular lattice of identical dots is seen as rows or as columns depending only on which neighbours are closer. Grouping direction is read off the ratio of vertical to horizontal spacing. Values are computed locally, not stored.

The Two Visual Streams

Beyond the early areas the cortical hierarchy splits into two great streams, and the division of labor between them reorganized how perception is understood. A ventral stream running from V1 into inferior temporal cortex codes the properties needed to identify what an object is, while a dorsal stream running into posterior parietal cortex codes the spatial and dynamic properties needed to act on it. The perception-action formulation of this distinction holds that the ventral stream constructs the conscious percept used for recognition and the dorsal stream computes the moment-to-moment visuomotor parameters used for reaching and grasping, so that the two streams serve different purposes rather than merely coding different attributes (Goodale & Milner, 1992). The strongest evidence is a double dissociation in neurological patients, in whom ventral damage can abolish the ability to report an object's size and orientation while the hand still scales its grip to that same object, and dorsal damage can do the reverse.

The two-stream account explains why some visual illusions that fool perceptual judgment leave the guidance of action comparatively unaffected, and it grounds the intuition that seeing-for-knowing and seeing-for-doing are partly separate achievements. It also connects vision to the rest of cognition: the ventral stream feeds recognition, memory, and the deliberate use of objects, whereas the dorsal stream feeds the fast, largely unconscious control of movement. The split is not absolute, and the streams interact extensively, but the framework remains the standard organizing map of high-level visual cortex and a reference point for questions about which visual computations reach awareness.

Object Recognition

Recognizing an object is the visual system's most demanding feat, because the same object casts wildly different images as its distance, pose, and lighting vary, and yet must be assigned a single identity. One influential class of theory solves this invariance problem with structural descriptions: an object is represented as an arrangement of simple volumetric parts in an object-centered frame, so that recognition is matching a viewpoint-independent description of parts and their relations rather than a stored image (Marr & Nishihara, 1978). Recognition-by-components made this concrete by proposing a small alphabet of generalized-cone primitives, called geons, recoverable from viewpoint-invariant properties of edges such as whether they are straight or curved and how they meet, so that a few geons in the right relations specify an object across most views (Biederman, 1987). Table 1 sets the leading families of recognition theory side by side.

Table 1. Leading families of object-recognition theory compared by their proposed representation and signature evidence.
Family How an object is represented Signature evidence
Structural description An object-centered arrangement of volumetric parts (geons) and their spatial relations, recovered from viewpoint-invariant edge properties Recognition robust to rotation when part structure is preserved; impaired when parts are occluded at their joins
Hierarchical feedforward A cascade of tuning and pooling stages that builds increasingly complex, tolerant features up to view-tolerant units in inferior temporal cortex Rapid recognition within about 100 ms; linearly decodable object identity in IT population activity
Generative inference A probabilistic model of how objects generate images, with recognition as inverting that model to find the most probable cause Prior-driven completion and illusions; reliability-weighted cue combination close to the statistical optimum

A second class of theory abandons explicit parts in favor of a feedforward hierarchy of feature detectors, in which alternating stages of tuning and pooling build units that are selective for complex forms yet tolerant of changes in position and scale, reaching view-tolerant object units at the top (Riesenhuber & Poggio, 1999). Physiological work supports a largely feedforward solution: core recognition is fast, and object identity can be read out by a simple classifier from the population activity of inferior temporal cortex, which untangles the identity information that is hopelessly interwoven at the retina (DiCarlo et al., 2012). The two classes are less opposed than they first appear, since single-unit studies of temporal cortex show neurons tuned to complex objects and views in ways both traditions must accommodate (Logothetis & Sheinberg, 1996). Figure 1 summarizes the growth of receptive-field size along the ventral hierarchy that any such account must implement.

Figure 1

Growth of Receptive-Field Size Along the Ventral Visual Hierarchy

Receptive-field size increases from V1 to inferior temporal cortex A bar chart with four ascending stages of the ventral visual hierarchy on the horizontal axis, labelled V1, V2, V4, and inferior temporal cortex, and approximate receptive-field diameter in degrees of visual angle on the vertical axis, from zero to fifty. The bars rise steeply from about one degree at V1 to about two degrees at V2, about four to five degrees at V4, and about forty to fifty degrees at inferior temporal cortex, illustrating the roughly order-of-magnitude growth in receptive-field size as processing ascends the hierarchy. 50 25 0 RF diameter (deg) Stage of the ventral hierarchy V1 V2 V4 IT ~1 ~2 ~5 ~45

Note. Representative receptive-field diameters at four stages of the ventral stream (form after Felleman & Van Essen, 1991, and Grill-Spector & Malach, 2004). Values are illustrative order-of-magnitude constants, not measured data, and vary with eccentricity.

Perception as Bayesian Inference

The idea that perception is inference, first put forward when Helmholtz described it as unconscious inference, has a modern formalization in Bayesian probability. On this account the visual system represents a likelihood, how probable the current image is given each possible scene, and a prior, how probable each scene is in the first place, and it combines them by Bayes' rule to obtain a posterior over scenes, perceiving the scene that maximizes it (Kersten et al., 2004). The framework explains constancies as the influence of priors that discount viewing conditions, explains illusions as priors misapplied to atypical stimuli, and above all explains cue combination: when several sources of information about a property are available, the optimal estimate weights each by its reliability, and human perceivers combine cues such as stereo and texture in close to this reliability-weighted way (Knill & Pouget, 2004).

Predictive coding gives the inference a plausible cortical implementation. In this scheme higher areas send predictions down the hierarchy through feedback connections, each level compares the prediction against its input, and only the residual prediction error is passed forward, so that the feedforward signal carries what the model failed to anticipate rather than the raw image; the extra-classical suppression of a V1 neuron by a predictable surround falls out of this arithmetic (Rao & Ballard, 1999). Perception on this view is the settling of the whole hierarchy into the state that best predicts its sensory input, which reframes the abundant feedback connections as the machinery of inference rather than mere modulation. The demonstration below makes the cue-combination principle concrete: the reader sets the estimate and reliability of two depth cues and reads off the reliability-weighted fused estimate, whose uncertainty is smaller than that of either cue alone.

Perception As Inference

Reliability-Weighted Cue Integration

Two cues, say binocular disparity and texture gradient, each estimate the depth of a surface, and each carries some uncertainty shown by the width of its curve. The visual system does not average them equally; it weights each by its reliability, so the fused estimate leans toward the sharper cue and is more precise than either. Adjust the estimates and their uncertainties and watch the fused curve.

05101520Estimated depth (arbitrary units)
Cue 1Cue 2Fused estimate
Cue 1 estimate10
Cue 1 uncertainty2.0
Cue 2 estimate13
Cue 2 uncertainty1.0
Cue 2 carries weight 0.80 and cue 1 weight 0.20, so the fused estimate is 12.40 with a standard deviation of 0.894, smaller than either cue's. Because cue 2 is the more reliable, the fused estimate is pulled toward it.
An illustrative implementation of Bayesian cue combination (form after Knill & Pouget, 2004; Kersten et al., 2004). Two independent cues to the same depth are fused by weighting each by its reliability, the reciprocal of its variance. The fused estimate is pulled toward the more reliable cue, and its uncertainty is smaller than that of either cue alone. Values are computed locally, not stored.

Visual Field Maps in the Cortex

The construction of a percept is laid out on the cortex in an orderly spatial format. Early visual areas are retinotopic: neighboring locations in the visual field map to neighboring patches of cortex, so that each area contains an internal map of the field, distorted to devote disproportionate territory to the high-acuity center of gaze. Functional imaging has traced more than a dozen such visual field maps across human occipital, parietal, and temporal cortex, and the boundaries between maps, where the representation of the visual field reverses, provide a reliable parcellation of visual cortex into distinct areas in the living brain (Wandell et al., 2007). These maps are the coordinate system on which the hierarchy's computations are performed, and their orderly layout is what makes early vision spatially precise.

Higher in the ventral stream the maps give way to a mosaic of regions selective for categories of stimuli, including patches that respond more strongly to faces, to places, or to printed words than to other objects (Grill-Spector & Malach, 2004). The transition from retinotopic maps of space to category-selective regions is the anatomical signature of the shift from coding where things are to coding what they are, and it dovetails with the two-stream architecture and with the untangling of object identity that recognition requires. Mapping these areas has also made high-level vision a testbed for the top-down influences discussed next, since the responses of category-selective cortex depend on attention, expectation, and task as well as on the stimulus.

What the Framework Does and Does Not Explain

The largest open question is how far vision is feedforward and how far it depends on feedback. The speed of recognition and the decodability of identity from inferior temporal cortex show that a single feedforward sweep accomplishes a great deal, and purely feedforward hierarchical models capture much of the behavior (DiCarlo et al., 2012). Yet the anatomy is dominated by feedback, and top-down signals carrying attention, expectation, task, and perceptual learning demonstrably reshape responses even in primary visual cortex, so that what a V1 neuron signals depends on the behavioral context and not only on the light in its receptive field (Gilbert & Li, 2013). A complete account must say when feedforward processing suffices and when recurrent computation is required, for example to segment cluttered scenes or resolve ambiguous figures.

The inferential frameworks face a complementary challenge. Casting perception as Bayesian inference is powerful and often quantitatively accurate, but a Bayesian model is only as constrained as the priors and likelihoods it assumes, and specifying those independently of the data they are meant to predict is not always possible (Yuille & Kersten, 2006). Predictive coding likewise remains partly a framework in search of decisive tests, since several distinct circuit schemes can implement approximate inference and the physiological signatures do not yet uniquely favor one. What is not in dispute is the reframing itself: the field has moved from asking how the image is transmitted to asking how the scene is inferred, and the enduring contribution of the last decades is the recognition that vision is a problem of statistical estimation performed by a hierarchical, recurrent cortical network.

Worked Example

Consider the receptive-field demonstration with its default constants. The modeled firing rate of an orientation-tuned cell is a Gaussian function of the difference between the stimulus orientation and the cell's preferred orientation, R equals R-max times the exponential of minus the squared orientation difference divided by twice the squared tuning width. With a maximum rate of 50 spikes per second, a preferred orientation of 0 degrees, and a tuning width of 20 degrees, a stimulus at 30 degrees gives an orientation difference of 30, so the exponent is minus 900 divided by 800, or minus 1.125, and the rate is 50 times the exponential of minus 1.125, which is 50 times 0.3247, or 16.2 spikes per second. A stimulus at the preferred orientation gives the full 50 spikes per second, and one 40 degrees away gives 50 times the exponential of minus 2.0, or 6.8 spikes per second, so the response falls to about an eighth of its peak within two tuning widths, the sharp orientation selectivity Hubel and Wiesel described.

The cue-integration demonstration reaches its numbers from the reliability-weighted combination rule. Each cue is weighted by its reliability, the reciprocal of its variance, and the weights are normalized to sum to one. Take a first depth cue estimating 10 units with a standard deviation of 2, so a reliability of one quarter, and a second cue estimating 13 units with a standard deviation of 1, so a reliability of 1. The normalized weights are 0.25 divided by 1.25, or 0.2, for the first cue and 1 divided by 1.25, or 0.8, for the second, so the fused estimate is 0.2 times 10 plus 0.8 times 13, which is 2 plus 10.4, or 12.4 units, pulled toward the more reliable cue. The variance of the fused estimate is the reciprocal of the summed reliabilities, 1 divided by 1.25, or 0.8, for a standard deviation of 0.894 units, smaller than either cue's, which is the central Bayesian result that combining information reduces uncertainty.

Discussion

Visual perception is best understood as inference on a hierarchy. The retina and early cortex analyze the image into local features through the oriented receptive fields Hubel and Wiesel discovered; intermediate stages group those features into surfaces and objects by principles that mirror the statistics of natural scenes; and higher stages, split into a ventral stream for identity and a dorsal stream for action, recognize objects by untangling their identity from the accidents of view. Across every stage the same logic recurs, that the image underdetermines the world and must be interpreted by combining evidence with prior knowledge, and the Bayesian and predictive-coding frameworks make that logic explicit and quantitative (Kersten et al., 2004; Rao & Ballard, 1999).

The framework's power is its unification. Constancies, illusions, cue combination, the growth of receptive fields, the two streams, and the feedforward speed of recognition are not a list of separate facts but consequences of a single idea, that the brain builds the most probable scene consistent with its input and its assumptions. The unresolved questions concern mechanism rather than principle: how much of the inference is completed in a feedforward sweep and how much requires recurrence, how priors are learned and stored, and how the abstract computation is realized in cortical circuits (Gilbert & Li, 2013; DiCarlo et al., 2012). That these are now the questions is itself the measure of progress, for the field has replaced the metaphor of the eye as a camera with the harder and more accurate picture of vision as the brain's best guess about what is there.

Glossary

Bayesian inference.
The combination of a likelihood, the probability of the image given a scene, with a prior over scenes, to obtain a posterior whose most probable scene is perceived.
Complex cell.
A visual cortical neuron that signals an oriented edge anywhere within its receptive field, showing position tolerance that a simple cell lacks.
Dorsal stream.
The pathway from primary visual cortex into posterior parietal cortex that codes spatial and dynamic information for the visual guidance of action.
Feature detector.
A neuron or filter that responds selectively to a specific image property, such as an edge of a particular orientation, at a particular location.
Geon.
In recognition-by-components, one of a small set of generalized-cone volumetric primitives from which objects are composed and by which they are recognized.
Gestalt grouping.
The organization of separate elements into perceived wholes according to principles such as proximity, similarity, common fate, and good continuation.
Inverse problem.
The ill-posed task of recovering the three-dimensional scene that produced a two-dimensional image, which has no unique solution from the image alone.
Object recognition.
The assignment of a single identity to an object across the varied images it casts as its view, distance, and lighting change.
Perceptual constancy.
The tendency to perceive a stable property of an object, such as its size, shape, or lightness, despite changes in the retinal image.
Predictive coding.
A scheme in which higher areas predict lower-level activity through feedback and only the residual prediction error is passed forward up the hierarchy.
Prior.
The probability the visual system assigns to a scene before the current image, encoding assumptions such as light coming from above or surfaces being smooth.
Receptive field.
The region of the visual field within which a stimulus alters a neuron's firing, together with the stimulus features to which the neuron is tuned.
Retinotopy.
The orderly mapping of the visual field onto cortex whereby neighboring field locations project to neighboring cortical locations within a visual area.
Simple cell.
A visual cortical neuron whose orientation preference is predicted by the layout of distinct on and off subregions in its receptive field.
Ventral stream.
The pathway from primary visual cortex into inferior temporal cortex that codes object identity and supports conscious recognition.
Visual field map.
A cortical area containing an orderly, if distorted, representation of the visual field, bounded by reversals in the field representation.

Key Researchers

David H. Hubel. Professor at Harvard Medical School until his death in 2013; with Wiesel he mapped the receptive fields of visual cortical neurons, discovering orientation selectivity and the columnar architecture of area V1, work that earned a share of the 1981 Nobel Prize. Wikipedia - Wikidata

Torsten N. Wiesel. Emeritus professor at Rockefeller University; with Hubel he established that cortical neurons are tuned to orientation and organized into columns whose development depends on early visual experience, sharing the 1981 Nobel Prize in Physiology or Medicine. Faculty Page - Wikipedia - Wikidata

David Marr. Vision scientist at the Massachusetts Institute of Technology until his death in 1980; he argued that vision must be understood at computational, algorithmic, and implementational levels, and with Nishihara proposed an object-centered, part-based representation of three-dimensional shape. Wikipedia) - Wikidata

Irving Biederman. Professor at the University of Southern California until his death in 2022; author of recognition-by-components, the theory that objects are recognized from a small alphabet of volumetric primitives, or geons, recoverable from viewpoint-invariant edge properties. ORCID - Wikipedia - Wikidata

Melvyn A. Goodale. Emeritus distinguished university professor at Western University; with Milner he proposed the perception-action model of the two visual streams, in which a ventral stream supports conscious perception and a dorsal stream supports the visual control of action. ORCID - Faculty Page - Google Scholar

James J. DiCarlo. Professor of neuroscience at the Massachusetts Institute of Technology; he characterized the ventral stream as a hierarchy that untangles object identity into a linearly readable representation in inferior temporal cortex, linking primate physiology to deep-network models of recognition. ORCID - Faculty Page - Google Scholar

Frequently Asked Questions

What is visual perception?
It is the process by which the brain constructs a representation of the surrounding world from the light imaged on the retina, actively inferring the surfaces, objects, and events that most probably produced the image rather than passively transcribing it (Kersten et al., 2004).

Why is visual perception called an inverse problem?
Because projecting a three-dimensional scene onto the two-dimensional retina destroys depth information, so infinitely many scenes could have produced the same image; recovering the scene is the inverse of projection and has no unique solution without added assumptions (Marr and Nishihara, 1978).

What did Hubel and Wiesel discover?
They found that neurons in primary visual cortex respond selectively to edges of a particular orientation rather than to diffuse light, distinguished simple from complex cells, and showed that these detectors are organized into orderly columns (Hubel and Wiesel, 1962).

What are the two visual streams?
A ventral stream running into inferior temporal cortex codes what an object is and supports conscious recognition, while a dorsal stream running into parietal cortex codes where it is and how to act on it, guiding movement (Goodale and Milner, 1992).

How does the brain recognize objects despite changes in view?
Theories propose either matching a viewpoint-invariant description of an object's parts and their relations, or a feedforward hierarchy that builds view-tolerant object units, with evidence that inferior temporal cortex represents identity in a form a simple classifier can read (Biederman, 1987; DiCarlo et al., 2012).

What does it mean to say perception is Bayesian inference?
It means the visual system combines the likelihood of the image under each possible scene with prior probabilities over scenes and perceives the most probable scene, which explains constancies, illusions, and the reliability-weighted combination of cues (Kersten et al., 2004; Knill and Pouget, 2004).

What is predictive coding?
It is a proposal that higher cortical areas predict the activity of lower areas through feedback, that each level passes forward only the error left unexplained by the prediction, and that perception is the state in which the hierarchy best predicts its input (Rao and Ballard, 1999).

Are the Gestalt grouping principles still accepted?
Yes, in updated form; proximity, similarity, common fate, and good continuation remain robust descriptions of how the visual system segments a scene, now understood as reflecting the statistical regularities of natural images (Wagemans et al., 2012).

References

Biederman, I. (1987). Recognition-by-components: A theory of human image understanding. Psychological Review, 94(2), 115-147. https://doi.org/10.1037/0033-295X.94.2.115

DiCarlo, J. J., Zoccolan, D., & Rust, N. C. (2012). How does the brain solve visual object recognition? Neuron, 73(3), 415-434. https://doi.org/10.1016/j.neuron.2012.01.010

Felleman, D. J., & Van Essen, D. C. (1991). Distributed hierarchical processing in the primate cerebral cortex. Cerebral Cortex, 1(1), 1-47. https://doi.org/10.1093/cercor/1.1.1

Gilbert, C. D., & Li, W. (2013). Top-down influences on visual processing. Nature Reviews Neuroscience, 14(5), 350-363. https://doi.org/10.1038/nrn3476

Goodale, M. A., & Milner, A. D. (1992). Separate visual pathways for perception and action. Trends in Neurosciences, 15(1), 20-25. https://doi.org/10.1016/0166-2236(92)90344-8

Grill-Spector, K., & Malach, R. (2004). The human visual cortex. Annual Review of Neuroscience, 27, 649-677. https://doi.org/10.1146/annurev.neuro.27.070203.144220

Hubel, D. H., & Wiesel, T. N. (1962). Receptive fields, binocular interaction and functional architecture in the cat's visual cortex. The Journal of Physiology, 160(1), 106-154. https://doi.org/10.1113/jphysiol.1962.sp006837

Kersten, D., Mamassian, P., & Yuille, A. (2004). Object perception as Bayesian inference. Annual Review of Psychology, 55, 271-304. https://doi.org/10.1146/annurev.psych.55.090902.142005

Knill, D. C., & Pouget, A. (2004). The Bayesian brain: The role of uncertainty in neural coding and computation. Trends in Neurosciences, 27(12), 712-719. https://doi.org/10.1016/j.tins.2004.10.007

Logothetis, N. K., & Sheinberg, D. L. (1996). Visual object recognition. Annual Review of Neuroscience, 19, 577-621. https://doi.org/10.1146/annurev.ne.19.030196.003045

Marr, D., & Nishihara, H. K. (1978). Representation and recognition of the spatial organization of three-dimensional shapes. Proceedings of the Royal Society of London. Series B, Biological Sciences, 200(1140), 269-294. https://doi.org/10.1098/rspb.1978.0020

Rao, R. P. N., & Ballard, D. H. (1999). Predictive coding in the visual cortex: A functional interpretation of some extra-classical receptive-field effects. Nature Neuroscience, 2(1), 79-87. https://doi.org/10.1038/4580

Riesenhuber, M., & Poggio, T. (1999). Hierarchical models of object recognition in cortex. Nature Neuroscience, 2(11), 1019-1025. https://doi.org/10.1038/14819

Wagemans, J., Elder, J. H., Kubovy, M., Palmer, S. E., Peterson, M. A., Singh, M., & von der Heydt, R. (2012). A century of Gestalt psychology in visual perception: I. Perceptual grouping and figure-ground organization. Psychological Bulletin, 138(6), 1172-1217. https://doi.org/10.1037/a0029333

Wandell, B. A., Dumoulin, S. O., & Brewer, A. A. (2007). Visual field maps in human cortex. Neuron, 56(2), 366-383. https://doi.org/10.1016/j.neuron.2007.10.012

Yuille, A., & Kersten, D. (2006). Vision as Bayesian inference: Analysis by synthesis? Trends in Cognitive Sciences, 10(7), 301-308. https://doi.org/10.1016/j.tics.2006.05.002