Abstract

Form perception, which the Medical Subject Headings classify under space perception, is the set of processes by which the visual system recovers the shapes, contours, and boundaries of objects from the retinal image. Because an outline must be extracted from local changes in luminance rather than read off directly, the system links oriented edge signals into continuous contours, assigns each border to a figure rather than its ground, and groups the parts into wholes by the Gestalt principles. Recovered shape is then coded in a form stable across position and size, matched to stored structure along the ventral cortical pathway, and completed across occluding surfaces. Whether shape is built from parts or from image statistics, and how far deep networks capture it, remains contested. Three interactive demonstrations model orientation tuning, contour integration, and illusory contours.

Keywords: form perception, contour integration, figure-ground organization, object recognition

Form perception refers to the recovery of the shape of an object, the layout of its bounding contour and the segregation of that object from its surroundings, from the distribution of light across the retina. The central problem is that a shape is not marked in the image: the retina registers only local intensities, and the boundary of an object appears as nothing more than a set of scattered luminance changes that the visual system must detect, link into continuous contours, and assign to one surface rather than another (Wertheimer, 1923). This article follows form perception from the oriented edge signals of early visual cortex, through the integration of those signals into contours, the assignment of borders and the perception of illusory edges, the Gestalt grouping of parts into wholes, the description of three-dimensional shape, the matching of shape to stored structure in the ventral pathway, and the completion of occluded objects, to the criticisms that current comparisons with deep neural networks have sharpened.

Key Takeaways
  • Form perception constructs the shape and boundary of an object from local luminance changes that do not, on their own, mark where one surface ends and the next begins.
  • Early visual cortex signals the orientation of local edges; these signals are then linked into continuous contours by an association field that favours smooth, co-circular paths.
  • Every contour must be assigned to a figure rather than its background, a decision the visual system can make even where no physical edge exists, as in illusory contours.
  • The Gestalt principles describe how parts are grouped into perceptual wholes, and recovered shape is matched to stored structure along the ventral cortical pathway.
  • Deep neural networks now rival human object recognition yet often classify by local texture rather than global shape, which has become a probe of what form perception really computes.

What Form Perception Is

Form perception is the construction of a representation of an object's shape from sensory data that do not specify it directly. The retina transduces the intensity and wavelength of light at each point, but shape is a relational property, defined by where luminance changes and how those changes curve and close, so the outline of an object is not given as such in the image and must be inferred from its traces (Wertheimer, 1923). The visual system resolves the resulting ambiguity by exploiting regularities that ordinarily hold in the world, such as the tendency of surfaces to be bounded by smooth continuous edges, of nearby similar elements to belong to the same object, and of occluding surfaces to interrupt rather than terminate the things behind them.

The problem has several connected parts. Local contrast must first be detected and its orientation measured; the resulting fragments must be linked into extended contours; each contour must be assigned to the figure it bounds rather than to the adjoining ground; the grouped contours must be organised into a shape; and that shape must be coded in a form stable enough to be recognised when the object moves, turns, or is partly hidden. A percept of form is therefore not a single image but the product of a sequence of inferences, each resolving one kind of ambiguity, and the sections that follow treat these in turn, beginning with the oriented edge signals from which everything else is built.

Types of Form Perception

Beyond being a subject in its own right, Form Perception is a formal category in the National Library of Medicine's Medical Subject Headings, which places it at tree position F02.463.593.778.435, beneath Space Perception, and hangs its recognised narrower kinds beneath it. These subtypes are a classification built to index the literature, not a claim about the mind's natural joints. Table 2 lists the direct children of the descriptor; neither has a dedicated article on this site yet, so each is described from its MeSH scope note rather than linked.

Table 2. Direct subtypes of Form Perception in the MeSH classification (tree F02.463.593.778.435).
Subtype In brief
Contrast Sensitivity The ability to detect sharp boundaries and slight changes in luminance where contours are faint, setting the lower limit on the contrast at which any form can be seen.
Stereognosis The perception of the shape and form of objects by touch and kinaesthesis, the haptic counterpart of visual form perception.

Two cautions keep this taxonomy in its place. It is a classification for indexing, built to organise the literature, not a theory asserting these subtypes are mutually exclusive or an exhaustive set of natural kinds. And a MeSH subtype is a narrower topic, not a component process: listing Contrast Sensitivity under Form Perception locates it in an index and says nothing, on its own, about the mechanisms this article describes.

From Edges to Oriented Contours

The building blocks of form are supplied by the primary visual cortex, where individual neurons respond selectively to the orientation of a local edge. Recording from the cat's striate cortex, Hubel and Wiesel found cells that fired vigorously to a bar or edge at one orientation within a small region of the visual field and fell silent as the bar was rotated away from that preferred angle, so each cell in effect measures the presence of a contour at a particular orientation and location (Hubel & Wiesel, 1962). These simple cells have receptive fields divided into adjacent excitatory and inhibitory zones, which is why an edge aligned with the boundary between the zones drives them best, and the population as a whole tiles the image with orientation detectors at every position.

Orientation selectivity is the first commitment the visual system makes about form: before any object is identified, the image has been recoded from a map of intensities into a map of local contour orientations. This recoding is powerful but strictly local, since each cell sees only a small patch and cannot tell whether the edge it signals belongs to a meaningful boundary or to noise. The first demonstration models the orientation tuning of such a cell, letting a reader rotate a bar across a simple cell's preferred orientation and watch the response rise and fall, and the Worked Example gives the idealised tuning curve behind it. The problem the next section takes up is how these local, ambiguous orientation signals are stitched together into the extended contours that actually bound objects.

From Edges to Contours

Orientation Tuning of a Simple Cell

A simple cell in the primary visual cortex responds best to an edge at one particular orientation, its preferred orientation, drawn here as vertical. Rotate the bar away from that orientation and the cell’s response falls, tracing the tuning curve on the right. The response is halved once the bar is turned 45 degrees from the preferred orientation.

preferredstimulus bar-90°-45°0°45°90°0.00.51.0response = cos²(θ)
Bar orientation (misalignment θ)0°
At a misalignment of 0° the normalised response is cos²(0°) = 1.00. Aligned with the preferred orientation, the cell fires maximally.
An illustrative model of the orientation tuning of a primary visual cortex simple cell, response = cos-squared of the misalignment theta (after Hubel & Wiesel, 1962). The receptive field's preferred orientation is drawn vertical; the response falls to half its peak at 45 degrees. The defaults reproduce the Worked Example. Values are computed locally, not stored.

Contour Integration and the Association Field

A boundary in a natural image is rarely a clean unbroken line; it is a chain of separate edge elements, often interrupted by noise, that the visual system must recognise as a single continuous contour. Field, Hayes and Hess showed that observers can pick out a path of oriented elements embedded in a field of randomly oriented distractors, but only when the elements along the path are aligned so that each is roughly tangent to a smooth curve passing through its neighbours, an arrangement they called the association field (Field, Hayes & Hess, 1993). Detection stays easy while successive elements turn through only a small angle and collapses as the path becomes more jagged, which shows that the linking process favours contours that are smooth and, more precisely, co-circular. Figure 1 shows the association field schematically.

Figure 1

The Association Field for Contour Integration

Preferential linking of oriented elements along a smooth curve A central oriented edge element sits at the middle of the panel. Neighbouring elements arranged along a gentle curve are oriented nearly tangent to that curve and are joined to the centre by solid links, showing strong grouping. Off the curve, elements oriented at sharp angles to the local tangent are joined by faint dashed links or none at all, showing weak grouping. The figure contrasts the smooth co-circular path that the visual system binds into a contour with the misaligned elements it leaves ungrouped. target element Aligned neighbours (amber) link strongly; misaligned ones (grey) do not

Note. Schematic of the association field. Elements whose orientation continues the target's along a smooth, co-circular curve are grouped into one contour (solid links); elements requiring a sharp turn are linked weakly or not at all (dashed). Illustrative geometry, not measured data.

The association field is a rule for grouping local orientation signals into global contours: an element links preferentially to neighbours whose orientation continues its own along a gentle curve, and links weakly or not at all to neighbours that would require a sharp turn. This is the perceptual expression of good continuation, one of the Gestalt grouping principles, implemented at the level of oriented edge detectors and their lateral connections. The second demonstration presents such a path in noise and lets a reader vary the angle by which successive elements turn, so that a smooth path pops out and a jagged one dissolves into the background. Once a continuous contour has been assembled, the visual system faces a further decision that the elements themselves do not settle: which side of the contour is the object.

Linking Edges into Contours

Contour Integration and the Association Field

A contour in a natural image is a chain of separate edge elements the visual system must recognise as one line. The amber elements below form a path through a field of random distractors. Increase the turn angle between successive path elements and watch the smooth contour break up: the association field links elements that continue one another along a gentle curve, not those that turn sharply.

Turn angle between elements (β)6°
Path elementsDistractors
With a turn angle of 6° between successive elements, the path’s salience is about 93%. The elements are nearly collinear, so the association field binds them into one smooth contour that pops out of the noise.
An illustrative model of contour integration (after Field, Hayes & Hess, 1993). A path of oriented elements (amber) is embedded in randomly oriented distractors; each path element is turned from the path direction by an alternating turn angle beta. Small angles keep the path smooth and salient, large angles make it jagged. The distractor field is generated with a fixed seed, so the first render is deterministic. The salience index is illustrative, not measured data.

Figure, Ground, and Illusory Contours

Every bounding contour is shared by two regions, and the visual system must decide which of them is the figure that owns the edge and which is the ground that merely abuts it. This assignment of border ownership is not dictated by the local image, since the same edge is compatible with either region owning it, yet the visual system commits to one interpretation and codes it early. Recording in monkey visual cortex, von der Heydt and colleagues found neurons that responded to a contour even when it was illusory, present in perception but absent from the image, as along the edges of a Kanizsa figure, and later work showed neurons signalling which side of an edge owns it (von der Heydt, Peterhans & Baumgartner, 1984). That cortical cells fire to a contour with no local luminance step demonstrates that form perception constructs boundaries rather than merely detecting them.

Illusory contours make the constructive character of figure-ground assignment vivid. When three disc-shaped inducers are each cut by a wedge and the wedges are aligned, observers see a bright triangle with crisp edges lying over three complete discs, although no edge is physically present between the inducers. The strength of the illusory edge grows with the support ratio, the fraction of each side that is physically specified by the inducers rather than interpolated across the gap, and the percept vanishes when the inducers are rotated so that their edges no longer align. The third demonstration lets a reader vary both the support ratio and the alignment of the inducers and watch the illusory figure appear and disappear, and the Worked Example gives the geometry of the support ratio it reports.

Constructing Boundaries

Kanizsa Illusory Contours and the Support Ratio

When three wedge-cut discs are aligned, a bright triangle with sharp edges appears to lie over them, though no edge is physically drawn between the discs. The illusory edge grows sharper as the inducers specify more of each side, the support ratio, and it disappears when the inducers are rotated so their mouths no longer face one another. Adjust the inducer size and their rotation.

Inducer radius (r)40
Inducer rotation (φ)0°
With an inducer radius of 40 on a side of length 200, the support ratio is 40/200 = 0.40. The mouths are aligned but weakly supported, so the illusory edge is faint.
An illustrative model of a Kanizsa figure (after von der Heydt et al., 1984). Three wedge-cut inducers evoke a bright triangle whose illusory edges strengthen with the support ratio s = 2r / L. Rotating the inducers so their mouths no longer align abolishes the percept. The geometry reproduces the Worked Example with L = 200. Values are computed locally, not stored.

Gestalt Principles of Perceptual Organization

The question of how local elements are organised into perceptual wholes is the oldest in the study of form, and the Gestalt psychologists gave the first systematic answer. Wertheimer set out the laws of perceptual organisation, the tendencies by which the visual field is grouped: elements close together, similar to one another, moving together, or arranged along a smooth continuous line are seen as belonging to the same figure, and the whole so formed has properties not present in any of its parts (Wertheimer, 1923). Grouping by proximity, similarity, common fate, good continuation, and closure is not a catalogue of curiosities but a description of the default parsing the visual system imposes whenever it segments a scene into objects.

Modern work has both formalised these principles and revised their foundations. Palmer and Rock argued that the classical laws presuppose a still more basic step, the formation of connected uniform regions, and proposed uniform connectedness, the grouping of any region of homogeneous luminance, colour, or texture into a single unit, as the entry-level operation on which proximity and similarity then act (Palmer & Rock, 1994). A century after Wertheimer, a comprehensive review confirmed that perceptual grouping and figure-ground organisation remain central to form perception and that the Gestalt agenda has been absorbed into quantitative vision science rather than superseded by it (Wagemans et al., 2012). Grouping settles what belongs to an object; the next problem is how the object's three-dimensional shape is described.

Structural Descriptions of Shape

Recognising an object requires a description of its shape that survives changes in viewpoint, and an influential tradition holds that this description is structural, specifying the object's parts and their spatial relations rather than a raw image. Marr and Nishihara proposed that shape is represented in an object-centred coordinate frame as a hierarchy of generalised cones, volumetric primitives whose axes and relations remain constant as the object is seen from different angles, so that recognition can proceed by matching this stable description rather than a viewpoint-dependent picture (Marr & Nishihara, 1978). The appeal of a structural description is that it factors the enormous variability of the retinal image into a small set of parts and relations that do not vary with pose.

Biederman developed this idea into a concrete account of everyday recognition. His recognition-by-components theory proposes that objects are parsed at regions of sharp concavity into a small alphabet of simple volumes he called geons, and that an object is recognised from the identities of its geons and their arrangement, which is why a few well-chosen parts often suffice and why degrading the vertices where parts meet is far more damaging than degrading the parts' midsegments (Biederman, 1987). On this view form perception delivers a part-based structural description that is largely invariant to viewpoint, size, and position. Whether the brain in fact recovers such explicit parts, or something more like a graded similarity to stored views, is a question the neural evidence bears on directly.

Shape in the Ventral Stream

The cortical machinery for object shape lies in the ventral visual pathway, running from the striate cortex into the temporal lobe, and human neuroimaging has localised a key stage of it. Grill-Spector, Kourtzi, and Kanwisher characterised the lateral occipital complex, a region that responds far more strongly to images of objects and coherent shapes than to scrambled controls or textures, and whose activity tracks whether an object is perceived rather than the low-level features of the image (Grill-Spector, Kourtzi & Kanwisher, 2001). Its responses generalise across changes in size and position and, to a degree, across the cues that define a shape, which marks it as a stage where form is coded more abstractly than in early cortex.

Two further findings frame how this region computes shape. Kourtzi and Kanwisher showed that the lateral occipital complex represents the perceived shape of an object rather than its physical contours, responding similarly to shapes that look alike even when their images differ and differently to identical images that are perceived as different shapes (Kourtzi & Kanwisher, 2001). On the computational side, Riesenhuber and Poggio proposed a hierarchical model in which alternating layers build selectivity for complex features and pool over position and scale to gain invariance, a design that reproduces the tuning and tolerance of ventral-stream neurons and that anticipated the architecture of the deep networks now dominant in machine vision (Riesenhuber & Poggio, 1999). The pathway thus climbs from local edges to invariant, perceived shape, but it must also cope with the everyday fact that objects are rarely seen whole.

Completion and Occluded Objects

In a cluttered world most objects are partly hidden, their contours interrupted where a nearer surface passes in front, yet observers perceive them as complete and continuous rather than as the fragments the retina actually receives. This filling in across a gap, amodal completion, is a routine achievement of form perception, and it is present very early in life. Kellman and Spelke showed that four-month-old infants treat two visible ends of a rod that moves together behind an occluder as a single connected object, revealing that the perception of partly occluded objects, and the interpolation of the hidden contour, does not wait on extensive visual learning (Kellman & Spelke, 1983).

Completion connects to the earlier stages of form perception in a direct way. The same bias toward smooth, co-circular contours that links visible edge elements into a path also governs how the visual system interpolates a boundary across an occluding gap or an illusory figure, so contour integration, illusory contours, and amodal completion can be seen as expressions of a single interpolation process operating under different conditions of visibility. Whether the completed contour is represented as sharply as a visible one, and where in the visual pathway the interpolation occurs, remain active questions, but the phenomenon establishes that form perception delivers whole objects, not the punctate sample the image provides.

Criticisms and Open Questions

Several strong claims in this area have been qualified, and the sharpest current pressure comes from comparisons with deep neural networks. Because such networks now match human accuracy on many object-recognition benchmarks, they were initially read as models of the ventral stream, and to a first approximation their intermediate representations do predict shape-sensitive responses (Kubilius, Bracci & Op de Beeck, 2016). But closer tests show a dissociation: standard networks often classify objects from local texture and small diagnostic features rather than from the global configuration of parts, so they can be highly accurate while failing on the very stimuli, such as silhouettes and outlines, that isolate global shape (Baker, Lu, Erlikhman & Kellman, 2018). This texture bias is a general property of the architecture rather than a quirk of one benchmark: Geirhos and colleagues found that networks trained on natural images label a cat's silhouette filled with elephant skin an elephant, following the texture against the shape, and that training the same networks toward shape instead improves both their robustness and their agreement with human judgements (Geirhos et al., 2019). This is precisely the aspect of form that structural-description theories placed at the centre.

Two further lines of evidence widen the gap and qualify the human account itself. Doerig and colleagues showed that visual crowding, in which clutter disrupts the recognition of a target, reveals fundamentally different local-versus-global processing in humans and machines, with human performance depending on global grouping that feedforward networks do not reproduce (Doerig et al., 2020). And on the human side, Jagadeesh and Gardner found that responses in high-level visual cortex to natural objects are surprisingly well captured by a texture-like statistical representation that discards much of the precise spatial arrangement of parts, suggesting that even the ventral stream may rely more on image statistics and less on explicit structural descriptions than the part-based tradition assumed (Jagadeesh & Gardner, 2022). The mature view treats form perception not as the delivery of a single structural model of an object but as a set of processes, some part-based and some statistical, whose balance is still being measured.

Worked Example

The first demonstration models the orientation tuning of a simple cell in the primary visual cortex, following the selectivity Hubel and Wiesel described. Let the cell's preferred orientation be zero degrees and let a bar be presented at orientation theta, so the misalignment is the angle theta itself. An idealised normalised tuning curve is the squared cosine of the misalignment, response equals cosine squared of theta, which is one when the bar is aligned and zero when it is orthogonal. At theta equal to zero the response is cosine squared of zero, which is one, the maximum. At thirty degrees it is cosine squared of thirty degrees; the cosine of thirty degrees is 0.8660, whose square is 0.75. At forty-five degrees the cosine is 0.7071 and its square is 0.50, so the response has fallen to half its peak, which makes forty-five degrees the half-width at half-maximum of this idealised curve. At sixty degrees the cosine is 0.5000 and its square is 0.25, and at ninety degrees the cosine is zero and the response is zero. Equal steps of orientation therefore produce unequal drops in response, steep near the flanks and shallow at the peak, which is the signature shape of an orientation-tuned cell.

The third demonstration models the support ratio of a Kanizsa illusory figure, the quantity that governs how strong the illusory edge appears. For one side of the illusory triangle of length L, each of the two inducers at its ends contributes a straight physically present edge of length equal to the inducer radius r, so the physically specified length along that side is two r and the remaining length L minus two r is interpolated across the gap. The support ratio is the physically specified fraction, two r divided by L. With a side of two hundred units and an inducer radius of forty, the support ratio is eighty divided by two hundred, which is 0.40. Increasing the radius to sixty raises it to one hundred and twenty divided by two hundred, which is 0.60, and a radius of one hundred gives two hundred divided by two hundred, which is 1.0, at which point the inducers meet and the triangle is fully drawn rather than illusory. The illusory contour is perceived as sharpest when the support ratio is high and fades as it falls toward zero, which is why the demonstration weakens the percept as the inducers are made small.

Discussion

Form perception is best understood not as a single achievement but as a graded sequence of inferences that turn a map of local intensities into a description of whole objects. Oriented edge detectors recode the image into local contour orientations; an association field links those orientations into continuous, smoothly curving contours; a decision about border ownership assigns each contour to a figure, sometimes constructing an edge where none exists; Gestalt grouping organises the parts into a shape; the ventral pathway codes that shape in a form tolerant of changes in size and position; and interpolation completes the shape across the surfaces that hide it. Each stage resolves an ambiguity the image leaves open, and each does so by assuming the regularities that hold in ordinary scenes. Table 1 sets the principal stages side by side with what each recovers and the evidence that most sharply characterises it.

Table 1. The principal stages of form perception compared across what each recovers and its signature evidence.
Stage What it recovers Signature evidence
Oriented edge detection The orientation of a local contour at each position Orientation-tuned simple cells in the primary visual cortex
Contour integration Continuous boundaries linked from aligned edge elements Detection of a smooth path in noise, the association field
Figure and ground Which side of a contour owns the edge, including illusory edges Cortical responses to illusory contours and border ownership
Perceptual organization The grouping of parts into a coherent whole Gestalt grouping and uniform connectedness
Invariant object shape Shape stable across size, position, and viewpoint Shape selectivity of the lateral occipital complex

Note. The stages share the goal of recovering object shape while differing in the scale, from a local edge to a whole object, at which each operates.

Read this way, form perception connects the physiology of a single cortical cell to the recognition of a whole object within one inferential story. What the visual system delivers is not a copy of the image but a description of the objects that most plausibly produced it, built by assuming that boundaries are smooth, that figures own their edges, that similar nearby elements cohere, and that hidden contours continue behind what occludes them. The current debate over whether that description is fundamentally part-based or statistical, sharpened by deep networks that recognise objects without recovering their global shape, does not overturn this picture so much as ask which of its inferences the brain actually performs, and in what form it stores the answer.

Glossary

Amodal completion.
The perception of a partly occluded object as complete and continuous behind the surface that hides it, with the missing contour interpolated across the gap.
Association field.
The pattern of preferential linking among oriented edge elements by which the visual system groups those that lie along a smooth, co-circular curve into a single contour.
Border ownership.
The assignment of a bounding contour to one of the two regions it separates, specifying which region is the figure that owns the edge and which is the ground.
Contour integration.
The linking of separate, aligned edge elements into a single continuous boundary, favouring paths that curve smoothly.
Figure-ground organization.
The segregation of a visual scene into a figure, seen as a bounded object in front, and a ground, seen as an unbounded surface behind it.
Geon.
In recognition-by-components, one of a small set of simple volumetric primitives, such as bricks, cylinders, and cones, from whose arrangement an object's shape is described.
Gestalt psychology.
The school of perception, founded on Wertheimer's work, holding that the visual field is organised into wholes with properties not reducible to their parts.
Illusory contour.
A perceived edge, such as the boundary of a Kanizsa figure, that is seen clearly although no corresponding luminance change is present in the image.
Lateral occipital complex.
A region of human ventral visual cortex that responds selectively to objects and coherent shapes, tracking perceived shape across changes in size and position.
Orientation tuning.
The selectivity of a visual neuron for the angle of an edge, firing maximally at a preferred orientation and less as the edge is rotated away from it.
Perceptual organization.
The processes that group the elements of a visual scene into figures and objects, described by the Gestalt principles of grouping.
Receptive field.
The region of the visual field within which a stimulus alters a neuron's firing, together with the stimulus features to which the neuron is tuned.
Recognition-by-components.
Biederman's theory that objects are recognised from the identities and arrangement of a small alphabet of volumetric parts, the geons.
Simple cell.
A neuron of the primary visual cortex with an oriented receptive field of adjacent excitatory and inhibitory zones, responding best to an edge at its preferred orientation and position.
Structural description.
A representation of an object's shape as a set of parts and the spatial relations among them, coded so as to remain stable across changes in viewpoint.
Support ratio.
In an illusory figure, the fraction of a contour that is physically specified by the inducers rather than interpolated across a gap, predicting how sharp the illusory edge appears.
Uniform connectedness.
Palmer and Rock's proposed entry-level grouping operation, by which a connected region of homogeneous luminance, colour, or texture is treated as a single perceptual unit.
Ventral stream.
The cortical visual pathway running from the striate cortex into the temporal lobe, specialised for the recognition of objects and their shapes.

Key Researchers

Irving Biederman (1939-2022). Formerly Harold Dornsife Professor of Neuroscience at the University of Southern California; proposed recognition-by-components, the geon theory of object recognition. ORCID - Wikipedia - Google Scholar

David J. Field. Professor of Psychology at Cornell University; co-authored the association-field account of contour integration and the theory of sparse coding of natural images. Faculty Page - Google Scholar

Kalanit Grill-Spector. Professor of Psychology at Stanford University; characterised the lateral occipital complex and its role in object recognition. ORCID - Faculty Page - Google Scholar - Wikipedia

Robert F. Hess. Professor of Ophthalmology at McGill University and director of McGill Vision Research; co-authored the association-field studies of contour integration and works on amblyopia and binocular vision. Faculty Page - Google Scholar

Rudiger von der Heydt. Professor Emeritus at the Zanvyl Krieger Mind/Brain Institute, Johns Hopkins University; discovered cortical responses to illusory contours and the neural coding of border ownership. ORCID - Faculty Page

David H. Hubel (1926-2013). Professor at Harvard Medical School; with Torsten Wiesel discovered orientation-selective receptive fields in the visual cortex, sharing the 1981 Nobel Prize in Physiology or Medicine. Wikipedia - Nobel Prize

Nancy Kanwisher (b. 1958). Walter A. Rosenblith Professor of Cognitive Neuroscience at the Massachusetts Institute of Technology; mapped category-selective regions of the ventral stream and the shape selectivity of the lateral occipital complex. ORCID - Faculty Page - Google Scholar - Wikipedia

Philip J. Kellman. Distinguished Professor of Psychology at the University of California, Los Angeles; studied contour interpolation, object completion, and the perception of partly occluded objects in infancy. Faculty Page - Google Scholar - Wikipedia

Tomaso A. Poggio (b. 1947). Eugene McDermott Professor at the Massachusetts Institute of Technology and a founder of the Center for Brains, Minds and Machines; proposed hierarchical models of object recognition in cortex. ORCID - Faculty Page - Google Scholar - Wikipedia

Johan Wagemans (b. 1963). Professor of Experimental Psychology at KU Leuven; led the modern reassessment of Gestalt perceptual organisation a century after Wertheimer. ORCID - Faculty Page - Google Scholar - Wikipedia

Max Wertheimer (1880-1943). Founder of Gestalt psychology, at the University of Frankfurt and later the New School for Social Research; set out the laws of perceptual organisation. Wikipedia

Torsten N. Wiesel (b. 1924). President Emeritus of the Rockefeller University; with David Hubel discovered orientation selectivity and the functional architecture of the visual cortex, sharing the 1981 Nobel Prize. Faculty Page - Wikipedia

Frequently Asked Questions

What is form perception?
It is the set of processes by which the visual system recovers the shape, contours, and boundaries of an object from the pattern of light on the retina, in which shape is not directly marked (Wertheimer, 1923).

Why is recovering shape a hard problem?
Because the retina registers only local intensities, so the boundary of an object appears as scattered luminance changes that must be detected, linked into continuous contours, and assigned to one surface rather than another before any shape is available (Wertheimer, 1923).

How does the brain detect the edges of objects?
Neurons in the primary visual cortex have oriented receptive fields and fire selectively to an edge at a particular orientation and location, recoding the image into a map of local contour orientations (Hubel & Wiesel, 1962).

How are separate edge fragments joined into a contour?
By an association field that links oriented elements preferentially when they lie along a smooth, co-circular curve, so a gently curving path stands out from randomly oriented distractors while a jagged one does not (Field, Hayes & Hess, 1993).

What are illusory contours?
They are edges that are perceived clearly although no luminance change is present, as along a Kanizsa figure, and cortical neurons respond to them, which shows that the visual system constructs boundaries rather than only detecting them (von der Heydt, Peterhans & Baumgartner, 1984).

What are the Gestalt principles of organization?
They are the tendencies, set out by Wertheimer, to group elements that are close together, similar, moving together, or arranged along a smooth line into a single figure whose properties exceed those of its parts (Wertheimer, 1923).

Where in the brain is object shape represented?
In the ventral visual pathway, and especially the lateral occipital complex, which responds selectively to objects and coherent shapes and tracks perceived shape across changes in size and position (Grill-Spector, Kourtzi & Kanwisher, 2001).

Do deep neural networks see shape the way people do?
Not fully; although they match human accuracy on many benchmarks, standard networks often classify by local texture rather than global shape, and they diverge from human perception on stimuli such as silhouettes and crowded displays (Baker, Lu, Erlikhman & Kellman, 2018).

References

Baker, N., Lu, H., Erlikhman, G., & Kellman, P. J. (2018). Deep convolutional networks do not classify based on global object shape. PLOS Computational Biology, 14(12), e1006613. https://doi.org/10.1371/journal.pcbi.1006613

Biederman, I. (1987). Recognition-by-components: A theory of human image understanding. Psychological Review, 94(2), 115-147. https://doi.org/10.1037/0033-295X.94.2.115

Doerig, A., Bornet, A., Choung, O. H., & Herzog, M. H. (2020). Crowding reveals fundamental differences in local versus global processing in humans and machines. Vision Research, 167, 39-45. https://doi.org/10.1016/j.visres.2019.12.006

Field, D. J., Hayes, A., & Hess, R. F. (1993). Contour integration by the human visual system: Evidence for a local association field. Vision Research, 33(2), 173-193. https://doi.org/10.1016/0042-6989(93)90156-Q

Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., & Brendel, W. (2019). ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In International Conference on Learning Representations (ICLR). https://doi.org/10.48550/arXiv.1811.12231

Grill-Spector, K., Kourtzi, Z., & Kanwisher, N. (2001). The lateral occipital complex and its role in object recognition. Vision Research, 41(10-11), 1409-1422. https://doi.org/10.1016/S0042-6989(01)00073-6

Hubel, D. H., & Wiesel, T. N. (1962). Receptive fields, binocular interaction and functional architecture in the cat's visual cortex. The Journal of Physiology, 160(1), 106-154. https://doi.org/10.1113/jphysiol.1962.sp006837

Jagadeesh, A. V., & Gardner, J. L. (2022). Texture-like representation of objects in human visual cortex. Proceedings of the National Academy of Sciences, 119(17), e2115302119. https://doi.org/10.1073/pnas.2115302119

Kellman, P. J., & Spelke, E. S. (1983). Perception of partly occluded objects in infancy. Cognitive Psychology, 15(4), 483-524. https://doi.org/10.1016/0010-0285(83)90017-8

Kourtzi, Z., & Kanwisher, N. (2001). Representation of perceived object shape by the human lateral occipital complex. Science, 293(5534), 1506-1509. https://doi.org/10.1126/science.1061133

Kubilius, J., Bracci, S., & Op de Beeck, H. P. (2016). Deep neural networks as a computational model for human shape sensitivity. PLOS Computational Biology, 12(4), e1004896. https://doi.org/10.1371/journal.pcbi.1004896

Marr, D., & Nishihara, H. K. (1978). Representation and recognition of the spatial organization of three-dimensional shapes. Proceedings of the Royal Society of London. Series B, Biological Sciences, 200(1140), 269-294. https://doi.org/10.1098/rspb.1978.0020

Palmer, S., & Rock, I. (1994). Rethinking perceptual organization: The role of uniform connectedness. Psychonomic Bulletin & Review, 1(1), 29-55. https://doi.org/10.3758/BF03200760

Riesenhuber, M., & Poggio, T. (1999). Hierarchical models of object recognition in cortex. Nature Neuroscience, 2(11), 1019-1025. https://doi.org/10.1038/14819

von der Heydt, R., Peterhans, E., & Baumgartner, G. (1984). Illusory contours and cortical neuron responses. Science, 224(4654), 1260-1262. https://doi.org/10.1126/science.6539501

Wagemans, J., Elder, J. H., Kubovy, M., Palmer, S. E., Peterson, M. A., Singh, M., & von der Heydt, R. (2012). A century of Gestalt psychology in visual perception: I. Perceptual grouping and figure-ground organization. Psychological Bulletin, 138(6), 1172-1217. https://doi.org/10.1037/a0029333

Wertheimer, M. (1923). Untersuchungen zur Lehre von der Gestalt. II. Psychologische Forschung, 4, 301-350. https://doi.org/10.1007/BF00410640