Abstract
Monocular vision is a form of visual perception: seeing with a single eye, or with the information available to one eye at a time. Although stereopsis needs two eyes, a single eye already receives a rich store of depth information, the monocular cues, from which the visual system recovers the three-dimensional layout of a scene. Pictorial cues such as occlusion, relative size, perspective, texture gradients, and shading specify depth in a static image; motion parallax and the kinetic depth effect add depth from a changing image; and defocus blur signals distance from the plane of focus. This article covers the pictorial cues, motion-based cues, shape from shading and cast shadows, how the visual system combines cues into one estimate, and monocular stereopsis, the vivid depth seen in a picture with one eye, with interactive demonstrations.
Keywords: monocular vision, monocular depth cues, motion parallax
Monocular vision is vision that uses the input of a single eye. It is the everyday condition of anyone who has lost an eye or covers one, and it is also the analytic starting point for depth perception, because most of the cues the visual system uses are available to each eye on its own. James Gibson made this the centre of his ecological account: a single moving eye is bathed in an ambient optic array whose structure, the way texture, edges, and surfaces project and transform, already specifies the layout of the world, so that depth need not be inferred from impoverished flat images but can be picked up from information the eye actually receives (Gibson, 2014). The catalogue of monocular cues is long, and each is reliable within its own range of distances, so that together they cover the whole span from the hand to the horizon (Cutting & Vishton, 1995). Monocular vision matters to cognitive psychology because it shows that binocular disparity, striking though it is, is one cue among many, and that the perception of a solid, three-dimensional world survives the loss of an eye largely intact (Linton, 2017). Direct evidence bears this out: patients with only one functioning eye judge large egocentric distances as accurately and as precisely as two-eyed observers, using the optic flow their own head movements generate to supply what the missing eye would have (Gao et al., 2023).
- Monocular vision uses one eye, and a single eye already carries a rich set of depth cues, so good depth perception does not require two eyes.
- Pictorial cues, occlusion, relative size, linear perspective, texture gradients, aerial perspective, and contrast, specify depth in a single static image.
- Motion parallax and the kinetic depth effect recover depth from the changing image produced by a moving observer or a moving object, and are geometrically as powerful as binocular disparity.
- Shading and cast shadows specify shape and depth under prior assumptions, such as that light comes from above.
- The visual system weights and combines cues into a single estimate, and a single eye viewing a picture can yield a vivid impression of depth, monocular stereopsis.
Pictorial Depth Cues
The pictorial cues are the depth information contained in a single static image, so called because painters have exploited them for centuries to make a flat canvas read as deep space (Table 1). When one surface hides part of another, occlusion, or interposition, gives an unambiguous order in depth: the covered object is farther, though the cue says nothing about how much farther. Relative size makes the smaller of two equal objects appear more distant; linear perspective converges the parallel edges of a receding surface toward a vanishing point; and a texture gradient, the progressive fining and compression of surface detail with distance, specifies the slant and recession of a surface directly, a cue Gibson placed at the heart of his account of surface perception (Gibson, 2014). Over large outdoor distances aerial perspective adds its own signal, as scattering by the atmosphere makes far objects fainter, bluer, and lower in contrast, and reduced contrast alone is sufficient to shift an object's apparent distance (O'Shea et al., 1994).
These cues differ not only in what they specify but in where they are useful, and their relative potency is orderly rather than arbitrary. Cutting and Vishton analysed the depth cues as a set of ranked, distance-dependent sources of information, showing that occlusion dominates at every distance while others, such as relative size, motion parallax, and binocular disparity, trade places in effectiveness as one moves from personal space near the body out to the horizon (Cutting & Vishton, 1995). What the pictorial cues ultimately support is the perception of three-dimensional shape, the layout of surfaces and their curvature, which the visual system recovers from these image regularities with considerable but imperfect fidelity (Todd, 2004).
| Cue | Type | What it specifies |
|---|---|---|
| Occlusion | Pictorial | Depth order only; the covered surface is farther, at any distance |
| Relative size and texture gradient | Pictorial | Relative distance and surface slant from the scaling of image detail |
| Linear and aerial perspective | Pictorial | Recession toward a vanishing point; greater distance from lost contrast and colour |
| Shading and cast shadows | Pictorial | Local surface shape and an object's height above a surface, under a light-from-above prior |
| Motion parallax and kinetic depth | Motion | Relative depth and three-dimensional structure from relative image motion |
| Defocus blur | Optical | Distance from the plane of focus, with the sign resolved by other cues |
Note. Every cue in the table is available to a single eye. Occlusion gives only depth order; the others scale with distance and so must be combined to yield a metric estimate.
Pictorial cues place an object in depth
Distance 3 of 10. Relative size makes the ball 28 px across; linear perspective and the texture gradient set its height in the scene; and it sits nearer than the post, which does not cover it.
Motion Parallax and Structure From Motion
A single eye that moves gains a powerful new source of depth. When an observer translates the head sideways, the images of near objects sweep across the retina faster than those of far objects, and this differential image motion, motion parallax, specifies relative depth in the same way that the two eyes' differing viewpoints specify disparity. Brian Rogers and Maureen Graham demonstrated that motion parallax is an independent cue: presenting observers with a moving random-dot display whose only depth information was the parallax linked to head movement, with no disparity, no perspective, and no shading, they found that people saw a clear corrugated surface in depth, proving that parallax alone is sufficient (Rogers & Graham, 1979). Because the effective baseline is the distance the head travels rather than the fixed separation of the eyes, motion parallax can in principle be made as sensitive as stereopsis, or more so.
Depth from motion is not limited to a moving observer. Hans Wallach and D. N. O'Connell showed that the shadow of a rotating wire form, a flat and meaningless outline when still, is instantly seen as a rigid three-dimensional object the moment it turns, the kinetic depth effect: the changing two-dimensional projection is enough for the visual system to recover the solid shape that would produce it (Wallach & O'Connell, 1953). The general problem of reconstructing three-dimensional structure from a sequence of changing images is structure from motion, and it is a monocular achievement, since one eye viewing the transforming image recovers the shape. Gibson placed exactly this kind of transforming optical structure, the flow of the optic array as an animal moves, at the foundation of perception, arguing that motion reveals the persistent layout of surfaces rather than obscuring it (Gibson, 2014).
Motion parallax: near and far objects move differently
Head centred: with no translation there is no parallax, and depth from this cue is unavailable until the head moves.
Shape From Shading and Cast Shadows
The gradation of light across a surface is a monocular cue to its shape, but one the visual system can read only by assuming where the light comes from. Vilayanur Ramachandran showed that observers perceive shape from shading rapidly and preattentively, and that they resolve the inherent ambiguity of a shaded patch, which could be a bump lit from one side or a dimple lit from the other, by assuming a single light source overhead: identical shaded disks are seen as convex when their bright side is up and as concave when it is down, and the percept flips when the image is inverted (Ramachandran, 1988). The light-from-above prior is a built-in assumption that turns an ambiguous image into a definite shape, an early and influential case of the visual system supplying a default where the data underdetermine the answer.
Cast shadows add a second, relational signal. The shadow an object throws onto a surface ties the object to that surface and specifies its height above it: Pascal Mamassian and colleagues showed that moving a shadow alone, with the object's retinal position unchanged, makes the object appear to move in depth and to lift off or settle onto the ground plane, and that the visual system reads shadows under a prior for a stationary, overhead light (Mamassian et al., 1998). Shadows are treated as a weak but genuine cue, readily overridden when they conflict with stronger information, which is why implausible or inconsistent shadows in a picture are often not consciously noticed even as they shape the perceived layout.
Shape from shading under a light-from-above prior
The visual system assumes light from above, so a disk brighter at the top is seen as a bump and one brighter at the bottom as a hollow. Here the effective light comes from above, so the disks read as convex bumps.
Combining Depth Cues
No single cue gives a complete metric description of depth, so the visual system must fuse the several estimates a scene provides into one. Michael Landy and colleagues framed this as a problem of cue combination and argued for modified weak fusion, in which each cue yields its own depth estimate and the estimates are averaged with weights proportional to their reliability, after a promotion step supplies the missing scale that cues such as motion parallax lack on their own (Landy et al., 1995). Marc Ernst and Martin Banks put this weighting to a direct test and found that observers combine cues, in their study visual and haptic estimates of an object's size, in a statistically optimal fashion, each cue weighted in inverse proportion to its variance so that the fused estimate is more reliable than either source alone (Ernst & Banks, 2002). Reliability weighting predicts that a cue counts for more when it is precise and for less when it is noisy or out of its useful range, which is why the ranking of cues shifts with viewing distance (Cutting & Vishton, 1995).
The optical blur of a defocused image is a case that shows how subtle the available cues can be. Robert Held, Emily Cooper, and Martin Banks demonstrated that retinal blur and binocular disparity are complementary cues to depth, blur carrying coarse distance information that helps scale the finer disparity signal (Held et al., 2012). That blur is informative at all reflects the structure of natural viewing: William Sprague and colleagues measured the pattern of blur across the visual field during everyday tasks and found that its statistics, jointly with the typical distances of surfaces, make defocus a systematic and usable signal for distance rather than mere degradation (Sprague et al., 2016). Blur is a purely monocular cue, present in every single-eye image, and its inclusion shows how thoroughly one eye's view is packed with depth information once the visual system knows how to read it (Figure 1).
Figure 1
Effective Range of Selected Depth Cues
Monocular Stereopsis
The richest demonstration that one eye suffices for solid depth is monocular stereopsis, the vivid, almost tangible three-dimensionality seen when a picture is viewed with a single eye through an aperture. Dhanraj Vishwanath and Paul Hibbard showed that observers report the same distinctive impression of real depth and object solidity from a two-dimensional image viewed monocularly through a small hole that they report from a binocular stereogram, even though the monocular image contains no disparity at all, evidence that this qualitative sense of depth is not the exclusive product of binocular stereopsis but can be triggered by the correct interpretation of pictorial cues alone (Vishwanath & Hibbard, 2013). The finding reframes stereopsis as a perceptual mode, an experience of scaled three-dimensional space, rather than a signal available only when two eyes cooperate.
The practical reach of monocular cues is clearest in virtual and augmented reality, where a display must conjure depth for a moving observer. Peter Scarfe and Andrew Glennerster reviewed how such displays exploit and sometimes misrepresent the cues the visual system expects, noting that pictorial and motion cues do much of the work of conveying a three-dimensional scene (Scarfe & Glennerster, 2019). Rachel Hornsey and Paul Hibbard measured the contributions directly, finding that pictorial cues carry a substantial share of perceived distance in a virtual environment alongside binocular cues, so that a convincing sense of depth does not depend on stereopsis alone (Hornsey & Hibbard, 2021).
Worked Example
Motion parallax obeys the same geometry as binocular disparity, and a short calculation shows why a single moving eye can rival two stationary ones. When an observer translates the head sideways by a distance b while fixating a point at distance z, a second object a small depth delta-z beyond the fixation point shifts its retinal position, relative to the fixation point, by an angle given to a good approximation by b times delta-z divided by the square of the distance: the parallax in radians is approximately b times delta-z over z squared. This is exactly the disparity formula, with the head's translation b playing the role that the interocular distance plays in stereopsis.
Fixate at z = 100 cm with a second object delta-z = 5 cm beyond it, and sway the head sideways by b = 10 cm. The parallax is 10 times 5 divided by 100 squared, which is 50 over 10,000, or 0.005 radians. Multiplying by 3,437.75 arc minutes per radian gives about 17.2 arc minutes, a large and easily seen shift. The decisive contrast with stereopsis is in the baseline. Two eyes are fixed about 6.5 cm apart, so the same depth step yields a binocular disparity of 6.5 times 5 over 10,000, or 0.00325 radians, about 11.2 arc minutes; a modest 10 cm head movement already exceeds it, and a wider sway of b = 20 cm produces 0.01 radians, roughly 34.4 arc minutes, three times the binocular signal. Because the observer chooses how far to move, the effective baseline for motion parallax is not capped at the width of the head, which is why a person who has lost an eye learns to rock the head slightly to read depth that stereopsis would otherwise have supplied (Rogers & Graham, 1979). The same inverse-square fall-off applies as for disparity: the parallax from a given head movement shrinks with the square of viewing distance, so motion parallax, like stereopsis, is precise up close and coarse across a room.
Discussion
Monocular vision reframes depth perception as the reading of many partly redundant cues rather than the operation of a single binocular mechanism. Gibson's insight that a single moving eye is immersed in structured information (Gibson, 2014) is borne out by the catalogue of cues that each eye receives: pictorial cues ordered by their potency across distance (Cutting & Vishton, 1995), motion parallax and the kinetic depth effect that recover structure from a changing image (Rogers & Graham, 1979; Wallach & O'Connell, 1953), shading and shadows read under built-in priors (Ramachandran, 1988; Mamassian et al., 1998), and even the defocus blur that ordinary vision might seem to waste (Held et al., 2012; Sprague et al., 2016).
The theme that unifies the modern work is that the visual system does not merely collect these cues but weights and combines them into a single estimate, promoting the ones that lack scale and trusting each in proportion to its reliability (Landy et al., 1995). That the product can be a full, solid percept from one eye is shown most directly by monocular stereopsis (Vishwanath & Hibbard, 2013) and by the effectiveness of pictorial and motion cues in virtual displays (Scarfe & Glennerster, 2019; Hornsey & Hibbard, 2021). The distinction worth keeping sharp is that binocular disparity adds precision and a compelling immediacy to near depth, but it is neither necessary nor sufficient for the perception of a three-dimensional world; the recovery of surface shape and layout is fundamentally a task the visual system can accomplish with a single eye (Todd, 2004; Linton, 2017).
Common Misconceptions
- A person with one eye sees the world as flat.
- Only stereopsis is lost with an eye; the many monocular cues, occlusion, perspective, motion parallax, shading, and blur, remain, and people with one eye judge depth well, especially once they use head movements to generate parallax (Cutting & Vishton, 1995).
- Monocular cues are learned tricks, weaker than real binocular depth.
- Motion parallax follows the same geometry as disparity and, because its baseline can exceed the width of the head, can match or surpass stereopsis in sensitivity (Rogers & Graham, 1979).
- A flat picture can never yield genuine three-dimensional depth.
- Viewing a picture with one eye through an aperture produces monocular stereopsis, a vivid impression of solid depth qualitatively like that of a binocular stereogram, from pictorial cues alone (Vishwanath & Hibbard, 2013).
Glossary
- Accommodation.
- The change in the lens's focal power that brings a target into focus; the effort involved is a weak monocular cue to the target's distance.
- Aerial perspective.
- The loss of contrast, detail, and colour saturation in distant objects caused by atmospheric scattering, which makes them appear farther away.
- Cast shadow.
- The shadow an object throws onto a surface, which ties the object to that surface and specifies its height above it.
- Defocus blur.
- The blur of an object away from the plane of focus, which grows with its distance from that plane and serves as a monocular cue to relative distance.
- Depth cue.
- Any source of information in the retinal image, or in its change over time, from which the visual system estimates distance or three-dimensional shape.
- Kinetic depth effect.
- The perception of a rigid three-dimensional shape from the changing two-dimensional projection of a rotating object, as in the shadow of a turning wire form.
- Light-from-above prior.
- The visual system's default assumption that illumination comes from overhead, used to resolve the convex or concave ambiguity of a shaded surface.
- Linear perspective.
- The convergence of parallel edges toward a vanishing point as a surface recedes, a pictorial cue to distance and surface orientation.
- Monocular cue.
- A depth cue available to a single eye, including all the pictorial, motion, and optical cues, as distinct from binocular cues that require two eyes.
- Monocular stereopsis.
- The vivid impression of solid three-dimensional depth obtained when a picture is viewed with a single eye through an aperture, from pictorial cues without disparity.
- Motion parallax.
- The differential image motion of near and far objects produced when the observer translates, specifying relative depth by the same geometry as binocular disparity.
- Occlusion.
- The partial hiding of one surface by another, also called interposition, giving an unambiguous depth order in which the covered surface is farther.
- Optic array.
- In Gibson's ecological optics, the structured pattern of light converging on a point of observation, whose structure and transformations specify the layout of surfaces.
- Optic flow.
- The pattern of image motion across the whole field of view produced by the observer's own movement, a rich monocular source of layout and heading information.
- Pictorial cue.
- A monocular depth cue present in a single static image, such as occlusion, relative size, perspective, texture gradient, aerial perspective, or shading.
- Relative size.
- The cue by which, among objects known or assumed to be the same physical size, the one projecting a smaller image is seen as more distant.
- Shape from shading.
- The recovery of a surface's local three-dimensional shape from the gradation of light across it, resolved by an assumption about the direction of illumination.
- Structure from motion.
- The reconstruction of three-dimensional shape from the changing image a moving object or moving observer produces, a monocular achievement.
- Texture gradient.
- The progressive fining and compression of surface detail with distance, which specifies the slant and recession of a surface directly.
Key Researchers
Martin S. Banks. Vision scientist at the University of California, Berkeley; he studies how the visual system combines depth cues and how retinal blur, the defocus that grows with an object's distance from the plane of fixation, serves as a monocular cue whose statistics match the natural environment. ORCID - Faculty Page - Google Scholar - Wikidata
James J. Gibson (1904-1979). Psychologist at Cornell University; he founded the ecological approach to visual perception, arguing that a single moving observer's ambient optic array, its texture gradients, optic flow, and occlusion, is rich in monocular information that specifies surface layout and depth without inference. Faculty Page - Wikipedia - Wikidata
Barbara Gillam. Emeritus professor at the University of New South Wales; she analysed how monocular cues, perspective, surface layout, and monocular occlusion configurations, specify depth and shape, showing that the visual system draws depth order and three-dimensional structure from single-image information alone. Faculty Page - Wikipedia - Wikidata
Pascal Mamassian. CNRS Director of Research at the Ecole Normale Superieure, Paris; he built Bayesian models of how the visual system reads pictorial depth cues, showing that shading and cast shadows are interpreted under prior assumptions, such as a single overhead light source, to recover shape and spatial layout. ORCID - Faculty Page - Google Scholar - Wikidata
Brian J. Rogers. Emeritus professor at the University of Oxford; he established motion parallax as an independent monocular cue to depth, showing with Maureen Graham that the relative image motion produced by a moving observer specifies three-dimensional structure on its own, without binocular disparity. ORCID - Faculty Page - Wikipedia - Wikidata
Dhanraj Vishwanath. Senior lecturer at the University of St Andrews; he showed that a single eye viewing a picture can produce a vivid impression of three-dimensional depth, monocular stereopsis, and developed theories of stereopsis and pictorial space that treat depth as an interpretation of the scene rather than a direct read-out of disparity. ORCID - Faculty Page
Frequently Asked Questions
What is monocular vision?
It is vision using a single eye, or the depth information available to one eye at a time; although it lacks stereopsis, a single eye still receives a rich set of monocular cues from which the visual system recovers a three-dimensional scene (Gibson, 2014).
What are monocular depth cues?
They are sources of depth information available to one eye, including the pictorial cues, occlusion, relative size, linear and aerial perspective, texture gradients, and shading, together with motion parallax and defocus blur (Cutting & Vishton, 1995).
Can someone see depth with one eye?
Yes; monocular cues support good depth perception, which is why losing an eye does not leave a person unable to judge distance, and a single eye viewing a picture can even yield a vivid impression of solid depth (Vishwanath & Hibbard, 2013).
What is motion parallax?
It is the differential image motion of near and far objects produced when the observer moves; it follows the same geometry as binocular disparity and by itself is sufficient to specify a three-dimensional surface (Rogers & Graham, 1979).
How does shading indicate depth?
The gradation of light across a surface specifies its shape once the visual system assumes a light direction; by default it assumes light from above, so a shaded patch bright at the top is seen as a bump and one bright at the bottom as a hollow (Ramachandran, 1988).
What is monocular stereopsis?
It is the vivid, scaled sense of three-dimensional depth obtained when a picture is viewed with one eye through an aperture, qualitatively like the depth of a binocular stereogram yet produced by pictorial cues without any disparity (Vishwanath & Hibbard, 2013).
Are monocular or binocular cues more important?
Neither dominates everywhere; occlusion and other monocular cues are potent at all distances while disparity is strongest up close, so the visual system weights each cue by its reliability at the current distance (Landy et al., 1995).
Do monocular cues work in virtual reality?
Yes; pictorial and motion cues carry a substantial share of the perceived depth in virtual environments, and displays rely heavily on them to convey a three-dimensional scene (Hornsey & Hibbard, 2021).
References
Cutting, J. E., & Vishton, P. M. (1995). Perceiving layout and knowing distances: The integration, relative potency, and contextual use of different information about depth. In W. Epstein & S. Rogers (Eds.), Perception of space and motion (pp. 69-117). Academic Press. https://doi.org/10.1016/b978-012240530-3/50005-5
Ernst, M. O., & Banks, M. S. (2002). Humans integrate visual and haptic information in a statistically optimal fashion. Nature, 415(6870), 429-433. https://doi.org/10.1038/415429a
Gao, L., Huang, Y., Zhang, Y., Zhang, X., Liu, Z., Pan, J. S., & Yu, M. (2023). Monocular information for perceiving large egocentric distance: A comparison between monocularly blind patients and normally sighted observers. Vision Research, 211, 108279. https://doi.org/10.1016/j.visres.2023.108279
Gibson, J. J. (2014). The ecological approach to visual perception (Classic ed.). Psychology Press. (Original work published 1979) https://doi.org/10.4324/9781315740218
Held, R. T., Cooper, E. A., & Banks, M. S. (2012). Blur and disparity are complementary cues to depth. Current Biology, 22(5), 426-431. https://doi.org/10.1016/j.cub.2012.01.033
Hornsey, R. L., & Hibbard, P. B. (2021). Contributions of pictorial and binocular cues to the perception of distance in virtual reality. Virtual Reality, 25(4), 1087-1103. https://doi.org/10.1007/s10055-021-00500-x
Landy, M. S., Maloney, L. T., Johnston, E. B., & Young, M. (1995). Measurement and modeling of depth cue combination: In defense of weak fusion. Vision Research, 35(3), 389-412. https://doi.org/10.1016/0042-6989(94)00176-M
Linton, P. (2017). The perception and cognition of visual space. Palgrave Macmillan. https://doi.org/10.1007/978-3-319-66293-0
Mamassian, P., Knill, D. C., & Kersten, D. (1998). The perception of cast shadows. Trends in Cognitive Sciences, 2(8), 288-295. https://doi.org/10.1016/S1364-6613(98)01204-2
O'Shea, R. P., Blackburn, S. G., & Ono, H. (1994). Contrast as a depth cue. Vision Research, 34(12), 1595-1604. https://doi.org/10.1016/0042-6989(94)90116-3
Ramachandran, V. S. (1988). Perception of shape from shading. Nature, 331(6152), 163-166. https://doi.org/10.1038/331163a0
Rogers, B., & Graham, M. (1979). Motion parallax as an independent cue for depth perception. Perception, 8(2), 125-134. https://doi.org/10.1068/p080125
Scarfe, P., & Glennerster, A. (2019). The science behind virtual reality displays. Annual Review of Vision Science, 5, 529-547. https://doi.org/10.1146/annurev-vision-091718-014942
Sprague, W. W., Cooper, E. A., Reissier, S., Yellapragada, B., & Banks, M. S. (2016). The natural statistics of blur. Journal of Vision, 16(10), 23. https://doi.org/10.1167/16.10.23
Todd, J. T. (2004). The visual perception of 3D shape. Trends in Cognitive Sciences, 8(3), 115-121. https://doi.org/10.1016/j.tics.2004.01.006
Vishwanath, D., & Hibbard, P. B. (2013). Seeing in 3-D with just one eye: Stereopsis without binocular vision. Psychological Science, 24(9), 1673-1685. https://doi.org/10.1177/0956797613477867
Wallach, H., & O'Connell, D. N. (1953). The kinetic depth effect. Journal of Experimental Psychology, 45(4), 205-217. https://doi.org/10.1037/h0056880