Skip to main content

Sensation & Perception V3: Chapter 7: Object Recognition

Sensation & Perception V3
Chapter 7: Object Recognition
  • Show the following:

    Annotations
    Resources
  • Adjust appearance:

    Font
    Font style
    Color Scheme
    Light
    Dark
    Annotation contrast
    Low
    High
    Margins
  • Search within:
    • My Notes + Comments
    • Notifications
    • Privacy
  • Project HomeSensation & Perception V3
  • Projects
  • Learn more about Manifold

Notes

table of contents
  1. Front Matter
  2. Preface
  3. Acknowledgements
  4. Chapter 1: Introduction to the Study of Sensation and Perception
  5. Chapter 2: Approaches to Studying Sensation and Perception
  6. Chapter 3: Receptors and Neural Processing
  7. Chapter 4: The Lateral Geniculate Nucleus (LGN) and Primary Visual Cortex (V1)
  8. Chapter 5: Higher-Level Visual Processing: Beyond V1
  9. Chapter 6: Attention and Visual Perception
  10. Chapter 7: Object Recognition
  11. Chapter 8: Color Vision
  12. Chapter 9: Depth Perception
  13. Chapter 10: Motion
  14. Chapter 11: Audition
  15. Chapter 12: Cutaneous Senses
  16. Chapter 13: Gustatory Senses
  17. Chapter 14: Olfaction
  18. Version History

Chapter 7: Object Recognition

Introduction

Object recognition is essential for identifying and categorizing the various objects and entities that populate our environment. In this chapter, we will explore how our brains process and recognize objects, touching upon key concepts and principles that help shape our perceptual experiences.

Gestalt Principles: Making Sense of Visual Information

The Birth of Gestalt Psychology

Max Wertheimer is credited with developing the Gestalt view of psychology. As the story goes, Wertheimer stumbled upon these groundbreaking ideas during a train journey in 1911. His encounter with a toy stroboscope, featuring alternating flashing lights, sparked a revelation. Instead of perceiving the lights as either stationary and flashing, or even as moving back and forth, Wertheimer perceived an invisible background element covering and uncovering the lights, creating the illusion of motion. This experience challenged the prevailing notion of structuralism, which asserted that sensations led to perceptions. In this case an invisible background object was perceived as moving in a manner to cover and uncover the lights. Wertheimer's observation hinted at something deeper - that our perceptual experience transcends sensory input (Wertheimer, 1912; Wagemans et al., 2012).

Gestalt Grouping Principles

Following this Gestalt psychologists sought to create rules or laws that would help to explain visual perception. Gestalt psychology introduced a series of principles that illuminate how we perceive objects as more than the sum of their individual parts (see Figure 7.1 for examples) (Wertheimer, 1912; Wagemans et al., 2012). These principles offer insights into how our brains naturally organize visual information:

  1. Proximity

Proximity suggests that we tend to group objects that are physically close to each other. For example, when presented with a set of dots, our brains instinctively group nearby dots, forming perceptual clusters. This principle often influences our perception of spatial relationships in scenes.

  1. Similarity

The principle of similarity highlights our tendency to group objects that share common attributes, such as shape, color, or size. When faced with a collection of elements, we naturally group those that exhibit similar characteristics, making it easier to identify patterns and objects within a scene.

  1. Good Continuation

Good continuation dictates that we prefer to perceive continuous and uninterrupted forms or patterns. When lines or contours intersect, we tend to follow the smoothest and most continuous path. This principle helps us distinguish objects from their backgrounds and identify cohesive shapes.

  1. Closure

Closure refers to our inclination to mentally complete or "close" incomplete shapes or forms. Even when presented with partial information, our brains strive to perceive whole objects. This phenomenon allows us to recognize objects even when some parts are obscured or missing.

  1. Common Region

Common region suggests that we group elements that appear within the same spatial region or boundary as belonging together. When objects share a common space, we perceive them as belonging to the same group, despite potential differences in their individual attributes (Palmer, 1992; Palmer, 2002).

  1. Connectedness

Connectedness occurs when objects are connected. When objects are physically connected by lines or other visual cues, we tend to group them together. This principle helps us perceive relationships and interactions between elements in a scene (Palmer & Rock, 1994).

  1. Figure-Ground Distinction

The figure–ground distinction is a fundamental aspect of object recognition. As we look at the world around us, our visual system continually separates the object of interest (the figure) from its surrounding environment (the ground). This process allows us to direct our attention to meaningful objects and perceive them as distinct from the background. Camouflage disrupts figure–ground organization by making the figure blend into its surroundings, making it more difficult to distinguish the object from the background (Figure 7.2) (Peterson & Gibson, 1994).

Figure 7.1

Six-panel figure illustrating six Gestalt principles of perceptual grouping. The panels are arranged in a 3×2 grid with large numbered titles.

Panel 1 – Proximity: Three clusters of dots are separated by wider spaces, demonstrating that nearby elements are perceived as belonging together. 
Panel 2 – Similarity of Shape: Columns of green circles, green squares, and green triangles show that objects with similar shapes are grouped together despite equal spacing. 
Panel 3 – Good Continuation: Two smooth black curves cross at the center, illustrating that people perceive continuous intersecting lines rather than four separate line segments.
Panel 4 – Closure: Four black Pac-Man-like circular shapes are positioned at the corners of an invisible square, causing viewers to perceive a white square that is not actually drawn. 
Panel 5 – Common Region: Red circles are enclosed within two rounded rectangles, demonstrating that elements inside the same boundary are perceived as a group. 
Panel 6 – Connectedness: Colored circles are linked by black lines into three distinct connected networks, illustrating that physically connected elements are perceived as belonging together.

Examples of six Gestalt principles of perceptual grouping. From left to right, top to bottom: proximity, similarity of shape, good continuation, closure, common region, and connectedness. Each panel illustrates how the visual system organizes individual elements into coherent groups based on these grouping cues. Panel 2 shows the principle of similarity of shape but this grouping variable could also be based on similarity of color.

"Gestalt grouping principles." by Kahan, T.A. is licensed under CC BY-NC-SA 4.0

White Arctic fox sitting in snow with scattered rocks. Because the fox's white fur closely matches the snowy background, its body blends into the environment, making it more difficult to separate the figure (the fox) from the ground (the snow). The image illustrates how camouflage can interfere with figure–ground organization.

Figure 7.2

The Arctic fox's white winter coat closely matches the snowy environment, reducing contrast between the figure and the background. As a result, the visual system has more difficulty separating the fox (the figure) from the surrounding snow (the ground), illustrating how camouflage interferes with figure–ground organization.

"Camouflage and figure–ground organization." by Kahan, T.A. is licensed under CC BY-NC-SA 4.0

Shading and Lighting: Perceiving Object Shape

Our perception of object shape and orientation is heavily influenced by shading and lighting cues. In our daily lives, we unconsciously assume that light comes from above, a bias likely shaped by our experience on Earth where sunlight illuminates objects from overhead. This assumption profoundly affects how we perceive the three-dimensional characteristics of objects.

Shading and Shape

Shading provides crucial information about an object's shape and depth. When an object is lit from above, it typically appears as if its upper surface is illuminated, while the lower parts remain in shadow. Ramachandran has found that this creates the impression of a raised or bulging form when a circle is lighter at the top and dark at the bottom. Conversely, when a circle is lighter at the bottom and darker at the top, it appears as though the upper surface is in shadow, creating the illusion of a depression or dimple (Ramachandran, 1988).

Figure 7.3

People have a bias to see the light source from above which creates the perception of a bump when the top portion of a circle is lighter and a dimple when the bottom portion is lighter.

"Perception of circle." by Kahan, T.A. is licensed under CC BY-NC-SA 4.0

Real-World Application

Humans typically assume that light comes from above when interpreting shading, allowing us to perceive three-dimensional shape from two-dimensional images. Designers have long recognized the importance of such perceptual assumptions in environments where orientation is difficult. For example, Soviet space stations used color-coded interiors and shading to reinforce a consistent sense of "up" and "down" in microgravity. These cues helped astronauts maintain a stable frame of reference. By contrast, the International Space Station (ISS) relies primarily on the consistent placement of equipment, displays, labels, and handrails rather than lighting schemes to establish orientation.

Recognition by Components (RBC) Theory

Recognition by Components (RBC) theory, developed by Irving Biederman, is a major theory of object recognition that provides valuable insights into how our brain processes visual information (Biederman, 1987). This theory suggests that the recognition of objects occurs by breaking them down into their fundamental components or geometric shapes, known as "geons." Here, we will delve into the details of Biederman's Recognition by Components theory.

Overview of Recognition by Components Theory

Biederman's theory postulates that the process of recognizing an object involves several stages (Biederman, 1987) (Figure 7.4):

  1. Edge Extraction: The initial step in object recognition is the extraction of edges from the visual input. When we perceive objects, our brain first identifies and processes the edges and contours that form the boundaries of those objects. This edge information is crucial for further analysis. This aligns with our understanding of visual processing in the brain, as regions like V1 (primary visual cortex) are known to respond strongly to edges in the visual field (Hubel & Wiesel, 1962; Hubel & Wiesel, 1968).
  2. Formation of Geons: Geons are basic geometric shapes that represent the building blocks of object recognition according to Biederman's theory. There are 36 such geons in total, and Biederman proposed that by combining these geons in various ways, we can recognize and represent all the diverse objects in the world (Biederman, 1987).
  • Geons can be thought of as three-dimensional shapes like cylinders, cones, cubes, and more. These shapes are simple yet versatile enough to compose complex objects.
  1. Object Recognition: By assembling and combining geons, we can recognize and identify different objects in our environment. For instance, putting together geons in a specific arrangement could represent a suitcase or a flashlight, demonstrating the versatility of this approach (Biederman, 1987).

Figure showing "Biederman's Recognition by Components (RBC) Theory" illustrating the three-stage process of object recognition. The diagram is arranged as a left-to-right flowchart with arrows connecting three panels.

Stage 1: Edge Extraction shows a photograph of a white coffee mug above a black image containing only its white outline, demonstrating how the visual system extracts edges from a 2D image.

Stage 2: Formation of Geons displays several simple three-dimensional geometric primitives (geons), including a cylinder, cone, sphere, cube, rectangular prism, and wedge. At the bottom, the mug is broken down into a cylinder and a curved handle, illustrating how objects can be represented as combinations of basic shapes.

Stage 3: Combining Geons into Objects shows a transparent cylinder with a curved handle attached, equated to a completed coffee mug.

Figure 7.4

Illustration of Biederman's theory of object recognition, showing the progression from edge extraction to the identification of basic geometric components (geons) and the combination of those geons into a structural description that supports object recognition.

"Biederman’s Recognition by Components." by Kahan, T.A. is licensed under CC BY-NC-SA 4.0

Challenges and Controversies in RBC Theory

While Recognition by Components theory has provided valuable insights into object recognition, it has also faced debates and challenges:

  • Contextual Influence: One major point of debate is the extent to which context influences object recognition. Biederman himself explored whether the context in which an object is presented affects perception (he did this using the object detection task described below). For instance, does seeing a deer in a forest make it easier to recognize than seeing it in an unexpected place like a bedroom? Biederman felt that scenes do provide a top-down influence that facilitates the recognition of congruent objects (Biederman et al., 1982).
  • Functional Isolation Model: Some researchers, such as Henderson and Hollingworth, have proposed a Functional Isolation Model of object recognition. This model suggests that scene context only plays a role after initial object recognition (Henderson & Hollingworth, 1999).

Object Detection Task and Contextual Influence

Biederman designed the Object Detection Task to investigate whether the context in which an object appears influences how easily it is recognized. For example, is a chair easier to recognize when it appears in a classroom than when it appears underwater or on a busy city street? If context helps object recognition, then objects that fit naturally within a scene should be identified more accurately than objects that appear in unusual or unexpected locations.

To test this idea, participants were first told the name of an object (e.g., fire hydrant) and then briefly shown a scene. Their task was simply to decide whether the named object was present in the picture. Biederman found that participants were more accurate at detecting an object when it appeared in a context in which it would normally be expected. For example, people were better at detecting a fire hydrant in a city street than when it appeared on the counter of a diner (Biederman et al., 1982).

However, these findings have been debated. Critics argued that participants may not actually have perceived the object more easily. Instead, they may simply have been more willing to respond "yes" when the object was appropriate for the scene and "no" when it seemed unlikely to appear there. In other words, the results may reflect response bias (guessing) rather than a genuine top-down influence of scene context on object recognition. Although this alternative explanation has been discussed extensively, there is still no definitive experiment that completely separates the effects of scene context on perception from its effects on decision making (Henderson & Hollingworth, 1999).

One possible explanation for these findings comes from research showing that people can extract the gist of a scene extremely quickly, often within a single brief glance (Potter, 1976; Greene & Oliva, 2009). Before we have had enough time to identify individual objects, we may already know whether we are looking at a kitchen, a classroom, a city street, or a beach. This rapid understanding of the overall scene could influence responses in the object detection task. For example, if a briefly presented image has the gist of a city street, participants may be more likely to report seeing a fire hydrant because it is an object that commonly belongs in that setting. Likewise, they may be more likely to respond that a fire hydrant was absent if the scene has the gist of a diner. If this interpretation is correct, scene context may influence participants' guesses about what is likely to be present rather than their perception of the object itself.

Interference Task and Contextual Influence

Kathy Mathis attempted to answer this question by conducting experiments using an interference task (Mathis, 2002). Participants had to categorize words as quickly as possible (categorizing the words as animal, food, furniture, or clothing) and these words would appear inside a picture of an object (e.g., a deer) or a nonobject (e.g., a blob that is object-like but does not have meaning). The idea was that it would be harder to categorize the word shirt as clothing when shown inside a deer relative to inside a blob since the deer, if recognized, would make people want to press the animal key. This is exactly what happened. Critically, these words and pictures were then inserted into scenes where the objects could be probable (e.g., a deer in a forest) or improbable (e.g., a deer in a bedroom).

  • In probable scenes (where the object fit the context), participants experienced more interference, indicating that the context facilitated recognition.
  • In improbable scenes (where the object did not fit the context), participants experienced no interference. This finding indicates that the scene affects object recognition but it raises questions about the timing of this contextual influence.
  • At first glance it may appear as though the object was fully ignored when it didn’t fit the scene, but in order to know the object mismatches with the scene the object and scene must first be recognized. So, when in the sequence of processing did the scene affect object recognition? Did it affect the online recognition of the object or did the scene only have an influence after the object was recognized?

Figure showing "Mathis (2002) Interference Task". The layout centers on a computer monitor displaying a forest scene with a deer. The word "shirt" appears in large white letters directly on the body of the deer. Participants are instructed to ignore the picture and scene and categorize the word as quickly as possible.

To the left, a blue instruction panel explains that the task is to decide which category the word belongs to and press the corresponding response key as quickly and accurately as possible. Below the monitor, four large response keys are labeled Animal, Food, Furniture, and Clothing. A pair of hands rests on the keyboard, with the right index finger pressing the Clothing key.

On the right, a green panel describes the example trial: the word "shirt" appears inside a picture of a deer in a forest scene, and the correct response is Clothing. A purple panel explains the purpose of the task: if the deer is recognized, it activates the competing animal category, creating response conflict and slowing classification of the word shirt even though the correct response is clothing. The figure illustrates how semantically related objects and scene context can interfere with word categorization, providing a measure of object recognition.
Figure 7.5

Participants categorized the centrally presented word into one of four semantic categories (animal, food, furniture, or clothing) while ignoring the object and background scene. Interference was measured to assess the extent to which the object was recognized when it appeared in scene contexts that were either compatible (e.g., a deer in a forest, shown here) or incompatible (e.g., a deer in a bedroom, not shown). Note: in the original study the objects and scenes were black and white line drawings rather than full-color photographs.

"Mathis’ Interference Task." by Kahan, T.A. is licensed under CC BY-NC-SA 4.0

Debate About Timing of Contextual Influence

The debate continues regarding whether the influence of scene context on object recognition occurs early in the processing stream (facilitating recognition) or late (affecting perception after recognition). Some argue that context influences early processing, while others suggest it might occur only after recognition of every object has taken place (Henderson & Hollingworth, 1999; Mathis, 2002; Oliva & Torralba, 2007).

Object Updating Theory

Introduction to Object Updating Theory:

The Object Updating Theory, developed by James Enns and Vincent DiLollo, is a fundamental concept in the field of visual perception (Enns & Di Lollo, 1997; Enns & Di Lollo, 2000). This theory elucidates how our brains process and update information about objects. Before delving into specific visual effects that might be explained by this theory (such as the Flash Lag Effect, Object Substitution Masking, and Change Blindness), let's explore the core principles of the Object Updating Theory.

A central idea of this theory is that our visual system is not simply taking a series of snapshots of the world. Instead, it attempts to maintain stable representations of objects over time by continually updating information about them. Most of the time this updating process is remarkably successful and allows us to perceive a stable, continuous visual world despite constant eye movements and changes in the environment. However, under certain conditions this otherwise helpful process can produce systematic perceptual errors. The Flash Lag Effect, Object Substitution Masking, and even Change Blindness have all been explained as consequences of the way the visual system updates these object representations over time (Kahneman et al., 1992; Enns & Di Lollo, 1997).

Principles of Object Updating Theory:

1. Object Files: In the context of this theory, "object files" are central constructs (Kahneman et al., 1992). These object files are essentially mental representations that our brains create for objects in our visual field. They contain vital information about an object's attributes, including its location and other relevant features.

2. Continuous Updating: One of the key tenets of the Object Updating Theory is the idea of continuous updating (Enns & Di Lollo, 1997; Di Lollo et al., 2000). When we perceive moving objects, our brains constantly update the information within these object files to track the objects' positions and attributes over time. This updating process relies on reentrant pathways in the brain, in which higher-level brain regions continuously compare visual information from lower-level regions with existing object files to determine whether an object should be updated or whether a new object file should be created.

3. Handling New Objects: As we encounter new objects in our visual field, our brains generate new object files for them. This allows us to seamlessly integrate these objects into our ongoing perceptual experience (Kahneman et al., 1992).

Taken together, these principles suggest that perception is an active process. Rather than creating a brand new representation every moment, the brain attempts to determine whether new visual information belongs to an existing object or represents the appearance of a new one. The answer to this question determines whether an existing object file is updated or whether a new object file is created (Enns & Di Lollo, 1997). As you will see, many visual illusions occur because the visual system makes the wrong decision about how objects should be updated.

The remainder of this section illustrates an important idea: the same mechanism—continuous updating of object files—can explain several seemingly unrelated visual phenomena.

The Flash Lag Effect

Understanding the Flash Lag Effect: The Flash Lag Effect is a captivating visual illusion that aligns with the principles of the Object Updating Theory. This phenomenon creates the perception that a flashed object lags behind a moving object, even when they are physically aligned (Eagleman & Sejnowski, 2000). In simpler terms, it makes the moving object appear ahead of the flashed object. As an example, imagine a dot moving in a clockwise manner around a central fixation point. If a square appears between the dot and the fixation when the dot hits the 3:00 position of a clock, the two (dot and square) will appear aligned if the dot stops moving (stopped motion condition) but the dot will appear to be ahead of the square (i.e., the square will lag behind) if the dot continues on its path toward the 4:00 position (continued motion condition) (Figure 7.6).

Diagram showing the flash lag effect in three conditions. In the continued motion condition, a dot moves past the location where a square briefly flashes, making the dot appear ahead of the square. In the stopped motion condition, the dot stops when the square flashes, and the two appear aligned. In the size change condition, the moving dot suddenly becomes much smaller at the time of the flash, and the small dot is perceived as a new object aligned with the flashed square; the moving large dot is perceived to be ahead. The figure illustrates how continued object updating produces the flash lag effect.

Figure 7.6

The flash lag illusion occurs when a moving dot continues after a flashed square appears. Stopping the dot or dramatically changing its size interrupts object updating, causing the dot and square to be perceived as aligned.

"Flash lag effect." by Kahan, T.A. is licensed under CC BY-NC-SA 4.0

Applying Object Updating Theory to the Flash Lag Effect: The Object Updating Theory offers an insightful explanation for the Flash Lag Effect. When we observe a moving object, our brains create an object file for it and continuously update its location. When a new object, such as a flashed square, appears, a new object file is generated. However, the updating process for the original moving object continues as the dot continues to move. As a result, when the dot continues moving and we are later asked where the moving dot had been when the dot was flashed people think the two were misaligned (Enns & Di Lollo, 1997; Moore & Enns, 2004).

Notice that the illusion does not occur because the visual system fails. Rather, it occurs because the visual system continues updating the object file for the moving dot after the flash has occurred. In everyday life, continuously updating moving objects is almost always beneficial because it allows us to accurately track moving people, animals, and vehicles. The flash lag effect is simply one situation in which this normally adaptive process produces an incorrect perceptual judgment.

Size Change and Object Perception: Moreover, the Object Updating Theory predicts that dramatic changes in an object's attributes, like size, can influence our perception. When an object undergoes a significant size change, our brains may treat it as a new object, leading us to perceive multiple objects instead of one object that changes over time (Moore et al., 2007). This aspect of the theory sheds light on why people perceive three distinct objects in a display involving a rapidly moving dot, a flashed square, and a dramatically resized dot. In this situation people perceive the small dot as a distinct object that is aligned with the square and central fixation at the moment the square was flashed (rather than perceiving the dot as appearing ahead of the square at the moment the square was flashed) (see size change condition in Figure 7.6).

In other words, the Object Updating Theory predicts that sufficiently large changes in an object's appearance can cause the visual system to stop updating the original object file and instead create a new one. Whether the brain updates an existing object file or creates a new one has important consequences for how we perceive events over time.

Object Substitution Masking:

Exploring Object Substitution Masking, which is also known as Common Onset Masking or Four Dot Masking[1]. In this effect, a briefly displayed target object, surrounded by four small dots, becomes challenging to identify when the dots persist on the screen after

the target vanishes (Di Lollo et al., 2000). However, if the dots disappear simultaneously with the target, identification becomes more manageable. For example, if a person is asked to recognize a shape (e.g., a triangle) that is surrounded by four small dots the person will have relatively little difficulty doing this if the display is shown rapidly and disappears. However, performance is much worse if the dots stay on the screen following the offset of the shapes (Di Lollo et al., 2000; Enns & Di Lollo, 2000) (see Figure 7.7).

Again, the central idea is that the visual system is continually updating object files. When the four dots remain after the target disappears, the brain treats the remaining dots as the continuation of the same object representation. As the object file is updated, information about the original target shape (triangle and dots) is effectively replaced by information about the dots alone, making the target difficult to identify afterward.

Diagram illustrating object substitution masking. A triangle briefly appears at the center of four small dots that mark its location. In one condition, the four dots disappear at the same time as the triangle, allowing the triangle to be identified easily. In the masking condition, the triangle disappears while the four dots remain visible for a short time, making the triangle much more difficult to identify. The figure demonstrates that the lingering four-dot mask can interfere with perception by replacing the target's object representation.

Figure 7.7

A briefly presented triangle is identified accurately when the surrounding four-dot mask disappears at the same time as the target. However, when the four dots remain visible after the triangle disappears (shown here), the triangle becomes difficult to identify because the visual system continues updating the existing object file. As later visual information contains only the four dots, the final object representation contains the dots but not the triangle, making the triangle difficult to identify.

"Object substitution masking." by Kahan, T.A. is licensed under CC BY-NC-SA 4.0

The Object Updating Theory provides a clear framework for understanding Object Substitution Masking. When the target and surrounding dots are presented rapidly and disappear together, our brains create an object file for the target and masking dots (e.g., triangle plus dots). If the dots disappear with the target then the person performs well when asked what shape had been shown. However, when the dots persist after the target's disappearance, the continuous updating process results in the object file containing the dots alone (without the triangle). Consequently, identifying the target becomes more difficult under these conditions. Notice that the theory does not propose that the visual system simply "forgets" the target. Instead, the target becomes difficult to identify because the object's representation has been updated. In other words, the original object file that initially contained both the target and the four dots is replaced by an updated object file containing only the dots. Like the Flash Lag Effect, this illusion arises because the visual system is continuously updating its representations of objects over time.

Change Blindness

The Object Updating Theory has also been used to explain change blindness (described in the previous chapter). When we view a scene, the visual system creates object files for the objects around us. As the scene changes, these object files are continually updated with new information. If attention is not directed toward a particular object at the moment it changes, the visual system may simply replace the old information with the new information without preserving a record of what was there previously. As a result, people often have the impression that nothing has changed because the updated object file represents only the current state of the object. Rather than comparing the old and new versions of the object, the visual system has simply overwritten the earlier representation with the updated one.

Summary

The examples in this chapter all illustrate the same general principle. The visual system continuously creates, maintains, and updates object files so that we experience a stable and coherent visual world. Most of the time this updating process is extremely useful because it allows us to keep track of moving objects and changing scenes. However, when the updating process does not perfectly match the physical events in the environment, it can produce visual illusions such as the Flash Lag Effect, Object Substitution Masking, and Change Blindness. Thus, the Object Updating Theory explains these phenomena not as failures of vision, but as consequences of a visual system that is constantly updating its internal representation of the world (Enns & Di Lollo, 2000).

Conclusion:

In conclusion, Gestalt grouping principles have demonstrated their influence on how individuals perceive and organize information within the visual environment, shaping the way elements are grouped into coherent and meaningful objects. Shading influences our perception of objects since humans have a bias to perceive the light source as coming from above. Concurrently, the Recognition by Components theory offers a framework that aligns with our understanding of brain processing, suggesting that objects are recognized by breaking them down into edges and then reassembling them into basic geometric shapes (geons) and combining these to form objects. Furthermore, the Object Updating theory posits that iterative processing in the brain plays a pivotal role in tracking objects as they move and change, shedding light on various visual phenomena such as the flash lag effect, object substitution masking, and change blindness. These theories collectively contribute to our comprehension of the intricate processes underlying object recognition and the dynamic nature of visual perception.

References

Biederman, I. (1987). Recognition-by-components: A theory of human image understanding. Psychological Review, 94(2), 115–147. https://doi.org/10.1037/0033-295X.94.2.115

Biederman, I., Mezzanotte, R. J., & Rabinowitz, J. C. (1982). Scene perception: Detecting and judging objects undergoing relational violations. Cognitive Psychology, 14(2), 143–177. https://doi.org/10.1016/0010-0285(82)90007-X

Di Lollo, V., Enns, J. T., & Rensink, R. A. (2000). Competition for consciousness among visual events: The psychophysics of object substitution. Journal of Experimental Psychology: General, 129(4), 481–507. https://doi.org/10.1037/0096-3445.129.4.481

Eagleman, D. M., & Sejnowski, T. J. (2000). Motion integration and postdiction in visual awareness. Science, 287(5460), 2036–2038. https://doi.org/10.1126/science.287.5460.2036

Enns, J. T., & Di Lollo, V. (1997). Object substitution: A new form of masking in unattended visual locations. Psychological Science, 8(2), 135–139. https://doi.org/10.1111/j.1467-9280.1997.tb00696.x

Enns, J. T., & Di Lollo, V. (2000). What's new in visual masking? Trends in Cognitive Sciences, 4(9), 345–352. https://doi.org/10.1016/S1364-6613(00)01520-5

Greene, M. R., & Oliva, A. (2009). Recognition of natural scenes from global properties. Quarterly Journal of Experimental Psychology, 62(1), 18–45.

Henderson, J. M., & Hollingworth, A. (1999). High-level scene perception. Annual Review of Psychology, 50, 243–271. https://doi.org/10.1146/annurev.psych.50.1.243

Hubel, D. H., & Wiesel, T. N. (1962). Receptive fields, binocular interaction and functional architecture in the cat's visual cortex. Journal of Physiology, 160(1), 106–154. https://doi.org/10.1113/jphysiol.1962.sp006837

Hubel, D. H., & Wiesel, T. N. (1968). Receptive fields and functional architecture of monkey striate cortex. Journal of Physiology, 195(1), 215–243. https://doi.org/10.1113/jphysiol.1968.sp008455

Kahneman, D., Treisman, A., & Gibbs, B. J. (1992). The reviewing of object files: Object-specific integration of information. Cognitive Psychology, 24(2), 175–219. https://doi.org/10.1016/0010-0285(92)90007-O

Mathis, K. M.  (2002). Semantic interference from objects both in and out of a scene context.  Journal of Experimental Psychology: Learning, Memory, and Cognition, 28, 171–182.   https://doi.org/10.1037/0278-7393.28.1.171

Moore, C. M., & Enns, J. T. (2004). Object updating and the flash-lag effect. Psychological Science, 15(12), 866–871. https://doi.org/10.1111/j.0956-7976.2004.00768.x

Moore, C. M., Mordkoff, J. T., & Enns, J. T. (2007). The path of least persistence: Object status mediates visual updating. Vision Research, 47(12), 1624–1630. https://doi.org/10.1016/j.visres.2007.01.030

Oliva, A., & Torralba, A. (2007). The role of context in object recognition. Trends in Cognitive Sciences, 11(12), 520–527. https://doi.org/10.1016/j.tics.2007.09.009

Palmer, S. E. (1992). Common region: A new principle of perceptual grouping. Cognitive Psychology, 24(3), 436–447. https://doi.org/10.1016/0010-0285(92)90014-s

Palmer, S. E. (2002). Perceptual grouping: It's later than you think. Current Directions in Psychological Science, 11(3), 101–106. https://doi.org/10.1111/1467-8721.00178

Palmer, S. E., & Rock, I. (1994). Rethinking perceptual organization: The role of uniform connectedness. Psychonomic Bulletin & Review, 1(1), 29–55. https://doi.org/10.3758/BF03200760

Peterson, M. A., & Gibson, B. S. (1994). Object recognition contributions to figure–ground organization: Operations on outlines and subjective contours. Perception & Psychophysics, 56(5), 551–564. https://doi.org/10.3758/BF03206951

Potter, M. C. (1976). Short-term conceptual memory for pictures. Journal of Experimental Psychology: Human Learning and Memory, 2(5), 509–522. https://doi.org/10.1037/0278-7393.2.5.509

Ramachandran, V. S. (1988). Perception of shape from shading. Nature, 331(6152), 163–166. https://doi.org/10.1038/331163a0

Wagemans, J., Elder, J. H., Kubovy, M., Palmer, S. E., Peterson, M. A., Singh, M., & von der Heydt, R. (2012). A century of Gestalt psychology in visual perception: I. Perceptual grouping and figure–ground organization. Psychological Bulletin, 138(6), 1172–1217. https://doi.org/10.1037/a0029333

Wertheimer, M. (1912). On perceived motion and figural organization (L. Spillmann, Trans.). MIT Press. https://doi.org/10.7551/mitpress/9222.001.0001


[1] These different terms have emerged over time because researchers have emphasized different aspects of the effect. Object substitution masking is the term most commonly used because it reflects the idea that the visual representation of the target is replaced or "substituted" by the later processing of the mask. Four-dot masking refers to a common experimental method in which the target is briefly surrounded by four small dots that remain visible after the target disappears. Common onset masking emphasizes that the target and mask begin at the same time, but the mask remains visible after the target has disappeared. Although these names highlight different characteristics of the procedure or theoretical interpretation, they all refer to the same masking phenomenon.

Annotate

Next Chapter
Chapter 8: Color Vision
PreviousNext
Powered by Manifold Scholarship. Learn more at
Opens in new tab or windowmanifoldapp.org