Skip to main content

Sensation & Perception V3: Chapter 11: Audition

Sensation & Perception V3
Chapter 11: Audition
  • Show the following:

    Annotations
    Resources
  • Adjust appearance:

    Font
    Font style
    Color Scheme
    Light
    Dark
    Annotation contrast
    Low
    High
    Margins
  • Search within:
    • My Notes + Comments
    • Notifications
    • Privacy
  • Project HomeSensation & Perception V3
  • Projects
  • Learn more about Manifold

Notes

table of contents
  1. Front Matter
  2. Preface
  3. Acknowledgements
  4. Chapter 1: Introduction to the Study of Sensation and Perception
  5. Chapter 2: Approaches to Studying Sensation and Perception
  6. Chapter 3: Receptors and Neural Processing
  7. Chapter 4: The Lateral Geniculate Nucleus (LGN) and Primary Visual Cortex (V1)
  8. Chapter 5: Higher-Level Visual Processing: Beyond V1
  9. Chapter 6: Attention and Visual Perception
  10. Chapter 7: Object Recognition
  11. Chapter 8: Color Vision
  12. Chapter 9: Depth Perception
  13. Chapter 10: Motion
  14. Chapter 11: Audition
  15. Chapter 12: Cutaneous Senses
  16. Chapter 13: Gustatory Senses
  17. Chapter 14: Olfaction
  18. Version History

Chapter 11: Audition

In this chapter, we examine auditory perception, the sense that allows us to perceive and interpret sound. Auditory perception plays a crucial role in our daily lives, from communicating with others to enjoying music and being aware of our environment. We will start by exploring the basics of sound, sound waves, and how our ears are finely tuned to detect and process auditory information.

Understanding Sound Waves

The Nature of Sound

Sound is a phenomenon that originates from vibrations or disturbances in the air, creating waves of air pressure that our ears can detect. These waves can be caused by various sources, from the vibration of a tuning fork to the rustling of leaves in the wind. Understanding the fundamentals of sound is essential to grasp the intricacies of auditory perception.

Frequency and Pitch

Frequency measures the number of cycles per second in a sound wave and determines the pitch of the sound. Human hearing is typically sensitive to frequencies ranging from approximately 20 hertz (Hz) to 20,000 Hz. Different animals have varying audible ranges based on their physiology.

Amplitude and Loudness

Amplitude refers to the size (height) of a sound wave, reflecting the changes in air pressure produced by the sound. This physical property is measured as sound intensity in decibels (dB). For sounds of the same frequency, increasing the intensity increases the perceived loudness. However, loudness is determined by more than intensity alone. The human ear is more sensitive to some frequencies than others, so sounds with the same intensity can be perceived as having different loudness depending on their frequency (Suzuki & Takeshima, 2004).

This relationship is illustrated by the threshold of hearing shown in Figure 11.1. The threshold of hearing is the minimum sound intensity required for a tone of a particular frequency to be detected. Because the ear is most sensitive to frequencies of approximately 2,000–5,000 Hz, sounds in this range can be heard at much lower intensities than very low- or very high-frequency sounds. Consequently, a tone presented at the same intensity (dB) may be inaudible at a very low frequency, clearly audible at a mid-range frequency, and become difficult or impossible to hear again at very high frequencies. For example, as shown in the figure, a 20 Hz tone presented at 20 dB falls below the threshold of hearing and would not be perceived, whereas a 1000 Hz tone presented at the same intensity lies above the threshold and would be heard. In addition to the threshold of hearing, the figure also shows the threshold for feeling, which represents the intensity at which sound vibrations become strong enough to be physically felt rather than only heard.

Graph showing the Thresholds of Hearing and Feeling and Typical Conversational Speech. The x-axis shows Frequency (Hz) on a scale from 20 to 20,000 Hz, and the y-axis shows Intensity Level (dB) from −20 to 140 dB. Two smooth black curves define the approximate limits of human hearing: a lower Threshold of Hearing curve that is lowest (most sensitive) at mid-range frequencies and rises toward both low and high frequencies, and an upper Threshold of Feeling curve near 100–130 dB. A shaded rounded rectangle labeled “Conversational Speech" shows that typical spoken conversation falls within the audible range.

Figure 11.1

The lower curve represents the approximate threshold at which sounds become audible, while the upper curve indicates the approximate intensity at which sound begins to be perceived as vibration or physical sensation. Human hearing is most sensitive at mid-range frequencies, where the threshold of hearing reaches its lowest values. The shaded region shows the approximate frequency and intensity range of conversational speech. The curves are illustrative and are intended to convey the general characteristics of human auditory sensitivity rather than exact clinical measurements.

"Human audible range." by Kahan, T.A. is licensed under CC BY-NC-SA 4.0

Auditory Anatomy

The Outer Ear

The outer ear consists of the visible external portion known as the pinna and the auditory canal. Its primary function is to capture and funnel sound waves towards the middle ear.

Figure 11.2


A tuning fork causes sound vibrations which stimulates the ear.

"Sound Waves and the Ear" by OpenStax is licensed under CC BY 4.0

The Middle Ear

The middle ear is home to the tympanic membrane (eardrum) and the ossicles, which include the malleus, incus, and stapes. These components work together to amplify and transmit sound vibrations from the eardrum to the inner ear.

The Inner Ear

Deep within the inner ear lies the cochlea, a coiled, fluid-filled structure responsible for transducing sound vibrations into neural signals. The cochlea contains specialized hair cells and membranes, including the basilar membrane and tectorial membrane, which are essential for sound perception.

The Process of Auditory Transduction

Cochlear Function

Sound vibrations transmitted through the ossicles eventually reach the oval window. When the stapes presses on the oval window this causes fluid within the cochlea to move. This movement stimulates hair cells along the basilar membrane, leading to auditory transduction.

Figure 11.3

Diagram of the ear, ossicles, and cochlea.

"Frequency Coding in the Human Ear and Cortex" by Chittka, Lars and Brockmann, Axel is licensed under CC BY 4.0

Cross-sectional diagram of the organ of the cochlea and organ of Corti. The basilar membrane forms the base of the structure, supporting one row of inner hair cells and three rows of outer hair cells. Sound-induced movement of the basilar membrane rubs the hair cells against the tectorial membrane, converting mechanical vibrations into neural signals.
Figure 11.4

Diagram showing a cross section of the cochlea. "Cochlea" by OpenStax is licensed under CC BY 4.0

Hair Cells

The inner hair cells are responsible for converting mechanical vibrations into neural signals, which are then transmitted via the auditory nerve to the brain. Outer hair cells actively amplify vibrations of the basilar membrane (Ashmore, 2008).

Diagram showing an unrolled cochlea. High frequency tones displace closer to the base and low frequency displace closer to the apex.

Figure 11.5

Diagram showing an unrolled cochlea. High frequency tones produce maximal displacement near the base of the cochlea, whereas low-frequency tones produce maximal displacement near the apex, reflecting the cochlea's tonotopic organization (Dallos, 1996).  "Schematic uncoil cochlea" by Kern A, Heid C, Steeb W-H, Stoop N, Stoop R is licensed under CC BY 2.5

Georg von Békésy & Place Theory

Georg von Békésy, born in Budapest, Hungary in 1899, was a renowned scientist whose groundbreaking research in auditory physiology earned him the Nobel Prize in Physiology or Medicine in 1961. His research fundamentally changed our understanding of the physiological mechanisms of hearing.

Early Life and Academic Pursuits

Békésy came from an educated family in Budapest. He initially studied engineering at the Budapest Technical University but soon shifted his interest towards physics, which led him to study at the University of Bern in Switzerland. There, he worked under the guidance of the eminent physicist Auguste Piccard, who encouraged his scientific endeavors.

Auditory Research: The Place Theory and the Cochlea

Békésy's most significant contributions came through his research on hearing and auditory perception. In the early 20th century, there was considerable debate about how humans perceive different sound frequencies. Békésy's pioneering work in the 1920s and 1930s helped resolve this mystery.

Békésy's work provided strong experimental support for the place theory of hearing, demonstrating how different sound frequencies produce maximal vibration at different locations along the basilar membrane. Although frequency theory explains the perception of very low-frequency sounds, place theory accounts for much of pitch perception across the audible range (Békésy, 1960). Békésy proposed that different sound frequencies are detected at specific locations along the cochlea, the spiral-shaped, fluid-filled structure of the inner ear. He developed an innovative technique that allowed him to directly observe the movement of the basilar membrane in response to sound. His experiments showed that sound produces a traveling wave along the basilar membrane, with the point of maximum displacement depending on the sound's frequency (Békésy, 1960).

According to place theory, high-frequency sounds produce their greatest displacement near the base of the cochlea, close to where the stapes presses on the oval window, whereas low-frequency sounds produce their greatest displacement near the apex, the far end of the cochlea (Békésy, 1960). This occurs because the basilar membrane is narrow and stiff at the base, making it most responsive to high frequencies, and wider and more flexible at the apex, making it most responsive to low frequencies.

Hair cells located at the point of maximum basilar membrane displacement convert these mechanical vibrations into neural signals. Because each location along the basilar membrane is tuned to a particular frequency, the auditory nerve preserves this tonotopic organization as sound information is transmitted to the brain. This frequency map is maintained throughout the auditory pathway and into the primary auditory cortex (A1) (Kaas & Hackett, 2000; Robles & Ruggero, 2001), allowing the brain to identify the pitch of sounds based on which neurons are activated.

The Glue Factory and the Elephant

One intriguing anecdote from Békésy's life involves an elephant skull. In his quest to understand hearing, he needed to study the cochlea of different animals to draw comparative conclusions. However, obtaining suitable specimens proved challenging.

Legend has it that Békésy had heard about an elephant skull at a glue factory in Zurich. He believed that studying the elephant cochlea could provide valuable insights since this would be larger and might be easier to study and photograph when sounds are played. Undeterred, he reportedly sent a graduate student to the glue factory with saws, knives, an axe, and other tools and in the end, Békésy obtained the elephant skull.

Illustration of a dog portrayed as Békésy's graduate student beside an elephant, with the dog saying, "This is fine." The cartoon references the anecdote that the student retrieved an elephant skull from a glue factory so it could be used to study the large cochlea.

Figure 11.6

Illustration inspired by the anecdote that, in pursuit of a larger cochlea for comparative hearing research, Georg von Békésy sent a graduate student to recover an elephant skull from a Zurich glue factory. The graduate student is depicted declaring, "This is fine." Note: AI refused to generate most of the prompts I wrote to depict this scene.

"Elephant glue story." by Kahan, T.A. is licensed under CC BY-NC-SA 4.0

Auditory Analysis

According to the place theory, our auditory system can analyze complex sounds, breaking them down into their component frequencies. This process, akin to Fourier analysis. A Fourier analysis is a mathematical process that decomposes a complex sound into its individual frequency components. The basilar membrane performs a similar function biologically by responding to different frequencies at different locations along its length, allowing the auditory system to analyze the frequency content of complex sounds before they reach the brain (Békésy, 1960; Robles & Ruggero, 2001).

Causes of Hearing Loss

Noise-Induced Hearing Loss

One of the primary causes of hearing loss is exposure to loud sounds, such as concerts or prolonged use of headphones at high volumes (Liberman & Kujawa, 2017). Noise-induced hearing loss can result in permanent damage to the auditory system.

Genetic Factors

A family history of hearing loss can increase the risk of inheriting hearing-related issues (Arnos & Pandya, 2003) . Understanding one's genetic predisposition to hearing loss can help in adopting preventative measures.

Other Factors

Hearing loss can also occur due to factors like ear trauma, illness, or damage to the auditory system's delicate components. In some cases, hearing loss can be accompanied by tinnitus, a persistent ringing in the ears (Baguley et al., 2013).

Differing Perspectives on Hearing Loss

Perspectives on hearing loss vary widely, particularly between cultural and medical viewpoints (Lane, 1992; Padden & Humphries, 1988; Napier & Leeson, 2016). Many members of the Deaf community view deafness not as a disability or deficit, but as a natural human difference and the basis of a rich linguistic and cultural identity centered around sign language and shared community experiences. From this perspective, deafness is something to be embraced rather than "fixed," and efforts to cure it are sometimes seen as overlooking the value of Deaf culture. In contrast, the medical model views hearing loss as a condition that limits auditory function and can often be treated or partially restored through interventions such as hearing aids or cochlear implants. Supporters of this perspective emphasize the potential benefits of improved access to spoken language, environmental sounds, and communication through hearing. Individual views on hearing loss and treatments such as cochlear implants are diverse, and many people with hearing loss hold perspectives that fall somewhere between these two viewpoints or vary depending on their personal experiences, cultural background, age of hearing loss, and communication preferences.

Gestalt Grouping

Auditory Grouping and Perceptual Processes

Auditory information can be grouped in ways that are similar to what has been found with visual information processing.

Similarity and Proximity in Auditory Grouping

Just as visual information can be grouped based on the similarity of color or spatial proximity, auditory information can be grouped based on similarity of pitch and temporal proximity.

Temporal Proximity: Manipulating Auditory Grouping

Temporal proximity is one of the primary cues the auditory system uses to organize sounds into meaningful perceptual groups. When alternating high- and low-pitched tones are presented slowly, listeners typically perceive a single stream that rises and falls in pitch. As the presentation rate increases, the shorter intervals between tones promote grouping by pitch similarity, causing the high and low tones to separate into distinct auditory streams. This process, known as auditory stream segregation, illustrates how temporal proximity and pitch similarity interact to organize complex auditory scenes (Bregman, 1990). Although the physical sequence of tones remains the same, changing only the timing between them can dramatically alter what listeners perceive.

Illustration demonstrating how the perceptual organization of alternating high- and low-pitched tones changes as presentation rate increases. In the slow alternation panel, evenly spaced alternating blue (low) and green (high) tones are perceived as a single auditory stream that rises and falls in pitch. In the moderate-speed panel, the same alternating pitch sequence is presented with shorter intervals between tones. The tones are now grouped as two separate streams, one consisting of the high-pitched tones and one consisting of the low-pitched tones, reducing the perception of a rising-and-falling melody.

Figure 11.7

The physical stimulus consists of alternating high- and low-pitched tones. When tones are presented slowly (top), listeners perceive a single stream that alternates between high and low pitches, producing a rising-and-falling percept. As the presentation rate increases (bottom), temporal proximity promotes grouping by pitch similarity, leading listeners to perceive two concurrent streams (one high, one low) rather than a single alternating sequence. This illustrates the interaction between pitch similarity and temporal proximity in auditory scene analysis.

"Diagram showing grouping by pitch and temporal proximity." by Kahan, T.A. is licensed under CC BY-NC-SA 4.0

Much like spatial proximity influences visual grouping, temporal proximity shapes how we perceive auditory information. When sounds occur close together in time, we tend to group them into a single auditory stream.

Think of it as organizing the sounds into distinct "streams" based on their sound similarities and temporal closeness, just as you might group visual objects together based on their color or spatial closeness.

Good Continuation in Auditory Perception

Another critical aspect of auditory perception is good continuation. This concept mirrors what happens in visual perception when our brains fill in gaps to create a continuous pattern or form.

In one auditory example, you hear a tone followed by silence, and the cycle repeats. Strangely, if static noise is added to the gaps that had been silent people tend to hear the tone as continuing behind the static (Warren et al., 1972). It feels as though it continues behind the static and isn't interrupted in much the same way as a visually presented line appears to continue even when it is blocked by an occluding object.

Figure illustrating the principle of good continuation in auditory and visual perception. The top half compares two repeating tone sequences. In the first example, blue tone bursts are separated by silent intervals, leading listeners to perceive a tone that repeatedly starts and stops. In the second example, identical tone bursts are separated by brief bursts of static noise. Although the tone is physically absent during the noise, listeners tend to perceive a single continuous tone continuing behind the static (auditory continuity illusion). The bottom half presents a visual analogy. The first example shows three separated purple line segments, which are perceived as discontinuous. The second shows the same lines but they are partially hidden behind submarine-shaped occluders, leading viewers to perceive one uninterrupted line extending behind the submarines. Together, the examples illustrate that the perceptual system often fills in missing information when an interruption can be attributed to an intervening object or masking sound.

Figure 11.8

In the auditory examples (top), a repeating tone separated by silence is perceived as a series of discrete tones that repeatedly start and stop. When the silent intervals are replaced with static noise, listeners typically perceive the tone as continuing behind the noise, even though it is physically absent during the masked intervals. The visual examples (bottom) demonstrate the same Gestalt principle. Separated line segments are perceived as discontinuous, whereas line segments hidden behind submarine-shaped occluders are perceived as extending uninterrupted behind the objects.  

"Diagram showing good continuation." by Kahan, T.A. is licensed under CC BY-NC-SA 4.0

Shepard scale and the Tritone Paradox

A Shepard tone is a sound created by combining several tones of different pitches (see Figure 11.9) (Shepard, 1964). Since multiple pitches are present at the same time, the brain cannot easily determine a single reference pitch, making the sound's overall pitch ambiguous.

A sequence of Shepard tones can be arranged to form a Shepard scale, which creates the compelling illusion that the pitch is continuously rising (or continuously falling) even though it never actually gets higher or lower. This is sometimes described as the auditory equivalent of a barber pole because, like the rotating stripes on a barber pole, it creates the impression of endless upward or downward motion. In the figure, time is shown on the x-axis and frequency on the y-axis. The darker bands represent greater amplitude (louder sound), and each of the 12 tones consists of multiple frequencies. If the tones are played from left to right, listeners perceive an ascending sequence. However, after the twelfth tone, the sequence simply repeats from the beginning, allowing the illusion of an endlessly rising scale to continue.

The tritone paradox is based on pairs of Shepard tones that are separated by a tritone (six semitones; for example, tone 1 followed by tone 6, or tone 3 followed by tone 8). Although every listener hears the same pair of tones, people often disagree about whether the sequence is ascending or descending in pitch. Some listeners perceive the second tone as higher, whereas others perceive it as lower. Diana Deutsch (1991) proposed that these differences are influenced by the language and dialect a person was exposed to during childhood, with listeners from different linguistic backgrounds tending to perceive the same tone pairs differently.

Time–frequency spectrogram illustrating a Shepard tone sequence. The horizontal axis is labeled Time, and the vertical axis is labeled Frequency, with higher frequencies at the top and lower frequencies at the bottom. Warmer colors (yellow–red) indicate greater intensity, while cooler colors (blue) indicate lower intensity. Three prominent frequency bands ascend diagonally across each 12-step sequence before repeating in a second identical sequence. Black rectangles highlight the first tone of each sequence, and red circles identify tones 1 and 6, which form a tritone pair used in the tritone paradox.

Figure 11.9

Time–frequency representation the Shepard scale and the tritone paradox. Color indicates spectral intensity, with warmer colors representing greater energy at a given frequency. Each numbered column corresponds to one Shepard tone in the sequence. The black rectangles mark the first tone of each sequence, which are identical. The red circles highlight tones 1 and 6, which are separated by a tritone (six semitones). When listeners hear tone 1 followed by tone 6 they will often disagree about whether the sequence appears to ascend or descend, demonstrating the perceptual ambiguity underlying the tritone paradox.

"Shepard scale and tritone paradox." by Kahan, T.A. is licensed under CC BY-NC-SA 4.0

Auditory Perception of Language

Finally, let's touch very briefly on how we perceive language. When we listen to spoken language, we tend to perceive individual words as distinct units. However, in reality, there aren't always clear pauses or silences between words.

Newborns have the remarkable ability to distinguish between various phonemes (the smallest units of sound in language) from all possible languages. However, as they grow, they start focusing on the specific phonemes relevant to their native language, filtering out those that aren't essential. This filtering process begins with vowels and later extends to consonants (Kuhl, 2004; Werker & Tees, 1984).

Language perception is complex, and our brains develop the ability to parse and understand spoken language over time. In essence, what seems like a series of isolated words is actually a continuous stream of sound, seamlessly processed by our brains (Saffran et al., 1996).

Wernicke's area is a critical region in the human brain that plays a pivotal role in the understanding of spoken language. Located in the posterior part of the left hemisphere, Wernicke's area is an integral component of the broader language processing network. It was first identified by the German neurologist Carl Wernicke in the late 19th century and has since become a focal point for research into language comprehension. The primary function of Wernicke's area is to process the auditory information received from the ears and convert it into meaningful linguistic representations (Binder et al., 2009). This region is especially important for the comprehension of spoken language, allowing individuals to decipher and make sense of the sounds and words they hear during conversations, lectures, or any form of oral communication.

Figure 11.10

Diagram showing Wernicke's area.

"Area Werinike" by Charlyzona is licensed under CC BY-SA 3.0

Pathway from the ear to A1

The auditory pathway carries sound information from the ear to the brain, where it is progressively analyzed and interpreted. The main pathway includes the cochlea, superior olivary nucleus (SON), inferior colliculus (IC), medial geniculate nucleus (MGN) of the thalamus, and finally the primary auditory cortex (A1). Rather than ending at A1, auditory processing continues into higher-order auditory cortical areas that support speech perception, music processing, and sound recognition.  

  1. Cochlea and Basilar Membrane: The pathway begins in the cochlea of the inner ear. Sound waves cause vibrations of the basilar membrane, where specialized hair cells convert mechanical vibrations into neural signals. The basilar membrane is tonotopically organized, meaning different regions respond best to different sound frequencies: high-frequency sounds are detected at the base of the cochlea, while low-frequency sounds are detected at the apex (Békésy, 1960). This frequency map is established at the very first stage of auditory processing.
  2. Superior Olivary Nucleus (SON): The auditory nerve carries these signals to the brainstem, where information reaches the superior olivary nucleus. The SON is the first structure to receive input from both ears simultaneously and is essential for sound localization (Grothe et al., 2010) . It compares small differences in the timing and intensity of sounds arriving at each ear, allowing the brain to determine where a sound is coming from in space
  3. Inferior Colliculus: From the superior olivary nucleus, auditory information is relayed to the inferior colliculus in the midbrain. The inferior colliculus integrates auditory information from multiple brainstem nuclei and helps orient attention toward important sounds. It also contributes to filtering or suppressing predictable self-generated sounds (c.f., Schneider et al., 2014), such as one's own speech, chewing, and breathing, allowing externally generated sounds to stand out.
  4. Medial Geniculate Nucleus (Thalamus): The information then travels to the medial geniculate nucleus in the thalamus. The thalamus acts as a relay center for various sensory inputs, including auditory information. It serves as a crucial gateway, regulating the flow of sensory data to the cortex (Bartlett, 2013).
  5. Auditory Cortex (A1): The primary auditory cortex (A1), located in the temporal lobe, is the first cortical area to receive auditory input. Here, sound is represented in a highly organized manner, with neurons responding to specific frequencies and temporal patterns (Kaas & Hackett, 2000). However, auditory processing does not stop at A1. Instead, A1 performs the initial cortical analysis of sound, and the information is then transmitted to higher-order auditory cortical areas for more complex processing, including speech comprehension, music perception, sound recognition, and interpretation of auditory scenes.

One of the key features of the auditory pathway is that the tonotopic organization established in the cochlea is preserved throughout the entire pathway, from the basilar membrane to the superior olivary nucleus, inferior colliculus, medial geniculate nucleus, and finally A1 (Kaas & Hackett, 2000). Neighboring neurons remain tuned to neighboring sound frequencies, creating a cortical map of pitch. Thus, high-frequency sounds activate one region of A1, while low-frequency sounds activate another. This organization is analogous to the retinotopic map of the visual system, where neighboring neurons represent neighboring locations in the visual field, allowing the brain to process sensory information with remarkable spatial precision.

Auditory Ambiguities

In the realm of perception, auditory processes share some intriguing similarities with visual perception, and we can encounter ambiguities in both sensory modalities.

Just as we discussed visual illusions such as the infamous dress that some people perceived as black and blue while others saw as gold and white, our auditory system can also be fooled by ambiguous sensory information. A striking example is the Laurel–Yanny illusion, which became widely shared on the internet because different listeners reported hearing completely different words from the same sound clip.

The recording contains a broad range of frequencies, with the lower-frequency components more closely matching the speech sounds in "Laurel" and the higher-frequency components more closely matching "Yanny." Which word a listener perceives depends on which parts of the frequency spectrum are most prominent to them. This can be influenced by several factors, including the quality of the recording, the playback device (such as speakers or headphones), the listening volume, age-related differences in hearing sensitivity, and the brain's interpretation of the ambiguous speech signal.

Researchers examined the spectrogram of the recording and confirmed that both sets of frequency cues are present. By boosting the lower frequencies, they could make most listeners hear "Laurel," whereas emphasizing the higher frequencies caused many listeners to hear "Yanny." Thus, although everyone hears the same recording, small differences in the frequencies that reach the ear—or in how the brain interprets those frequencies—can lead to two very different perceptual experiences (Pressnitzer et al., 2018).

Synesthesia: A Blend of Senses

Although synesthesia is not strictly an auditory phenomenon, it is included here because several forms involve hearing, such as music-color synesthesia, in which musical notes or sounds automatically evoke vivid color experiences. More broadly, synesthesia is a fascinating perceptual phenomenon in which stimulation of one sensory or cognitive pathway automatically and consistently triggers an additional sensory experience. Rather than processing information through a single sensory modality, the brain of a synesthete links two or more sensory systems, creating experiences that combine multiple senses. Although the term literally means "joined perception" (from the Greek syn, meaning "together," and aisthesis, meaning "perception"), synesthesia is not a hallucination or a mental disorder. Instead, it represents a naturally occurring variation in the way the brain organizes sensory information for some individuals.

Many researchers suggest that all humans possess some degree of cross-modal perception (Ward, 2013). For example, people often associate high-pitched sounds with lighter colors or smaller objects, and low-pitched sounds with darker colors or larger objects. However, these associations are metaphorical or learned. In true synesthesia, the additional sensory experience occurs automatically, involuntarily, and consistently throughout a person's lifetime.

Common Types of Synesthesia

More than 80 different forms of synesthesia have been documented. The most common forms include:

Music-color synesthesia:

Musical notes, chords, or even individual instruments evoke vivid colors or moving visual patterns. A trumpet might consistently appear as golden-yellow, while a cello produces deep shades of blue.

Grapheme-color synesthesia

Letters, numbers, or words automatically evoke specific colors. For example, one synesthete may always perceive the letter A as bright red and the number 5 as green. These color experiences remain remarkably stable over time (Eagleman et al., 2007); decades later, the same individual will report the identical color associations.

Spatial-sequence synesthesia:

Numbers, months, or days of the week are perceived as occupying fixed locations in three-dimensional space. A synesthete might "see" the months of the year arranged in a circular pattern surrounding their body.

Although the specific experiences differ among individuals, one defining characteristic of synesthesia is consistency. Unlike imagination or memory, synesthetic perceptions occur automatically and remain stable across time.

The Neural Basis of Synesthesia

Research by neuroscientist V. S. Ramachandran and colleagues have provided important insights into the neurological mechanisms underlying synesthesia (Ramachandran & Hubbard, 2001). One influential explanation is known as the cross-activation theory, which proposes that neighboring regions of the brain become unusually interconnected.

In grapheme-color synesthesia, two neighboring regions in the fusiform gyrus of the temporal lobe appear to interact more strongly than usual:

  • The Visual Word Form Area (VWFA) specializes in recognizing letters and numbers.
  • Area V4, located nearby, plays a critical role in processing color.

Because these regions are anatomically adjacent, increased neural connectivity may allow activation of one area to automatically stimulate the other (Rouw & Scholte, 2007). As a result, viewing the letter "A" not only activates the brain region responsible for recognizing the letter but also simultaneously activates the color-processing region, producing the vivid experience of seeing the letter as red.

A second explanation, known as the disinhibited feedback theory, suggests that synesthesia may result not from extra anatomical connections but from reduced inhibition between existing neural pathways (Grossenbacher & Lovelace, 2001). Normally, inhibitory processes prevent information from one sensory system from activating another. If these inhibitory mechanisms are weaker, activity may spread between sensory regions, allowing cross-modal experiences to emerge.

Current evidence suggests that both increased structural connectivity and altered patterns of neural communication may contribute to different forms of synesthesia (Hubbard & Ramachandran, 2005; Ward, 2013).

Figure illustrating color–grapheme synesthesia. On the left, black letters and numbers (A, 7, B, 3, k, Z, g, 2, ?) represent the visual stimulus. An arrow points to the silhouette of a person's head with a brain containing the same graphemes shown in different colors, illustrating that each letter or number is automatically experienced as a specific color. A large red letter "A" on the right represents the perceived synesthetic color experience. Additional panels explain that these color associations are automatic and consistent over time and provide an example chart assigning arbitrary colors to various letters and numbers.

Figure 11.11

Conceptual illustration of color–grapheme synesthesia. In individuals with color–grapheme synesthesia, viewing letters or numbers automatically and consistently evokes specific color experiences. The particular color associated with each grapheme varies across individuals but remains highly stable within an individual over time. The example color associations shown are illustrative rather than representative of any universal pattern.

"Color-grapheme synesthesia." by Kahan, T.A. is licensed under CC BY-NC-SA 4.0

Conclusion

Audition is far more than the passive detection of sound—it is an active process in which the nervous system transforms patterns of air pressure into meaningful perceptual experiences. In this chapter, we explored how sound waves are characterized by frequency and amplitude, how the anatomy of the ear converts these physical vibrations into neural signals, and how the cochlea's tonotopic organization forms the foundation of pitch perception. We examined Békésy's place theory, the auditory pathway from the cochlea to the primary auditory cortex, and the mechanisms that allow us to localize sounds and perceive speech. We also saw that perception is shaped by Gestalt principles, enabling the brain to organize complex auditory scenes, while auditory illusions such as the Laurel–Yanny recording and the Shepard scale demonstrate that perception is an active interpretation rather than a direct reflection of reality. Finally, synesthesia illustrates how variations in neural connectivity can produce remarkable perceptual experiences, providing valuable insight into the flexibility of the human brain. Together, these topics reveal that hearing is not simply the registration of sound but the brain's sophisticated construction of a rich and meaningful auditory world.

References

Arnos, K. S., & Pandya, A. (2003). Advances in the genetics of deafness. In M. Marschark & P. E. Spencer (Eds.), Oxford handbook of deaf studies, language, and education (pp. 392–405). Oxford University Press.

Ashmore, J. (2008). Cochlear outer hair cell motility. Physiological Reviews, 88(1), 173–210. https://doi.org/10.1152/physrev.00044.2006

Baguley, D., McFerran, D., & Hall, D. (2013). Tinnitus. The Lancet, 382(9904), 1600–1607. https://doi.org/10.1016/S0140-6736(13)60142-7

Bartlett, E. L. (2013). The organization and physiology of the auditory thalamus and its role in processing acoustic features important for speech perception. Brain and Language, 126(1), 29–48. https://doi.org/10.1016/j.bandl.2013.03.003

Békésy, G. von. (1960). Experiments in hearing. McGraw-Hill.

Binder, J. R., Desai, R. H., Graves, W. W., & Conant, L. L. (2009). Where is the semantic system? A critical review and meta-analysis. Cerebral Cortex, 19(12), 2767–2796. https://doi.org/10.1093/cercor/bhp055

Bregman, A. S. (1990). Auditory scene analysis. MIT Press.Dallos, P. (1996). Overview: Cochlear neurobiology. In P. Dallos, A. N. Popper, & R. R. Fay (Eds.), The Cochlea (pp. 1–43). Springer. https://doi.org/10.1007/978-1-4612-0757-3

Deutsch, D. (1991). The tritone paradox. Music Perception, 8(3), 213–239. https://doi.org/10.2307/40285514

Eagleman, D. M., Kagan, A. D., Nelson, S. S., Sagaram, D., & Sarma, A. K. (2007). A standardized test battery for the study of synesthesia. Journal of Neuroscience Methods, 159(1), 139–145. https://doi.org/10.1016/j.jneumeth.2006.07.012

Grossenbacher, P. G., & Lovelace, C. T. (2001). Mechanisms of synesthesia: Cognitive and physiological constraints. Trends in Cognitive Sciences, 5(1), 36–41. https://doi.org/10.1016/S1364-6613(00)01571-0

Grothe, B., Pecka, M., & McAlpine, D. (2010). Mechanisms of sound localization in mammals. Physiological Reviews, 90(3), 983–1012. https://doi.org/10.1152/physrev.00026.2009

Hubbard, E. M., & Ramachandran, V. S. (2005). Neurocognitive mechanisms of synesthesia. Neuron, 48(3), 509–520. https://doi.org/10.1016/j.neuron.2005.10.012

Kaas, J. H., & Hackett, T. A. (2000). Subdivisions of auditory cortex. Current Opinion in Neurobiology, 10(4), 433–443. https://doi.org/10.1073/pnas.97.22.11793

Kuhl, P. K. (2004). Early language acquisition. Nature Reviews Neuroscience, 5(11), 831–843. https://doi.org/10.1038/nrn1533

Lane, H. (1992). The mask of benevolence: Disabling the Deaf community. Alfred A. Knopf.

Liberman, M. C., & Kujawa, S. G. (2017). Cochlear synaptopathy in acquired sensorineural hearing loss. Hearing Research, 349, 138–147. https://doi.org/10.1016/j.heares.2017.01.003

Napier, J., & Leeson, L. (2016). Sign Language in Action. In Sign Language in Action (pp. 50–84). Palgrave Macmillan UK. https://doi.org/10.1057/9781137309778_3

Padden, C., & Humphries, T. (1988). Deaf in America: Voices from a culture. Harvard University Press.

Pressnitzer, D., Graves, J., Chambers, C., de Gardelle, V., & Egré, P. (2018). Auditory Perception: Laurel and Yanny Together at Last. Current Biology, 28(13), R739–R741. https://doi.org/10.1016/j.cub.2018.06.002

Ramachandran, V. S., & Hubbard, E. M. (2001). Synaesthesia—A window into perception, thought and language. Journal of Consciousness Studies, 8(12), 3–34.

Robles, L., & Ruggero, M. A. (2001). Mechanics of the mammalian cochlea. Physiological Reviews, 81(3), 1305–1352. https://doi.org/10.1152/physrev.2001.81.3.1305

Rouw, R., & Scholte, H. S. (2007). Increased structural connectivity in grapheme-color synesthesia. Nature Neuroscience, 10(6), 792–797. https://doi.org/10.1038/nn1906Saffran, J. R., Aslin, R. N., & Newport, E. L. (1996). Statistical learning by 8-month-old infants. Science, 274(5294), 1926–1928. https://doi.org/10.1126/science.274.5294.1926

Schneider, D. M., Nelson, A., & Mooney, R. (2014). A synaptic and circuit basis for corollary discharge in the auditory cortex. Nature, 513(7517), 189–205. https://doi.org/10.1038/nature13724

Shepard, R. N. (1964). Circularity in judgments of relative pitch. The Journal of the Acoustical Society of America, 36(12), 2346–2353. https://doi.org/10.1121/1.1919362

Suzuki, Y., & Takeshima, H. (2004). Equal-loudness-level contours for pure tones. The Journal of the Acoustical Society of America, 116(2), 918–933. https://doi.org/10.1121/1.1763601

Ward, J. (2013). Synesthesia. Annual Review of Psychology, 64(1), 49–75. https://doi.org/10.1146/annurev-psych-113011-143840

Warren, R. M., Obusek, C. J., & Ackroff, J. M. (1972). Auditory induction: Perceptual synthesis of absent sounds. Science, 176(4039), 1149–1151. https://doi.org/10.1126/science.176.4039.1149

Werker, J. F., & Tees, R. C. (1984). Cross-language speech perception. Infant Behavior and Development, 7(1), 49–63. https://doi.org/10.1016/S0163-6383(84)80022-3

Annotate

Next Chapter
Chapter 12: Cutaneous Senses
PreviousNext
Powered by Manifold Scholarship. Learn more at
Opens in new tab or windowmanifoldapp.org