RECA-293 - How Music Rocks and Rolls
Cognitive Neuroscience of Music: Emotion, Meaning, and the Mind

MODULE 04B

 

cognitive processing in music listening
Experiencing Musical Time: Attention
Memory Limits - Categories - Accents - Schemas - Complexity
Implicit & Explicit Knowledge - Analytic & Synthetic Listening
Melody
   
Melodic Units - Pitch & Temporal Contours - Rhythm
    Harmony - Tonality
Cognitive Aspects of Loudness, Pitch, Timbre, & Source Localization

 


  

Melody - Pitch & Temporal Contours - Rhythm - Harmony - Tonality

 

 

(common melodic "seeds" or "cells"; after Fuentes, 2020)

 

 


Example of melody generation using the LSTM* and Markov Chain modeling (after Bihani et al., 2023)
[*LSTM: Long Short-Term Memory, a type of recurrent neural network (RNN) designed to handle long-term dependencies in current sequential data.]

 
Melodic Units

The units of a melody are not individual notes. During a piece of music, perception produces a coded version of the continuous flow of sound based on how our implicit and explicit rules interact to identify accents and creatively identify local or global features and boundaries. 

A melodic unit refers to a distinct section of a melody; a small, self-contained set of notes that contributes to the overall melodic line.
A motif (or motive), a basic such unit, is a short, identifiable pattern of notes; a "motivating idea...the small cell out of which the music evolves." [ in Berry, 1986 ]. Out of the many short melodic gestures in a piece of music only the few that figure prominently in its growth are designated as motifs.
The figure to the left illustrates 24 common melodic motifs that can be combined in various time organizations.

Larger melodic entities include "phrases" and "periods," made out of multiple motifs that often end with a cadence (i.e. resting point). Listen to 20 iconic film-music motifs.

Analogously to language phrases, melodic 'phrases' are organized in structural units partly in terms of breathing restrictions (3-5 seconds per phrase). 
In contrast to language phrases, melodic 'phrases' are also organized in terms of patterning & redundancies (repetitions).
Redundancy fosters predictability which, in turn, gives rise to expectation.
Consequently, examining melodies as sets of distinct, isolated notes is inadequate because it does not allow for the patterning that supports the experience/prediction/expectation/game-of-expectations affective cycle (more in Module 6). 

When we break a whole down to a series of isolated bounded units (known in mathematics as a Markov chain* of order 0) we destroy the level of meaning that is based on syntax; that is, on the relationship among units that binds them into a larger, single, bounded event.

In relevant experiments, participants have been asked to contribute 1, 2, 3, etc.. words to a text after reading 0, 1, 2, 3, etc. of the already existing words respectively. [ e.g. Miller & Selfridge, 1950 - a similar experiment was conducted by Davies (1978), using notes instead of words. ]
The more words the participants were able to read (i.e. the more developed the syntax & context) the more coherent and meaningful were their contributed text and the resulting phrase.
In other words, the more we are able to reflect on the past, act in the present, and project on the future, the more meaningful our actions.

Such experiments
a) illuminate the statement "the whole is different than the sum of its parts"; the meaning of phrases is not just an aggregate of the meaning of the included words;
b) highlight the feedback relationship between experiencing time and experiencing music (explored later in the course), as it relates to memory and expectation.
[ We experience 'time' as a tension between what has been (longing/regret for the past) and what might be (excitement/fear for the future), in negotiation with the 'present' experienced by us and others. ]

In both language and music, meaning is not simply lexical (i.e. it does not simply depend on the words/notes used); it is also syntactical. It depends on the way the words/notes are patterned together; on how they relate to their past (previous words/notes) and what they anticipate as their future (words/notes to come).

Music syntax outlines tension-resolution patterns that are creatively followed by listeners, as a piece of music unfolds in time, and support the (largely implicit) assembly of individual notes into logical structures (i.e. chunking). Syntactical relationships may be set up by the composer but are always mediated by what the listener/reader brings to the experience.

Melodies are not random collections of notes [ i.e. not Markov chains of order 0 ] but larger units, tied together by boundary-providing accents and syntactical rules, set up by the composer AND configured by the listener(s). The bounded musical units are outlined by accents and incorporate redundancies, with periodic repetition being key to musical syntax communication and perception.

[ * A Markov chain is a mathematical concept that outlines the probability of a system to move towards a particular new state, given its current state, named after 19-20th century Russian mathematician A.A. Markov. The concept has been instrumental to major scientific and technological developments, including nuclear power, search-engine functionality, and artificial intelligence. Watch this fascinating historical account. More here. ]

Sound events in time are organized in terms of pitch, duration, etc. contours that track accents and outline potential musical units.

 

Archetypal pitch contour examples

 


 


Descending major scale vs. "Joy to the World" 

 

 


 
Melody - Pitch & Duration Contours - Rhythm

Melody can be thought of as the superimposition (layering/stratification) and interaction of pitch (or melodic) and duration (or temporal) contours.

Pitch (melodic) contour describes the pattern of pitch direction changes within a melody: the pitch can either go up, down, or remain the same. Points of change in pitch direction (i.e. pitch contour inflections) are perceptually salient. They correspond to contrasts, which become the accents that outline bounded melodic units.
Pitch-contour periodicity outlines the pitch-contour clock.

Pitch contour similarity corresponds to similarity in the sequence of pitch contour inflections, even if the interval sizes differ. Pitch changes in a melody that do not alter the pitch contour are perceived as melodic variations rather than new melodies.
On first hearing, it is the pitch contour that is stored in memory, with the exact pitches being subsequently 'assigned' to this contour [ e.g. Graves et al., 2019 ].

Several archetypal contours are illustrated to the left.
[ Additional contour archetypes such as "gap-fill" and "changing-note" (examples in Meyer, 1973; Rosen & Meyer, 1982) are directly related to gestalt principles of perception, addressed later in the course. ]

If a pitch contour includes large pitch leaps (e.g. leaps of two or more scale steps), the resulting melody is considered disjunct. Otherwise (e.g. for single-step leaps in the scale), the resulting melody is considered conjunct.
Disjunct melodies tend to be more difficult to remember and to reproduce. Melodies with a balance between conjunct and disjunct portions are usually preferred and are judged as more interesting/exciting [ e.g. Brantingham, 2013 ].
For a balanced example, listen to "Here, There and Everywhere" by The Beatles

Duration (temporal) contour describes the pattern of sound and silence duration changes: a sound/silence event can become longer, shorter, or remain the same. Points of change in 'sound-event' duration (i.e. duration contour inflections) are perceptually salient. They correspond to contrasts, which become the accents that outline bounded temporal units.
Duration contour periodicity outlines the duration contour clock.

Duration contour similarity corresponds to similarity in the sequence of duration contour inflections, even if the actual durations differ. Duration changes in a melody that do not alter the duration contour are perceived as melodic variations rather than new melodies.

According to Dowling's melodic contour theory [ in Dowling, 1978 ], at first hearing melodies are coded based on their pattern of accents and are remembered in terms of the sonic motion implied by their pitch and duration contours. NOTE: contour inflections become accents after they have occurred, pointing to the importance of the temporal aspects of musical organization.
 

Beat / Meter / Tempo
Every duration contour outlines a beat (i.e. underlining pulse of a piece of music). This is organized into a meter (i.e. repeating accented patterns) that reflects the interaction among duration contour shape, clock, and beat.
Tempo describes the number of beats per unit time. The average spontaneous motor tempo (SMT) has been measured at 100 beats per minute (bpm) but with considerable individual variations [ e.g. Fraisse, 1982, in Wearden, 2024 ].
 

Duration vs. pitch contour salience

Numerous studies suggest that duration contours are more salient than pitch contours and may be more important in melody coding [ e.g. Monahan et al., 1987; Kendall & Carterette, 1990; Palmer, 1996; Schmuckler & Moranis, 2023 ].

For example, a descending diatonic scale and the opening of "Joy to the World" have identical pitch contours but can be easily recognized as two different 'tunes' because of their different duration contours.

Consider "America the Beautiful" (by Ward & Bates - Ray Charles's rendition).
The duration and pitch contours are aligned (in 3s). If we superimpose a syncopated duration contour in 2s (tango), the piece becomes unrecognizable, even if the precise sequence of pitches remains unchanged (after Kendall & Carterette, Unpublished).
Changing the duration contour of most melodies will likely result in what is perceived as new melody, even if its pitch contour is kept intact, (e.g. the "Star Spangled Banner" modification presented in class).

Pitch contour changes may also be sufficient to create a new melody (even if the duration contour is kept intact), particularly when the changes result in a:
a) significantly more disjunct/conjunct melody and/or
b) misalignment between pitch and temporal contour clocks.
(e.g. "Star Spangled Banner" vs. "Happy Birthday"). 

 
Rhythmm
is a collective property of a piece of music, emerging out of the combination and interaction of all available sonic-contrast patterns and corresponding contours (i.e. pitch, duration, dynamic, and/or timbral) within the piece.
[ see Bronzini, 2024 for an outline of rhythm from a music theory perspective ]
[ see Ravignani et al., 2017 and McAuley, 2010 for outlines of rhythm from a music cognition perspective ]

Theme variation often involves changes in some of the contours while keeping the others relatively unchanged. This results in a variety of melodic phrases that are still understood as parts of the same composition.

During musical performance, micro-deviations from the written/precise contours function as a performer's means to communicate 'expressive intent.'
[ Listen to the verse's snare-tom pattern in "Ticket To Ride" by The Beatles; Live at the Hollywood Bowl, 1965. ]

The importance of duration, dynamic, and other contour micro-variations to musical expression has been documented extensively.
[ e.g. Gabrielsson, 1988; Kendall & Carterette, 1990; Palmer, 1996, Juslin, 2000 ].

Kendall & Carterette (1990), for example, first illustrate that musical expression works like all communication: performers shape sound in ways that listeners can reliably interpret. They then reveal that temporal micro-variations —tiny speeding up, slowing down, and stretching of notes— are more significant in communicating expression than small changes in pitch or loudness. Even when other sound qualities were simplified, listeners could still largely detect differences in expressiveness based on timing patterns alone. In other words, expression in music seems to be carried primarily by how notes are timed.
Drawing an analogy to speech communication (i.e. to communication via linguistic performance), how something unfolds in time (e.g. pauses, pacing, hesitation, rushing, lingering) often communicates more emotion than raw content alone (e.g. exact words/notes).

 

 

 

 

 

Pitch & Duration Contour Interaction

It is easier to remember melodies whose pitch and duration contour clocks 'line up'.
[ e.g. Monahan et al., 1987 ]

'Frére Jacques' (French children's song) opens with a simple and conjunct pitch contour, that has a regular clock (4 notes per repetition).
The duration contour is also relatively 'flat', with only minor duration changes (most pitches are quarter-notes with no rests inserted), and the simplest possible periodic contour clock (1 note per repetition).
The resulting melody involves no conflict between pitch and duration contours. Pitch and time accent structures align, resulting in a simple and, at some level, uninteresting melody.

Similarly, 'Three blind mice,' incorporates a conjunct pitch contour with regular pitch contour clock (3 notes per repetition) that aligns perfectly with the duration contour clock (also 3 notes per repetition), resulting in another example of a simple (and dull) melody.

Such pieces, along with several nursery and pop tunes, are extreme examples of simple melodies. In most other cases there is no perfect contour alignment or clock periodicity, at least not for the entire duration of a piece or theme, avoiding a level of redundancy high enough to make a piece not only easily codified but also uninteresting (remember the relationship between complexity and preference).

 

 

The "Star Spangled Banner", for example (Whitney Huston's 1991 Super Bowl performance), layers a pitch contour clock in 2s over a duration contour clock in 3s, resulting in a relatively complex and interesting melody. [ Watch Jimmy Hendrix in his iconic 1969 Woodstock performance ].

"Yesterday," by The Beatles, involves a more complicated relationship between pitch and duration contours, with clocks that are not perfectly periodic and melody lines that are frequently disjunct (i.e. increased complexity). Nonetheless, it is possible for listeners to trace an overall arch-like contour that becomes the piece's signature (i.e. decreased complexity). This high/low complexity balance helps sustain both high interest and high preference levels.

Layering multiple non-aligned pitch and duration contours results in polyrhythms.
Musical pieces in Jazz, renaissance, 20th century, and several non-western traditions employ polyrhythms of varied degrees of complexity.
Given that contour repetition facilitates memory [ e.g. Monahan et al., 1987 ], pieces with complicated contour relationships compensate for the information overload through repetition.
[ e.g. "Star Spangled Banner"; "Yesterday" (The Beatles); "Theseus and Minotauros" (Daedalus Project); "In the Mood" (Joe Garland); "A Night in Tunisia" (Dizzy Gillespie) ]

 
Perceptual dimensions of melodies, rhythms, and entire pieces

_ Single melodies can be understood in terms of primarily duration and pitch (and secondarily dynamic and timbral) contour layering and
_ Musical pieces can be understood as the layering of one or more melodies and/or rhythms (or, more generally, one or more "sonic structures'), with contour clocks
   of various degrees of periodicity [ see Patel, 2007; section on Melody Statistics and Contours ].

Theorists have proposed several hierarchy-based music analysis methods to address single- and multi-layered musical syntax [ e.g. Schenker, Forte, & Deutch; in Tan et al., 2010a ]. From the perceptual and cognitive perspectives, it is all about the interaction among pitch, duration, dynamic, and timbral contours. 
[ Detailed discussions on melody perception in Butler, 1992 and Tan et al., 2010b ]
.

Identifying the perceptual boundaries (accents) that help define contours depends in part on previous learning and experience. For example, a musical listening task that appears simple and clear cut to a native of Indonesia may appear complicated and unorganized to a Western trained listener, and vice versa
[ E.g. 'Jaya Semara'. Indonesian Gamelan for Kebyar gong. UCLA Gamelan Ensemble ].

Whether a pitch or temporal inflection will constitute an accent (i.e. a salient point) depends not only on physical contrasts but also on implicit rules of data organization that utilize personal, stylistic, or cultural schemas acquired through experience.
In addition, when faced with melodies that exhibit recognizable patterns on several levels, listeners will organize them by focusing on the most salient level, within the given context.

Implicit rules assist us in coding accent patterns, resulting in a melody being remembered in terms of the contout-implied melodic and rhythmic motions.

Motion indicates more than shape (contour); it indicates time. As previously noted, contour inflections become accents after they have occurred, highlighting the importance of the syntactical and temporal aspects of musical organization.

Stimuli used to assess neural response to harmonic anomalies (after Maess et al., 2001).
 (A) A sequence of five in-key consonant chords (key of C), with the fifth chord highlighted in green.
 (B) A sequence of the first four chords in (A), followed by a fifth chord that contains two in-key notes
(F and E) and two out-of-key notes (A flat and D flat), highlighted in red.
The four chords preceding the fifth chord set up a harmonic (syntactic) expectancy in the listener, which the fifth chord fulfills in the case of (A) but violates in the case of (B).
 

   

 

 

 

  


 

 

 

 

(Jazz harmony details here)

    
Harmony

Musical harmony refers to the simultaneous combination of two or more notes and to the progression of such combinations in time. Two-note combinations are called dyads or harmonic intervals. Combinations of three or more notes are called chords (they include triads, four-part harmony, etc.).

Listener experience and culture-dependent music theory rules, outline chord relationships that support the creation of chord progressions with specific harmonic perceptual impact.
As chord progressions unfold in time, they set up expectations that composers/performers can partially or completely fulfill/violate. This process communicates patterns of tension and release that, at a basic level, constitute music's intrinsic meaning.

Cultural musical norms and the music theory rules that codify them also outline pitch relationships, which support melodies whose implied chord progressions have harmonic perceptual impact as well.

Harmony perception involves cognitive processes that, along with the auditory periphery, engage brain regions responsible for prediction, emotional response, and learning.

At the auditory periphery level, harmonies and harmonic progressions are analogous to isolated timbres and timbre contours, respectively, and are processed as time-variant spectra.

At the neural level, music interval and chord recognition engages the brain's auditory cortex AND the dorsolateral prefrontal cortex, which is involved in cognitive control and the balancing of emotional and deliberative responses.

Harmonic perception relies heavily on working memory to retain previous chords and anticipate upcoming ones. The prefrontal cortex and hippocampus are involved in storing and updating harmonic progressions in real time, helping generate musical expectations and emotional responses based on the outcome of these predictions.

Harmonic structures also activate reward systems in the brain, reflecting whether harmonies meet or defy expectations. The dopaminergic system responds to these predictions, producing pleasurable sensations when musical surprises (e.g. unexpected modulations or cadences) are resolved satisfactorily/plausibly.

At the cognitive level, harmony perception relies on pattern recognition and tonal schemas, developed as internal representations of common tonal structures (e.g. major or minor scales). When chords and harmonies match these schemas, the brain processes them easier. It able to predict the temporal unfolding of chord progressions based on familiarity with musical styles (e.g. a dominant chord leading to a tonic resolution).
Departures from established schemas violate stylistic expectations (see the figure top-left) and elicit emotional responses that are manifested and confirmed both behaviorally and neurologically [ e.g. Limb, 2006 ].
 

The cognitive effort required to process harmony depends on the musical structure's perceived complexity.
Complexity can arise from:
   _ dissonance,
   _ ambiguous tonal centers,
   _ unpredictable modulations or progressions, or
   _ some other unexpected harmonic feature.
The degree of perceived complexity largely depends on a given listener's context (e.g. personal and cultural previous experience).

Simple harmonic structures rely on a limited menu of triadic chords (e.g. tonic, dominant, and subdominant - in the simplest case: C, G, and F triads in their root or inverted versions, and common progressions with predictable resolutions. Such features impose minimal cognitive demands, even for untrained listeners. By aligning with learned schemas, they typically evoke familiar and comforting/dull emotions.

As expected from our discussion on the relative nature of preferred complexity level, listeners without formal musical training or sufficiently varied listenings tend to prefer structures that accommodate their harmonic cognitive bandwidth.
Most harmonically simple songs that enjoy widespread success make up for this simplicity with complex arrangements, varied contours, and skillful, "signature" performances, raising the average complexity of the experience.
[ e.g. "Get Back" by The Beatles;  "Free Fallin" by Tom Petty;  "Songbird" by Oasis;  "La Bamba" by Richie Valence / performed by Los Lobos ].

Complex harmonic structures use an extensive menu of chords (e.g. 7ths, 9ths, diminished, hybrid, modal), chromaticism, modulations across tonally distant keys, and ambiguous/unstable tonal centers. Processing such features requires greater cognitive effort.
Working memory must track multiple harmonic layers, and the listeners are called to make predictions in vaguer contexts.

Complex harmonies can evoke more nuanced emotional responses that include tension, surprise, or even discomfort. Experienced listeners may find pleasure in resolving such emotions while novices may feel overwhelmed or alienated.
Most harmonically complex songs that enjoy widespread success balance this complexity with repetition, less cognitively demanding contours, etc..
[ e.g. "Maybe I'm Amazed" by Paul McCartney; "Josie" by Steely Dan; "Never Gonna Let You Go", by Sergio Mendez ].
Occasionally, songs succeed in capturing the listeners' attention and imagination in spite of maintaining complexity at multiple levels.
[ e.g. "Strawberry Fields Forever" by The Beatles ]

A listener's cultural tradition supports the development of cognitive frameworks that shape harmonic perception. Perceived degree of musical complexity influences cognitive and emotional responses, always within these frameworks.

As is the case with most aspects of human experience, the more varied a listener's musical exposure & practice, the broader the gamut of musical experiences that may be considered interesting, preferable, and pleasurable.

 

 

 

 

 
Tonality / Consonance-Dissonance / Non-Western Music

Tonal harmony provides a learned context that guides perceptual organization of melodic and harmonic passages and supports operation of the various gestalt principles of perception (addressed later in the course). The concepts of tonality and harmony are perceptually interdependent and weave an important framework for music making and listening.

Tonality describes the organization of pitch relationships around a central tone (key/tonic).
It outlines how likely or unlikely a note is to be included in a melodic or harmonic passage, given the notes that have already been played/heard [ Patel, 2007; Melody Statistics and Contours ].
[ Review of studies exploring the neural basis of tonal processing in Asano et al., 2022. ]

The complete set of rules that guides tone relationships within a given key is unclear but there is consensus on some of them (e.g. tonic-dominant relationship; major/minor modalities; musical consonance/dissonance contrasts; cadences).

Geometric models of pitch relationships such as Shepard's pitch spiral, Krumhansl's pitch cone, or various versions of Pythagoras's circle of fifths (see the figure bottom-left) represent attempts to describe tonality's organizational principles in terms of tonal hierarchies. Within the context of tonality, the term harmony refers to
a) the range of melodic and harmonic expectations outlined within a given key &
b) the specific, key-based melodic and harmonic implications of what has already
    been performed.

Musical pieces that follow tonality rules give a sense of direction/progression, particularly but not exclusively to listeners familiar with the underlying tonal framework, supporting the game of expectations that seems to be at the basis of the affective potential of all experience, musical or otherwise.

Musical consonance (representing stability/release) and dissonance (representing instability/tension) depend on many variables and are often defined in the context of tonal hierarchy. In harmonic passages, consonance also describes the degree of a harmonic interval's/chord's pleasantness, fittingness, and/or perceptual smoothness.

Musical consonance/dissonance judgments map onto tension/release judgments.
Within the Western musical tradition, musical tension/release judgments are linked to contrasts in several aspects, including:
_ tonal center (e.g. key);
_ sensory consonance/dissonance (degree of beating & roughness sensations);
_ dynamics; pitch; rhythm; timbre; orchestration;
_ performance techniques.

Numerous musical traditions base melodic and harmonic development on musical consonance/dissonance contrasts, with musical pieces structured around musical consonance/dissonance 'contours.'

For example, in the opening of Leonard Bernstein's "Maria" (from West Side Story), the voice moves from a tritone melodic interval (dissonance) to a 5th (consonance). Later on in the song, the voice follows a similar pattern, moving from a major 2nd (dissonance) to a major 6th (consonance). One can think of these passages as having a similar consonance/dissonance 'contour.'

 
When listening to unknown pieces, listeners trained in the Western musical tradition make tonal assessments swiftly and often implicitly, determining tonality based on the absence as well as presence of certain intervals. Tonally knowledgeable listeners are able to extrapolate tonal harmony information from incomplete tonal cues. Any tone one hears may suffice as a tonal center, until the listener is probed by additional tonal evidence to opt for a more plausible choice.

Several studies on melodic motion and tonality judgments provide convergent evidence that the combination of consonant and dissonant harmonic intervals conveys a tonal center more efficiently than the use of consonant intervals alone [ e.g. Butler & Brown, 1984 ].

In atonal contexts,* intervalic similarity among chords does not translate to perceptual similarity. This observation is consistent with the experienced dissimilarity between major and minor triads in tonal music, both of which have the same intervalic content, just in different distributions.
Atonal chord similarity/dissimilarity judgments are likely based on a combination of musical context, previous experience, and the spectral distribution of the chord-signals in question, as atonal music has no hierarchical tone structure on which to base similarity comparisons.
[ *atonal context: melodic/harmonic context where all available notes are equally likely to occur ]

Most musical traditions in the world employ hierarchical tone structures analogous to, even if quite different from, Western tonal harmony [e.g. Indian ragas; Arabic maqams; Indonesian gamelan ]. Tone hierarchies within each individual musical tradition are more salient to members of the same tradition than to "outsiders." At the same time, different traditions make equally strong claims of "good" musical organization for widely differing musical structures.
Tone hierarchies, musical scale systems, perceptual similarities/differences, music pattern recognition, tension/release judgments, aesthetic judgment standards, etc. may therefore be largely culture-dependent, although they are processed through common cognitive principles.

Globalization has increased exposure to diverse musical styles. Many listeners now enjoy music that combines harmonic, melodic, tonal, and rhythmic elements from different traditions. However, agreement on the intellectual and affective response to hybrid styles still relies on the degree of shared cultural familiarity and exposure.

  

Cognitive Aspects of Loudness, Pitch, Timbre, & Source Localization

 


LOUDNESS COGNITION

Auditory Masking

The term masking describes the ability of one tone (masker) to cover (i.e. render inaudible or raise the audibility threshold of) a second tone (maskee), depending on their relative and absolute levels: the more intense tone may mask the less intense tone and the stronger the intense tone the broader the masking bandwidth.
The smaller the frequency difference and the larger the level difference between two tones, the more likely it is for masking to occur. In addition, the masking effect is asymmetrical with respect to frequency: the masking bandwidth extends more into frequencies above the masker rather than below.

Simultaneous Masking occurs when two tones close in frequency are presented simultaneously. It has unambiguous physiological roots in the auditory periphery.

In Forward Masking, a signal masks a tone that comes 0-200ms after it. Forward masking does not produce the broad masking effects of simultaneous masking. It increases with masker duration and level and occurs most effectively for masker-signal delays ~20-30ms. Forward masking depends on the signals used and may be due to masker activity persisting at some level in the auditory system, impacting signal perception.

In Backward Masking, a signal masks a tone that came 0-50ms before it. Backward masking effects are even slighter than those of forward masking and are more prominent in untrained listeners. Backwards masking depends on the signals used and may be due to higher level cognitive processing, involving an interaction between analytic & synthetic listening operations.

Backward and forward masking examples.
Auditory Masking Section - LMU: RECA220

(simultaneous masking)


(backward and forward masking)

Backward masking may be due to higher level cognitive processing, analogously to the above examples (source)

 
PITCH COGNITION

Our discussion on analytic & synthetic listening addressed several examples where auditory sensation and contextual information interact to support unexpected cognitive responses to pitch changes [ e.g. Smoorenburg's pitch direction experiment and Deutsch's scale illusion experiment ].
Additional examples include the perception of Absolute Pitch, Pitch Circularity, and the Tritone Paradox.

Absolute/Perfect Pitch vs. Relative Pitch

Absolute pitch (AP) refers to the ability to identify (passive AP or AP recognition) or even produce (active AP or AP recall) a specific pitch in isolation (i.e. in the absence of any reference pitch and independently from any other attribute of the sound, such as timbre). This ability is not necessary to musical activities and may even be disruptive to performers / listeners who possess it. [ e.g. Miyazaki, 2007 -- Watch a video outline of AP by Rick Beato ]

AP is an example of implicit musical processes at play, describing an ability that is partially genetic and partially learned. There is converging evidence suggesting that it may reflect a universal genetic propensity, evolved for/through language processing and encouraged by environmental factors.
[ e.g. Zattore, 2003; Deutch et al., 2004 ]
Studies exploring the neural activity accompanying AP perception indicate that AP possessors engage the same cortex regions for all tonal tasks.
[ e.g. Schulze et al., 2009 ]


Constructivist psychologists have argued that the AP ability develops during the preoperational stage of cognitive development (~2-7yrs; one of the 4 developmental stages postulated by Piaget), when the automatic acquisition window of language is still open. Formal and consistent contact with sounds from the Western (or any other) tuning system during this stage is thought to mentally imprint absolute pitch relationships.

The idea that, while the window of language acquisition is open, implicit knowledge is extracted automatically from the environment is explored by the Suzuki approach to children's music learning. In his approach, children do not learn music through symbols or formal rules. Rather, along with their parents, they learn to express themselves musically through interaction and feedback. Suzuki's training method has proven rather successful, and children trained with this method are very quick to later acquire the relevant formal knowledge.

Relative pitch refers to the ability to identify a specific pitch based on its distance (interval) from a reference pitch, which may be concurrently present or recalled from short-term memory. Relative pitch is an example of an explicit musical process and is necessary to musical activities.

For harmonic complex tones, relative pitch may be a special case of spectral contrast recognition. Abilities that resemble passive AP, but are only demonstrated for specific instrument(s) per musician, may be based on recognition of memorized timbral differences among different pitches for a given instrument. 
In the case of pure tones (spectra with only one component), aural harmonics (i.e. harmonics introduced to the sound by the way the inner ear responds to sine signal inputs) may enrich the spectra of the pure tones enough to facilitate spectral similarity/difference comparisons.

 

(Pitch Circularity by D. Deutsche)
 

                           (source)                                           ("Ascending-Descending" - Escher, 1960)  

 
Pitch Circularity

Pitch circularity is an auditory illusion where a series of tones (or a continuous frequency glide) appear to endlessly rise/drop in pitch. It is also referred to as the barber pole illusion, drawing an analogy to the visual illusion of a rotating barber pole, where the lines appear to move endlessly towards one direction.

Pitch circularity reveals that pitch is multidimensional with at least three dimensions, involving pitch height (one dimension: frequency) and pitch chroma/class (two dimensions, separating pitches within an octave and linking pitches an octave apart).
 

Within a single octave, the octave interval represents the largest physical distance in terms of pitch height but the smallest perceptual distance in terms of pitch chroma, whose dimensions capture the circularity in pitch perception. 
Pitch perception wraps on the octave, with scales defining sets of different pitch chromas that repeat at different pitch heights for each new octave. This is best represented by a pitch spiral (versus a pitch scale).

The perceptual circularity of pitch is explored in Shepard-tone scales and Risset pitch slides. Named after American cognitive scientist Roger Shepard and French composer Jean-Claude Risset, respectively, they present the paradox of a continuously ascending/descending pitch.
Listen to two pitch spiral examples (after Houtsma et al., 1987).
Pitch spiral scales/slides are the auditory analog of the continuously ascending/descending staircases, explored conceptually by mathematicians Lionel and Rodger Penrose and artistically by Escher.

The spectral distributions of the complex tones used in pitch circularity demonstrations include octave-separated frequency components (e.g. 1f, 2f, 4f, 8f, 16f) that are shaped by a fixed, inverted-u spectral envelope and are gradually shifted up or down in frequency. As they shift, some components exit one side of the spectral envelope while new ones enter.

Pitch circularity is a 'cognitive geometry' that does not exist in the physical world. In the context of music listening, this cognitive geometry constitutes a 'cognitive reality,' where implicit knowledge structures appear to track both, pitch height and pitch chroma.

[ The same cognitive geometry can also explain the psychoacoustic phenomenon of the pitch-shift effect. ]
 

_ Watch this pitch circularity demonstration/explanation.
_ Risset extended the concept to create an analogous tempo circularity illusion.
_ Watch Shepard illusion examples in sound design and experiment with this Shepard tone generator.
_ Watch this Escher waterfall model in action[ what's the trick? ]

 

(tritone spectral envelopes, after Malek, 2018)


Tritone Paradox

The tritone paradox is a musical illusion that occurs when two tones separated by a half octave (also called a tritone, an augmented fourth, or a diminished fifth), are played in sequence. Some hear the tones as ascending, while others hear them as descending. Linguistic background of the listener seems to be a factor, with Californians and people from the south of England consistently hearing the tritone paradox differently.

Diane Deutsch at UCSD was the first to describe the trirone paradox in 1975. She and colleagues have since explored the issue extensively, arguing that the speech patterns to which we have been exposed can influence how music is perceived.

It's important to note that the illusion occurs when the stimuli are Shepard tones, fact that may explain the pitch movement ambiguity observed in the tritone paradox [ explanation challenged by Malek, 2018 ]. Consider that:

  • the tritone interval is positioned at the center of the pitch chroma cycle explored in Shepard scales

  • the illusory perception of continuous rising/dropping pitch in Shepard scales implies that, at some stage in the accompanying spectral shift, the pitch has to jump down/up to maintain the continuous rise/drop illusion.

  • the central point in the spectral shift is the most ambiguous perceptually and the best candidate for the pitch jump necessary for the Shepard illusion to work;

  • the pitch shift effect also includes an analogous ambiguity at the mid-octave point
    [ e.g. Vassilakis, 1998 ].

 

Further Reading

Butler, D. (1992). "Cognitive Aspects of Pitch in Music" (book chapter)

 

 
 
TIMBRE COGNITION

Categorical vs. Continuous Timbre Perception

Categorical Timbre Perception

The same instrument may produce notes with widely differing spectral envelope shapes, depending on performance technique, intensity level (see figure to the left), or register, but will most likely retain its timbral identity, suggesting that timbre perception may be categorical.

For example, the signals of low, middle, and high pitched notes on a violin have very different spectral envelopes but will, in general, continue to be identified as belonging to a single instrumental timbre category, that of a violin. 
Conversely, imposing the same spectral envelope and overall spectral distribution across the playing range of a single instrument results in tones that cannot convincingly convey the instrument's identity.

Listen to the sound of a violin playing C4.
Now listen to the same sound transposed up to C5 or transposed down to C3.
The examples of the two new tones retain the C4 spectral envelope. 
Do the transposed tones convincingly convey the violin's timbral identity?
Spectral evolution video of a similar example on the bassoon.

 
When we listen to a familiar type of musical ensemble, for instance, we are able to identify whether a sound is coming from this or that instrument, even if all instruments perform in a similar pitch range and/or in sync. Successful categorization is based less on explicit analysis of acoustic features (e.g. spectral envelopes) and more on learned associations and conceptual representations of each instrument's unique sound identity, influenced by cultural exposure and individual experience (musicians do develop finer discrimination abilities for the instrument(s) they've been trained on).

Listeners in different musical traditions may categorize instruments or sound features differently. For example, in Western classical music, timbre categories are often associated with families of instruments (strings, woodwinds, brass, etc.), while in other cultures, organizing principles may be based on performance techniques (blown, plucked, struck, etc.) or contexts (work, war, ceremonies, etc.).

Continuous Timbre Perception

Several studies that gradually morph signals from one instrumental spectral distribution to another have shown that, perceptually, the timbre does not abruptly move from the first instrument to the second at some fixed point in the morphing stage but appears to transform perceptually in a gradual manner, as does the physical stimulus. This suggest that timbre perception may be continuous

In this sound morphing example, a C4 tone played on a French Horn gradually (in 10 steps) morphs into a C4 tone played on a Bb Clarinet
[ after Butler, 1992 ]
.  Does the transition from French Horn to Clarinet seem gradual or abrupt? 

Based on such observations, the previously discussed ability to group together the widely different spectral distributions of different notes on the violin under a single "violin" timbre category must be based on higher level cognitive processing, guided by our experience with an instrument's sound throughout its playing range. This claim is supported by studies that show a larger decline in timbre identification with changes in pitch register for unfamiliar versus familiar instrument sounds. 

So, timbre perception appears to be more continuous for unfamiliar sounds, based on explicit rules, and more categorical for familiar sounds, based on implicit rules. Furthermore, even during categorical timbre perception, continuous timbre perception allows us to perceive and explicitly track subtle variations (e.g. vibrato, articulation, spectral time variance) within a given timbral category.

 

 
Multidimensionality of Timbre

In a seminal timbral similarity study, Grey (1975) revealed three primary physical dimensions along which timbral judgments are made:
    a) narrow vs. wide spectra;
    b) coherent/synchronous vs. independent/asynchronous spectra; and
    c) low- vs. high-centroid attack spectra.
Instrument identification experiments support the timbral clustering in the figure to the left [ Grey, 1975; in Butler, 1992 ], but reveal asymmetries in identification confusion (e.g. the bassoon is confused for a French horn and the saxophone is confused for an English horn but not the other way around).

Experimental studies that involve melodic and harmonic contexts [ e.g. Grey, 1977 ] suggest that perceptual strategies for timbral recognition and discrimination are varied and depend on:
    a) whether spectral or temporal characteristics of a tone are more pronounced
        (e.g. tones corresponding to continuous versus impulse signals),
    b) whether or not other instruments/sources are present, and
    c) attention shifts and larger musical/sonic context.

[ Look over this set of timbre spaces based on a combination of psychoacoustic and cognitive studies. For creative explorations of timbre spaces see "spectral music" ]
 

Musical Context and Timbre

A series of studies by Kendall and colleagues [ e.g. Kendall et al., 1999 ] provide further support to Grey's claim that the perception of timbre also depends on musical context.

Different strategies are being employed depending on the types of tones in question and timbre recognition / discrimination is based on different acoustical cues depending on whether the tones in question are in solo and static, versus multi-instrumental and time-variant contexts. [ e.g. McAdams, 2013 ]

[ Outline of Timbre Spaces & Semantics (printable copy) ]

 
Musical Texture / Timbre / Emotion

Timbre plays a crucial role in shaping musical texture, contributing to the layering and interaction of sounds in a piece. Analogously to timbre, sonic textures appear to be linked to spectral information gleaned from the auditory periphery via statistical processes [ e.g. McDermott & Simoncelli, 2011 ].
Timbre helps define a musical texture's complexity (via timbral density and variety) and expressiveness (via timbral associations and/or contrast contours), playing a key role in how we perceive and emotionally respond to music [ e.g. McAdams et al., 2017 ].

Timbre's precise contribution to music's affective impact remains elusive due to the challenge of isolating it from other musical parameters. Studies recording brain responses' ERPs (event-related potentials) suggest that the brain’s capacity to process emotional information from musical textures and structure is implicit: automatic and pre-attentive [ e.g. Spreckelmeyer et al., 2013 ].
 

Further Reading

Butler, D. (1992). "Cognitive Aspects of Timbre"  (book chapter)
McAdams, S. (2013). "Musical Timbre Perception." (book chapter)
Siedenburg et al. (2016). "Acoustic and Categorical Dissimilarity of Musical Timbre: Evidence from Asymmetries Between Acoustic and Chimeric Sounds."
Siedenburg et al. (2019). "The Present, Past, and Future of Timbre Research."
Wei et al. (2022). "A Review of Research on the Neurocognition for Timbre Perception."


(source)

 
SOURCE LOCALIZATION COGNITION

Precedence Effect (Haas effect / Franssen effect)

The precedence effect describes a learned strategy employed implicitly by listeners in order to address conflicting or ambiguous localization cues occurring in environments where sound wave reflections play an important role (e.g. all rooms other than anechoic environments). This strategy is believed to develop through exposure to reverberant listening contexts and is eventually applied automatically to all listening contexts, leading to possible auditory localization "illusions."

According to this effect, listeners make their localization judgments based on the earliest arriving sound onset. The term "precedence" is used to indicate that the direct sound, with presumably accurate localization information, is given precedence over the subsequent reflections and reverberation, which usually convey inaccurate localization information. In fact, in reflective/reverberant environments, both IPD and ILD cues are diffused to such an extent that they become unusable, making the precedence effect a necessary localization strategy.

The Haas effect describes the observation that, for arrival-times differences up to 30–40 milliseconds, the precedence effect persists (i.e. listeners localize the sound source based on the earlier arriving signal) even if the delayed signal is up to 12dB stronger than the first (see the image to the left and watch this demonstration).

The Franssen effect describes the observation that, an intense, fast-rising tone presented in a reverberant space via a single speaker will continue to be perceived as originating from that speaker even after it has gradually moved to a second loudspeaker, 45 degrees away.

Overview of the precedence effect in Hartmann, 1999.
More details in Litovsky et al., 1999.
Review of more recent studies in Brown et al., 2015.

 

Auditory Scene Analysis

Auditory scene analysis is a term coined by Canadian cognitive psychologist Albert Bergman in the 1980s to describe the process by which complex concurrent and sequential acoustic events are analyzed and organized into coherent auditory objects or “streams.”  Auditory scene analysis postulates models of perceptual mechanisms and cognitive abilities that segment auditory input into manageable chunks, integrate portions of related layers into streams, and segregate different streams, allowing attention to focus on each of them separately. The process has analogues in the visual system's organization of visual scenes but is more challenging to model due to hearing sense's transient and less localized nature.

Integration / Segregation

Integration refers to the process of grouping sounds that originate from the same source into distinct auditory streams. For example, it helps the brain group together sounds from the same voice or instrument into a coherent percept.
Segregation refers to the process by which the brain separates sonic streams that originate from different sources. For example, it helps distinguish a person’s voice from background noise, or one instrument from others in an orchestra.

Both mechanisms support auditory stream formation, a process by which the brain creates distinct perceptual “streams” of sound and allows us to track different sources of sound over time. They explain the cocktail party effect (discussed later in the course) and rely on a combination of:

  • bottom-up auditory cues, processed at the periphery [ primary auditory cortex ] and

  • top-down cognitive strategies, processed at more central brain regions [ the superior temporal sulcus (source identification), the planum temporale (sound sequence organization), and the prefrontal cortex (relevance determination) ].

(refresh your memory on "attentional capacity")

 
Auditory Cues
supporting auditory stream formation

Spectral cues:
   _ degree of harmonicity/inharmonicity;
   _ average and time-variant spectral envelope;
   _ formant structure; etc.
Temporal cues
   _ degree of spectral component attack synchrony;
   _ temporal coherence (e.g. periodicity/rhythm)
   _ temporal gaps (e.g. phrase segmentation)
   _ simultaneous and forward masking
Spatial cues:
   _ static (for location) and dynamic (for movement) interaural level / phase / spectral differences;
   _ degree of reverberation/echo
   _ perceived source distance changes

Cognitive Strategies supporting auditory stream formation

_ backward masking
_ Gestalt principles of perception (discussed later in the course)
_ top-down processing based on previous experience and expectations about how a source
   will sound within a given context.

 

Further Reading

Calcus, A. (2024). "Development of auditory scene analysis: a mini-review."
McAdams, S. & Bregman, A.S. (1979). "Hearing Musical Streams."
Bregman, A.S. (1990). "Auditory Scene Analysis: The Perceptual Organization of Sound."
Plack, C.J. (2005). "The Auditory Scene."
Shinn-Cunningham, B.G. (2020). "Brain Mechanisms of Auditory Scene Analysis."
Shamma, S.A. et al. (2011). "Temporal coherence and attention in auditory scene analysis."
Sussman, E.S. (2017). "Auditory Scene Analysis: An Attention Perspective." (review article)
Noble, J. (2019). "Auditory Scene Analysis." (blog - summary)

 

 


  

Loyola Marymount University - School of Film & Television