RECA-293 - How Music Rocks and Rolls
Cognitive Neuroscience of Music: Emotion, Meaning, and the Mind

MODULE 04A

 

cognitive processing in music listening
Experiencing Musical Time: Attention
Memory Limits - Categories - Accents - Schemas - Complexity
Implicit & Explicit Knowledge - Analytic & Synthetic Listening
Melody
    Melodic Units - Pitch & Temporal Contours - Rhythm
    Harmony - Tonality
Cognitive Aspects of Loudness, Pitch, Timbre, & Source Localization

 


  

Experiencing Musical Time: Attention

 

  

 

 

 

 

 

(adapted from Polti et al., 2018)

 

 


(attention guided by schemas)

 
Experiencing Time

Our experience of time seems to often proceed independently from clock/chronological time. Musical organization of sounds offers us means to creatively explore this difference.

The fleeting nature of all non-linguistic sound events (in music, film, etc.), due to their potential absence of a referent, highlights the broader challenge of defining the perceptual 'present.'

What we understand as the present (i.e. our sense of 'now') is shaped by our selective attention to (aspects of) events and can stretch/contract to degrees that significantly impact duration perception.
Our experience of present events is always mediated by the memory of past events and the expectation of future ones, while reshaping both our memories and our expectations.

The dialectic nature of the way we experience the present, past, and future is at the core of the human experience of time and the topic of extensive research from within philosophy in general and phenomenology in particular.

The ever-moving present is the only existentially 'real' dimension of what we experience as time, since it is the the only temporal dimension in which we may act.
An essentially temporal field of action, the present is experientially defined by the productive tension between our memory of past events and our anticipation of future events, in negotiation with the actions we and others take.
    Philosophy/Phenomenology Resources

    _ Ricoeur, P. (1998). "Critique and Conviction." Columbia University Press.
    _ Ricoeur, P. (1984-85). "Time and Narrative" (3 volumes). University of Chicago Press.
    _ Gadamer, H.G. (1989). "Truth and Method." Continuum Press.
    _ Gadamer, H.G. (1986). "The Relevance of the Beautiful." Cambridge University Press
    _ Kramer, J. (1988). "The Time of Music." Schirmer Books.
    _ Additional bibliography (Stanford Encyclopedia of Philosophy)

Human perception compresses time frames:
    _ Recent events are perceived as more distant and vice versa (telescoping effect).
    _ Shorter intervals are overestimated and, more so, vice versa (Vierordt's law).
    _ Working towards a desired goal shrinks perceived duration and vice versa
       (linked to the goal-gradient effect: motivation & effort increase with goal proximity).

 

Storage-Size (computational) Theories

Memory storage needs allegedly influence our estimates of time: a percept containing a large amount of information requires more storage capacity in short-term memory, generating the impression of greater elapsed time [ e.g. Ornstein, 1969; Sheth et al., 2000 ].

The idea is that our brain uses memory capacity as a metric for time duration perception. Information-rich events take up more "mental space" and, consequently, feel longer. According to the theory, events with high complexity --like musical pieces with busy arrangement, many details, and little repetition-- tend to be perceived as longer than simple events, due to the increased memory storage required [ e.g. Matthews et al., 2014 ].

In a version of this approach, the Scalar Expectancy Theory (SET) proposes that our perception of time is based on the accumulation of pulses within an internal "clock." The number of stored pulses represents the perceived duration, essentially acting as a storage size counter for time information
[ e.g. Gibbon, 1977 ].
Partially consistent with this model is the argument that melodic expectations operate within a limited-width time-window, restricted by short-term memory storage capacity limits. [ e.g. Dowling et al., 1987 ]

Drawbacks

The theory is challenged by difficulties in reliably measuring the "storage size" of memory, which is necessary to the theory's testing. It has also been criticized as simplistic, in that it does not account for factors like previous experience, attention, and emotional state.

More importantly, research on how event density influences perceived duration points to the opposite direction. Essentially, the more events happening within a timeframe, the higher the event density, which can often lead to a perception that the time passed more quickly and the events' perceived duration was shorter [ e.g. Lambrou-Kokolaki et al., 2024 ].

Regularity (i.e. reduction in new information) correlates with the perception of longer durations for listeners familiar with a musical style/piece [ e.g. Sasaki & Yamada, 2017 ] but may have the opposite effect on unfamiliar listeners (see Berger, 2014).
Slower tempos and irregular meters/rhythms make it difficult to accurately predict how much time has passed, leading to duration over/underestimations, experienced as 'losing track of time' (see Why music can make us literally lose track of time).

 

Attentional Capacity and Emotional State (cognitive) Theories

Assuming that perceptual activities compete for attentional center stage, the more we attend to some events -or aspects of events- the less we will attend to time itself, leaving our 'cognitive clock' with the impression of less elapsed time [ e.g. Hicks et al., 1976; Block, 1978; Polti et al., 2018 ].

Some models hypothesize an "attentional gate" that regulates the flow of temporal information into our internal clock. When attention is diverted, the gate narrows, reducing the amount of temporal information that can be processed [ e.g. Zakay & Block, 1997 ].

More recently, the impact of emotional state and context on time perception has been appreciated and systematically studied. Results have shown that anxiety, arousal, task difficulty, and skill level (musical or otherwise) are important factors.

For example, when people are excited or are pursuing a goal, they may feel like time is passing quickly. At the same time, being presented with the opportunity to earn a reward can make time seem longer [ e.g. Dawson & Sleek, 2018; see also Lehockey et al., 2018 ].

Issues of attention are central to our ability to partition our sonic environment into distinct streams and shift focus among them. The perceptual strategy (auditory scene analysis) and manifestations (e.g. the cocktail party effect) of such processes will be addressed later in the course. 

Suggested approaches

Attention operates on the stimulus (limited channel capacity; filter model).
Memory has a limited capacity; perceptual filters only admit portions of incoming stimuli, blocking the remainder. Blocked stimuli never reach long term memory
[ Broadbent, 1958 is one of the early monographs exploring this approach ]
.
Data related to the 'cocktail party effect' seem to support this theory. Evidence of future recall of 'unattended' information does not [ e.g. Hutmacher & Kuhbandner, 2020 ].

Attention dynamically controls the flow of information (parallel-access theory).
Attention facilitates a perceptual zigzag among several continuous streams of information, supported by expectations built on previous experience and on emerging patterns that imply future direction(s). 

_ At any given moment, focusing on some streams comes at the cost of attenuating others, which are still processed to some degree. This flexibility allows for dynamic shifts in attention based on context
[ attenuation & feature integration models; e.g. Treisman & Gelade, 1980 ].

_ Alternatively, observers can simultaneously attend to multiple streams of information at a cost (time, strain, etc.) that increases with stream complexity
[
multitasking; divided attention model; e.g. Kahneman, 1973 ].

Attention operates on memory rather than on stimulus (late selection models).
All information may be submitted to memory and processed by cognition, with focus operating at a later stage, supporting more nuanced engagements with the environment
[ e.g. Deutsch & Deutsch, 1963 ].

Attention is guided by accents and schemas that allow us to anticipate contextually important features of incoming streams of stimuli (more later in the module). It can be seen as a manifestation of a hypothesized interaction between explicit* and implicit** perceptual rules that helps us parse incoming information in a meaningful, to us, manner, with different attentional mechanisms being activated depending on context.

[*Explicit perceptual rules: perception-guiding strategies that we're consciously aware of and can express in words, such as the rules of grammar, the rules outlining what is and is not allowed in a game of chess, the rules of music theory, etc..]
[**Implicit perceptual rules: perception-guiding strategies that we're not consciously aware of and cannot express in words, such as the sensorimotor rules guiding the performance of an instrument, the rules underlying the communication of musical expression, etc..]

 

Experiencing Musical Time

The perception of musical time depends on:
    _ endogenous factors, such as attention, motivation, and physiological state;
    _ exogenous factors, such as tempo, meter, salience of sonic features, degree of regularity
       vs. complexity, and musical or broader spatiotemporal context.  

Listening to music appears to be a perceptual scanning process, where attention zigzags, shifts focus, and allows us to track multiple simultaneous streams of information through time. This perceptual 'zigzag' picks out only a single or a small subset of information streams at a time, out of a complex sonic environment. Implicit rules fill in the resulting gaps to create the 'illusion' of simultaneous, continuous perception of several -only partially attended- streams.

The scanning process varies for different listeners and for the same listener at different times, since implicit rules, explicit rules, and their interaction vary according to previous experience and context (more later the module).

The efficiency of the scanning process is improved by:
    _ repetition;
    _ gradual introduction of musical layers;
    _ selective variation of one or two layers at a time, while the rest remain unchanged;
    _ selective attenuation/removal and amplification/reintroduction of layers in time;
    _ assignment of different contours/timbres/registers per layer, etc..

Musical pieces that make use of such devices seem to retain their 'freshness' even after many hearings. For example, the chorus in "Good Vibrations" by the Beach Boys layers 6 different melodic lines that we seem to be able to follow 'simultaneously.'
[ recording session clip ]

  • Each melodic line is introduced successively, giving listeners the opportunity to learn each line on its own. The contrast created every time a new line is introduced constitutes an accent that focuses attention on the new line.

  • Shifting focus is facilitated by contour differences among melodic lines.

  • Most of the melodic lines occupy relatively different pitch ranges, further facilitating selective attention and focusing. Even at points where melodic lines cross one another, they are still discernable because of the combination of the other factors:
    a) previous knowledge, b) different sonic contours*, etc..
    [*Sonic Contour: pattern of sonic changes, outlined by inflection points.]

Eventually, listeners follow the chorus by shifting focus to different lines, as time goes by, and implicitly filling-in any gaps, based on their previous knowledge of each individual line.
The statistical improbability that this process can be repeated in exactly the same way may explain the sustained freshness of and interest in such pieces.

 

Further Reading

Butler, D. (1992). "Cognition and Time in Music." (book chapter)
Tan et al. (2010). "Perception of musical time." (book chapter)
Allman M.J. et al. (2016). "A brief history of 'The Psychology of Time Perception.'"
Maniadakis, M. & Trahanias, P. (2014). "Time models and cognitive processes: a review."

  

Memory limits - Categories - Accents - Schemas - Complexity

 

 
All music making and listening is, at some level, an act of communication. The complexity level of communicated messages is limited by the versatility of our information-encoding process.

According to information theory, a perceptual system's efficiency is based on its ability to reduce the 'size ' of a message without significantly (i.e. noticeably) altering it; that is, on its ability to identify patterns, reduce uncertainty, and increase predictability, while, at the same time, maintaining variety [ e.g. Shannon & Weaver, 1949/1963 ].
The musical systems we invent appear to involve parameters that satisfy the need to reduce our cognitive load by effecting a data reduction, through processes that also allow for maximum variety (i.e. maximum possible combinations) within the resulting set of reduced data.

Data reduction: Process that reduces the infinite variability of the world of experience so that it can be workable in short-term memory.
'Data reduction' refers less to actual elimination of individual pieces of information and more to reduced complexity/randomness/uncertainty, and increased predictability, based on identifying redundancies. It refers to the organization / patterning of random information into larger units we call chunks and categories.

Categories: Psychological constructs that depend on long-term previous learning and experience, connected to the idea of data reduction.
Large amount of data can be enveloped in a single category which is then understood in terms of the criterial attributes [ refresh your memory ] that link this data into a unit (e.g. timbre, pitch, scales, harmony, tonality). Categories can be multidimensional, incorporating many interrelated criterial attributes, and may therefore vary in one or more attributes simultaneously without losing their identity.

Chunks: While categories are large-scale organizations that depend on long-term previous learning and experience, chunks are rather temporary perceptual units that depend on immediate context and implicit rules (e.g. melodic theme; chord progression; rhythmic motif). Perceptual chunking can lead to the formation of a category or can organize several existing categories into a single chunk that may become a new, higher-level category (e.g. musical style/genre), if the same chunking process is repeated sufficiently often.

Categorization and data reduction are fundamental to survival.
Every moment and aspect of our life is framed by categorization choices that include value judgments (i.e. judgments on the relative benefit/appropriateness/value of the available options) and discrimination (i.e. selection of this over that option).

 

  

(early, algorithmic AI efforts in music genre identification did not take into account musical schemas)

  
Two contrasting definitions of Information:
     a) number and variance/difference among the possible events within a system and
     b) anything that reduces uncertainty.

The human brain's hypothesized limited channel capacity imposes a limit to the amount of information that can be processed before there is 'overflow.' Theoretically, the limit to the rate of information that short-term memory can deal with is expressed by Miller's 7+-2 rule: A maximum of only 7+-2 unrelated / un-patterned / random events out of a one-dimensional continuum can be stored in short-term memory [ Miller, 1956 ].

It follows that the complexity of the outside world needs to somehow be reduced in order to become perceptually meaningful.
Implicit rules help us deal with 'input data overload' by focusing on accents to abstract criterial attributes, relative to existing schemas, and accomplish the necessary data reduction.

Attention focuses on and is guided by accents (i.e. points of perceptual salience/importance). 
Accents are created through contrasts; that is through changes (in shape, color, brightness, direction, speed, size, pitch, loudness, timbre, rhythm, etc.), that help us identify boundaries and parse incoming stimuli into meaningful units (e.g. chunks). Empirical studies confirm that perceptually salient portions (i.e. accents) of a melody are attended to more often and are recalled more easily than other portions.

Accents help outline schemas (i.e. super-ordinate knowledge structures, based on prior experience).
In cognitive psychology, schemas denote cognitive frameworks that help us categorize objects and events, based on common characteristics, to support interpretation and prediction of new objects/events.

Schemas develop through early immersion in one's culture (musical or otherwise). They are abstractions that reduce cognitive load via
    i) identifying redundancies,
    ii) recognizing/imposing patterns, and occasionally 
    iii) discarding information that diverges from an implied pattern.
Expertise significantly influences schematizing ability, reduces the amount of discarded information, and helps develop highly refined schemas that support faster and more accurate encoding and retrieval of information.

Schemas influence how events are perceived, interpreted, and remembered. Existing schema templates are activated by specific informational instances (schema instantiation) and modulate early perceptual processing, enhancing but possibly also distorting mnemonic action
[ e.g. Gilboa & Marlatte, 2017 ]
.

In musical contexts, listeners draw on familiar patterns, genres, and styles to form schemas that facilitate chunking and categorization. Attention facilitates the integration of musical chunks into coherent sequences and helps listeners organize sense perceptions into cognitions (i.e. into meaningful wholes).

Pattern learning and recognition relies on repeated exposure to musical elements that supports the development of predictive neural pathways and sensitivity to specific intervals, rhythms, and harmonies.
The brain's ability to recognize patterns is enhanced through familiarity with musical styles, leading to the formation of schemas that outline expectations about how musical phrases typically unfold.

For example, a listener familiar with blues might chunk a complex improvisation into recognizable motifs that align with common blues patterns. Analogous processes support perceptual chunking of melodic contours, intervals, chords, chord progressions, or tonality [ e.g. Radocy & Boyle, 2012 ]. The brain continuously updates its predictions, based on auditory input, contributing to the experience of surprise, tension, or resolution (more later in the course).

Intra-cultural consistency and inter-cultural differences in the way listeners complete unfinished melodic [ e.g. Carlsen, 1981 ] and harmonic [ e.g. Bharucha & Stoeckig, 1986 ] passages support the view that musical expectations, perceptual chunking, and the associated schemas are culture dependent. Musical pieces that do not conform to a listener's cultural expectations are more difficult to process [ e.g. Unyk & Carlsen, 1987; Pearce & Wiggins, 2006 ].

Schematization, chunking, and pattern recognition are interrelated processes that feed back on one-another.

Schema formation supports effective chunking, which facilitates pattern recognition and the other way round. Similarly, recognizing patterns allows for the formation of larger chunks, while identifying chunks aids in the recognition of overarching musical structures.
For example, listeners may recognize the verse and chorus sections of a song as distinct chunks, while also chunking thematic materials and their transformations within those sections.

[ Note: In 1700s, the term 'schema' denoted a 'stock musical phrase' used to construct melodic, harmonic, & rhythmic scaffolds for Galant-style compositions; e.g. Gotham & Shaffer, 2023 ]

Notation of the major scale from C4 to C5 showing the intervals in semitones.
All 12 intervals are possible, indicating that this specific scale (subset of the twelve tones in a semitone-divided octave) satisfies Miller's rule while, at the same time, allows for maximal intervalic variety (in Kendall & Carterette, 1996).

 

 
Musical Scales as Perceptual Categories

The octave interval is an example of a category based on data reduction.
Consistent with Miller's rule, the pitch systems used to construct melodies employ a maximum of 7±2 individual pitches within the octave, in practically all music cultures.

This is not to say that only 7±2 pitches are available. The Western musical system has 12 pitches available within an octave and other musical traditions have even more. No musical system, however, violates Miller's 7±2 rule in terms of the number of functional pitches used in melodic units.

Functional pitches usually belong to pitch subsets called Musical scales: Systems prescribing the frequency relationships among pitches within an octave.
Musical scales can be arbitrary and reflect compositional / aesthetic choices. Such choices, more often than not, are guided by the properties of the musical instruments available and by cultural standards, rather than being determined by the harmonic series, the construction of the ear, or some other 'natural' factor.

Major and minor diatonic scales contain 7 out of the 12 available pitches, distributed so that they include five whole-tone intervals and two semitone intervals. As was demonstrated mathematically [ Balzano, 1980 ], this pitch configuration supports maximum variety (i.e. maximum interval-combination possibilities) from only 7 pitches out of an octave.

Diatonic scales divide the octave in 12 log-frequency units [ 21/12 - equal temperament ], select 7, and distribute them in a way that facilitates reproduction of all intervals available in the 12-tone chromatic scale (see image to the left).
The principle of coherence (i.e. no interval between any two successive scale notes/degrees should be larger than the sum of any two successive intervals) is closely related to maximum intervalic variety and the possibility of developing fully functional harmony within a scale.

The unique sonic character or 'signature sound' of different pitch intervals outlines musical categories that support prediction during listening. Combinations of these signature sounds can create cognitively recognizable patterns, which outline higher level perceptual groupings in the form of larger scale categories or schemas (e.g. chord progressions).

The key features of western art music tuning and scale systems (i.e. octave divided in 12 log-equal units and organized diatonically) attempt to optimally address:
     _ cognitive limits (short term memory/Miller's rule),
     _ data reduction/pitch circularity (more later in the course), and
     _ maximum variety (interval distribution within the major/minor scales).
  

    


Postulated Relationships among preference, interest, and complexity [ after Berlyne, 1971 ].

 
Complexity / Prediction / Preference / Interest

W. Wundt [ 19-20th century German physiologist, psychologist, and philosopher ] postulated a relationship among complexity, preference, and interest, illustrated to the left. He proposed that emotional responses are driven by perceived degree of stimulus intensity, complexity, and novelty.
    _ Moderate levels of complexity produce positive emotions.
    _ Too little complexity can lead to boredom and disinterest.
    _ Excessive complexity can result in overstimulation and negative emotions.

Consequently, preference decreases when there is extreme simplicity (redundancy; little information; maximal order; low entropy), as well as when there is extreme complexity (randomness; too much information; minimal order; high entropy).

A system has maximum entropy when all possible events within the system are equally likely. If such a system involves a large number of possible events, then prediction becomes difficult, resulting in a high level of uncertainty. 
A musical example of maximum entropy is strict dodecaphonic, serial music which -to some- does not even qualify as music (e.g. Excerpt from 'Symphony No. 21 - Opening' by Anton Webern).

Maximum preference corresponds to an optimum level of complexity, which is -in general- different for each individual and changes with previous knowledge of the system in question. Listeners, for example, prefer music with an intermediate level of predictability; not so predictable that it's boring, but not so unpredictable that it's unintelligible.
 
Interest follows a similar relationship but with maximum interest corresponding to an optimum level of complexity that is -in general- higher than that for maximum preference. That is, we remain interested in levels of complexity that are on average higher than we'd prefer.

In a sense, successful / good art strikes a balance between low and high entropy, but always within a context. A system's degree of entropy and its relationship to interest and preference is relative to the observer. Analogously to the processes of schematizing and categorizing, familiarity with a certain style, artist, or work of art influences what level of complexity a listener is likely to consider optimal. This may explain individual response differences when experiencing works of art in general and music in particular, as well as individual differences on what constitutes successful / good music.

Repeated exposure to preferred works does not seem to reduce their enjoyment. However,
“…it seems likely that there may be changes in an individual’s preferred complexity level with the passage of time. For example, a piece of music which is initially too complex for an individual to like, may, with repeated playings, move down to a lower complexity level at which liking may begin to emerge.” [ in J.B. Davies's 1978, now classic monograph, "The Psychology of Music" ]

 
Musical complexity generally refers to
_ the structural intricacy of a piece, in terms of rhythm, harmony, melody, texture, and/or form, and
_ the degree of unpredictability, density, or richness of musical information.

  • Rhythmic complexity: syncopation; irregular meters; time signature changes; etc.

  • Harmonic complexity: dissonance; unconventional chord progressions; atonality; etc.

  • Melodic complexity: wide intervals; intricate ornamentation; thematic & tonal variety; etc.

  • Textural complexity: rhythmic, harmonic, melodic, and timbral layering; textural contrasts; etc.

  • Formal complexity: departure from standard or repetitive forms; etc..

Musical predictability refers to information content that helps reduce uncertainty and counter musical complexity, achieving a degree of balance appropriate to a given listener and context. The devices that improve the scanning efficiency of complex (i.e. multi-stream) musical pieces, mentioned earlier, also support predictability and include:

  • Repetition;

  • Gradual introduction of musical layers;

  • Selective variation of one or two layers at a time, while the rest remain unchanged;

  • Selective attenuation/removal and amplification/reintroduction of layers in time;

  • Assignment of different contours/timbres/registers per layer, etc..

For example, North & Hargreaves (1995) explored the complexity/familiarity/enjoyment relationship with 75 students listening and rating 60 excerpts from popular (at the time) songs. Artists included David Bowie, Simple Minds, Duran Duran, Enya, Japan, Tangerine Dream, and many others, covering a wide range of what is considered popular music. When rating 'complexity,' listeners were provided with a broad definition, based on: "how easy it is to predict what the music will do next" vs. "how many surprises the music contains." The results were compatible with Wundt's model, showing, highest preference for intermediate levels of complexity.

Overall, listeners seem to reliably prefer music of intermediate complexity, with preference for more predictability rising with context uncertainty [ e.g. Gold et al., 2019 ].
 

 

Example of simplicity/complexity balance: "Wouldn't It Be Nice" by the Beach Boys (1999 remaster and official video - first released in 1966)

  

Implicit & Explicit Knowledge - Analytic & Synthetic Listening

 


Based on Ellis et al., 2009

 


In Kendall & Carterette,1990: 135.

 

 

 

       
Rotating Point                                                          Receding Circle

 


The McGurk Effect

 
Implict & Explicit Levels of Understanding

The chunking, categorization, and schematizing processes that guide perception reveal a non-isomorphic relationship between input (external world) and output (understanding).
There are two interacting levels of knowledge involved, implicit and explicit, with the implicit level dealing with knowledge units that we do not have direct access to.

The explicit level deals with symbols and involves rules that we're consciously aware of, can be expressed in words, and are applied deliberately. E.g.:
_ rules of grammar;
_ rules outlining what is and is not allowed in a game of chess;
_ rules in sports;
_ rules of music theory.
The implicit level deals with knowledge units (meta-symbols / schemas) that we may not have direct access to, but which are crucial to our eventual awareness and actions. This level involves rules that we're not consciously aware of, cannot be expressed in words, and are applied 'automatically.' E.g.:
_ the sensorimotor rules guiding repeated physical activity
   (e.g. sports; the performance of an musical instrument)
_ the communication of musical expression.

Development and enrichment of human consciousness is accompanied by a concurrent growth of implicit knowledge. Consequently, our comprehension of the external world (environment) necessarily also relies on what our implicit knowledge avails to us from the external world.
Our explicit understanding of and communication in and about the external world is always mediated by implicit rules, which we cannot access but which effect a translation from short-term to long-term memory.

This mediation is approached as
_ "noise in message transmission," within Information Theory, where inability to explicitly
   analyze contributes to disorganization;
_ "a necessary message-shaping step," within Cognitive Psychology, where ability to
   identify patterns, reduce uncertainty, increase predictability, and reduce a message's
   'size' without altering it enhances our perceptual system's efficiency and effectiveness.
   [ per Shannon & Weaver, 1949 ]
.

Awareness of the world is the result of synthesis. Input from the external world is mediated at the implicit level before it reaches our explicit awareness and is processed again at both the explicit and implicit levels before it can be expressed as a response, reaction, or understanding. 
 

Implicit & explicit processes in meaning construction

Processes at the implicit level allow us to predict and give rise to expectations, based on our fundamental intuition that, in organized systems, not all events are equally likely.

  • An external input is first partitioned into events.
    Event partitioning (i.e. 'chunking') involves identification of boundaries. The process assisted by contrasts that produce accents on which attention can focus [ the term 'contrast' brings to the forefront issues of similarity versus dissimilarity, identity, etc. ].
     

  • Event partitioning is followed by event categorization.
    Event categorization into meaningfully grouped units relies on existing schemas, which may be innate or acquired through experience and which are based on our natural urge to identify and store patterns.

If, for example, we suppose that frequency and pitch events proceed on a continuum (e.g. from low to high), categorization effectuates a data reduction by breaking (or 'quantizing') this continuum into bounded units, within which all events represent (and are represented by) a single category. Melodic motion is not a motion on a continuum from low to high but a motion along discrete note categories, organized through the ordering of discrete note durations.

  • Chunking and Categorization allow for a pattern synthesis that makes incoming data meaningful, in that it helps us anticipate what may come next (i.e. determine a 'most likely outcome').

Deliberate reflection on a given action/thought or a conflict between explicit and implicit rules may reveal their interaction and interrelationship. Such contexts support the occasional transformation of implicit knowledge into explicit statements that can permit communication and manipulation of what eventually becomes a set of explicit rules. Such transformations provide insights on the interpretive strategies we employ to modify/override explicit rules, implicit rules, or both. Such conflicts are usually resolved through the transformation of implicit .

For example (see below and to the left):

In the Rotating Point clip, a dot bounces as it moves from left to right, followed by a dot moving in the same direction along a straight line. When both dots are presented together, a joint motion emerges that, due to context, is perceived as different from the combination of the individual motions.

In the Receding Circle clip (see to the left), a circle diminishes in size while the intensity of the accompanying sound increases, decreases or stays the same. The circle's motion appears to influence perceived loudness, particularly in the case of an accompanying sound with fixed intensity, where there is a conflict between
_ explicit rules (constant intensity corresponds to constant loudness) and
_ implicit rules (in a context that calls upon our experience of perspective, constant intensity corresponds to increasing, constant, or decreasing loudness, depending on whether the source of sound appears to move away from us, stand still, or move towards us respectively - see a visual analog).
[ For research exploring an analogous effect see Kitagawa & Ichihara, 2002 ]
.

The McGurk Effect illustrates the interaction among conflicting visual/aural stimuli (explicit) and speech/timbral expectations (implicit), resulting in a constructed perception that resolves the conflict but corresponds to neither stimulus.

In a set of pilot experiments at Columbia College Chicago, subjects rated the degree of sonic 'warmth' of acoustical spaces. Wood-color-painted cement surfaces were rated almost as 'warm' as actual wooden surfaces and less 'warm' than unpainted cement surfaces and cement-color-painted wooden surfaces.
[ Vassilakis, 2014 -  Columbia College Chicago, Audio Arts & Acoustics ].

The phenomenon of the "missing fundamental," mentioned earlier in the course, provides an additional example of implicit processes at work: Implicit rules synthesize a pitch that does not correspond to a frequency present in the physical frame of reference.

 
The phenomenon of chromaesthesia
describes the consistent association of pitches with colors, shapes, and/or movement, in a behavior that appears to largely rely on implicit knowledge. It is a rare phenomenon and the way it operates is still not fully understood [ e.g. van Campen, 1997a ].
Although not all individuals who exhibit chromaesthesia associate the same colors with the same pitches, every individual is internally consistent.

During the second half of the 19th century, chromaesthesia fascinated a number of composers who created works believed to have intrinsic relationships to color compositions [ e.g. Scriabin: Poem of Ecstasy; Rimsky-Korsakov: Scheherazade; Kandinsky: The Music of Colors ].
     For a discussion of Scriabin's and Kandinsky's chromaesthesia artistic explorations, see van Campen (1997b).

Chromaesthesia is a special case of the phenomenon of synesthesia, the perceptual cross-modal connection of the senses.
     For reviews see Hochel & Milan (2008) & Spence, (2011). Comprehensive accounts in Cytowic (2018) & Sathian & Ramachandran (2020), Ch 13.

  

While the processing of musical emotions at the implicit or explicit levels appears to engage distinct neural circuits [ e.g. Bogert et al., 2016 ], the implicit/explicit distinction is not rigid and does not imply clear-cut boundaries. There is a constant interaction between perception rules (explicit & implicit) and the external world. THis may may explain why "the whole is different from the sum of its parts," statement that holds true not only for music and art but for all experience. Whether an event will be processed implicitly or explicitly depends on previous knowledge, experience, motivation, purpose, context, attention, and mental set.

In a, now classic study, Schachter & Singer (1962) demonstrated that the behavior of an individual at any given time depends on their mental set as much as (if not more than) it depends on their physiological state. Subjects were given a drug or a placebo and were put in a room with an actor that had supposedly taken the same substance and was instructed to act either drugged or normal. Subjects acted drugged or normal based more on the attitude of the actor than on the administered substance (simplified description). The study formed the basis for the authors' two-factor theory of emotions, according to which emotional experiences reflect the combination of some physiological arousal/response and its cognitive interpretation and labeling (video outline).

Explicit and implicit rules may be shared or unshared during communication (module 1 refresher).
Introducing the idea of unshared explicit and implicit rules to understanding helps explain how we can validly have different interpretations of the same event (or musical piece), without resorting to subjective preference or whim.

     ADDITIONAL PERCEPTS

(based on Deutsch, 1995)

 

 

 

 

 
Analytic & Synthetic Listening

Analytic listening is based largely on explicit knowledge and rules;
Synthetic listening is based largely on implicit knowledge and rules.

EXAMPLES

Two 2-component complex tones, presented in succession:
(a)
800Hz+1000Hz     (b) 750Hz+1000Hz. [ adapted from Smoorenburg, 1970 ].

When moving from (a) to (b):

  • Many listeners hear the pitch going down by following the frequency change of the first component in each tone:
    800Hz 750Hz ~1 semitone drop.
    Explicit rules are employed to track the physical attributes of the tones and determine the pitch motion.
    This is considered an example of analytic listening.
     

  • Some listeners hear the pitch going up, by reconstructing the motion of the (missing) fundamental implied by the two complex tones:
    200Hz    250Hz ~4 semitone (i.e. a major third) rise.
    Implicit rules
    are employed to synthesize a physical attribute that is implied by the rest of each tone's attributes, guiding pitch motion. This is considered an example of synthetic listening.

Listen to a version with two 5-component complex tones, presented in succession:
(a) 800Hz+1000+1200+1400+1600        (b) 750Hz+1000+1250+1500+1750.   
In this case, more listeners respond synthetically and hear the pitch going up by a major 3rd, perceiving the motion of the (missing) fundamental: 200Hz 250Hz. (Why?)

Intermittent fixed frequency & frequency modulated tones, with noise bursts filling the gaps
[ fixed frequency & frequency modulated tones without noise bursts - after Dannenbring & Bregman, 1976 ]:

  • Listeners synthesize a sensation of steady tones, overlaid with noise bursts, in an example synthetic/holistic listening, largely based on implicit rules.

  • Listeners alerted to the gaps in the tones may be able to perceive them by directing their attention to separate portions of the combined stimulus. This 'directed,' analytic form of listening is largely based on explicit rules.

Diane Deutsch's scale illusion published in 1975 (see figures to the left) consists of a major scale with successive tones alternating from ear to ear. The scale is played simultaneously in both ascending and descending form; however when a tone from the ascending scale is in the right ear a tone from the descending scale is in the left ear, and vice versa [ figure (a) ]. The tones are equal-amplitude sine waves, and the sequence is played repeatedly without pause at a fast rate.

Even though panning alternates with each scale note, listeners reorganize perceptually the stimulus and assign each melodic line to a single ear (left / right) as in (c). This perception persists even when both panning and timbre alternate for each scale note in (a), with each melodic line being assigned to a single ear and a single timbre. Implicit rules are employed to process incoming information and send a synthesized version of 'reality' to conscious awareness.
(video explanation of the illusion)

If the timbre difference between the left and right channels becomes extreme, the illusion breaks down and listeners cannot synthesize the conjunct melodies of (c). Instead, they hear what is actually happening: constant leaps of pitch and timbre in each ear, as in (a). In other words, after a dissimilarity threshold between the two ears has been crossed, we shift from synthetic (c) to analytic (a) listening.

Whether an actual shift from implicit to explicit rules will occur depends strongly on context and previous knowledge. For example, although non-musicians and Western-trained musicians usually hear the conjunct melodies in (c), musicians skilled in the 12-tone system often hear the actual pitch leaps in (a), even without the help of timbral dissimilarity. They can attend to them by listening analytically with the help of explicit and implicit rules able to handle disjunct melodies, developed through explicit training and organic experience.

The McGurk effect, the receding circle, and the rotating point examples already discussed also illustrate the general link of analytic listening to explicit knowledge and of synthetic listening to implicit knowledge. However, distinctions such as explicit/implicit and analytic/synthetic are not rigid and do not imply clear-cut boundaries. Both implicit and explicit rules are involved in analytic as well as synthetic listening.

All examples discussed so far illustrate that perceptions are the result of an experience-guided synthesis of information from our environment. Rather than necessarily matching the stimuli, responses to a set of often conflicting stimuli reflect the best way such stimuli fit to previous experience, concurrent context, and the associated expectations. Removing or exaggerating conflicts supports a switch between synthetic and analytic perception.

_ In the McGurk effect example:
When staring at the speaker's lips, the synthetic listening effect is so robust that, even when alerted to the conflict, listeners/viewers are unable to listen to the a/v composite analytically. Listening with eyes closed presents no ambiguity and supports analytic listening.
_ In Deutsch's scale illusion experiments:
When the timbral dissimilarity between the audio feeds per ear is exaggerated, the synthetic, conjunct melody perception gives way to the analytic, disjunct melody perception.

As another example, a study presented musicians and non-musicians with stretched and compressed versions of a perfect fifth see Siegel & Siegel, 1977 ]:

_ Musicians heard the intervals not as different but as out-of-tune versions of a single interval. They grouped the stimuli in terms of an explicit category (perfect 5th), formed through years of applying implicit rules in formal training/practice,.
_ Non-musicians heard the detuned intervals as representing separate perceptual entities, because they had not developed through training explicit interval categorization rules or any clear pitch category boundaries.

The explicit category of the perfect 5th interval employed by musicians is based on implicit knowledge. At the same time, the inability of non-musicians to utilize the explicit category of the perfect fifth is accompanied by their ability to utilize the explicit rule that assigns larger (smaller) frequency values to higher (lower) pitches.

 
Analytic vs. Synthetic Operations & Brain Area Specialization

The recurring contrast between analytic & synthetic or explicit & implicit has been addressed physiologically in terms of brain-area specialization (see to the left). While the infant brain shows a hemispheric specialization in processing music as early as after the first postnatal hours [ e.g. Perani et al., 2010 ], brain-area specialization can, to some extent, be retrained.

Analytic operations are processed mostly on the left brain hemisphere.
Scientists are often thought of as left-brain oriented. In ~95% of right-handers (~60-70% of left-handers), the left side of the brain is dominant for language. When musicians listen to music analytically (employing mostly explicit rules of organization: i.e. music theory, explicit categories, etc.) they are using their left brain hemisphere more than their right.
Synthetic operations are processed mostly on the right brain hemisphere.
Artists are often thought of as right-brain oriented. When musicians listen to music synthetically/holistically (employing mostly implicit rules of organization: i.e. general patterns, gestalt rules, etc.) they are using more the right brain hemisphere.

Brain area specialization's impact on music perception is reviewed in Peretz & Zatorre (2005).

 

 

 


  

Loyola Marymount University - School of Film & Television