|




(adapted from
Polti
et al., 2018)

(attention guided by schemas) |
Experiencing Time
Our experience of time seems to often
proceed independently from clock/chronological time. Musical organization of
sounds offers us means to creatively explore this difference.
The
fleeting nature of all non-linguistic sound events (in music, film, etc.),
due to their potential absence of a referent, highlights the broader challenge of defining the perceptual
'present.'
What we understand as
the present (i.e.
our sense of 'now') is shaped by our
selective attention to (aspects of) events
and can stretch/contract
to degrees that significantly impact duration perception.
Our experience of
present events is always mediated by the memory of past events and
the expectation of future ones, while reshaping
both our memories and our expectations.
The dialectic
nature of the way we experience the present, past, and future is
at the core of the human experience of time and the topic of extensive research from within philosophy
in general and phenomenology in particular.
The
ever-moving present is the
only existentially 'real' dimension of what we experience as
time, since it is the the only temporal dimension in which we
may act.
An essentially temporal field of action,
the present is experientially defined by the
productive tension between our memory of past events and our
anticipation of future events, in negotiation
with the actions we and others take.
Philosophy/Phenomenology Resources:
_ Ricoeur, P. (1998). "Critique and Conviction." Columbia
University Press. _ Ricoeur, P. (1984-85). "Time and
Narrative" (3 volumes). University of Chicago Press.
_ Gadamer, H.G.
(1989). "Truth and Method." Continuum Press.
_ Gadamer, H.G.
(1986). "The Relevance of the Beautiful." Cambridge
University Press _ Kramer, J. (1988). "The Time of Music." Schirmer Books.
_
Additional bibliography
(Stanford Encyclopedia of Philosophy)
Human
perception compresses time frames:
_
Recent events are perceived
as more distant and vice versa (telescoping effect).
_ Shorter intervals are overestimated and, more so, vice versa
(Vierordt's law).
_ Working towards a desired goal shrinks perceived duration and vice
versa
(linked to the goal-gradient effect:
motivation & effort increase with goal proximity).
Storage-Size
(computational) Theories
Memory storage needs allegedly influence
our estimates of time: a percept containing a large amount of
information requires more storage capacity in short-term
memory, generating the impression of greater elapsed time [ e.g.
Ornstein, 1969; Sheth
et al., 2000 ].
The idea is that our brain
uses memory capacity as a metric for time duration
perception. Information-rich events
take up more "mental space" and, consequently, feel longer.
According to the theory, events with high complexity --like
musical pieces with busy arrangement, many details, and
little repetition-- tend to be perceived as longer than
simple events, due to the increased memory storage required
[ e.g.
Matthews et al., 2014
].
In a version of this
approach, the
Scalar Expectancy Theory (SET) proposes that our
perception of time is based on the accumulation of
pulses within an internal "clock." The number of stored
pulses represents the perceived duration, essentially
acting as a storage size counter for time information
[ e.g.
Gibbon, 1977 ].
Partially consistent with this model is the argument
that melodic expectations operate
within a limited-width time-window, restricted by short-term
memory storage
capacity limits. [ e.g.
Dowling
et al., 1987 ]
Drawbacks
The theory is challenged by difficulties in reliably measuring
the "storage size" of memory, which is necessary to the theory's
testing. It has also been criticized as simplistic,
in that it does not account for factors like
previous experience, attention, and emotional state.
More importantly,
research on how event density influences
perceived duration points to the opposite direction.
Essentially, the more events happening within a timeframe, the
higher the event density, which can often lead to a perception
that the time passed more quickly and the events' perceived duration was
shorter [ e.g.
Lambrou-Kokolaki
et al., 2024 ].
Regularity (i.e.
reduction in new information) correlates with the perception
of longer durations for listeners familiar with a musical style/piece
[ e.g.
Sasaki & Yamada, 2017 ]
but may have the opposite effect on unfamiliar
listeners (see
Berger, 2014). Slower tempos and irregular meters/rhythms make it
difficult to accurately predict how much time has passed,
leading to duration over/underestimations, experienced as
'losing track of time' (see
Why music can make us literally lose track of time).
Attentional Capacity
and Emotional State (cognitive)
Theories
Assuming that perceptual activities
compete for attentional center stage, the more we attend to some events
-or aspects of events- the less we will attend to time itself,
leaving our 'cognitive clock' with the impression of less elapsed
time [ e.g. Hicks
et
al., 1976; Block, 1978;
Polti
et al., 2018 ].
Some models hypothesize an "attentional
gate" that regulates the flow of temporal information
into our internal clock. When attention is diverted, the
gate narrows, reducing the amount of temporal information
that can be processed [ e.g.
Zakay &
Block, 1997 ].
More recently, the impact of emotional state
and context on time perception
has been appreciated and systematically studied. Results have
shown that
anxiety, arousal, task difficulty, and skill
level (musical or otherwise) are
important factors.
For example, when people are excited or
are pursuing a goal, they may feel like time is passing quickly.
At the same time, being presented with the opportunity to earn a
reward can make time seem longer [ e.g.
Dawson & Sleek, 2018;
see also
Lehockey
et al., 2018 ].
Issues of attention are central to our ability to partition our
sonic environment into distinct streams and shift focus among
them. The perceptual strategy (auditory scene
analysis) and manifestations (e.g. the cocktail party effect) of
such processes will be addressed later in the course.
Suggested
approaches
Attention operates on the stimulus
(limited channel capacity; filter model). Memory has
a limited capacity; perceptual filters only admit
portions of incoming stimuli, blocking the remainder.
Blocked stimuli never reach long term memory
[ Broadbent,
1958 is one of the early monographs exploring this
approach ].
Data related to the 'cocktail party
effect' seem to support this theory. Evidence of future recall of
'unattended' information does not [ e.g.
Hutmacher &
Kuhbandner, 2020 ].
Attention
dynamically controls the flow of
information (parallel-access theory). Attention
facilitates a perceptual zigzag among several continuous streams
of information, supported by expectations built on previous
experience and on emerging patterns that imply future direction(s).
_ At any given moment,
focusing on some streams comes at the cost of
attenuating others, which are still processed to some
degree. This flexibility allows for dynamic shifts in
attention based on context
[ attenuation
& feature
integration models; e.g.
Treisman & Gelade, 1980 ].
_ Alternatively,
observers can simultaneously attend to multiple streams
of information at a cost (time, strain, etc.) that increases with
stream complexity
[ multitasking; divided
attention model; e.g.
Kahneman, 1973 ].
Attention
operates on memory rather than on stimulus (late selection models).
All information may be submitted to memory
and processed by cognition, with focus operating
at a later stage, supporting more nuanced engagements with the
environment
[ e.g.
Deutsch &
Deutsch, 1963 ].
Attention
is guided by accents and schemas that allow us to
anticipate contextually important features of incoming streams of
stimuli (more later in the module).
It can be seen as a manifestation of a hypothesized
interaction between explicit* and implicit** perceptual
rules
that helps us parse incoming information in a meaningful, to us,
manner, with different attentional mechanisms being
activated depending on context.
[*Explicit perceptual
rules: perception-guiding strategies that we're
consciously aware of and can express in words, such as the
rules of grammar, the rules outlining what is and is not
allowed in a game of chess, the rules of music theory, etc..]
[**Implicit perceptual rules:
perception-guiding strategies that we're not consciously
aware of and cannot express in words, such as the sensorimotor rules guiding
the performance of an instrument, the rules underlying the
communication of musical expression, etc..]
|
|


|
Experiencing Musical Time
The perception of musical time depends
on:
_ endogenous factors, such
as attention, motivation, and physiological state;
_
exogenous factors, such as tempo, meter, salience of sonic features, degree of regularity
vs.
complexity, and musical or broader
spatiotemporal context.
Listening to music
appears to
be a perceptual scanning process, where attention
zigzags, shifts focus, and allows us to track multiple simultaneous
streams of information through time. This perceptual 'zigzag' picks
out only a single or a small subset of information streams at a
time, out of a complex sonic environment. Implicit rules fill in the
resulting gaps to create
the 'illusion' of simultaneous, continuous perception of several
-only
partially attended- streams.
The scanning process varies for
different listeners and for the same listener at different times,
since implicit rules, explicit rules, and their interaction
vary according to previous experience and context (more
later the module).
The efficiency of the scanning
process is improved by: _ repetition; _ gradual introduction of
musical layers; _ selective variation of one or two
layers at a time, while the rest remain unchanged; _
selective attenuation/removal and
amplification/reintroduction of layers in time; _ assignment of different contours/timbres/registers per
layer, etc..
Musical
pieces that make use of such devices seem to retain
their 'freshness' even after many hearings. For
example, the chorus in
"Good Vibrations"
by the Beach Boys
layers 6
different melodic lines that we seem to be able to follow
'simultaneously.'
[ recording session clip ]
-
Each melodic line is
introduced successively, giving listeners the opportunity to
learn each line on its own. The contrast created every time
a new line is introduced constitutes an accent that focuses
attention on the
new line.
-
Shifting focus is facilitated by
contour differences among melodic lines.
-
Most of the melodic lines occupy
relatively different pitch ranges, further facilitating
selective attention and focusing. Even at points where melodic lines cross
one another, they are still discernable
because of the combination of the other factors:
a) previous
knowledge, b) different sonic contours*, etc..
[*Sonic Contour:
pattern of sonic changes, outlined by inflection points.]
Eventually, listeners follow the
chorus by shifting focus to different lines, as time goes by,
and implicitly filling-in any gaps, based on their previous
knowledge of each individual line.
The statistical improbability that this process can be repeated
in exactly the same way may explain the
sustained freshness of and interest in such pieces.
Further Reading
Butler,
D. (1992). "Cognition and Time in Music."
(book chapter)
Tan et al. (2010). "Perception of musical time." (book
chapter)
Allman
M.J. et al. (2016). "A brief history of 'The
Psychology of Time Perception.'"
Maniadakis,
M. & Trahanias, P. (2014). "Time models and
cognitive processes: a review."
|
All music making and
listening is, at some level, an act of communication. The
complexity level of communicated messages is limited by
the versatility of our information-encoding process.
According to information theory, a perceptual system's efficiency is based on its
ability to reduce the 'size ' of a message without
significantly (i.e. noticeably) altering it; that is,
on its ability to identify patterns, reduce uncertainty,
and increase predictability, while, at the same
time, maintaining variety [ e.g.
Shannon & Weaver, 1949/1963 ].
The
musical systems we invent appear to involve parameters that
satisfy the need to reduce our cognitive load by
effecting a data reduction,
through processes that also
allow for maximum variety
(i.e. maximum possible combinations) within the resulting
set of reduced data.
Data reduction: Process that reduces the infinite
variability of the world of experience so that it can be
workable in short-term memory. 'Data reduction' refers
less to actual elimination of individual pieces
of information and more to reduced complexity/randomness/uncertainty, and
increased
predictability, based on identifying redundancies. It
refers to the organization / patterning of random
information into larger units we call chunks
and categories.
Categories:
Psychological constructs that depend on
long-term previous learning and experience, connected to
the idea of data reduction. Large amount of data can be
enveloped in a single category which is then
understood in terms of the criterial attributes [
refresh
your memory ] that
link this data into a unit (e.g. timbre, pitch,
scales, harmony, tonality). Categories can be multidimensional, incorporating many
interrelated criterial attributes, and may therefore vary in one or
more attributes simultaneously without losing their
identity.
Chunks: While
categories are large-scale organizations that
depend on long-term previous learning and
experience, chunks are rather temporary
perceptual units that depend on immediate
context and implicit rules (e.g. melodic
theme; chord progression; rhythmic motif).
Perceptual chunking can lead to the formation of
a category or can organize several existing categories into a single chunk
that may become a new, higher-level category (e.g. musical style/genre), if the same
chunking process
is repeated sufficiently often.
Categorization
and data reduction are fundamental to survival.
Every moment and aspect of our life is framed by
categorization choices that include value judgments (i.e.
judgments on the relative
benefit/appropriateness/value of the available
options) and discrimination (i.e. selection
of
this over that option).
|
|


(early, algorithmic AI efforts in music genre identification did not take
into account musical schemas) |
Two contrasting
definitions of Information:
a) number and variance/difference among the possible events
within a system and
b) anything that reduces uncertainty.
The human brain's
hypothesized limited channel capacity imposes a limit to the
amount of information that can be processed before there
is 'overflow.' Theoretically, the limit to the rate
of information that short-term memory can deal with is
expressed by Miller's 7+-2
rule:
A maximum of only
7+-2 unrelated /
un-patterned / random events out of a one-dimensional
continuum can be stored in short-term memory [ Miller,
1956 ].
It follows that the complexity of the outside world
needs to somehow be reduced in order to become
perceptually meaningful.
Implicit rules help us
deal with 'input data overload' by focusing on
accents to abstract
criterial attributes, relative to existing
schemas, and accomplish the
necessary data
reduction.
Attention focuses on
and is guided by accents (i.e. points of perceptual salience/importance). Accents are
created through contrasts; that is through changes (in shape,
color, brightness, direction, speed, size, pitch,
loudness, timbre, rhythm, etc.), that help us identify
boundaries and parse incoming stimuli into meaningful units (e.g.
chunks). Empirical studies confirm that
perceptually salient portions (i.e. accents) of a
melody are attended to more
often and are recalled more easily than other portions.
Accents
help outline schemas
(i.e. super-ordinate knowledge structures, based on prior experience). In cognitive psychology,
schemas denote cognitive
frameworks that help us categorize objects and events, based on
common characteristics, to support interpretation and prediction
of new objects/events.
Schemas develop through early
immersion in one's culture (musical or otherwise). They are
abstractions that reduce cognitive load
via i) identifying
redundancies, ii) recognizing/imposing patterns,
and occasionally iii) discarding information that diverges from an implied pattern.
Expertise significantly influences schematizing ability, reduces the
amount of discarded information, and helps develop highly
refined schemas that support faster and more accurate encoding and retrieval of
information.
Schemas
influence how
events are perceived, interpreted, and remembered.
Existing schema templates are activated by specific
informational instances (schema instantiation)
and modulate early perceptual processing, enhancing
but possibly also distorting mnemonic action
[ e.g.
Gilboa &
Marlatte, 2017 ].
In musical contexts, listeners
draw on familiar patterns, genres, and styles to form schemas
that facilitate chunking and categorization. Attention facilitates
the integration of musical chunks into coherent sequences and
helps listeners organize sense perceptions into cognitions (i.e.
into meaningful wholes).
Pattern learning and recognition relies on repeated exposure to musical elements
that supports the development of predictive neural
pathways and
sensitivity to specific intervals, rhythms, and harmonies.
The brain's ability to recognize patterns is enhanced through
familiarity with musical styles, leading to the formation of
schemas that outline expectations about how musical phrases typically unfold.
For example, a
listener familiar with blues might chunk a complex
improvisation into recognizable motifs that align
with common blues patterns. Analogous processes
support perceptual chunking of melodic contours,
intervals, chords, chord progressions, or tonality [
e.g.
Radocy & Boyle,
2012 ]. The brain continuously updates its predictions, based on auditory
input, contributing to the experience of surprise, tension, or
resolution (more later in the course).
Intra-cultural consistency and
inter-cultural differences in the way listeners complete
unfinished melodic [ e.g.
Carlsen, 1981 ] and harmonic
[ e.g.
Bharucha & Stoeckig, 1986 ] passages support the view that
musical
expectations, perceptual chunking, and the associated schemas are culture dependent.
Musical pieces that do not conform to a listener's
cultural expectations are more difficult to process
[ e.g.
Unyk
& Carlsen, 1987;
Pearce & Wiggins, 2006 ].
Schematization, chunking, and pattern
recognition are interrelated processes that feed back on one-another.
Schema formation supports effective chunking, which facilitates pattern recognition
and the other way round. Similarly, recognizing patterns allows for the formation of
larger chunks, while identifying chunks aids in the recognition of
overarching musical structures. For example, listeners may recognize the
verse and chorus sections of a song as distinct chunks, while
also chunking thematic materials and their transformations
within those sections.
[
Note: In 1700s, the term 'schema' denoted a 'stock
musical phrase' used to construct melodic, harmonic,
& rhythmic scaffolds for Galant-style compositions; e.g.
Gotham & Shaffer, 2023 ]
|
|

Notation of the major scale from C4 to C5
showing the intervals in semitones.
All 12 intervals are
possible, indicating that this specific scale (subset of
the twelve tones in a semitone-divided octave) satisfies
Miller's rule while, at the same time, allows for
maximal intervalic variety (in Kendall & Carterette,
1996).
|
Musical Scales as
Perceptual Categories
The octave interval is an
example of a category based on data reduction. Consistent with
Miller's rule, the
pitch systems used to construct melodies employ a maximum of
7±2 individual
pitches within the octave, in practically all music
cultures.
This is not to say that only
7±2 pitches are
available. The Western musical system has 12 pitches
available within an octave and other musical traditions
have even more. No musical system, however, violates
Miller's 7±2 rule in terms of the number of
functional pitches used in melodic units.
Functional pitches usually belong to pitch subsets
called Musical scales: Systems
prescribing the frequency relationships among pitches
within an octave. Musical scales can be
arbitrary and reflect compositional / aesthetic
choices. Such choices, more often than not, are guided by the properties of the
musical instruments available and by cultural
standards, rather than being determined by the harmonic series, the
construction of the ear, or some other 'natural'
factor.
Major and minor
diatonic scales
contain 7 out of the 12 available pitches, distributed so
that they include five whole-tone intervals
and two semitone intervals.
As was
demonstrated mathematically [
Balzano, 1980
], this pitch
configuration supports maximum
variety (i.e. maximum interval-combination possibilities)
from only 7 pitches out of an octave.
Diatonic scales divide the octave in 12 log-frequency units
[ 21/12 - equal temperament ], select 7,
and distribute them in a way that
facilitates reproduction of all intervals
available in the 12-tone chromatic scale (see image to the left).
The principle of coherence (i.e. no interval
between any two successive scale notes/degrees should be
larger than the sum of any two successive intervals) is closely related to maximum intervalic variety
and the possibility of developing fully functional
harmony within a scale.
The unique sonic
character or 'signature sound' of different pitch
intervals outlines musical categories that support
prediction during listening. Combinations of these
signature sounds can create cognitively recognizable
patterns, which outline higher level perceptual
groupings in the form of larger scale categories
or schemas (e.g. chord progressions).
The key features of western
art music tuning
and scale systems (i.e. octave divided in 12 log-equal units
and organized diatonically) attempt to optimally
address: _ cognitive limits (short term memory/Miller's rule),
_ data reduction/pitch circularity (more later
in the course), and _ maximum variety (interval distribution within
the major/minor
scales).
|
|


Postulated Relationships among preference, interest, and complexity [ after
Berlyne, 1971 ].
|
Complexity / Prediction / Preference
/ Interest
W. Wundt [ 19-20th century
German physiologist, psychologist, and philosopher ]
postulated a relationship among complexity, preference, and
interest, illustrated to the left. He proposed that emotional
responses are driven by perceived degree of stimulus intensity,
complexity, and novelty. _ Moderate levels of complexity produce positive emotions.
_ Too little complexity can lead to boredom and
disinterest. _ Excessive complexity can result in
overstimulation and negative emotions.
Consequently, preference decreases when
there is extreme simplicity (redundancy; little information;
maximal order; low entropy), as well as when there is
extreme complexity (randomness; too much information;
minimal order; high entropy).
A system has maximum entropy when all possible events within
the system are equally likely. If such a system involves a
large number of possible events, then prediction becomes
difficult, resulting in a high level of uncertainty. A musical example of maximum entropy is strict
dodecaphonic, serial music which -to some- does not even qualify as
music (e.g.
Excerpt from 'Symphony No. 21 - Opening' by Anton Webern).
Maximum
preference corresponds to an optimum level of
complexity, which is -in general- different for each
individual and changes with previous knowledge of
the system in question. Listeners, for example, prefer music with an
intermediate level of predictability; not so predictable that
it's boring, but not so unpredictable that it's unintelligible.
Interest follows a similar relationship but with maximum
interest corresponding to an optimum level of complexity
that is -in general- higher than that for maximum
preference. That is, we remain interested in levels of
complexity that are on average higher than we'd prefer.
In a sense, successful / good art strikes a balance between
low and high entropy, but always within a context. A
system's degree of entropy and its relationship to interest
and preference is relative to the observer. Analogously to
the processes of schematizing and categorizing, familiarity
with a certain style, artist, or work of art influences what level of
complexity a listener is likely to consider optimal. This may explain individual response
differences when experiencing works of art in general and
music in particular, as well as individual differences on
what constitutes successful / good music.
Repeated exposure to preferred works does not seem to
reduce their enjoyment. However,
“…it seems
likely that there may be changes in an individual’s preferred
complexity level with the passage of time. For example, a piece
of music which is initially too complex for an individual to
like, may, with repeated playings, move down to a lower
complexity level at which liking may begin to emerge.” [ in
J.B. Davies's 1978, now classic monograph, "The Psychology
of Music" ]
|
Musical complexity generally
refers to
_ the structural intricacy of a piece, in terms of
rhythm, harmony, melody, texture, and/or form, and
_ the degree
of unpredictability, density, or richness of
musical information.
-
Rhythmic complexity: syncopation; irregular
meters; time signature changes; etc.
-
Harmonic complexity: dissonance; unconventional chord progressions;
atonality; etc.
-
Melodic complexity: wide intervals;
intricate ornamentation; thematic & tonal variety; etc.
-
Textural complexity:
rhythmic, harmonic, melodic, and timbral layering; textural
contrasts; etc.
-
Formal complexity: departure
from standard or repetitive forms; etc..
Musical predictability refers
to information content that helps reduce uncertainty and
counter musical complexity, achieving a degree of balance
appropriate to a given listener and context. The devices
that improve the scanning efficiency of complex (i.e.
multi-stream) musical pieces, mentioned earlier, also
support predictability and
include:
-
Repetition;
-
Gradual introduction of
musical layers;
-
Selective variation of one or two
layers at a time, while the rest remain unchanged;
-
Selective attenuation/removal
and
amplification/reintroduction of layers in time;
-
Assignment of different contours/timbres/registers per
layer, etc..
For example,
North & Hargreaves (1995) explored the
complexity/familiarity/enjoyment relationship with
75 students listening and rating 60 excerpts from popular
(at the time)
songs. Artists included David Bowie, Simple
Minds, Duran Duran, Enya, Japan, Tangerine Dream, and many
others, covering a wide range of what is considered popular
music. When rating 'complexity,' listeners were provided with
a broad definition, based on: "how easy it is to predict what the
music will do next" vs. "how many surprises the music
contains." The results were compatible with Wundt's
model, showing, highest preference for intermediate levels
of complexity.
Overall, listeners seem
to reliably
prefer music of intermediate
complexity, with
preference for more predictability rising with context
uncertainty
[
e.g. Gold
et al., 2019 ].
|
|

Example of simplicity/complexity balance:
"Wouldn't It Be Nice" by the
Beach Boys (1999 remaster and
official video - first released in 1966) |