
(common melodic "seeds" or "cells"; after Fuentes, 2020)

Example of melody generation using the
LSTM* and Markov Chain modeling (after
Bihani et al., 2023)
[*LSTM: Long Short-Term Memory, a type of recurrent neural network (RNN)
designed to handle long-term dependencies in current sequential data.] |
Melodic Units
The units of a melody are not individual notes. During a piece of music, perception
produces a coded version of the continuous flow of sound based on how our
implicit and explicit rules interact to identify accents and
creatively identify local or global features and boundaries.
A melodic unit refers to
a distinct section of a
melody; a small, self-contained set of notes that contributes to the overall melodic line.
A motif (or motive), a basic such unit, is a short,
identifiable pattern of notes; a "motivating idea...the small
cell out of which the music evolves." [ in Berry, 1986 ]. Out of the
many short melodic gestures in a piece of music only the few
that figure prominently in its growth are designated as motifs.
The figure to the left illustrates 24 common melodic motifs that
can be combined in various time organizations.
Larger melodic
entities include
"phrases" and "periods," made out of
multiple motifs that often end with a cadence (i.e. resting point). Listen to
20 iconic
film-music motifs.
Analogously to language phrases,
melodic 'phrases' are organized in structural units partly in terms of
breathing restrictions (3-5 seconds per phrase).
In
contrast to language phrases, melodic 'phrases' are also organized
in terms of patterning & redundancies (repetitions).
Redundancy fosters predictability which, in turn, gives rise to
expectation.
Consequently, examining melodies as sets of distinct, isolated notes is inadequate because it does not
allow for the patterning that supports the experience/prediction/expectation/game-of-expectations
affective cycle (more in Module 6).
When we break a whole down to a series of isolated bounded units
(known in mathematics as a Markov chain* of order 0) we destroy the level of
meaning that is based on syntax; that is, on the relationship among units that
binds them into a larger, single, bounded event.
In relevant experiments,
participants have been asked
to contribute 1, 2, 3, etc.. words to a text after reading 0, 1,
2, 3, etc. of the already existing words respectively.
[
e.g.
Miller & Selfridge, 1950
- a
similar experiment was conducted by Davies (1978), using notes
instead of words. ] The more words the participants were able to read (i.e.
the more developed the syntax & context) the more coherent and
meaningful were their contributed text and the resulting phrase. In other words,
the more we are able to reflect on the past,
act in the present, and project on the future, the
more meaningful our actions.
Such experiments
a) illuminate the statement "the whole is
different than the sum of its parts"; the meaning of phrases
is not just an aggregate of the meaning of the included
words;
b) highlight the feedback relationship between experiencing time
and experiencing music (explored later in the course),
as it relates to memory and expectation.
[ We
experience 'time' as a tension between
what has been (longing/regret for the past) and what might be
(excitement/fear for the future), in negotiation with the
'present' experienced by us and others. ]
In both
language and music, meaning is not simply lexical (i.e.
it does not simply depend on the words/notes used); it is
also syntactical. It depends on the way the words/notes are
patterned together; on how they relate to their past
(previous words/notes) and what they anticipate as their
future (words/notes to come).
Music syntax outlines
tension-resolution patterns that are creatively followed by listeners, as a piece of
music unfolds in time, and support the (largely implicit) assembly of
individual notes into
logical structures (i.e. chunking). Syntactical relationships
may be set up by the composer but are always mediated by what the
listener/reader brings to the experience.
Melodies are not random collections of notes
[ i.e. not Markov chains
of order 0 ] but larger units, tied together by
boundary-providing accents and syntactical
rules, set up by the composer AND configured by the listener(s). The bounded musical units
are outlined by accents
and incorporate redundancies, with periodic repetition being
key to musical syntax communication and perception.
[ * A Markov
chain is a
mathematical concept that outlines the probability of a system to
move towards a particular new state, given its current state, named
after 19-20th century Russian mathematician A.A. Markov. The concept
has been instrumental to major scientific and technological
developments, including nuclear power, search-engine functionality,
and artificial intelligence.
Watch this fascinating historical account.
More
here.
]
Sound events in time are organized in terms of pitch,
duration, etc. contours that track accents and outline potential musical
units.
|

Archetypal pitch contour examples

Descending major scale vs. "Joy to the World"

 |
Melody - Pitch &
Duration Contours - Rhythm
Melody
can be thought of as the superimposition
(layering/stratification) and interaction of pitch (or melodic) and
duration (or temporal) contours.
Pitch
(melodic) contour describes the pattern of pitch
direction changes within a melody: the pitch can either go up, down,
or remain the same. Points of change in pitch direction (i.e. pitch
contour inflections) are perceptually salient. They correspond
to contrasts, which become the accents that outline bounded
melodic units.
Pitch-contour periodicity outlines the pitch-contour clock.
Pitch contour similarity
corresponds to similarity in the sequence of pitch
contour inflections, even if the interval sizes differ.
Pitch changes in a melody that do not alter the pitch
contour are perceived as melodic variations rather than new melodies.
On
first hearing, it is the pitch contour that is stored in memory,
with the exact pitches being subsequently 'assigned' to this
contour [ e.g.
Graves
et al., 2019 ].
Several archetypal contours are
illustrated to the left.
[ Additional contour archetypes such as "gap-fill" and
"changing-note" (examples in Meyer, 1973; Rosen & Meyer,
1982) are directly related to gestalt principles
of perception, addressed later in the course. ]
If a pitch contour includes large pitch leaps (e.g. leaps of two or more scale steps), the
resulting melody is considered disjunct.
Otherwise (e.g. for single-step leaps in the scale), the resulting melody is considered conjunct.
Disjunct
melodies tend to be more difficult to remember and to
reproduce. Melodies with a balance between conjunct and
disjunct portions are usually preferred and are judged
as more interesting/exciting [ e.g.
Brantingham, 2013 ].
For a balanced example, listen to
"Here, There and Everywhere" by The Beatles.
Duration
(temporal) contour describes the pattern of sound and
silence duration changes: a sound/silence event can become
longer, shorter, or remain the same. Points of change in
'sound-event' duration (i.e. duration contour inflections) are
perceptually salient. They correspond to contrasts, which become
the accents that outline bounded temporal units.
Duration contour periodicity outlines the duration contour
clock.
Duration contour similarity
corresponds to similarity in the sequence of
duration contour inflections, even if the actual durations
differ. Duration changes in a melody that do not alter the
duration contour are perceived as melodic variations rather than
new melodies.
According to Dowling's melodic contour theory
[ in
Dowling, 1978 ],
at first hearing melodies are coded based on their pattern of
accents and are remembered in terms of the sonic motion implied
by their pitch and duration contours. NOTE: contour inflections become accents after they have occurred, pointing to the importance of the temporal aspects of musical organization.
Beat /
Meter / Tempo Every duration contour outlines a
beat (i.e. underlining
pulse of a piece of music). This is organized into
a meter (i.e. repeating accented patterns)
that reflects the
interaction among duration contour shape, clock, and beat.
Tempo describes the number of beats per unit time. The average spontaneous motor tempo (SMT) has been measured at
100 beats per minute (bpm) but with considerable individual
variations [ e.g. Fraisse,
1982, in
Wearden, 2024 ].
Duration vs. pitch contour
salience
Numerous studies suggest
that duration contours are more salient than pitch contours and may be more important in melody coding [ e.g. Monahan
et al., 1987;
Kendall & Carterette, 1990;
Palmer, 1996;
Schmuckler & Moranis, 2023 ].
For example, a descending
diatonic scale and the opening of "Joy to the World"
have identical pitch contours but can be easily
recognized as two different 'tunes' because of their
different duration contours.
Consider "America the Beautiful" (by Ward & Bates - Ray Charles's rendition).
The duration and pitch contours
are aligned (in 3s). If we superimpose a syncopated duration
contour in 2s (tango), the piece becomes unrecognizable, even if
the precise sequence of pitches remains unchanged (after Kendall
& Carterette, Unpublished).
Changing the duration contour of most melodies will likely result in what is
perceived as new melody, even if its
pitch contour is kept intact, (e.g. the "Star Spangled Banner"
modification presented in class).
Pitch
contour changes may also be sufficient to create a new
melody (even if the duration contour is kept intact), particularly when
the changes result in a:
a) significantly more disjunct/conjunct melody
and/or
b) misalignment between pitch and temporal contour
clocks.
(e.g. "Star Spangled Banner" vs. "Happy Birthday").
|
Rhythmm
is a collective property of a piece of music, emerging out of
the combination and interaction of all available sonic-contrast patterns and corresponding contours (i.e.
pitch,
duration, dynamic, and/or timbral) within the piece.
[ see
Bronzini, 2024 for an
outline of rhythm from a music theory perspective ] [ see
Ravignani et al., 2017
and
McAuley, 2010 for outlines of rhythm from a music cognition
perspective ]
Theme
variation often involves changes in some of the contours
while keeping the others relatively unchanged. This results
in a variety of melodic phrases that are still understood as
parts of the same composition.
During musical performance,
micro-deviations from the written/precise contours
function as a performer's means to communicate 'expressive
intent.' [ Listen to the verse's snare-tom pattern in
"Ticket To Ride" by The Beatles; Live at the
Hollywood Bowl, 1965. ]
The importance of
duration, dynamic, and other contour micro-variations
to musical expression has been documented extensively.
[ e.g.
Gabrielsson, 1988;
Kendall & Carterette, 1990;
Palmer, 1996,
Juslin, 2000 ].
Kendall &
Carterette (1990), for example, first illustrate that musical
expression works like all communication:
performers shape sound in ways that listeners
can reliably interpret. They then reveal that
temporal micro-variations —tiny speeding up,
slowing down, and stretching of notes— are more
significant in communicating expression than
small changes in pitch or loudness. Even when
other sound qualities were simplified, listeners
could still largely detect differences in
expressiveness based on timing patterns alone.
In other words, expression in music seems to be
carried primarily by how notes are timed.
Drawing an analogy to speech communication (i.e.
to communication via linguistic performance),
how something unfolds in time (e.g.
pauses, pacing, hesitation, rushing, lingering)
often communicates more emotion than raw content
alone (e.g. exact words/notes).
|
|

 |
Pitch & Duration Contour Interaction
It is easier to remember
melodies whose pitch and duration contour clocks 'line
up'.
[ e.g. Monahan
et al., 1987 ]
'Frére Jacques'
(French children's song) opens with a simple and conjunct
pitch contour, that has a regular
clock (4 notes per repetition).
The duration contour is also relatively 'flat', with only minor
duration changes (most pitches are quarter-notes
with no rests inserted), and the simplest possible periodic
contour clock
(1 note per repetition).
The resulting melody involves no conflict between pitch and duration contours.
Pitch and time accent structures align, resulting in a
simple and, at some level, uninteresting melody.
Similarly, 'Three blind
mice,' incorporates a conjunct
pitch contour with regular pitch contour clock (3 notes per
repetition) that aligns perfectly with the duration contour
clock (also 3 notes per repetition), resulting in another
example of a simple (and dull) melody.
Such pieces, along with
several nursery and pop tunes, are extreme examples of
simple melodies. In most other cases there is no perfect
contour alignment or clock periodicity, at least not for the
entire duration of a piece or theme, avoiding a level of redundancy high enough to make a piece not
only easily codified but also uninteresting (remember the
relationship between
complexity and preference).
|
 |
The
"Star Spangled Banner", for example
(Whitney Huston's 1991 Super Bowl performance), layers a
pitch contour clock in 2s over a duration contour clock in
3s, resulting in a relatively complex and interesting
melody. [ Watch
Jimmy
Hendrix in his iconic
1969 Woodstock performance ].
"Yesterday," by
The
Beatles, involves a more complicated relationship
between pitch and duration contours, with clocks that are
not perfectly periodic and melody lines that are frequently disjunct (i.e. increased complexity). Nonetheless, it is possible for
listeners to trace an overall arch-like contour that becomes
the piece's signature (i.e. decreased complexity). This
high/low complexity balance helps sustain both high interest and
high preference levels.
Layering multiple non-aligned
pitch and duration contours results in polyrhythms.
Musical
pieces in Jazz, renaissance, 20th century, and several
non-western traditions employ polyrhythms of varied degrees
of complexity.
Given that contour repetition
facilitates memory
[ e.g. Monahan
et al., 1987 ], pieces with complicated contour
relationships compensate for the information overload
through repetition.
[ e.g. "Star Spangled Banner";
"Yesterday" (The Beatles);
"Theseus and Minotauros" (Daedalus
Project);
"In the Mood" (Joe Garland); "A Night in
Tunisia" (Dizzy Gillespie) ]
|
Perceptual dimensions of melodies,
rhythms, and entire pieces
_ Single melodies
can be understood
in terms of primarily duration and pitch (and secondarily
dynamic and timbral) contour layering and
_ Musical pieces
can be understood as the layering of one or more melodies and/or
rhythms (or,
more generally, one or more "sonic structures'), with
contour clocks of various degrees of periodicity
[ see Patel, 2007; section on
Melody Statistics and Contours ].
Theorists have proposed several hierarchy-based music analysis methods
to address single- and multi-layered musical syntax
[ e.g. Schenker, Forte,
& Deutch; in
Tan et al., 2010a ]. From the perceptual and cognitive perspectives, it is all about
the interaction among pitch, duration, dynamic, and timbral
contours.
[ Detailed discussions on melody perception in
Butler, 1992
and
Tan et al., 2010b ].
Identifying the
perceptual boundaries (accents) that help define contours
depends in part on previous learning and experience. For example, a
musical listening task that appears simple and clear cut to a native of Indonesia may appear
complicated and unorganized to a Western trained listener, and
vice versa
[ E.g.
'Jaya Semara'. Indonesian Gamelan for Kebyar gong. UCLA Gamelan
Ensemble ].
Whether a pitch or temporal inflection will
constitute an accent (i.e. a salient point) depends not
only on physical contrasts but also on implicit rules of
data organization that utilize personal, stylistic, or
cultural schemas acquired through experience.
In
addition, when faced with melodies that exhibit
recognizable patterns on several levels, listeners will
organize them by focusing on the most salient level,
within the given context.
Implicit
rules assist us in coding accent patterns, resulting in a
melody being remembered in terms of the contout-implied melodic and rhythmic motions.
Motion
indicates more than shape (contour); it
indicates time. As previously noted, contour inflections become
accents after they have occurred, highlighting the
importance of the syntactical and temporal aspects of musical
organization.
|
|

Stimuli used to assess neural
response to harmonic anomalies (after
Maess et al., 2001).
(A) A sequence of five in-key consonant chords (key of C),
with the fifth chord highlighted in green.
(B) A sequence of the first four chords in (A), followed by a fifth
chord that contains two in-key notes
(F and E) and two out-of-key notes (A
flat and D flat), highlighted in red.
The four chords preceding the fifth chord set up a harmonic (syntactic) expectancy in the listener, which
the fifth chord fulfills in the case of (A) but violates in the case of (B).


(Jazz harmony details here) |
Harmony
Musical harmony refers to
the simultaneous combination of two or more notes and to the
progression of such combinations in time. Two-note combinations
are called dyads or harmonic intervals. Combinations of three or
more notes are called chords (they include triads, four-part
harmony, etc.).
Listener experience and
culture-dependent music theory rules, outline chord
relationships that support the creation of chord progressions
with specific harmonic perceptual impact. As chord
progressions unfold in time, they set up expectations
that composers/performers can
partially or completely fulfill/violate. This process
communicates patterns of tension and release that, at a basic level, constitute music's
intrinsic meaning.
Cultural musical norms
and the music theory rules
that codify them also outline pitch relationships, which support
melodies whose implied chord progressions have harmonic perceptual impact as
well.
Harmony
perception involves cognitive processes that, along with the auditory periphery,
engage brain regions
responsible for prediction, emotional response, and learning.
At
the auditory periphery level, harmonies and harmonic
progressions are analogous to isolated timbres and timbre contours,
respectively, and
are processed as
time-variant spectra.
At the neural level,
music interval
and chord recognition engages the brain's auditory cortex
AND the dorsolateral
prefrontal cortex, which is involved
in cognitive control and the balancing of emotional and
deliberative responses.
Harmonic perception relies heavily on working memory
to retain previous chords and anticipate
upcoming ones. The prefrontal cortex and
hippocampus are involved in storing and updating
harmonic progressions in real time, helping generate musical
expectations and emotional responses based on the outcome of these predictions.
Harmonic
structures also activate reward systems in the
brain, reflecting whether harmonies meet or defy expectations. The dopaminergic
system responds to these
predictions, producing pleasurable sensations when musical
surprises (e.g. unexpected modulations or cadences) are
resolved satisfactorily/plausibly.
At the cognitive level,
harmony perception relies on pattern recognition and tonal
schemas, developed as internal
representations of common tonal structures (e.g. major or minor
scales). When chords and harmonies match these schemas, the
brain processes them easier. It able to predict the
temporal unfolding of chord progressions based on
familiarity with musical styles (e.g. a dominant chord leading
to a tonic resolution).
Departures from established schemas violate stylistic expectations (see the figure
top-left) and elicit emotional responses that are manifested
and confirmed both behaviorally and neurologically [ e.g.
Limb, 2006 ].
The cognitive effort required to process harmony
depends on
the musical structure's perceived complexity. Complexity
can arise from:
_ dissonance,
_ ambiguous tonal centers,
_ unpredictable modulations or progressions, or
_ some other unexpected
harmonic feature.
The degree of perceived complexity largely depends on
a given listener's context (e.g. personal and cultural previous
experience).
Simple harmonic structures
rely on a limited menu of triadic chords (e.g. tonic,
dominant, and subdominant - in the simplest case: C, G, and F
triads in their root or inverted versions,
and common progressions with predictable resolutions. Such
features impose minimal
cognitive demands, even for
untrained listeners. By aligning with learned schemas, they
typically evoke familiar and comforting/dull emotions.
As expected from our discussion on the relative
nature of preferred complexity level, listeners without formal musical training
or sufficiently varied listenings tend to prefer structures
that accommodate their harmonic cognitive bandwidth. Most
harmonically simple songs that
enjoy widespread success make up for this
simplicity with complex arrangements,
varied contours, and skillful, "signature" performances, raising the average
complexity of the experience.
[ e.g.
"Get Back"
by The Beatles;
"Free Fallin"
by Tom
Petty;
"Songbird"
by Oasis;
"La
Bamba" by Richie Valence / performed by Los Lobos ].
Complex harmonic structures
use an extensive menu of chords (e.g. 7ths, 9ths,
diminished, hybrid, modal), chromaticism, modulations across
tonally distant keys, and ambiguous/unstable tonal centers. Processing
such features
requires greater cognitive effort.
Working memory must track
multiple harmonic layers, and the listeners are called to make
predictions in vaguer contexts.
Complex harmonies can evoke more
nuanced emotional responses that include tension, surprise, or even
discomfort. Experienced listeners may find pleasure in resolving
such emotions while novices may feel overwhelmed or alienated.
Most harmonically complex songs that
enjoy widespread success
balance this complexity with repetition, less cognitively demanding
contours, etc..
[ e.g.
"Maybe
I'm Amazed" by Paul McCartney;
"Josie"
by Steely Dan;
"Never
Gonna Let You Go", by Sergio Mendez ].
Occasionally, songs succeed in capturing the listeners'
attention and imagination in spite of maintaining complexity at
multiple levels.
[ e.g.
"Strawberry Fields Forever"
by The Beatles ].
A listener's cultural tradition
supports the development of cognitive frameworks that shape
harmonic perception. Perceived degree of musical complexity influences cognitive and emotional responses,
always within these frameworks.
As is
the case with most aspects of human experience, the more
varied a listener's musical exposure & practice, the broader
the gamut of musical experiences that may be considered
interesting,
preferable, and pleasurable.
|

 |
Tonality / Consonance-Dissonance /
Non-Western Music
Tonal harmony provides a learned
context that guides perceptual organization of melodic and
harmonic passages and supports operation of the various
gestalt principles of perception (addressed
later in the course). The concepts of tonality and harmony are
perceptually interdependent and weave an important framework for music making
and listening.
Tonality
describes the organization of pitch relationships around a central tone
(key/tonic).
It outlines how likely or unlikely a note is to be
included in a melodic or harmonic passage, given the
notes that have already been played/heard [ Patel, 2007;
Melody Statistics and Contours ].
[ Review of studies exploring the
neural basis of tonal processing in
Asano et al., 2022.
]
The complete set of
rules that guides tone relationships within a given
key is unclear but there is consensus on some
of them (e.g. tonic-dominant
relationship; major/minor modalities; musical
consonance/dissonance contrasts; cadences).
Geometric models
of pitch relationships such as Shepard's pitch spiral,
Krumhansl's pitch cone, or various versions of Pythagoras's circle of fifths
(see the figure bottom-left)
represent attempts to describe tonality's organizational
principles in terms of tonal hierarchies. Within the context of tonality,
the term harmony refers to
a) the range of melodic and harmonic expectations
outlined within a given key &
b) the specific, key-based melodic and harmonic implications of what has already
been performed.
Musical pieces that follow tonality
rules give a sense of direction/progression, particularly
but not exclusively to listeners familiar with the
underlying tonal framework, supporting the game of
expectations that seems to be at the basis of the affective
potential of all experience, musical or otherwise.
Musical consonance (representing
stability/release) and dissonance (representing
instability/tension) depend on many variables
and are often defined in the context of tonal hierarchy. In harmonic
passages, consonance also describes the degree of a harmonic
interval's/chord's pleasantness, fittingness, and/or perceptual
smoothness.
Musical consonance/dissonance judgments map
onto tension/release judgments. Within the
Western musical tradition, musical tension/release judgments
are linked to contrasts in
several aspects, including:
_ tonal center (e.g. key);
_ sensory consonance/dissonance (degree of
beating & roughness sensations);
_ dynamics; pitch; rhythm; timbre; orchestration;
_ performance techniques.
Numerous musical traditions base melodic
and harmonic development on musical consonance/dissonance contrasts,
with musical pieces structured around
musical consonance/dissonance 'contours.'
For example, in the opening of Leonard Bernstein's
"Maria" (from
West Side Story), the voice moves from a tritone melodic interval (dissonance) to a 5th
(consonance). Later on in the song, the voice follows a
similar pattern, moving from a major 2nd (dissonance) to
a major 6th (consonance). One can think of these
passages as having a similar consonance/dissonance
'contour.'
|
When listening to unknown pieces, listeners
trained in the Western musical tradition make
tonal assessments swiftly and often implicitly,
determining tonality based on the absence as
well as presence of certain intervals. Tonally knowledgeable listeners are able to extrapolate
tonal harmony information from incomplete tonal cues. Any tone one hears may suffice as a tonal center, until the
listener is probed by additional tonal evidence to opt for a
more plausible choice.
Several studies on
melodic motion and tonality judgments provide convergent
evidence that the
combination of consonant and dissonant
harmonic intervals conveys a tonal center
more efficiently than the use of consonant
intervals alone [ e.g. Butler &
Brown, 1984 ].
In
atonal contexts,* intervalic
similarity among chords does not translate to perceptual
similarity. This observation is consistent with the experienced
dissimilarity between major and minor triads in tonal music, both of
which
have the same intervalic content, just in different distributions.
Atonal chord similarity/dissimilarity
judgments are likely based on a combination of
musical context, previous experience, and the
spectral distribution of the chord-signals in
question, as atonal music has no hierarchical
tone structure on which to base similarity
comparisons. [ *atonal
context: melodic/harmonic context where all available notes are
equally likely to occur ]
Most
musical traditions in the
world employ hierarchical tone structures analogous to, even if
quite different from, Western tonal harmony [e.g.
Indian
ragas;
Arabic maqams;
Indonesian gamelan ].
Tone
hierarchies within each individual musical tradition are more
salient to members of the same tradition than to "outsiders."
At the same time, different traditions make equally strong claims of "good"
musical organization for widely differing musical structures.
Tone
hierarchies, musical scale systems, perceptual similarities/differences,
music pattern recognition, tension/release judgments,
aesthetic judgment standards, etc. may
therefore be
largely culture-dependent, although they are processed through common
cognitive principles.
Globalization has increased exposure to diverse musical
styles. Many listeners now enjoy music that combines
harmonic, melodic, tonal, and rhythmic elements from different traditions.
However, agreement on the intellectual and affective response to
hybrid styles still relies on the degree of shared cultural familiarity and exposure.
|