What is
tone of voice?
How pitch, texture and timing shape what we hear in a voice.
Emphasize “blue” in “I asked for the blue one,” and the sentence contrasts colors. Emphasize “I,” and it contrasts people.
We call differences like this “tone of voice.” We also use the phrase for a voice that sounds excited, tired, formal or hesitant. It is a useful umbrella term, but it bundles together several properties of speech and the impressions we form from them.
To understand those connections, we ran a listening experiment. Across more than 205,000 usable ratings of over 61,000 clips, we asked which acoustic properties were associated with listeners’ emotion and speaking-style judgments. The results connect familiar impressions to measurable parts of the sound, while showing how much an individual cue leaves unexplained.
A sentence has a shape.
Prosody is the patterning of pitch, duration and intensity across speech. It creates melody, rhythm, emphasis and phrasing. These patterns can express emotion or social meaning, and they also organize the message itself.
Jennifer Cole’s review of prosody in context describes two central aspects of this organization: prominence, which makes parts of speech stand out, and phrasing, which groups speech into units. A pitch movement, a longer sound, greater intensity or a pause can help establish those patterns. What the pattern means depends on the words and the conversation.
Changing the emphasis in our sentence makes this audible distinction visible. The words stay the same; their relative prominence changes.
i
Drag left or right to move the pitch prominence; drag upward to strengthen it. The six words remain fixed. This ribbon is a schematic contour, not measured speech. Strand depth is decorative. Real prominence can also change duration and intensity. Emphasizing “I” can contrast speakers; emphasizing “blue” can contrast colors.
The ingredients of a voice.
Pitch height, pitch variation, brightness and timing describe different things. A high voice can have a narrow pitch span. Two sounds can share a pitch and still differ in brightness. Separating these dimensions helps us ask what each contributes to a listener’s impression.
The explorer below isolates the acoustic ideas. Its shapes are illustrations, not recordings or measured effects from our experiment.
A generated sound, unfolded across time and frequency. Height is the square root of power; the same axes and vertical scale stay fixed. Drag horizontally and vertically, or use the two sliders.
Spectrum: pitch moves harmonic spacing; brightness changes the balance of harmonic power. Its readout is the model spectrum’s power-weighted centroid at the median pitch, not an estimate from the full sound. Rhythm: pulses change the number of amplitude swells; quiet changes how much of the envelope sits below 15% amplitude. Contour: span widens the pitch path in semitones (p10–p90); harmonicity sets the generated periodic-to-noise power ratio.
Listen plays four seconds of this synthetic setting at a limited level. Changing a control stops playback. The visual shows a smooth noise floor; the sound uses low-pass random noise at the same component power ratio. Neither represents a speech recording, a listener judgment, or a model prediction.
Pitch and melody
In voiced speech, the main acoustic correlate of pitch is fundamental frequency, or F0: the repetition rate of vocal-fold vibration, measured in hertz. Marc Garellek’s account of voice production explains how that vibration produces the sound. Pitch level describes where F0 typically sits; pitch span describes how widely it varies.
A pitch contour traces the rises and falls over time. Intonation concerns how that melody is organized linguistically, including where pitch movements occur and how phrases end. Pierrehumbert and Hirschberg’s account of intonational meaning connects that organization to how an utterance fits into a conversation.
Brightness and texture
A voiced sound contains harmonics at whole-number multiples of F0. The spectrum describes how energy is distributed across frequencies; the spectral envelope is its broad shape. We measured brightness using the spectral centroid, an energy-weighted average frequency. Think of a balance point that shifts upward when more energy lies at higher frequencies.
Brightness also reflects the sounds being spoken and the recording conditions, so it does not identify a particular vocal gesture. For periodic regularity, we measured harmonics-to-noise ratio, or HNR. Higher HNR means more periodic energy relative to noise, not higher pitch. It captures one part of voice quality, which includes qualities such as breathiness, creakiness and a pressed sound.
Timing and energy
The amplitude envelope follows broad changes in signal strength over time. We counted distinct peaks in that outline per second: detectable swells of speech energy. Related work by de Jong and Wempe uses intensity and voicing to detect syllable nuclei. Our peak measure is different and does not directly count syllables or pauses.
We also measured the fraction of a clip at or below an energy threshold. This includes silence and quiet speech. Intensity itself is acoustic power per unit area; perceived loudness also depends on frequency content and listening conditions. These measurements capture rough pieces of temporal organization rather than its full linguistic structure.
What listeners heard.
Participants heard speech clips and selected the emotion and style labels that fit. More than one label could apply. After excluding ineligible responses, the listening dataset contained 205,537 ratings of 61,040 clip IDs, from 3,031 participant sessions.
For each clip, we calculated the fraction of eligible ratings that selected each label. This preserves disagreement and overlapping impressions instead of forcing every voice into one category. We analyzed ten judgments, including excited, neutral, hesitant and formal.
The acoustic analysis used a smaller set with complete measurements and sufficient clips per speaker key. We kept speaker keys separate across training, validation and test sets. The main test contained 7,513 clips from 267 keys. A key sometimes identifies a speaker within one recording, rather than a globally unique person.
The connections below show each cue’s association with a judgment after accounting for the other five cues, task controls and stable speaker-key differences. Each clip received equal weight. Select a cue or judgment to follow its connections. The initial view shows the 19 associations that passed the multiple-comparison correction and retained their direction across sensitivity checks.
Info
Each ribbon is one reported partial correlation between an acoustic cue and a listener judgment, controlling for the other five cues and the study’s covariates. These are small associations, not individual predictions or causal effects. Width encodes absolute r on a fixed 0–.15 scale; blue is negative and rose is positive. Curvature and node positions are layout only.
Values are the two-decimal estimates transcribed from Irene’s supplied plot. Rounded zeros use dotted gray hairlines; they are not precise zero estimates. Uncertainty intervals are not available for every pair in this visualization. Filtering uses the displayed rounded values.
Supported means Benjamini–Hochberg q < .05 across 60 comparisons, with matching directions in the alternative speaker-key split and the subset with at least two raters. Nineteen pairs met this criterion. The four starred judgments were exploratory additions. Quiet means low-energy fraction; Pulses means energy-envelope peaks per second, not syllables. HNR measures harmonicity.
Click or tap a name to pin its connections; click a ribbon to inspect its value. Click again or press Escape to clear. Keyboard users can tab through names and ribbons and press Enter or Space. Use arrow keys on the filter.
Several cues suggest animated delivery.
Higher pitch, brighter spectra and more envelope peaks per second were each associated with excited and energetic judgments. All six associations passed the correction across 60 comparisons and retained their direction in the sensitivity checks. The same cues contributed to both impressions; none uniquely identified either label.
Lower pitch was associated with neutral, deadpan and formal judgments. These effects were small. Even the strongest association, between pitch level and neutral judgments, was r = −.121, with a 95% interval of −.148 to −.092.
Brightness added some predictive information beyond pitch alone. Adding it reduced held-out prediction error by a further 0.82% for excited and 1.23% for energetic. Both intervals excluded zero in the main speaker split, but included zero in an alternative split. That additional benefit was modest and split-dependent.
Four explicit cue interactions, including pitch × brightness and low-energy fraction × envelope peaks, did not consistently improve on the six-cue additive model. The results do not establish a particular formula listeners use to combine acoustic cues.
Hesitation has a timing pattern.
It is tempting to assume that more silence makes a voice sound more hesitant. Our results point to a more specific temporal association: clips with fewer energy-envelope peaks per second received more hesitant judgments.
This association remained after accounting for pitch level, pitch span, brightness, low-energy fraction, harmonicity, task controls and stable speaker-key differences. The conditional correlation was −.088, with a 95% interval of −.109 to −.065. It survived the correction for 60 comparisons.
How to read this
The weave is synthetic. Dragging changes the number of energy swells in a four-second illustration while keeping the low-energy fraction at exactly 50%. It does not predict an emotion or simulate the study’s recordings. The moving seam marks playback position.
The ribbon below shows a measured conditional correlation with hesitant judgments. Its ends are the actual 95% speaker-cluster bootstrap interval, and the dot is the estimate. Its decorative folds do not encode a distribution. Fewer envelope peaks were associated with more hesitant selections; low-energy fraction did not survive correction across 60 comparisons (q = .086), even though its pointwise interval narrowly excludes zero.
Estimates adjust for the other five cues, task controls, and stable speaker-key differences. Test sample: 7,513 clips and 267 speaker keys. Neither measure directly counts pauses or syllables; low energy includes quiet speech.
Fewer peaks mean fewer detectable energy pulses per second. Those pulses can reflect syllabic organization, articulation and stretches of quieter speech. They do not tell us exactly which linguistic behavior listeners responded to.
The evidence for low-energy fraction was weaker. Its association with hesitant judgments did not survive the multiple-comparison correction (q = .086) and was unstable across sensitivity checks. The supported result concerns the rate of energy swells. We cannot summarize it as “more silence means more hesitation.”
The sound still needs a context.
The patterns extend beyond animated delivery and hesitation. Narrower pitch span and less bright spectra were associated with deadpan judgments. Those cues, together with fewer energy peaks, were associated with tired judgments. Formal judgments involved lower pitch, more energy peaks and higher harmonicity; confident judgments had a small association with more energy peaks. Higher harmonicity was also associated with neutral judgments.
These overlapping associations explain why “tone of voice” covers so much. A listener hears acoustic properties together with the words, the surrounding speech and what they know about the speaker. The same sound can acquire different social meanings in different settings.
Our experiment identifies small associations between measurable properties and listeners’ judgments. It does not establish causes: listeners heard complete speech, and we did not manipulate individual cues. Six acoustic summaries cannot fully explain an interpretation. Understanding tone of voice means asking which properties support an impression, and how the conversation makes that impression possible.