Skip to content

What is
tone of voice?

How pitch, texture and timing shape what we hear in a voice.

A voice has more than one dimension. Conceptual artwork.

Emphasize “blue” in “I asked for the blue one,” and the sentence contrasts colors. Emphasize “I,” and it contrasts people.

We call differences like this “tone of voice.” We also use the phrase for a voice that sounds excited, tired, formal or hesitant. It is a useful umbrella term, but it bundles together several properties of speech and the impressions we form from them.

To understand those connections, we ran a listening experiment. Across more than 205,000 usable ratings of over 61,000 clips, we asked which acoustic properties were associated with listeners’ emotion and speaking-style judgments. The results connect familiar impressions to measurable parts of the sound, while showing how much an individual cue leaves unexplained.

A sentence has a shape.

Prosody is the patterning of pitch, duration and intensity across speech. It creates melody, rhythm, emphasis and phrasing. These patterns can express emotion or social meaning, and they also organize the message itself.

Jennifer Cole’s review of prosody in context describes two central aspects of this organization: prominence, which makes parts of speech stand out, and phrasing, which groups speech into units. A pitch movement, a longer sound, greater intensity or a pause can help establish those patterns. What the pattern means depends on the words and the conversation.

Changing the emphasis in our sentence makes this audible distinction visible. The words stay the same; their relative prominence changes.

Synthetic
i

Drag left or right to move the pitch prominence; drag upward to strengthen it. The six words remain fixed. This ribbon is a schematic contour, not measured speech. Strand depth is decorative. Real prominence can also change duration and intensity. Emphasizing “I” can contrast speakers; emphasizing “blue” can contrast colors.

The ingredients of a voice.

Pitch height, pitch variation, brightness and timing describe different things. A high voice can have a narrow pitch span. Two sounds can share a pitch and still differ in brightness. Separating these dimensions helps us ask what each contributes to a listener’s impression.

The explorer below isolates the acoustic ideas. Its shapes are illustrations, not recordings or measured effects from our experiment.

Pitch171 HzBrightness520 Hz
Synthetic

Pitch and melody

In voiced speech, the main acoustic correlate of pitch is fundamental frequency, or F0: the repetition rate of vocal-fold vibration, measured in hertz. Marc Garellek’s account of voice production explains how that vibration produces the sound. Pitch level describes where F0 typically sits; pitch span describes how widely it varies.

A pitch contour traces the rises and falls over time. Intonation concerns how that melody is organized linguistically, including where pitch movements occur and how phrases end. Pierrehumbert and Hirschberg’s account of intonational meaning connects that organization to how an utterance fits into a conversation.

Brightness and texture

A voiced sound contains harmonics at whole-number multiples of F0. The spectrum describes how energy is distributed across frequencies; the spectral envelope is its broad shape. We measured brightness using the spectral centroid, an energy-weighted average frequency. Think of a balance point that shifts upward when more energy lies at higher frequencies.

Brightness also reflects the sounds being spoken and the recording conditions, so it does not identify a particular vocal gesture. For periodic regularity, we measured harmonics-to-noise ratio, or HNR. Higher HNR means more periodic energy relative to noise, not higher pitch. It captures one part of voice quality, which includes qualities such as breathiness, creakiness and a pressed sound.

Timing and energy

The amplitude envelope follows broad changes in signal strength over time. We counted distinct peaks in that outline per second: detectable swells of speech energy. Related work by de Jong and Wempe uses intensity and voicing to detect syllable nuclei. Our peak measure is different and does not directly count syllables or pauses.

We also measured the fraction of a clip at or below an energy threshold. This includes silence and quiet speech. Intensity itself is acoustic power per unit area; perceived loudness also depends on frequency content and listening conditions. These measurements capture rough pieces of temporal organization rather than its full linguistic structure.

What listeners heard.

Participants heard speech clips and selected the emotion and style labels that fit. More than one label could apply. After excluding ineligible responses, the listening dataset contained 205,537 ratings of 61,040 clip IDs, from 3,031 participant sessions.

For each clip, we calculated the fraction of eligible ratings that selected each label. This preserves disagreement and overlapping impressions instead of forcing every voice into one category. We analyzed ten judgments, including excited, neutral, hesitant and formal.

205,537usable ratings
61,040clip IDs
3,031participant sessions

The acoustic analysis used a smaller set with complete measurements and sufficient clips per speaker key. We kept speaker keys separate across training, validation and test sets. The main test contained 7,513 clips from 267 keys. A key sometimes identifies a speaker within one recording, rather than a globally unique person.

The connections below show each cue’s association with a judgment after accounting for the other five cues, task controls and stable speaker-key differences. Each clip received equal weight. Select a cue or judgment to follow its connections. The initial view shows the 19 associations that passed the multiple-comparison correction and retained their direction across sensitivity checks.

−+
Follow a connection19 / 19
0|r| .15
Touch a name or a ribbon. Drag to filter.
Info

Each ribbon is one reported partial correlation between an acoustic cue and a listener judgment, controlling for the other five cues and the study’s covariates. These are small associations, not individual predictions or causal effects. Width encodes absolute r on a fixed 0–.15 scale; blue is negative and rose is positive. Curvature and node positions are layout only.

Values are the two-decimal estimates transcribed from Irene’s supplied plot. Rounded zeros use dotted gray hairlines; they are not precise zero estimates. Uncertainty intervals are not available for every pair in this visualization. Filtering uses the displayed rounded values.

Supported means Benjamini–Hochberg q < .05 across 60 comparisons, with matching directions in the alternative speaker-key split and the subset with at least two raters. Nineteen pairs met this criterion. The four starred judgments were exploratory additions. Quiet means low-energy fraction; Pulses means energy-envelope peaks per second, not syllables. HNR measures harmonicity.

Click or tap a name to pin its connections; click a ribbon to inspect its value. Click again or press Escape to clear. Keyboard users can tab through names and ribbons and press Enter or Space. Use arrow keys on the filter.

Several cues suggest animated delivery.

Higher pitch, brighter spectra and more envelope peaks per second were each associated with excited and energetic judgments. All six associations passed the correction across 60 comparisons and retained their direction in the sensitivity checks. The same cues contributed to both impressions; none uniquely identified either label.

Lower pitch was associated with neutral, deadpan and formal judgments. These effects were small. Even the strongest association, between pitch level and neutral judgments, was r = −.121, with a 95% interval of −.148 to −.092.

Brightness added some predictive information beyond pitch alone. Adding it reduced held-out prediction error by a further 0.82% for excited and 1.23% for energetic. Both intervals excluded zero in the main speaker split, but included zero in an alternative split. That additional benefit was modest and split-dependent.

Four explicit cue interactions, including pitch × brightness and low-energy fraction × envelope peaks, did not consistently improve on the six-cue additive model. The results do not establish a particular formula listeners use to combine acoustic cues.

Hesitation has a timing pattern.

It is tempting to assume that more silence makes a voice sound more hesitant. Our results point to a more specific temporal association: clips with fewer energy-envelope peaks per second received more hesitant judgments.

This association remained after accounting for pitch level, pitch span, brightness, low-energy fraction, harmonicity, task controls and stable speaker-key differences. The conditional correlation was −.088, with a 95% interval of −.109 to −.065. It survived the correction for 60 comparisons.

4energy swells
4 seconds50% low energy
Illustration
Drag the weave. Change the rhythm.
Measured association
−0.088
95% CI −0.109 to −0.065Supported · q < .001
partial r
How to read this

The weave is synthetic. Dragging changes the number of energy swells in a four-second illustration while keeping the low-energy fraction at exactly 50%. It does not predict an emotion or simulate the study’s recordings. The moving seam marks playback position.

The ribbon below shows a measured conditional correlation with hesitant judgments. Its ends are the actual 95% speaker-cluster bootstrap interval, and the dot is the estimate. Its decorative folds do not encode a distribution. Fewer envelope peaks were associated with more hesitant selections; low-energy fraction did not survive correction across 60 comparisons (q = .086), even though its pointwise interval narrowly excludes zero.

Estimates adjust for the other five cues, task controls, and stable speaker-key differences. Test sample: 7,513 clips and 267 speaker keys. Neither measure directly counts pauses or syllables; low energy includes quiet speech.

Fewer peaks mean fewer detectable energy pulses per second. Those pulses can reflect syllabic organization, articulation and stretches of quieter speech. They do not tell us exactly which linguistic behavior listeners responded to.

The evidence for low-energy fraction was weaker. Its association with hesitant judgments did not survive the multiple-comparison correction (q = .086) and was unstable across sensitivity checks. The supported result concerns the rate of energy swells. We cannot summarize it as “more silence means more hesitation.”

The sound still needs a context.

The patterns extend beyond animated delivery and hesitation. Narrower pitch span and less bright spectra were associated with deadpan judgments. Those cues, together with fewer energy peaks, were associated with tired judgments. Formal judgments involved lower pitch, more energy peaks and higher harmonicity; confident judgments had a small association with more energy peaks. Higher harmonicity was also associated with neutral judgments.

These overlapping associations explain why “tone of voice” covers so much. A listener hears acoustic properties together with the words, the surrounding speech and what they know about the speaker. The same sound can acquire different social meanings in different settings.

Our experiment identifies small associations between measurable properties and listeners’ judgments. It does not establish causes: listeners heard complete speech, and we did not manipulate individual cues. Six acoustic summaries cannot fully explain an interpretation. Understanding tone of voice means asking which properties support an impression, and how the conversation makes that impression possible.

Written by
How we measured it

Eligible responses

We included only sessions that explicitly passed the server attention check. We excluded test sessions; practice, attention and repeat probes; failed or flagged audio; and ratings without usable labels. The resulting listening dataset contained 205,537 ratings across 61,040 clip IDs from 3,031 participant sessions.

Acoustic measurements

Pitch level
Median voiced fundamental frequency (F0), transformed to semitones.
Pitch span
The distance between the 10th and 90th percentiles of voiced F0, in semitones. Twelve semitones is a doubling of frequency.
Spectral brightness
Median power-weighted spectral centroid across windows above the energy threshold.
Envelope peaks per second
Distinct peaks in the smoothed amplitude envelope, divided by the full clip duration.
Low-energy fraction
The proportion of short windows at or below the energy threshold.
Harmonicity
Median valid harmonics-to-noise ratio (HNR), using Praat’s cross-correlation method.

Modeling and uncertainty

We required complete acoustic measurements and at least ten eligible clips per speaker key, removed exact audio duplicates and excluded reserved reference clips. The final modeling sample contained 31,192 clips across 1,065 speaker keys. Training, validation and test sets were separated by speaker key. The main test contained 7,513 clips from 267 keys and 19,377 ratings.

Some keys identify speakers within recordings rather than globally unique people. Models weighted clips equally. Prediction comparisons tested whether cues added information beyond source corpus, duration, rating count and task-version composition. Tuning used only training and validation data.

We estimated conditional cue-judgment associations in the test subset, adjusting for the other five cues, task controls and stable speaker-key differences. We estimated uncertainty using 2,000 speaker-key resamples, checked an alternative speaker split and clips with at least two distinct raters, and applied a Benjamini–Hochberg correction across 60 comparisons. The 19 marked associations passed q < .05 and retained their direction in those sensitivity checks. A separate check accounting for both speaker and listener clustering retained the same 19 associations.

The ten judgments were excited, energetic, neutral, deadpan, hesitant, tired, confident, formal, sincere and sarcastic. The last four were added for this exploratory follow-up. Conditional correlations describe adjusted associations; they are not label probabilities or causal effects. The conceptual figures illustrate acoustic ideas, while the association weave and uncertainty intervals report the study’s results.

References
  1. Jennifer ColeProsody in context: a review
  2. Marc GarellekThe phonetics of voice
  3. Janet Pierrehumbert & Julia HirschbergThe meaning of intonational contours in the interpretation of discourse
  4. University College LondonIntroduction to intensity perception
  5. Nivja de Jong & Ton WempePraat script to detect syllable nuclei and measure speech rate automatically
  6. Praat documentationSpectrum: Get centre of gravity
  7. Praat documentationHarmonicity

A voice. Many meanings.

How a sound acquires social meaning.

Pitch contours and perceived emotion

The interpretations that follow a changing melody.

What speech encoders hear

Where speech models represent voice and emotion.

← More from the labBack to top ↑