Skip to content
OrukResearch

Language & perception

A voice.
Many meanings.

How small differences in speech acquire social meaning.

When we hear someone speak, we hear more than words.

We also hear social information. We form impressions of where someone is from, their gender, their education, or the kind of person they might be. Those impressions can be mistaken, but they feel available to us because we have learned how the social world “sounds”, such as how people from different backgrounds speak, and how a person’s speech changes across situations.

Speakers have some awareness of this. We may change how we sound to fit into a group, take a particular stance, or present a different side of ourselves. Sociolinguists study these connections between linguistic variation and social meaning, or the inferences people make about speakers through the ways they speak.

01 / Meaning

A sound points somewhere.

A central idea here is indexicality, the capacity of a linguistic form to point toward something in the social world. A pronunciation can suggest a stance, a recognizable kind of person, or membership in a group. Where it points depends on the history of the form and on the experiences of both speaker and listener. A southern U.S. accent might sound “southern” to another American, but simply “American” to a British listener.

Penelope Eckert’s 2008 account of the indexical field describes a constellation of related meanings that can become relevant in different uses. The connections come from shared ideas about who speaks in particular ways. They can change as speakers combine features differently and listeners reinterpret them.

Consider a released /t/, or the audible burst of air you might hear at the end of an emphatic “cat.” Across different styles, it can suggest articulateness, precision, refinement, or emphasis. These possibilities help explain how the same feature appears in the speech of quite different social groups.

The indexical field

One feature.
Many possibilities.

Change the context

Learnedness and authorityBenor · 2001

About this figure ↗

A conceptual illustration of Eckert’s indexical field. A released /t/ can participate in different styles; none of these meanings is fixed to the sound. Brightness and position are visual emphasis, not measured probabilities. The context groupings illustrate the literature, rather than simulate any individual listener. Scholarship refers to Benor’s study of Orthodox Jewish English and the value of Torah learning. The public-speech findings varied with speaker identity and listener expectations.

Eckert (2008), Variation and the indexical field. Community-specific findings are described in the accompanying figures and source notes.

01 Select a context to explore possible meanings of released /t/. The connections are conceptual, not measured probabilities.

02 / The sound of /s/

One measure, different social worlds.

Following an individual phonetic feature through the literature makes this easier to see. Take /s/. Its noisy sound contains energy spread across frequencies. One common measure, the spectral center of gravity (COG), summarizes where that energy is concentrated as a weighted average frequency.

A higher COG generally goes with a more fronted, higher-frequency /s/. This is different from the fundamental frequency of vocal-fold vibration, which contributes to the pitch of voiced speech. Move the energy in the figure below to see what the spectral measure captures.

The acoustics of /s/

A center of gravity.

Spectral center of gravity
5.97kHz
An illustrative /s/ spectrumEnergy is distributed across frequencies. Its energy-weighted mean is currently 5.97 kilohertz. Move the slider to shift this distribution. The repeated contours give the same spectrum visual depth; they are not separate observations. This spectral measure is distinct from vocal-fold vibration, or F0.024681012Frequency · kHzRelative energyIllustrative spectrum
The measure, not the meaning

Center of gravity is a weighted mean of the frequencies in a spectrum. Here the weights are synthetic power values: Σ(frequency × power) / Σ(power). Studies may use amplitude or power weights, and their analysis settings matter.

This is an illustrative /s/ spectrum, not a recording or a dataset. The repeated contours add depth to a single distribution. Spectral center of gravity describes where a sound’s energy sits; it is not F0, the repetition rate of vocal-fold vibration, and does not identify a speaker’s gender or identity.

Context: Calder & King, Whose Gendered Voices Matter? (2022), as discussed in Irene’s accompanying article.

02 An illustrative spectrum. The center of gravity follows the energy distribution; it does not assign a social identity.

What does that measurement mean socially? The spectral characteristics of /s/ have well-documented associations with gender. Higher center of gravity is conventionally associated with femininity, and lower center of gravity with masculinity, although these meanings depend on the surrounding voice and style (Zimman, 2017). Higher-frequency /s/ can also contribute to perceptions of a male speaker as gay, with experimental evidence showing that these judgments depend on listeners’ beliefs about gender and sexuality (Levon, 2014). These associations are therefore neither mutually exclusive nor fixed. As Zimman (2017:362) explains, a high-frequency /s/ can acquire different gendered meanings when combined with a lower-pitched voice, contributing to a gay or queer masculine persona.

The feature of /s/, nevertheless, has different interpretations across different communities, changing due to a particular context’s history or circumstance. In Calder and King’s study of Bakersfield, California, White women produced higher-COG /s/ than White men, but among African American speakers, the researchers found no significant gender split. Treating higher COG as a uniform measure of femininity would misdescribe this community.

The authors interpret the difference through the local association of backed /s/ with White country masculinity. African American men may avoid a pronunciation tied to that historically oppressive local identity. This is an interpretation of a production pattern, rather than an experiment establishing what listeners hear in each voice. Even within one city, the same acoustic dimension does not organize gender in the same way.

Farther north, Podesva and Van Hofwegen’s Redding study found that older country-oriented women and several lesbian speakers produced comparably low-COG /s/. The authors argue that retracted /s/ can suggest a country orientation alongside associations with masculinity, non-normative femininity, or lesbian identity. The local setting helps explain which of those associations matters.

In San Francisco, J Calder followed /s/ through drag performers’ visual transformations. As the transformation progressed, COG decreased and /s/ became shorter, while intensity increased. The pattern involved several acoustic dimensions at once. Calder argues that increasing intensity helped project fierceness alongside the performers’ changing appearance. Reading COG alone would miss how the voice and visual presentation worked together.

These studies follow the same sound through different communities. A measurable difference in /s/ participates in each community’s social distinctions, without carrying one fixed meaning between them.

03 / Released /t/

Even a burst of air has context.

Released /t/ offers another compact example. Eckert discusses its place in styles associated with intellectual precision, professional articulateness, a meticulous persona, or even “prissiness”. It can also contribute to emphasis, and the duration and strength of a release burst matter too.

In D’Onofrio and Eckert’s experiment on affect, listeners heard acoustically manipulated versions of an audio clip saying “Todd” and evaluated a hypothetical speaker’s feelings about working with him. A longer release burst increased perceived negativity in the louder condition, but not in the quieter one. Even this small acoustic change depended on another part of the signal.

The figure below shows a schematic release burst; change its duration and intensity to explore how the two features affected listeners’ impressions in the experiment.

The release burst

A few milliseconds.
A different impression.

/t/
Release burst “Todd”
A 64-millisecond release burst at 75 decibelsA schematic sculpture of a speech waveform. Selecting a longer burst extends it horizontally. Selecting the louder setting makes it taller. The illustration is not an experimental recording.RELEASE64 msTIME →
Duration
Intensity
When the burst gets longer
More negativeAt the louder setting
Sources & context

D’Onofrio & Eckert (2021), Affect and iconicity in phonological variation ↗. Conditions and the qualitative comparison are summarized from Irene Yi’s account of the experiment: lengthening the burst from 25 to 64 ms increased perceived negativity at 75 dB, but not at 65 dB.

This schematic does not reproduce the stimuli, rating distributions, or effect sizes. “Negative” refers to listeners’ evaluations in the reported experiment, not a universal meaning of /t/.

03 Change duration and intensity to explore the reported comparison. The waveform is schematic; the outcome describes listeners’ impressions in the experiment.

An audible /t/ release can give an impression of care, but what that care signifies depends on the setting. In Benor’s study of Orthodox Jewish English, phrase-final /t/ release participates in the expression of learnedness and authority. Religious scholarship gives learnedness a particular social value in this community. Sounding knowledgeable and sounding entitled to make a claim are also distinguishable, and the pronunciation alone does not specify which interpretation applies.

Podesva, Reynolds, Callier, and Baptiste combined recordings of six prominent U.S. politicians with a perception experiment using manipulated speech. Released /t/ did not produce a uniform impression. Word-medial releases carried stronger social meanings than word-final ones, and listeners’ interpretations varied with the politician and their customary release rate. Associations with intelligence, education, and articulateness depended on who was speaking and what listeners expected.

A released stop cannot simply be assigned the meaning “careful” or “nerdy,” any more than a longer burst can always be assigned “negative.” The linguistic surroundings, the social setting, and the speaker all help determine the interpretation.

04 / Pitch & place

When a higher voice sounds tougher.

Fundamental frequency, or F0, measures the repetition rate of vocal-fold vibration and contributes strongly to perceived pitch. Researchers can manipulate it while preserving much of a recording, then ask whether listeners’ impressions change.

Average differences in vocal anatomy contribute to perceived gendered patterns in voice, where longer, more massive vocal folds generally support lower fundamental frequencies, while longer vocal tracts produce lower resonance frequencies. Adult men, on average, have lower F0 and lower formant frequencies than adult women (Titze, 1989; Pisanski et al., 2016).

These population-level associations, however, do not determine the social meaning of pitch in a particular setting. In a matched-guise study in Hawaiʻi with ethnically diverse speakers and listeners, Drager and colleagues found that changing pitch could alter impressions of body size in a direction that runs against the familiar association between low pitch and largeness, where some speakers sounded larger at a higher pitch. The effect depended on other characteristics listeners attributed to them as well.

For one speaker, Kent, higher F0 elicited descriptions of toughness, whereas lower F0 evoked a more easygoing, affectionate persona. The authors hypothesize that these perception patterns reflect a locally meaningful Native Hawaiian persona called moke, which is associated with Pidgin, surfing, rural life, pickup trucks, and toughness. M. Joelle Kirtley’s (2015) dissertation, Language, Identity, and Non-Binary Gender in Hawai‘i, similarly discusses how speakers in Hawaiʻi use high pitch to construct toughness.

Listening in context

Who is speaking?
And who is listening?

These are just three of the many ways acoustic features take on social meanings. Ethnographic work helps explain how features acquire significance within communities and styles. Perception experiments show how a change in sound can affect what listeners hear. Acoustic analysis becomes more explanatory when we connect it to both.

The question grows beyond “What does this sound mean?” It includes who is speaking, what they are doing, how a feature combines with other cues, and what histories speakers and listeners bring to the encounter. A pronunciation can have strong conventional associations and still take on a more specific, or competing, meaning in use. Social meaning helps explain how a measurable difference in sound becomes part of a social world.