Skip to content
oruk

Speech acoustics / Interactive essay

What a smile does to a voice

How mouth shape changes speech, and what 546 clips reveal about the happy labels inside our models.

An acoustic smile, drawn as a field of spectral contours. The ripples follow rising resonances in a simple voice model, with the source pitch held fixed. A stylized illustration of how mouth shape can change sound.

You can hear a smile in a whisper.

Ordinary whispering has no regular vocal-fold vibration. Yet in experiments by Tartter and Braun, listeners still heard smiled speech as happier. Pitch alone cannot explain that result.

A smile changes the instrument that makes a voice. We wanted to know how that physical change relates to the happy labels in our speech models. We started with the acoustics, then followed Fourier’s happy tags through two oruk encoders.

Both encoders retained a distinction that a simple classifier could read. Removing one direction from Fourier’s internal activity lowered its happy score. But half a second of silence also changed 161 of 546 happy-tag decisions. To explain the model, we have to account for both results.

The mouth changes which frequencies stand out

In voiced speech, your vocal folds supply a sound source. Your throat and mouth shape it. The repetition rate of the folds is the fundamental frequency, or F0. The resonances of the vocal tract are formants: F1, F2, F3, and so on. Moving the pitch and moving the formants change different parts of the sound.

Pull the lips back and you change the opening at the end of the tract. You can also shorten its effective acoustic length. A shorter tube has higher resonances. In 1980, John Ohala demonstrated the idea with plasticine models driven by an artificial sound source. A tube with its mouth corners pulled back behaved acoustically more like a shorter tube.

A human mouth has a tongue, a jaw, and several connected cavities. The effect depends on the vowel and on what the other articulators do. The simple tube explains why a smile can affect resonance. It cannot tell us exactly how much a particular person’s voice will change.

01 / The sound source and the mouthInteractive illustration

Move the pitch. Then move the resonances.

Harmonics: multiples of pitchFilter: resonances of the tract
Pitch and resonance can change separatelyIllustrative spectrum with a 140 hertz pitch and resonances shifted by 0 percent. Blue harmonic lines move with pitch. Teal resonance peaks move with the filter control. Relative amplitudes are schematic.01,0002,0003,000F1F2F3Hz

Pitch sets the spacing of the blue lines. The filter sets where the teal peaks sit.

The blue lines are harmonics. The teal curve is an illustrative spectral envelope. Change either control while holding the other still. These peaks are drawn to explain source and filter; they are not a measured vowel or a simulation of a particular smile.

Vocal-tract anatomy also sets a baseline. Serrurier and Neuschaefer-Rube built a vocal-tract model from MRI data covering 41 speakers. Their acoustic simulations linked longer tracts to lower formants. That study concerned anatomical differences across people, not smiling. It explains why we should compare expressions within speakers before attributing a difference to the expression.

Kim and colleagues watched the tract move during emotional speech. Ten actors repeated a digit sequence inside an MRI scanner. Happy delivery usually had a shorter estimated tract than angry or sad delivery. These scans capture the combined movements of the speech system. They do not isolate the contribution of a smile.

Pitch and mouth shape both contribute

Robson and Mackenzie Beck trained eleven speakers to spread their lips using anatomical instructions that named no emotion. Video confirmed wider mouth openings. Fifteen listeners compared those recordings with the same speakers’ ordinary versions. They usually chose the spread-lip version as the one that sounded more smiled. Listeners heard the two versions back to back.

To separate pitch from mouth shape, Simon Stone and colleagues used an articulatory synthesizer in 2022. They generated 15 German sentences with a higher pitch, altered vocal-tract geometry, both changes, or neither. The geometry change combined less lip protrusion with a raised larynx and adjustments to keep the vowels recognizable.

Fifty-seven listeners judged whether each voice sounded smiled. Either change increased smile judgments. Together, they increased them further. The individual pitch and geometry conditions did not differ significantly.

Smile judgments: baseline 20.9 percent, raised pitch 38.1 percent, altered tract 39.8 percent, both changes 61.4 percent.
Published result from Stone et al. (2022), Table 4. Pitch rose by two semitones. Values are smile judgments across 57 listeners and 15 sentences. Participant-level responses were unavailable for this redraw, so no uncertainty intervals are shown.

The effect varies in conversation. Barthel and Quené examined 1,242 word tokens from filmed Dutch conversations. They matched words within speakers and marked visible smiles. The second formant rose for the rounded vowel in their /oː/ category, by about 93 Hz. They did not find the same effect for neutral or spread vowels. Pitch rose significantly in ordinary words, but not in the brief acknowledgment ja.

Keough and colleagues also found that the vowel mattered in English. Ten speakers widened their lips and raised their pitch while smiling, without a significant change in measured larynx height. F1 and F2 rose significantly in “caw.” Neither showed a significant change in “key” or “coo.”

The filter can also be changed after a voice is recorded. Pablo Arias and colleagues altered the spectral envelope of recorded voices while preserving pitch, content, and speed. Listeners heard more smiling. Their validation also increased ratings of both joy and irony. The same acoustic cue can support more than one reading.

A voice that sounds smiled need not come with a visible smile. Drahota and colleagues compared listeners’ judgments with facially coded recordings. Some acoustic cues were heard as smiles even when a smile was not coded on the face. That gap matters when we move from measuring a gesture to assigning a label to its sound.

Hold the sentence still

The studies above connect speech with listeners’ judgments. Our model analysis starts from the model’s own happy tag. We use 546 CREMA-D clips from 91 actors. Every clip contains the same sentence: “It’s eleven o’clock.” The actors deliver it in different ways. This holds the words constant while allowing the acoustics to change.

We use the same Fourier emotion checkpoint as the pinned production version. We group clips by whether its native output includes happiness: 247 clips do, and 299 do not. A clip without that tag can sound angry, sad, or something else. Here, “happy” means the model’s tag. We have not coded the actors’ faces.

The label is broad. Fourier tags 74 of the 91 acted-happy clips as happy. It also tags 71 of the 91 acted-neutral clips. Across all deliveries, 162 clips receive both happy and neutral tags. These are overlapping outputs from a model with 15 emotion labels. The actors’ instructed deliveries are a separate annotation. The actor instructions and model tags divide the recordings differently.

Of 91 clips per acted delivery, Fourier tagged happy: 74 happy, 71 neutral, 37 fear, 23 disgust, 21 anger, and 21 sadness.
Each row contains one delivery from each of 91 actors. Bars show how often Fourier’s score crossed its native happy threshold of 0.375. The row labels describe the actors’ instructions, a separate annotation from the model’s tags.

The label is readable near the input

We read the model’s activity at successive depths. A small classifier, called a linear probe, tries to recover the happy-label distinction from each representation. We split the actors into five groups. Each probe trains on four groups and is tested on the fifth. No actor appears on both sides of a probe split.

The test measures how well each representation supports Fourier’s own decision. It does not independently validate the tag. Training overlap is also unverified: an actor held out of our probe may have appeared in the encoder’s original training data.

The probe reaches 0.925 ROC-AUC at the input state and 0.971 after the fourth encoder block. At the model’s native pooling stage, it reaches 0.976. ROC-AUC measures ranking: 0.5 is chance and 1.0 is perfect ranking of the happy-tagged clips above the others. The learned input state already supports a strong readout.

Fourier ROC-AUC rises from 0.925 at its input state to 0.971 at block four and 0.976 at native pooling. Seven acoustic features score 0.847; the log-mel summary scores 0.683. Chance is 0.5.
Switch models to inspect their representations. Both probes predict Fourier’s tags on the same 546 clips with the same five actor-disjoint folds. Intervals use 500 actor-bootstrap resamples, conditional on fitted probes. “Native pool” is Fourier’s attention-weighted summary. Blue baseline rows use audio features. Models differ in training, architecture, and feature dimensions.

Fourier’s tags are also predictable from Orukeet, a transcription model. The probe reaches 0.899 AUC at its speech encoder’s final layer. The highest observed AUC is 0.941 at block seven.

Some of that distinction is easy to extract from the audio. Seven features, including pitch, loudness, spectral brightness, and duration, reach 0.847. In a follow-up check, duration alone reaches 0.844. Happy-tagged clips have a median length of 2.00 seconds; the others, 2.44 seconds.

To test what those features explain, we used them to predict each hidden-state summary and subtracted that linear prediction. We fitted the correction on training actors only. The fourth-block probe fell from 0.971 to 0.822 AUC. The remaining signal could reflect other acoustic features or relationships this correction missed.

Change the activity, then read the answer again

A probe can read information that the original model does not use. We therefore edited one direction associated with the happy tag and checked the model’s own output. For each training actor with clips in both groups, we subtracted the average untagged representation from the average happy-tagged representation. Averaging those differences gave us one direction in the model’s 384-dimensional activity.

On actors outside that training group, we removed the component along this direction from each time frame, before the model pooled the sequence into its final summary. Then we ran the original output head again. The audio stayed identical.

Among the 247 originally happy-tagged clips, the mean happiness score fell from 0.660 to 0.315. That is a drop of 0.345 on the model’s zero-to-one score scale. The 95% actor-bootstrap interval for the drop runs from 0.325 to 0.365. About 70 percent of those clips lost the happy tag.

We then moved each frame by exactly the same distance in 64 random directions. Those controls lowered the score by 0.011 on average. Their mean tag-removal rate was about 3 percent. The learned direction had a much larger effect than the 64 random directions we tested.

Removing the learned direction decreases the mean happy score by 0.345, with a 95 percent interval from 0.325 to 0.365. Sixty-four matched random perturbations decrease it by 0.011 on average.
Each gray dot is the mean score change for one random direction across the same 247 clips. Vertical jitter only separates the dots. The blue diamond is their average. The teal interval uses 4,000 actor-bootstrap resamples. All directions were applied to probe-held-out actors; controls match the target’s perturbation norm at every frame.

Other labels move too. On those same clips, the neutral score falls by 0.192 and disgust rises by 0.188. The edit also changes how the model weights time frames.

Half a second of silence changes the label

The happy tag also changed when we left every speech sample intact. Adding half a second of digital silence before each recording flipped 161 of the 546 happy-tag decisions: 153 clips lost the tag and eight gained it. Moving the same silence to the end flipped only 32.

Halving the waveform’s amplitude flipped 104 decisions, with 103 clips gaining the tag. The original vocal gesture stayed the same; the recording level changed.

Of 546 clips: leading silence removes 153 happy tags and adds 8; trailing silence removes 3 and adds 29; half amplitude removes 1 and adds 103.
Exploratory controls on all 546 clips. Silence consists of 8,000 zero samples at 16 kHz, placed before or after the original waveform. The gain control multiplies samples by 0.5, a reduction of about 6 dB. We rerun the complete Fourier input pipeline and compare its native happy tag with the original decision.

The missing measurement is the face

A physical smile can leave an audible trace. Our encoders retain information about Fourier’s happy tag, and one learned direction changes Fourier’s output. Connecting that direction to the face requires measured facial movement.

We would record the same speakers saying the same words, track their lips, and compare matching vowels. We would keep the silence around the speech and the recording level fixed, then shift pitch and the spectral envelope separately. Those comparisons would test which measured movements and acoustic changes move the representation along the direction we found.

Methods and reproducibility

We froze the primary protocol before examining these outputs. The sample contains one CREMA-D IEO clip per actor and acted delivery: 91 actors × six deliveries. Non-neutral clips use the corpus’s medium-intensity setting; neutral clips have unspecified intensity. The Fourier checkpoint is fourier-gabor-api15-v1-s43, with 15,481,373 parameters. Its SHA-256 begins 564d1e42e4cf0b93. Local CPU inference used FP32 and evaluation mode. All 546 happy-tag decisions matched saved production API receipts from September 7; small score differences remain between inference precisions.

Probes use time means and standard deviations over valid frames, or native attention pooling, followed by train-only standardization and logistic regression with C = 0.01 and balanced class weights. We report pooled out-of-fold ROC-AUC. Feature counts and representations differ, so these baselines do not isolate the effect of architecture. Intervals resample actors with all their clips; they do not capture retraining uncertainty. The input state is a learned convolutional projection, not raw audio.

We derived each intervention axis from within-actor differences among training actors with clips in both groups, centered on their untagged mean. The random controls reuse the target axis’s signed projection coefficient while changing its direction. They are matched perturbations, not projections that erase random axes. We recomputed attention pooling and all 15 outputs. Repeated extraction was identical, and replaying the native head agreed to numerical precision.

Follow-up checks were exploratory. With 64 within-actor label shuffles and the entire probe refit each time, mean native-pool AUC was 0.533. Restricting the task to 85 happy-without-neutral and 80 neutral-without-happy clips gave native-pool AUC 0.992. Other emotion tags can still co-occur. This subset excludes 162 clips carrying both tags and 219 carrying neither. Using the source’s acted-happy versus acted-neutral categories instead gives fourth-block AUC 0.969, versus 0.569 from duration alone. Acoustic residualization subtracts a fitted linear prediction; it is not a causal adjustment. We did not measure vowel formants or facial movements.

Orukeet features come from all 24 encoder blocks of its staged NeMo checkpoint, using the identical audio IDs and probe folds. We exclude padding and summarize each block with time means and standard deviations. This run used CPU FP32; repeated extraction was identical. Model-training overlap is unverified for both checkpoints. The silence and gain controls apply to Fourier’s full local feature-extraction and inference path, so they do not locate the responsible component.

Download the aggregate results and provenance. The local research package includes the frozen protocol, input hashes, extraction code, per-clip results, and reproducible figure scripts.

Sources

  1. Tartter & Braun (1994). Hearing smiles and frowns in normal and whisper registers. Smile perception in whispered speech.
  2. Ohala (1980). The acoustic origin of the smile. Physical tube demonstration. The proposed evolutionary origin is a hypothesis.
  3. Serrurier & Neuschaefer-Rube (2023). Morphological and acoustic modeling of the vocal tract. MRI anatomy and simulated acoustics.
  4. Kim et al. (2020). Vocal tract shaping of emotional speech. Real-time MRI during acted emotional delivery.
  5. Robson & Mackenzie Beck (1999). Hearing smiles: perceptual, acoustic and production aspects of labial spreading. Anatomically instructed lip spreading and paired listening judgments.
  6. Stone, Abdul-Hak & Birkholz (2022). Perceptual cues for smiled voice. The four-condition experiment plotted above.
  7. Barthel & Quené (2015). Acoustic-phonetic properties of smiling revised. Visually coded smiles in conversation.
  8. Keough et al. (2015). Acoustic and articulatory qualities of smiled speech. Ten-speaker production study; preliminary conference report.
  9. Arias, Belin & Aucouturier (2018). Auditory smiles trigger unconscious facial imitation. Controlled spectral transformation.
  10. Drahota, Costall & Reddy (2008). The vocal communication of different kinds of smile. Facial coding compared with audio-only judgments.
  11. Cao et al. (2014). CREMA-D: Crowd-sourced Emotional Multimodal Actors Dataset. Source of the acted speech recordings.
  12. Hewitt & Liang (2019). Designing and Interpreting Probes with Control Tasks. Why probe performance needs controls.
  13. Elazar et al. (2021). Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals. Separating readable information from its effect on model behavior. Our single-direction edit is a different procedure.