Research
Notes from the speech lab
- 7 min read
When the words lie: measuring sarcasm in speech
Sarcasm inverts the words. On MUStARD, transcript sentiment drops below a constant classifier on unseen shows, while acoustic scores keep most of their ground.
Read → - 9 min read
A score you can check
The full accounting of how our speech emotion models score: 64 systems on our own benchmark with the conflicts stated, nine commercial APIs on IEMOCAP with identical audio, and sarcasm on MUStARD — with the methodology and every system’s failure modes, including ours.
Read → - 8 min read
The shape of a voice
Run a sentence through a speech model and it becomes a shape. We traced thousands of real recordings through four models in rotatable 3-D: emotions separate into threads nobody trained them to draw, four architectures agree on the layout, and Whisper rebuilds the vowel chart phoneticians drew a century ago.
Read → - 7 min read
Pitch contours and perceived emotion
A pitch contour can shift the odds of several interpretations without belonging to any one emotion. In a pilot study, six unsupervised contour families coincided with different distributions of listener labels, but no family pointed to a single answer.
Read → - 4 min read
Speech models learn physics
We examined convolution filters across six speech systems and found that training repeatedly produces Gabor-like shapes predicted by signal theory.
Read → - 3 min read
What speech encoders hear
We mapped every layer of four speech encoders and found that emotion, speaker identity, and sentence content settle at different depths.
Read → - 1 min read
Hello, world
Why oruk is starting a public speech research journal, what we plan to publish, and the standard we want the work to meet.
Read →
