Skip to content

Try Resonance 2 early.

Tell us a little about yourself and what you’re building.

A brief description of your use case.

We’ll use these details to review your application and contact you about access. Privacy policy

Explore oruk

Speech emotion

What is speech emotion recognition?

Speech emotion recognition (SER) estimates expressed emotion from the soundof a voice. It can distinguish vocal patterns in two recordings with the same words, such as a calm complaint and an irritated delivery. The result is a model annotation of how speech sounds, not a direct observation of someone’s feelings.

What emotion recognition models listen to

Prosody
Pitch contours, speaking rate, pauses, and stress patterns can help distinguish vocal expressions. Their meaning depends on the speaker and context.
Voice quality
Breathiness, vocal tension, and changes in loudness contribute to how a voice sounds. They do not establish whether an expression is sincere.
Temporal dynamics
Changes across a recording can reveal where delivery shifts. Timed annotations let a reviewer hear those passages in context.
Linguistic context
The words and preceding conversation help a listener interpret delivery. A flat "great, thanks" can sound different in different conversations; an acoustic label alone does not establish sarcasm or intent.

How speech emotion recognition is measured

SER systems are evaluated on labeled corpora such as IEMOCAP and MELD, usually reported as accuracy and weighted or macro F1 across emotion classes. Published numbers are hard to compare because papers use different label sets, splits, and scoring code. Start with the model and recordings you plan to use, and keep errors and abstentions in the reported sample.

Our September 7, 2026 Fourier API evaluation used 546 CREMA-D recordings and measured 55.7% accuracy and 0.542 macro F1 on six acted labels. It publishes input hashes, saved responses, an error analysis, and a scorer you can run without an API key. Training overlap has not been audited. This dated result does not establish accuracy on spontaneous conversations or for the current Resonance endpoint.

The separate historical speech-emotion-bench comparison uses the full 64,384 clips for open models across ~20 languages and 7 emotion classes; closed/API and audio-LLM systems use a fixed 5,000-clip subset. All 64 rows use the same label mapping and scorer, but not identical audio.

In the July 14, 2026 emotion run, oruk Resonance 1 measured 77.8% accuracy and 0.816 macro F1. The Oruk entry was trained in-distribution; the compared systems were evaluated zero-shot across corpora. This is a historical checkpoint comparison, not a current API accuracy ranking. The accuracy guide explains how speaker splits, label choices, and calibration change what a score can tell you.

Where speech emotion AI is used

Contact-center teams can score completed calls to surface moments of frustration for human review instead of relying on keyword matches alone. Voice-agent teams can analyze prerecorded evaluation calls to find interactions where the agent’s response should be revised. Health and research teams can review affect labels across recorded sessions while keeping interpretation with qualified people.

oruk exposes acoustic emotion recognition through one API

Emotion is one output of oruk’s speech understanding models, alongside English transcription and a separate 16-label speaking-style task. API v1 accepts prerecorded audio files and can return emotion alone or combine every output in unified analysis.

Primary sources

The benchmark discussion above draws on the original IEMOCAP database paper and the MELD dataset paper. For sarcastic delivery, context, and multimodal incongruity, see the MUStARD paper. oruk’s own sample counts, label mapping, and scorer are published in the benchmark methodology.