Skip to content
BenchmarksMethodology

August 2026 report · evaluated 2026-07-11

Speech emotion recognition benchmark

speech-emotion-bench scores 64 systems — commercial APIs, audio LLMs, open models, and text-only LLMs working from transcripts — with one seven-class label mapping and scoring implementation. Open models use the full 64,384-clip evaluation across ~20 languages; closed/API and audio-LLM systems use a fixed 5,000-clip stratified subset. oruk Resonance 1 reaches 77.8% accuracy and 0.816 macro F1; the best non-oruk open model (emotion2vec+ seed) reaches 68.7%, and the highest-scoring commercial API snapshot in this release (Hume AI legacy prosody) reaches 49.6%.

See BeSimple’s September 23 VocalAffectBench leaderboard, including Resonance 2 and the evaluation conditions for its published result.

For a more recent API result, replay the September 17 Resonance-2 diagnostic on 546 fixed CREMA-D recordings. The September 7 Fourier evaluation also includes pinned inputs, response receipts and scoring code.

For transcription serving speed, see the September 12 Orukeet API latency and throughput evaluation. It includes request-level timings, an offline results verifier, and instructions to reconstruct the twelve base audio inputs.

Systems scored
64
Held-out clips
64,384
Languages
~20
External snapshots
IEMOCAP · MUStARD

One caveat stated plainly: the oruk entry is trained in-distribution, while other systems are evaluated zero-shot. Protocol, sample counts, exclusions, and limitations are on the methodology page. Closed/API and audio-LLM systems are scored on a fixed 5,000-clip stratified subsample (marked below); rescoring open models on the same subsample shifts results by less than two points.

These results describe historical evaluation checkpoints. The benchmark’s seven classes and approximately 20 dataset languages do not define the production API’s outputs or supported languages. For an integration, use the current model catalog and endpoint capabilities.

Product-snapshot note: the Hume row is the legacy 48-dimension Expression Measurement prosody endpoint, not Hume’s current Tagger or real-time Prosody products. The current products have not been evaluated in this release. See Hume’s official current product page.

September 23, 2026 · Published leaderboard

BeSimple VocalAffectBench

280 recordings · 7 emotions · English speech

49.3%

Highest listed accuracy · Resonance 2

This Resonance 2 revision was adapted using all 280 benchmark clips for training and selection. Its published score is not a held-out generalization result.

BeSimple’s published VocalAffectBench results, ordered by reported accuracy. The Oruk checkpoint was adapted on the benchmark recordings.
ModelAccuracyCorrect / scored
Oruk Resonance 2Adapted on these recordings49.3%138 / 280
Gemini 3.8 Flash45.7%128 / 280
Gemini 3.5 Flash44.3%124 / 280
Hume Prosody38.0%106 / 279
Qwen3.5 Omni+37.9%106 / 280
Voxtral Small34.3%96 / 280
Thinking Machines Inkling32.1%90 / 280
Inworld Voice28.6%80 / 280
OpenAI Realtime27.9%78 / 280
Scoring and evaluation details

Accuracy is the share of scored recordings whose top emotion matches the target label. The dataset has 40 acted recordings per emotion. Oruk selects the highest of seven continuous emotion scores, mapping “scared” to “fearful.” Its published run scores all 280 recordings; Hume’s published score uses 279 scored responses.

These are the model versions in BeSimple’s September 23 snapshot. The ranking describes reported scores under their respective evaluation conditions, including Oruk’s training overlap; it does not establish a held-out advantage over the other systems.

Dataset and protocolOruk run metadata

September 17, 2026 · Recorded API diagnostic

Replay every Resonance-2 prediction

We sent the same sentence, acted with six intended emotions by 91 CREMA-D actors, to the Resonance-2 API. The primary scoring rule matched 337 of 546 intended labels. Every recording stays in the denominator, including the 170 scoring abstentions. The archive contains every response, the input hashes and the model and calibration revisions.

Two preregistered scoring rules applied to the same 546 saved Resonance-2 responses.
Scoring ruleCorrect / totalAccuracy · 95% intervalScoring abstentions
Primary: threshold-selected emotions337 / 54661.7%58.1–65.9%170
Sensitivity check: all emotion scores346 / 54663.4%59.9–67.2%153

The primary result includes 39 wrong target labels and 170 cases with an empty selected-emotion list or a winning label outside the six classes. All 546 requests returned HTTP 200; a scoring abstention is different from a failed request. The sensitivity check ranks all 15 continuous emotion scores before mapping the winner. Both rules were registered before inference.

Training overlap has not been audited. This single-sentence acted diagnostic does not establish accuracy on unseen conversations or customer recordings. The intervals resample whole actors 2,000 times and describe uncertainty within this saved run. They do not account for training exposure.

Check the result without an API key

Run these commands in a macOS or Linux terminal with Python 3.12 or newer. After the archive downloads, the replay runs offline using only the Python standard library. It checks file hashes, all 546 predictions, aggregate scores and bootstrap intervals.

replay_dir=$(mktemp -d)
curl --fail --location --output "$replay_dir/replay.zip" \
  https://oruk.ai/research/resonance-2/reproduction-2026-09-19.zip &&
python3 -m zipfile -e "$replay_dir/replay.zip" "$replay_dir" &&
cd "$replay_dir/reproduction-2026-09-19" &&
python3 replay.py --check

Expected output includes "correct": 337, "n": 546 and all archive hashes, predictions, metrics and intervals match. This replays saved outputs; it does not rerun inference, download the audio or verify the historical competitor rows.

Historical Resonance 1 evaluation · July–August 2026

Results on unseen audio

The public emotion result uses in-distribution training for Oruk and zero-shot evaluation for the alternatives. A separate set of 3,088 recordings was held back from every model in this comparison. On that unseen set, Resonance 1 scores 58.39% accuracy, compared with 77.76% on the public set. These results describe the dated Resonance 1 checkpoint. They do not establish the same performance for subsequent builds or Fourier.

Both sets use the same scoring pipeline. Every row below was evaluated on all 3,088 unseen clips; models without that run are omitted. These are descriptive point estimates. Per-clip predictions needed for paired comparisons and speaker-clustered uncertainty are not published.

Historical emotion accuracy on the public set and the held-back unseen-audio set. Rows are ordered by unseen-audio accuracy.
ModelPublic setUnseen audio
oruk Resonance 177.76%58.39%
SpeechBrain w2v239.20%45.95%
WavLM Vox-Profile45.28%37.21%
emotion2vec+ seed68.69%31.57%
emotion2vec+ base68.52%31.12%
wav2vec2 SUPERB31.43%25.49%
Whisper SER (firdhokk)32.50%23.41%
HuBERT Large SUPERB40.88%22.31%
XLS-R SER (firdhokk)25.98%20.79%
oruk (previous generation)77.56%20.14%
emotion2vec finetuned63.61%17.13%
WavLM Odyssey MSP44.59%16.97%
emotion2vec+ large68.59%16.94%
SenseVoice Small55.69%12.47%
XLS-R (hatman)40.24%12.47%
wav2vec2 emotion (dpngtm)38.18%10.49%
XLS-R SER (hughlan1214)44.36%9.29%

speech-emotion-bench

Leaderboard: all 64 systems

Sorted by accuracy on the shared seven-class mapping (anger, happiness, sadness, fear, disgust, surprise, neutral). Systems marked 5k were scored on the stratified subset rather than the full 64,384 clips, so this is a descriptive release table, not a claim that every system saw identical audio.

#SystemTypeAccuracy %Macro F1Sample
1oruk Resonance 1oursoruk77.80.816full
2emotion2vec+ seedOpen model68.70.680full
3emotion2vec+ largeOpen model68.60.677full
4emotion2vec+ baseOpen model68.50.683full
5emotion2vec finetunedOpen model63.60.616full
6EmotionThinkerAudio LLM60.50.5045k
7SenseVoice SmallOpen model55.70.469full
8emotion2vec base (probe)Open model55.10.489full
9Kimi-Audio 7B InstructAudio LLM52.40.4475k
10Hume AI legacy prosodyCommercial API49.60.4145k
11Qwen2-Audio 7B InstructAudio LLM48.90.4025k
12Step-Audio 2 miniAudio LLM46.50.3815k

speech-emotion-bench

Per-emotion F1

oruk Resonance 1 against the best frontier multimodal API result (Gemini 3 Flash Preview) on each of the seven classes.

Emotionoruk Resonance 1Gemini 3 FlashDelta
Disgust0.9050.185+0.720
Fear0.8640.355+0.509
Surprise0.8580.268+0.590
Anger0.8060.462+0.344
Neutral0.7560.529+0.227
Sadness0.7460.313+0.433
Happiness0.7370.502+0.235

August 2026 snapshot · evaluated 2026-08-01

IEMOCAP test split: nine commercial systems

The 2,008-clip IEMOCAP test split (acted dyadic emotional speech), sent as identical audio to each vendor’s public API in one harness. The common footing across all nine systems is valence polarity: gold categorical labels mapped to negative (angry, frustrated, sad), neutral, and positive (happy, excited), and each system’s native output mapped to the same three classes. Chance is 33%. The grey ± is half the 95% Wilson interval; several vendors stopped early at free-tier credit or rate limits, so rows with a smaller n carry a wider ±. Transcript-sentiment rows are speech-to-text products whose sentiment is computed from the transcript text, not the acoustics — included because they are commonly used as emotion signals, not because they claim to hear prosody.

#SystemTypeValence accuracy %n
1oruk Resonance 1oursoruk71.5±2.01,947
2ImentivVoice emotion63.6±8.6118
3Smallest.ai PulseVoice emotion59.3±2.11,947
4Mistral Voxtral SmallVoice emotion54.4±2.21,947
5Rev AI sentimentTranscript sentiment53.5±5.0381
6Hume EVI prosodyVoice emotion51.3±7.8154
7Deepgram nova-3 sentimentTranscript sentiment46.3±2.21,947
8AssemblyAI sentimentTranscript sentiment36.0±2.21,917
9Gladia sentimentTranscript sentiment35.7±16.828
  • Imentiv: Stopped at free-tier credit limit.
  • Smallest.ai Pulse: Taxonomy has no neutral class.
  • Mistral Voxtral Small: Prompted audio-LLM classification, voxtral-small-latest.
  • Rev AI sentiment: Stopped at free-tier credit limit.
  • Hume EVI prosody: EVI websocket top-1 prosody snapshot; stopped at free-tier credit limit. Not a measurement of other current Hume products.
  • AssemblyAI sentiment: 33 clips returned no sentiment.
  • Gladia sentiment: Free tier rate-limited almost immediately; interval too wide to rank.

Four-class emotion, acoustic systems only

Angry / happy+excited / sad / neutral for the five systems that emit acoustic emotion labels. Predictions outside the four classes count as errors; chance is 25%. UAR weights all four classes equally, and its intervals are bootstrap.

SystemAccuracy %UAR %n
oruk Resonance 1ours65.3±2.566.3±2.51,384
Imentiv67.4±9.465.8±10.092
Mistral Voxtral Small44.0±2.644.7±2.71,384
Smallest.ai Pulse44.7±2.646.0±2.41,384
Hume EVI prosody37.3±8.634.3±8.9118

Imentiv’s accuracy point estimate sits above oruk’s, but on 92 clips its interval fully contains oruk’s; on UAR oruk ranks first. oruk is measured on 1,384 scoreable clips — the full test split minus clips whose gold label falls outside the four classes.

August 2026 snapshot · evaluated 2026-08-01

Sarcasm probes: MUStARD, Oruk vs Deepgram

The full official MUStARD set (ACL 2019): 690 sitcom utterances, 345 sarcastic and 345 not, target utterance audio only. Neither system ships a sarcasm classifier, so both are scored identically: a logistic probe over the scores each API returns — oruk emotion and speaking-style scores from /v1/audio/affect, Deepgram nova-3 transcript sentiment — with identical folds and macro-F1 on out-of-fold predictions.

orukaudio affectDeepgramtranscript sentiment
always “not sarcastic” · 0.333macro-F1 · 95% CI
0.562±0.038
0.517±0.038
0.510±0.039
0.395±0.037
0.462±0.053
0.299±0.024

Standard 5-fold

speaker-dependent

Speaker-independent

grouped 5-fold

Cross-show

BBT+GG → Friends

Oruk’s probe has the higher point estimate on each split, but the confidence intervals overlap; these results do not establish a reliable lead. On the cross-show split, the transcript-sentiment probe’s point estimate is below the trivial always-“not sarcastic” baseline. That observation alone does not establish why it failed. Deepgram’s input is transcript sentiment, while Oruk’s input is acoustic affect scores. The original paper’s full-feature SVM uses weighted F1 and a different setup, so its score is not directly comparable to this macro-F1 table. Scripted sitcom audio with laugh tracks and the small dataset limit generalization.

Historical Resonance 1 evaluation · July 2026

English transcription: four-set average

Average word error rate across LibriSpeech clean/other · Common Voice · TED-LIUM. Lower is better. The measured panel contains three checkpoints; Open ASR A and B are anonymized internal registry entries, with no public reproducible checkpoint identifiers in this release. These results do not evaluate subsequent API builds or Fourier.

Confidence intervals are unavailable: this release contains aggregate WER without the per-clip errors and sample counts needed to estimate uncertainty. The 0.03-point difference between the first two measured rows does not establish a significant lead. Vendor-published references below come from separate evaluations and are not part of the measured ranking.

Measured in the July 2026 panel

Measured in the July 2026 panel: English word error rate and recorded source scope.
SystemWERRecorded scope
oruk Resonance 14.44%2026-07 evaluation release
Open ASR A4.47%Internal checkpoint registry, pinned 2026-07-11
Open ASR B6.53%Internal checkpoint registry, pinned 2026-07-11

Separately published references

Separately published references: English word error rate and recorded source scope.
SystemWERRecorded scope
Azure Speech-to-Text5.5%Published four-set average accessed 2026-07-12
Whisper Large V35.7%Published four-set average accessed 2026-07-12
Whisper Medium6.1%Published four-set average accessed 2026-07-12
Whisper Small7.0%Published four-set average accessed 2026-07-12
Google Speech-to-Text8.9%Published four-set average accessed 2026-07-12
Picovoice Leopard9.7%Vendor-published value accessed 2026-07-11
Picovoice Cheetah10.1%Vendor-published value accessed 2026-07-11

The JSON report preserves the model versions and source categories. See the evaluation protocol or run the current transcription API on representative audio to measure your own error rate.

Vendor comparisons

Each comparison above has its own stated protocol. A product decision usually needs more than a score — which capabilities exist at all, what streaming and language support look like, what each vendor charges. These pages cover that, and each states the date its product claims were checked.

All comparisons

How to cite

This page is the canonical version of the speech-emotion-bench results. Cite the evaluation date rather than the access date: the harness run is pinned to 2026-07-11, and rows added later carry their own dates on the methodology page.

@misc{oruk2026speechemotionbench,
  title        = {speech-emotion-bench: 64 speech emotion recognition
                  systems under one label mapping and scorer},
  author       = {{oruk}},
  year         = {2026},
  howpublished = {\url{https://oruk.ai/benchmarks}},
  note         = {Evaluation run 2026-07-11. Open models scored on the full
                  64,384-clip set; closed/API and audio-LLM systems on a fixed
                  5,000-clip stratified subset}
}

Evaluate Oruk on your audio

This report records the July–August 2026 Resonance 1 evaluation checkpoint. The current catalog lists available models and task support. Try the API on representative recordings before choosing a model; subscriptions include audio minutes shared across tasks.