August 2026 report · evaluated 2026-07-11
Speech emotion recognition benchmark
speech-emotion-bench scores 64 systems — commercial APIs, audio LLMs, open models, and text-only LLMs working from transcripts — with one seven-class label mapping and scoring implementation. Open models use the full 64,384-clip evaluation across ~20 languages; closed/API and audio-LLM systems use a fixed 5,000-clip stratified subset. oruk Resonance 1 reaches 77.8% accuracy and 0.816 macro F1; the best non-oruk open model (emotion2vec+ seed) reaches 68.7%, and the highest-scoring commercial API snapshot in this release (Hume AI legacy prosody) reaches 49.6%.
See BeSimple’s September 23 VocalAffectBench leaderboard, including Resonance 2 and the evaluation conditions for its published result.
For a more recent API result, replay the September 17 Resonance-2 diagnostic on 546 fixed CREMA-D recordings. The September 7 Fourier evaluation also includes pinned inputs, response receipts and scoring code.
For transcription serving speed, see the September 12 Orukeet API latency and throughput evaluation. It includes request-level timings, an offline results verifier, and instructions to reconstruct the twelve base audio inputs.
- Systems scored
- 64
- Held-out clips
- 64,384
- Languages
- ~20
- External snapshots
- IEMOCAP · MUStARD
One caveat stated plainly: the oruk entry is trained in-distribution, while other systems are evaluated zero-shot. Protocol, sample counts, exclusions, and limitations are on the methodology page. Closed/API and audio-LLM systems are scored on a fixed 5,000-clip stratified subsample (marked below); rescoring open models on the same subsample shifts results by less than two points.
These results describe historical evaluation checkpoints. The benchmark’s seven classes and approximately 20 dataset languages do not define the production API’s outputs or supported languages. For an integration, use the current model catalog and endpoint capabilities.
Product-snapshot note: the Hume row is the legacy 48-dimension Expression Measurement prosody endpoint, not Hume’s current Tagger or real-time Prosody products. The current products have not been evaluated in this release. See Hume’s official current product page.
September 23, 2026 · Published leaderboard
BeSimple VocalAffectBench
280 recordings · 7 emotions · English speech
49.3%
Highest listed accuracy · Resonance 2
This Resonance 2 revision was adapted using all 280 benchmark clips for training and selection. Its published score is not a held-out generalization result.
| Model | Accuracy | Correct / scored |
|---|---|---|
| Oruk Resonance 2Adapted on these recordings | 49.3% | 138 / 280 |
| Gemini 3.8 Flash | 45.7% | 128 / 280 |
| Gemini 3.5 Flash | 44.3% | 124 / 280 |
| Hume Prosody | 38.0% | 106 / 279 |
| Qwen3.5 Omni+ | 37.9% | 106 / 280 |
| Voxtral Small | 34.3% | 96 / 280 |
| Thinking Machines Inkling | 32.1% | 90 / 280 |
| Inworld Voice | 28.6% | 80 / 280 |
| OpenAI Realtime | 27.9% | 78 / 280 |
Scoring and evaluation details
Accuracy is the share of scored recordings whose top emotion matches the target label. The dataset has 40 acted recordings per emotion. Oruk selects the highest of seven continuous emotion scores, mapping “scared” to “fearful.” Its published run scores all 280 recordings; Hume’s published score uses 279 scored responses.
These are the model versions in BeSimple’s September 23 snapshot. The ranking describes reported scores under their respective evaluation conditions, including Oruk’s training overlap; it does not establish a held-out advantage over the other systems.
September 17, 2026 · Recorded API diagnostic
Replay every Resonance-2 prediction
We sent the same sentence, acted with six intended emotions by 91 CREMA-D actors, to the Resonance-2 API. The primary scoring rule matched 337 of 546 intended labels. Every recording stays in the denominator, including the 170 scoring abstentions. The archive contains every response, the input hashes and the model and calibration revisions.
| Scoring rule | Correct / total | Accuracy · 95% interval | Scoring abstentions |
|---|---|---|---|
| Primary: threshold-selected emotions | 337 / 546 | 61.7%58.1–65.9% | 170 |
| Sensitivity check: all emotion scores | 346 / 546 | 63.4%59.9–67.2% | 153 |
The primary result includes 39 wrong target labels and 170 cases with an empty selected-emotion list or a winning label outside the six classes. All 546 requests returned HTTP 200; a scoring abstention is different from a failed request. The sensitivity check ranks all 15 continuous emotion scores before mapping the winner. Both rules were registered before inference.
Training overlap has not been audited. This single-sentence acted diagnostic does not establish accuracy on unseen conversations or customer recordings. The intervals resample whole actors 2,000 times and describe uncertainty within this saved run. They do not account for training exposure.
Check the result without an API key
Run these commands in a macOS or Linux terminal with Python 3.12 or newer. After the archive downloads, the replay runs offline using only the Python standard library. It checks file hashes, all 546 predictions, aggregate scores and bootstrap intervals.
replay_dir=$(mktemp -d)
curl --fail --location --output "$replay_dir/replay.zip" \
https://oruk.ai/research/resonance-2/reproduction-2026-09-19.zip &&
python3 -m zipfile -e "$replay_dir/replay.zip" "$replay_dir" &&
cd "$replay_dir/reproduction-2026-09-19" &&
python3 replay.py --checkExpected output includes "correct": 337, "n": 546 and all archive hashes, predictions, metrics and intervals match. This replays saved outputs; it does not rerun inference, download the audio or verify the historical competitor rows.
Historical Resonance 1 evaluation · July–August 2026
Results on unseen audio
The public emotion result uses in-distribution training for Oruk and zero-shot evaluation for the alternatives. A separate set of 3,088 recordings was held back from every model in this comparison. On that unseen set, Resonance 1 scores 58.39% accuracy, compared with 77.76% on the public set. These results describe the dated Resonance 1 checkpoint. They do not establish the same performance for subsequent builds or Fourier.
Both sets use the same scoring pipeline. Every row below was evaluated on all 3,088 unseen clips; models without that run are omitted. These are descriptive point estimates. Per-clip predictions needed for paired comparisons and speaker-clustered uncertainty are not published.
| Model | Public set | Unseen audio |
|---|---|---|
| oruk Resonance 1 | 77.76% | 58.39% |
| SpeechBrain w2v2 | 39.20% | 45.95% |
| WavLM Vox-Profile | 45.28% | 37.21% |
| emotion2vec+ seed | 68.69% | 31.57% |
| emotion2vec+ base | 68.52% | 31.12% |
| wav2vec2 SUPERB | 31.43% | 25.49% |
| Whisper SER (firdhokk) | 32.50% | 23.41% |
| HuBERT Large SUPERB | 40.88% | 22.31% |
| XLS-R SER (firdhokk) | 25.98% | 20.79% |
| oruk (previous generation) | 77.56% | 20.14% |
| emotion2vec finetuned | 63.61% | 17.13% |
| WavLM Odyssey MSP | 44.59% | 16.97% |
| emotion2vec+ large | 68.59% | 16.94% |
| SenseVoice Small | 55.69% | 12.47% |
| XLS-R (hatman) | 40.24% | 12.47% |
| wav2vec2 emotion (dpngtm) | 38.18% | 10.49% |
| XLS-R SER (hughlan1214) | 44.36% | 9.29% |
speech-emotion-bench
Leaderboard: all 64 systems
Sorted by accuracy on the shared seven-class mapping (anger, happiness, sadness, fear, disgust, surprise, neutral). Systems marked 5k were scored on the stratified subset rather than the full 64,384 clips, so this is a descriptive release table, not a claim that every system saw identical audio.
| # | System | Type | Accuracy % | Macro F1 | Sample |
|---|---|---|---|---|---|
| 1 | oruk Resonance 1ours | oruk | 77.8 | 0.816 | full |
| 2 | emotion2vec+ seed | Open model | 68.7 | 0.680 | full |
| 3 | emotion2vec+ large | Open model | 68.6 | 0.677 | full |
| 4 | emotion2vec+ base | Open model | 68.5 | 0.683 | full |
| 5 | emotion2vec finetuned | Open model | 63.6 | 0.616 | full |
| 6 | EmotionThinker | Audio LLM | 60.5 | 0.504 | 5k |
| 7 | SenseVoice Small | Open model | 55.7 | 0.469 | full |
| 8 | emotion2vec base (probe) | Open model | 55.1 | 0.489 | full |
| 9 | Kimi-Audio 7B Instruct | Audio LLM | 52.4 | 0.447 | 5k |
| 10 | Hume AI legacy prosody | Commercial API | 49.6 | 0.414 | 5k |
| 11 | Qwen2-Audio 7B Instruct | Audio LLM | 48.9 | 0.402 | 5k |
| 12 | Step-Audio 2 mini | Audio LLM | 46.5 | 0.381 | 5k |
speech-emotion-bench
Per-emotion F1
oruk Resonance 1 against the best frontier multimodal API result (Gemini 3 Flash Preview) on each of the seven classes.
| Emotion | oruk Resonance 1 | Gemini 3 Flash | Delta |
|---|---|---|---|
| Disgust | 0.905 | 0.185 | +0.720 |
| Fear | 0.864 | 0.355 | +0.509 |
| Surprise | 0.858 | 0.268 | +0.590 |
| Anger | 0.806 | 0.462 | +0.344 |
| Neutral | 0.756 | 0.529 | +0.227 |
| Sadness | 0.746 | 0.313 | +0.433 |
| Happiness | 0.737 | 0.502 | +0.235 |
August 2026 snapshot · evaluated 2026-08-01
IEMOCAP test split: nine commercial systems
The 2,008-clip IEMOCAP test split (acted dyadic emotional speech), sent as identical audio to each vendor’s public API in one harness. The common footing across all nine systems is valence polarity: gold categorical labels mapped to negative (angry, frustrated, sad), neutral, and positive (happy, excited), and each system’s native output mapped to the same three classes. Chance is 33%. The grey ± is half the 95% Wilson interval; several vendors stopped early at free-tier credit or rate limits, so rows with a smaller n carry a wider ±. Transcript-sentiment rows are speech-to-text products whose sentiment is computed from the transcript text, not the acoustics — included because they are commonly used as emotion signals, not because they claim to hear prosody.
| # | System | Type | Valence accuracy % | n |
|---|---|---|---|---|
| 1 | oruk Resonance 1ours | oruk | 71.5±2.0 | 1,947 |
| 2 | Imentiv | Voice emotion | 63.6±8.6 | 118 |
| 3 | Smallest.ai Pulse | Voice emotion | 59.3±2.1 | 1,947 |
| 4 | Mistral Voxtral Small | Voice emotion | 54.4±2.2 | 1,947 |
| 5 | Rev AI sentiment | Transcript sentiment | 53.5±5.0 | 381 |
| 6 | Hume EVI prosody | Voice emotion | 51.3±7.8 | 154 |
| 7 | Deepgram nova-3 sentiment | Transcript sentiment | 46.3±2.2 | 1,947 |
| 8 | AssemblyAI sentiment | Transcript sentiment | 36.0±2.2 | 1,917 |
| 9 | Gladia sentiment | Transcript sentiment | 35.7±16.8 | 28 |
- Imentiv: Stopped at free-tier credit limit.
- Smallest.ai Pulse: Taxonomy has no neutral class.
- Mistral Voxtral Small: Prompted audio-LLM classification, voxtral-small-latest.
- Rev AI sentiment: Stopped at free-tier credit limit.
- Hume EVI prosody: EVI websocket top-1 prosody snapshot; stopped at free-tier credit limit. Not a measurement of other current Hume products.
- AssemblyAI sentiment: 33 clips returned no sentiment.
- Gladia sentiment: Free tier rate-limited almost immediately; interval too wide to rank.
Four-class emotion, acoustic systems only
Angry / happy+excited / sad / neutral for the five systems that emit acoustic emotion labels. Predictions outside the four classes count as errors; chance is 25%. UAR weights all four classes equally, and its intervals are bootstrap.
| System | Accuracy % | UAR % | n |
|---|---|---|---|
| oruk Resonance 1ours | 65.3±2.5 | 66.3±2.5 | 1,384 |
| Imentiv | 67.4±9.4 | 65.8±10.0 | 92 |
| Mistral Voxtral Small | 44.0±2.6 | 44.7±2.7 | 1,384 |
| Smallest.ai Pulse | 44.7±2.6 | 46.0±2.4 | 1,384 |
| Hume EVI prosody | 37.3±8.6 | 34.3±8.9 | 118 |
Imentiv’s accuracy point estimate sits above oruk’s, but on 92 clips its interval fully contains oruk’s; on UAR oruk ranks first. oruk is measured on 1,384 scoreable clips — the full test split minus clips whose gold label falls outside the four classes.
August 2026 snapshot · evaluated 2026-08-01
Sarcasm probes: MUStARD, Oruk vs Deepgram
The full official MUStARD set (ACL 2019): 690 sitcom utterances, 345 sarcastic and 345 not, target utterance audio only. Neither system ships a sarcasm classifier, so both are scored identically: a logistic probe over the scores each API returns — oruk emotion and speaking-style scores from /v1/audio/affect, Deepgram nova-3 transcript sentiment — with identical folds and macro-F1 on out-of-fold predictions.
Standard 5-fold
speaker-dependent
Speaker-independent
grouped 5-fold
Cross-show
BBT+GG → Friends
Oruk’s probe has the higher point estimate on each split, but the confidence intervals overlap; these results do not establish a reliable lead. On the cross-show split, the transcript-sentiment probe’s point estimate is below the trivial always-“not sarcastic” baseline. That observation alone does not establish why it failed. Deepgram’s input is transcript sentiment, while Oruk’s input is acoustic affect scores. The original paper’s full-feature SVM uses weighted F1 and a different setup, so its score is not directly comparable to this macro-F1 table. Scripted sitcom audio with laugh tracks and the small dataset limit generalization.
Historical Resonance 1 evaluation · July 2026
English transcription: four-set average
Average word error rate across LibriSpeech clean/other · Common Voice · TED-LIUM. Lower is better. The measured panel contains three checkpoints; Open ASR A and B are anonymized internal registry entries, with no public reproducible checkpoint identifiers in this release. These results do not evaluate subsequent API builds or Fourier.
Confidence intervals are unavailable: this release contains aggregate WER without the per-clip errors and sample counts needed to estimate uncertainty. The 0.03-point difference between the first two measured rows does not establish a significant lead. Vendor-published references below come from separate evaluations and are not part of the measured ranking.
Measured in the July 2026 panel
| System | WER | Recorded scope |
|---|---|---|
| oruk Resonance 1 | 4.44% | 2026-07 evaluation release |
| Open ASR A | 4.47% | Internal checkpoint registry, pinned 2026-07-11 |
| Open ASR B | 6.53% | Internal checkpoint registry, pinned 2026-07-11 |
Separately published references
| System | WER | Recorded scope |
|---|---|---|
| Azure Speech-to-Text | 5.5% | Published four-set average accessed 2026-07-12 |
| Whisper Large V3 | 5.7% | Published four-set average accessed 2026-07-12 |
| Whisper Medium | 6.1% | Published four-set average accessed 2026-07-12 |
| Whisper Small | 7.0% | Published four-set average accessed 2026-07-12 |
| Google Speech-to-Text | 8.9% | Published four-set average accessed 2026-07-12 |
| Picovoice Leopard | 9.7% | Vendor-published value accessed 2026-07-11 |
| Picovoice Cheetah | 10.1% | Vendor-published value accessed 2026-07-11 |
The JSON report preserves the model versions and source categories. See the evaluation protocol or run the current transcription API on representative audio to measure your own error rate.
Vendor comparisons
Each comparison above has its own stated protocol. A product decision usually needs more than a score — which capabilities exist at all, what streaming and language support look like, what each vendor charges. These pages cover that, and each states the date its product claims were checked.
How to cite
This page is the canonical version of the speech-emotion-bench results. Cite the evaluation date rather than the access date: the harness run is pinned to 2026-07-11, and rows added later carry their own dates on the methodology page.
@misc{oruk2026speechemotionbench,
title = {speech-emotion-bench: 64 speech emotion recognition
systems under one label mapping and scorer},
author = {{oruk}},
year = {2026},
howpublished = {\url{https://oruk.ai/benchmarks}},
note = {Evaluation run 2026-07-11. Open models scored on the full
64,384-clip set; closed/API and audio-LLM systems on a fixed
5,000-clip stratified subset}
}Evaluate Oruk on your audio
This report records the July–August 2026 Resonance 1 evaluation checkpoint. The current catalog lists available models and task support. Try the API on representative recordings before choosing a model; subscriptions include audio minutes shared across tasks.