Skip to content
orukBenchmarks Results

August 2026 evaluation report

Benchmark methodology

This page separates measurements run in one oruk harness from vendor-published references, records the sample and metric behind each figure, and states known limitations directly.

Measured panels

Speech emotion

Seven-class emotion, 64 systems

Sample
64,384 held-out clips; closed/API and audio-LLM systems on a fixed 5,000-clip stratified subsample
Metric
Accuracy across the shared seven-class mapping; macro F1 in the downloadable report
Protocol
Open models use the full 64,384 clips; closed/API and audio-LLM systems use the fixed 5,000-clip subset. Every row uses the same seven-class mapping and scorer within its assigned sample; refusals and API errors count as neutral predictions.
Exclusions and scope
Audio Flamingo 3 is excluded: it returned unparseable output on every clip in two environment configurations.
Confidence intervals
Not reported for this panel because per-example bootstrap artifacts are not available for every compared system.

IEMOCAP

Valence polarity, nine commercial APIs

Sample
IEMOCAP test split, 2,008 clips of acted dyadic emotional speech
Metric
Valence polarity accuracy (negative / neutral / positive); four-class accuracy and UAR for acoustic systems
Protocol
Identical WAV files sent to each vendor public API in one harness with retries, backoff, and per-clip checkpointing. Native outputs mapped to the shared three-class valence footing; 95% Wilson intervals on accuracy.
Exclusions and scope
Several vendors stopped early at free-tier credit or rate limits; rows report the clip count that produced a usable prediction. Transcript-sentiment products are included as commonly used emotion proxies, not as acoustic emotion claims.
Confidence intervals
95% Wilson intervals on accuracy; bootstrap intervals on UAR. Smaller n widens the interval.

Sarcasm

MUStARD, oruk vs Deepgram

Sample
Full official MUStARD set (ACL 2019): 690 utterances, 345 sarcastic / 345 not
Metric
Macro-F1 of a logistic probe over out-of-fold predictions
Protocol
Target-utterance audio only. Identical probe and folds for oruk affect scores (/v1/audio/affect) and Deepgram nova-3 transcript sentiment. Three setups: standard 5-fold, speaker-independent, and cross-show.
Exclusions and scope
Neither system ships a sarcasm classifier; this measures shipped product scores under identical treatment, not dedicated sarcasm-model ceilings. Scripted sitcom audio with laugh tracks; MUStARD known limitations apply.
Confidence intervals
95% bootstrap intervals on macro-F1 for each setup.

Transcription

English public four-set average

Sample
Four public English test sets
Metric
Corpus word error rate; lower is better
Protocol
LibriSpeech clean and other, Common Voice, and TED-LIUM with one standard English normalization path.
Exclusions and scope
Hatched vendor rows use official published values and are not represented as same-harness measurements.
Confidence intervals
Not reported for this panel because per-example bootstrap artifacts are not available for every compared system.

Pinned release

Evaluation date
2026-07-11
API version
v1
Transcription release
oruk-spectra-1
Unified release
oruk-resonance

The downloadable report pins an evaluation snapshot for every row and labels identities retained only in the internal checkpoint registry. It does not disclose private architecture or training implementation.

Open emotion benchmark

The separate speech-emotion-bench evaluation contains 64,384 clips, seven classes, approximately 20 languages, and 64 systems. Audio is normalized to 16 kHz mono and capped at 16 seconds. Closed APIs use a fixed 5,000-clip stratified subset; refusals and errors are counted rather than silently removed.

Fairness note: the oruk entry is trained in-distribution; the compared public and API systems are evaluated zero-shot across corpora. The result is useful for task performance, but it is not a pure zero-shot comparison.

Hume snapshot note: the 49.6% row is the legacy 48-dimension Hume Expression Measurement prosody endpoint, scored on the fixed 5,000-clip subset in this dated release. It is not a measurement of Hume’s current Tagger or real-time Prosody products. See Hume’s current product page.

Reporting rules

  • Solid bars denote results measured by oruk with the stated shared harness.
  • Hatched bars denote official vendor-published references from a different evaluation source.
  • Public API identifiers are reported; private architecture and training implementation are not.
  • Raw audio is not redistributed where source licenses prohibit redistribution.
  • Evaluation dates and pricing versions are fixed in downloadable artifacts.
  • Rankings retain each row's sample marker; mixed 64,384-clip and 5,000-clip rows are never represented as identical-audio evaluations.

Data licensing and access

The benchmark code is available under Apache-2.0. That license does not apply to the source audio. The evaluation set draws from corpora with different licenses, and oruk does not redistribute their audio. Researchers must obtain each source corpus under its own terms. The public manifest contains source identifiers and sample counts so the composition can be inspected without republishing restricted data.