Skip to content
orukBenchmarks Results

July 2026 evaluation report

Benchmark methodology

This page separates measurements run in one oruk harness from vendor-published references, records the sample and metric behind each figure, and states known limitations directly.

Measured panels

Speech emotion

Seven-class emotion, 64 systems

Sample
64,384 held-out clips; closed/API and audio-LLM systems on a fixed 5,000-clip stratified subsample
Metric
Accuracy across the shared seven-class mapping; macro F1 in the downloadable report
Protocol
Open models use the full 64,384 clips; closed/API and audio-LLM systems use the fixed 5,000-clip subset. Every row uses the same seven-class mapping and scorer within its assigned sample; refusals and API errors count as neutral predictions.
Exclusions and scope
Audio Flamingo 3 is excluded: it returned unparseable output on every clip in two environment configurations.
Confidence intervals
Not reported in this release because per-example bootstrap artifacts are not available for every compared system.

Transcription

English public four-set average

Sample
Four public English test sets
Metric
Corpus word error rate; lower is better
Protocol
LibriSpeech clean and other, Common Voice, and TED-LIUM with one standard English normalization path.
Exclusions and scope
Hatched vendor rows use official published values and are not represented as same-harness measurements.
Confidence intervals
Not reported in this release because per-example bootstrap artifacts are not available for every compared system.

Pinned release

Evaluation date
2026-07-11
API version
v1
Transcription release
oruk-spectra-1
Unified release
oruk-resonance

The downloadable report pins an evaluation snapshot for every row and labels identities retained only in the internal checkpoint registry. It does not disclose private architecture or training implementation.

Open emotion benchmark

The separate speech-emotion-bench evaluation contains 64,384 clips, seven classes, approximately 20 languages, and 64 systems. Audio is normalized to 16 kHz mono and capped at 16 seconds. Closed APIs use a fixed 5,000-clip stratified subset; refusals and errors are counted rather than silently removed.

Fairness note: the oruk entry is trained in-distribution; the compared public and API systems are evaluated zero-shot across corpora. The result is useful for task performance, but it is not a pure zero-shot comparison.

Hume snapshot note: the 49.6% row is the legacy 48-dimension Hume Expression Measurement prosody endpoint, scored on the fixed 5,000-clip subset in this dated release. It is not a measurement of Hume’s current Tagger or real-time Prosody products. See Hume’s current product page.

Reporting rules

  • Solid bars denote results measured by oruk with the stated shared harness.
  • Hatched bars denote official vendor-published references from a different evaluation source.
  • Public API identifiers are reported; private architecture and training implementation are not.
  • Raw audio is not redistributed where source licenses prohibit redistribution.
  • Evaluation dates and pricing versions are fixed in downloadable artifacts.
  • Rankings retain each row's sample marker; mixed 64,384-clip and 5,000-clip rows are never represented as identical-audio evaluations.

Data licensing and access

The benchmark code is available under Apache-2.0. That license does not apply to the source audio. The evaluation set draws from corpora with different licenses, and oruk does not redistribute their audio. Researchers must obtain each source corpus under its own terms. The public manifest contains source identifiers and sample counts so the composition can be inspected without republishing restricted data.