This page separates measurements run in one oruk harness from vendor-published references, records the sample and metric behind each figure, and states known limitations directly.
One acted sentence and six intended labels; empty or out-of-taxonomy answers count as errors. Training overlap unaudited. Separate development holdouts were reused; neither establishes independent unseen accuracy.
How does a different inference runtime change speed?
Historical development audio; preloaded-waveform timings exclude HTTP, model loading and prewarming. Aggregate evidence only, not a full replay package or a production API speed claim.
What did the Fourier API return?
September 7 replay: 546 input hashes, response receipts, errors and actor-bootstrap intervals.
Six acted labels; training overlap unaudited. Operator artifact hash, without an API-reported immutable revision.
Developer-run, reused integration sample; unseen status unverified. The r3 protocol discloses checkpoint selection and adaptation overlap. Fresh inference covers the ONNX CPU pair only; other runtimes retain historical evidence and different timing boundaries.
Research checkpoints, with Oruk trained in-distribution and other systems evaluated zero-shot. Not a current product ranking.
September 7, 2026 · Live API evaluation
Inspect and replay a real API evaluation
We sent 546 distinct CREMA-D recordings through oruk-fourier at /v1/audio/emotions. The input set was fixed before the run: all 91 actors, one IEO sentence, five acted emotions at medium intensity and neutral at its unspecified intensity. Each label has 91 recordings. The original WAV bytes are hash-checked and unmodified.
The reference is the emotion the actor was asked to portray, as recorded in the filename. Training overlap has not been audited. This measures one acted-speech protocol; it does not establish performance on spontaneous conversations, customer audio or unseen training sources. It is separate from the historical 64,384-clip, seven-class evaluation below.
Accuracy
55.7%
95% interval: 51.8%–59.2%
Macro F1
0.542
95% interval: 0.502–0.579
Valid API responses
546 / 546
0 errors; no retries
Prediction coverage
99.6%
Out-of-taxonomy winners abstain
Six-class current API evaluation, with all registered recordings retained.
Intended emotion
Clips
Precision
Recall
F1
anger
91
56.8%
92.3%
0.703
disgust
91
48.8%
64.8%
0.557
fear
91
41.5%
56.0%
0.477
happiness
91
53.7%
24.2%
0.333
neutral
91
82.1%
60.4%
0.696
sadness
91
75.0%
36.3%
0.489
Where this run made mistakes
Of 546 recordings, 304 matched the intended label, 240 received another target label, and 2 abstained. There were 0 API errors. The largest confusions were sadness labeled as fear (37 recordings) and happiness labeled as anger (34 recordings). These counts describe the saved September 7 run; they do not explain why an error occurred.
Full confusion matrix, including abstentions and errors
Rows are intended acted labels; columns are the scored API predictions. Every row totals 91 recordings.
Reference
anger
disgust
fear
happiness
neutral
sadness
Abstained
Error
anger
84
4
1
1
0
1
0
0
disgust
16
59
11
0
2
1
2
0
fear
10
19
51
2
4
5
0
0
happiness
34
17
14
22
4
0
0
0
neutral
3
4
9
16
55
4
0
0
sadness
1
18
37
0
2
33
0
0
This error analysis was added on September 9 after the results were known. It changes no inputs, predictions, thresholds or aggregate scores. The downloadable tool rechecks the original audio and saved response scores before building the matrix. Download all confusion counts and abstention receipts.
The scorer takes the highest returned label score, applies fixed spelling mappings, and abstains if the winning label is outside the six-class taxonomy. It does not renormalize scores or treat them as calibrated probabilities. Errors and abstentions stay in the denominator. Confidence intervals resample whole actors 2,000 times with a fixed seed.
Successful-response latency was 1.19s at the median and 1.60s at the 95th percentile, measured sequentially from one client and including upload and network time. The client also performed other local work during part of the run, so these are observed request times, not an isolated performance benchmark or an SLA. The run processed 20.78 minutes of audio through an internal evaluation account; it is excluded from customer conversion reporting.
Recompute the result without an API key
Download the source bundle and saved predictions, fetch the pinned source recordings, then run the scorer. Python 3.10 or later and its standard library are sufficient. This replays the recorded results; the bundle also includes the separate, explicitly invoked command for a new API run.
The public API returns a model alias, without an immutable revision field. Our operator checked the deployed emotion artifact before and after this run; its SHA-256 is 564d1e42e4cf0b936d911dbc71e30fa858f674ff60f587e6029b2e14a1643bbc. This is operator evidence, not independent verification. A later API deployment can change predictions; the saved receipts and scorer preserve this dated result.
What an independent result still needs
A third-party dataset can still overlap a model’s training. Our next evaluation needs permitted human recordings, model-blind listener labels, and exclusion checks against training, pretraining, validation, calibration and model selection, bound to the actual served artifact. Those inputs and provenance are not yet verified, so no independent result is claimed here.
The prospective protocol fixes the planned six-class selection, speaker limits, annotations, failure accounting and uncertainty rules. Its readiness template lists the evidence required before registering a run. Neither document is a completed evaluation.
Source audio: Cao and colleagues’ CREMA-D (2014), under ODbL 1.0 / DbCL 1.0. Audio is fetched from its pinned upstream source; it is not bundled here. See the bundle’s attribution and license notices. To test your own workload, start with the emotion API documentation.
Orukeet API latency and throughput
The September 12, 2026 serving evaluation records 1,905 requests, including warmups, with zero recorded errors. It covers REST and PCM streaming, sustained REST load, and separate optional-task runs from one client against one warm A100 in the US. The base runs reuse twelve Earnings22 recordings; request count is not a count of distinct recordings.
The reproduction guide recalculates the published aggregates without an API key and reconstructs all twelve base WAVs with exact hash checks in the documented Linux environment. The separate synthetic optional-task recording is not included. These measurements describe serving performance, not accuracy, training-data independence, or latency from every region.
August 2026 historical evaluation panels
Speech emotion
Seven-class emotion, 64 systems
Sample
64,384 held-out clips; closed/API and audio-LLM systems on a fixed 5,000-clip stratified subsample
Metric
Accuracy across the shared seven-class mapping; macro F1 in the downloadable report
Protocol
Open models use the full 64,384 clips; closed/API and audio-LLM systems use the fixed 5,000-clip subset. Every row uses the same seven-class mapping and scorer within its assigned sample; refusals and API errors count as neutral predictions.
Exclusions and scope
Audio Flamingo 3 is excluded: it returned unparseable output on every clip in two environment configurations.
Confidence intervals
Approximate 95% Wilson intervals calculated from published, rounded accuracy and clip counts. They assume independent clips, do not adjust for repeated speakers or training overlap, and are not a paired significance test. Macro-F1 intervals are unavailable.
IEMOCAP
Valence polarity, nine commercial APIs
Sample
IEMOCAP test split, 2,008 clips of acted dyadic emotional speech
Metric
Valence polarity accuracy (negative / neutral / positive); four-class accuracy and UAR for acoustic systems
Protocol
Identical WAV files sent to each vendor public API in one harness with retries, backoff, and per-clip checkpointing. Native outputs mapped to the shared three-class valence footing; 95% Wilson intervals on accuracy.
Exclusions and scope
Several vendors stopped early at free-tier credit or rate limits; rows report the clip count that produced a usable prediction. Transcript-sentiment products are included as commonly used emotion proxies, not as acoustic emotion claims.
Confidence intervals
95% Wilson intervals on accuracy; bootstrap intervals on UAR. Coverage differs by system, and smaller n widens the interval.
Sarcasm
MUStARD, oruk vs Deepgram
Sample
Full official MUStARD set (ACL 2019): 690 utterances, 345 sarcastic / 345 not
Metric
Macro-F1 of a logistic probe over out-of-fold predictions
Protocol
Target-utterance audio only. Identical probe and folds for oruk affect scores (/v1/audio/affect) and Deepgram nova-3 transcript sentiment. Three setups: standard 5-fold, speaker-independent, and cross-show.
Exclusions and scope
Neither system ships a sarcasm classifier; this measures shipped product scores under identical treatment, not dedicated sarcasm-model ceilings. Scripted sitcom audio with laugh tracks; MUStARD known limitations apply.
Confidence intervals
95% bootstrap intervals on out-of-fold macro-F1 for each setup. Overlapping intervals do not establish a reliable lead.
Transcription
English public four-set average
Sample
Four public English test sets
Metric
Corpus word error rate; lower is better
Protocol
LibriSpeech clean and other, Common Voice, and TED-LIUM with one standard English normalization path.
Exclusions and scope
Hatched vendor rows use official published values and are not represented as same-harness measurements.
Confidence intervals
Confidence intervals are unavailable: this release contains aggregate WER without the per-clip errors and sample counts needed to estimate uncertainty.
Pinned release
Harness pinned
2026-07-11
API version
v1
Evaluation checkpoint
oruk-resonance-1
External evaluation
2026-08-01
Resonance 1, July–August 2026 evaluation checkpoint. Results apply to the recorded build and protocol; they are not a new evaluation of subsequent API releases or Fourier. The emotion run was recorded on July 14 after the July 11 harness pin. Row-level notes identify later corrections and external evaluation dates.
Open emotion benchmark
oruk-bench is Oruk’s evaluation toolkit, separate from the production API and its client SDK. Its seven-class label mapping and roughly 20-language dataset describe the benchmark protocol, not the API’s supported languages or native outputs. Whisper, WavLM, HuBERT, and Conformer research comparisons are not the served Oruk model catalog. See current API models and language support.
The separate speech-emotion-bench evaluation contains 64,384 clips, seven classes, approximately 20 languages, and 64 systems. Audio is normalized to 16 kHz mono and capped at 16 seconds. Closed APIs use a fixed 5,000-clip stratified subset; refusals and errors are counted rather than silently removed.
Fairness note: the oruk entry is trained in-distribution; the compared public and API systems are evaluated zero-shot across corpora. The result is useful for task performance, but it is not a pure zero-shot comparison.
Hume snapshot note: the 49.6% row is the legacy 48-dimension Hume Expression Measurement prosody endpoint, scored on the fixed 5,000-clip subset in this dated release. It is not a measurement of Hume’s current Tagger or real-time Prosody products. See Hume’s current product page.
Verify the published report
Save this as audit_benchmarks.py and run python3 audit_benchmarks.py. It uses Python’s standard library, requires no API key, and downloads only the two public JSON artifacts. It checks the source counts and label mapping, prints the dated Oruk result, and separates measured rows from vendor references in each panel.
import json
from urllib.request import Request, urlopen
ORIGIN = "https://oruk.ai"
def read_json(path):
request = Request(ORIGIN + path, headers={
"User-Agent": "oruk-benchmark-audit/1.0 (+https://oruk.ai/benchmarks/methodology)",
"Accept": "application/json",
})
with urlopen(request, timeout=30) as response:
return json.load(response)
report = read_json("/api/benchmarks/report")
manifest = read_json("/benchmarks/emotion-eval-manifest.json")
emotion = report["open_emotion_benchmark"]
sources = manifest["sources"]
if sum(row["samples"] for row in sources) != manifest["sample_count"]:
raise ValueError("Source counts do not add up to the manifest total")
if len(sources) != manifest["source_count"]:
raise ValueError("Source count does not match the manifest")
if manifest["sample_count"] != emotion["clips"]:
raise ValueError("Report and manifest sample counts disagree")
if manifest["labels"] != emotion["emotionClasses"]:
raise ValueError("Report and manifest label mappings disagree")
print(f"Harness pin: {report['evaluation_date']}")
print(f"Report revision: {report['report_revision']}")
print(f"{manifest['sample_count']:,} clips across {len(sources)} sources")
for row in emotion["results"]:
if row.get("ours"):
n = emotion["subsampleClips"] if row.get("subsample") else emotion["clips"]
print(f"{row['model_id']}: {row['acc']}% accuracy, {row['mf1']} macro F1; n={n:,}")
for panel in report["panels"]:
measured = sum(row["measured_in_oruk_harness"] for row in panel["results"])
reference = sum(row["vendor_published_reference"] for row in panel["results"])
print(f"{panel['id']}: {measured} measured rows, {reference} vendor references")
The manifest should total 64,384 clips from 50 source entries. The printed checkpoint is oruk-resonance-1; its evaluation date and the report revision are different fields. A newer report revision can correct presentation without recording a new model evaluation.
What this check can establish
It checks the consistency of published aggregates. It does not recompute accuracy or F1 from predictions, verify an audio split, or run the current API. The downloaded report contains aggregate results, and the manifest contains source counts rather than clip-level audio and predictions. Reproducing an inference run requires the licensed input audio, exact held-out split, evaluated model snapshot, label mapping, preprocessing, and scoring configuration. The public oruk-bench repository provides evaluation and scoring code; it does not redistribute the evaluation audio. Its synthetic fixture exercises the pipeline, not the reported accuracy.
The largest manifest entry, RijoSLal/SpeechEmotionClassification, contributes 26,326 clips (40.9% of the full set). Accuracy counts each clip equally, so larger sources contribute more to the total. Macro F1 gives each emotion class equal weight; it does not give each source or language equal weight. The manifest counts alone cannot establish performance for a particular source, accent, or recording environment.
Reporting rules
Solid bars denote results measured by oruk with the stated shared harness.
Hatched bars denote official vendor-published references from a different evaluation source.
Public API identifiers are reported; private architecture and training implementation are not.
Raw audio is not redistributed where source licenses prohibit redistribution.
Evaluation dates and pricing versions are fixed in downloadable artifacts.
Rankings retain each row's sample marker; mixed 64,384-clip and 5,000-clip rows are never represented as identical-audio evaluations.
Data licensing and access
The benchmark code is available under Apache-2.0. That license does not apply to the source audio. The evaluation set draws from corpora with different licenses, and oruk does not redistribute their audio. Researchers must obtain each source corpus under its own terms. The public manifest contains source identifiers and sample counts so the composition can be inspected without republishing restricted data.