Skip to content

Emotion API

Emotion Recognition API

Score 15 calibrated emotions from prerecorded English speech through one synchronous REST call.

$50 trial credit · No card · From $0.0060 per audio minute

Create your account

$50 trial credit · No card · Sign up, then create a key.

Already have an account? Sign in

curl https://speech-api.oruk.ai/v1/audio/emotions \
  -H "Authorization: Bearer $ORUK_API_KEY" \
  -F file=@call.wav -F model=oruk-spectra-1
{
  "model": "oruk-spectra-1",
  "emotions": [
    { "label": "frustrated", "score": 0.83 },
    { "label": "worried",    "score": 0.61 }
  ],
  "usage": { "audio_seconds": 42.7, "estimated_cost_usd": 0.0043 }
}

Emotion recognition from audio, calibrated for decisions

This API recognizes emotion in recorded speech — contact-center calls, meetings, interviews, voicemails. It does not analyze faces, images, video, or text sentiment. Each request returns multilabel scores for 15 emotions with thresholds calibrated on held-out audio, so you can act on a label without inventing your own cutoffs. Full endpoint reference in the API docs.

Measured in the open, stated with caveats

speech-emotion-bench

77.6% accuracy

Top measured result in our July 2026 release. Open models use the full 64,384 held-out clips; closed/API and audio-LLM systems use a fixed 5,000-clip subset. The highest-scoring non-oruk row is emotion2vec+ seed at 68.7%. The caveat, stated plainly: the oruk entry is trained in-distribution while other systems are evaluated zero-shot. Protocol and downloadable results are on the methodology page.

What a response means

A score of 0.83 for frustrated means the audio crossed a threshold calibrated on held-out speech — it describes how the clip sounds, not what the speaker meant or felt inside. Scores are probabilistic acoustic signals and should never be the sole basis for a consequential decision.

Try it yourself first: run the live analyzer on your own voice, no account needed.

Data handling: audio and outputs are used to produce the response and are not retained after it is returned; oruk never trains on customer content. Full commitments in the privacy policy.

Get an API key

$50 trial credit. No card required.

Exactly what you are buying

The limits below are stated here so you do not discover them after integrating.

Speech only
Emotion is recognized from audio recordings. Facial expression, image, video, and text-sentiment analysis are out of scope.
Input
Prerecorded English audio files only. No production multilingual support in API v1.
No real-time streaming
The public contract is file-based. You upload a file and get the full result back; there is no websocket or streaming endpoint.
File limits and formats
Up to 30 MB and 60 minutes per file. WAV, FLAC, MP3, M4A, OGG, and WebM are supported.
Request behavior
Synchronous REST: one POST, one response. Every response reports its measured duration and estimated cost.
Pricing
Emotion on Spectra 1 costs $0.0060 per audio minute (pricing version 2026-07-11), metered per measured second with a one-second minimum. Full rates on the pricing page.
Data retention and training
Audio and outputs are not retained after the response is returned; billing and audit metadata is kept. oruk never trains on customer content.
What scores are — and are not
Emotion scores are probabilistic acoustic signals calibrated on held-out audio. They are not intent, truth, medical, employment, or psychological judgments, and must not be the sole basis for consequential decisions. See responsible use.
Model and version stability
Models are versioned; Spectra 1 and Resonance are stable, Spectra 2 is in preview. Changes are announced on the changelog.
Support
Email access@oruk.ai; the production guide covers retries, limits, and error handling.
Get an API key

$50 trial credit. No card required.

Start with $50 in trial credit

Create an account, verify your email, and make your first request in minutes. No card required.