Skip to content

Contact center

Voice call sentiment analysis

Score the emotion and speaking style of a recorded call from its audio — per sentence, not one averaged label. One synchronous REST call per recording, priced by the second.

$50 trial credit · No card · From $0.0060 per audio minute

Create your account

$50 trial credit · No card · Sign up, then create a key.

Already have an account? Sign in

curl https://speech-api.oruk.ai/v1/audio/emotions \
  -H "Authorization: Bearer $ORUK_API_KEY" \
  -F file=@call.wav -F model=oruk-spectra-1
{
  "model": "oruk-spectra-1",
  "emotions": [
    { "label": "frustrated", "score": 0.83 },
    { "label": "worried",    "score": 0.61 }
  ],
  "usage": { "audio_seconds": 42.7, "estimated_cost_usd": 0.0043 }
}

A sentiment arc, not one score per call

Most call sentiment tooling averages a whole conversation into a single positive-or-negative verdict, which hides the moment things turned. oruk returns time-local segments, each with its own emotion and speaking-style scores, so a call reads as an arc: where frustration climbed, where it settled, and which sentence sat at the turn. Scoring runs on the audio itself, so tone that never reaches the transcript is still measured.

Where teams apply it

The same endpoint backs several call-analytics programs, each covered in depth on its own page: contact center analytics for queue-wide scoring and QA, voice of customer analytics for trending how customers sound over time, and sales call analysis for coaching against specific moments in a deal.

Measured in the open, stated with caveats

speech-emotion-bench

77.6% accuracy

Top measured result in our July 2026 release. Open models use the full 64,384 held-out clips; closed/API and audio-LLM systems use a fixed 5,000-clip subset. The highest-scoring non-oruk row is emotion2vec+ seed at 68.7%. The caveat, stated plainly: the oruk entry is trained in-distribution while other systems are evaluated zero-shot. Protocol and downloadable results are on the methodology page.

What a response means

A score of 0.83 for frustrated means the audio crossed a threshold calibrated on held-out speech — it describes how the clip sounds, not what the speaker meant or felt inside. Scores are probabilistic acoustic signals and should never be the sole basis for a consequential decision.

Try it yourself first: run the live analyzer on your own voice, no account needed.

Data handling: audio and outputs are used to produce the response and are not retained after it is returned; oruk never trains on customer content. Full commitments in the privacy policy.

Get an API key

$50 trial credit. No card required.

Exactly what you are buying

The limits below are stated here so you do not discover them after integrating.

Post-call, not live agent assist
The API analyzes completed recordings. Real-time call monitoring and live agent assist are out of scope for API v1.
Input
Prerecorded English audio files only. No production multilingual support in API v1.
No real-time streaming
The public contract is file-based. You upload a file and get the full result back; there is no websocket or streaming endpoint.
File limits and formats
Up to 30 MB and 60 minutes per file. WAV, FLAC, MP3, M4A, OGG, and WebM are supported.
Request behavior
Synchronous REST: one POST, one response. Every response reports its measured duration and estimated cost.
Pricing
Emotion on Spectra 1 costs $0.0060 per audio minute (pricing version 2026-07-11), metered per measured second with a one-second minimum. Full rates on the pricing page.
Data retention and training
Audio and outputs are not retained after the response is returned; billing and audit metadata is kept. oruk never trains on customer content.
What scores are — and are not
Emotion scores are probabilistic acoustic signals calibrated on held-out audio. They are not intent, truth, medical, employment, or psychological judgments, and must not be the sole basis for consequential decisions. See responsible use.
Model and version stability
Models are versioned; Spectra 1 and Resonance are stable, Spectra 2 is in preview. Changes are announced on the changelog.
Support
Email access@oruk.ai; the production guide covers retries, limits, and error handling.
Get an API key

$50 trial credit. No card required.

Built on an emotion recognition API

Underneath, this is one emotion recognition API call. The same models power our speech emotion recognition API; this page covers how that signal applies to recorded calls. Prefer to see it first? Try the captioning demo or read the API docs.

Get an API key

$50 trial credit. No card required.

Frequently asked questions

What is voice call sentiment analysis?
Voice call sentiment analysis measures the emotion and tone of a spoken call from its audio — not just the words. The oruk Speech API returns a transcript plus calibrated emotion and speaking-style scores for each sentence, so you can see how a caller sounded and how that shifted through the conversation.
Can you analyze sentiment and emotion per sentence in a call?
Yes. Long audio is returned as time-local segments, each carrying its own emotion and speaking-style scores, so a single call breaks into a sentence-by-sentence arc rather than one averaged label.
Can I use this for live agent assist?
Not with API v1. The public contract is file-based: you upload a completed recording and get the full result back, and there is no streaming or websocket endpoint. Teams building agent assist use oruk after the call — for scoring, coaching, and quality review — rather than during it.
How accurate is emotion scoring on call audio?
Scores are probabilistic acoustic signals calibrated on held-out audio, and published benchmark results are on the benchmarks page. Accuracy depends on your microphones, codecs, accents, and call types, so validate on your own recordings before relying on the output.
How is this different from an emotion recognition API?
It is the same emotion recognition API. This page covers how that signal applies to recorded calls; the emotion recognition API page documents the raw endpoint and the label space.

Start with $50 in trial credit

Create an account, verify your email, and make your first request in minutes. No card required.