Skip to content

Ranked by measurement

Speech emotion APIs ranked in our July 2026 benchmark

Correction — July 20, 2026: Hume’s legacy benchmark row is not its current product. Hume’s current Expression Measurement page advertises offline Tagger and real-time Prosody; neither has been evaluated here.

Most “best emotion API” lists are written from vendor marketing pages. This one is ranked by measured accuracy — every system uses the same seven-class mapping and scorer. Open models use the full 64,384-clip evaluation; closed/API and audio-LLM rows use a fixed 5,000-clip stratified subset. Full disclosure: oruk builds the top measured entry, trained it in-distribution, and publishes the complete 64-system results and methodology so you can check the ranking yourself.

  1. 01oruk Speech API

    77.6% measured

    Top measured result in our July 2026 benchmark

    15 calibrated multilabel emotion labels plus 16 speaking styles, transcription, and unified analysis from one synchronous REST call. English prerecorded audio, priced per second (emotion from $0.0060/min), $50 trial credit. The caveat we state everywhere: the oruk benchmark entry is trained in-distribution while other systems are zero-shot.

  2. 02emotion2vec+ (self-hosted)

    68.7% measured

    Best open-weights option

    The highest-scoring non-oruk and open-weights row in this release, and free apart from compute. You take on GPU serving, calibration, thresholding, and scaling — the right trade for ML-mature teams with steady volume.

  3. 03Gemini 3 Flash Preview

    46.0% measured

    Highest-scoring frontier multimodal API in this release

    Prompt an audio-capable LLM and parse the answer. Flexible and convenient if you already run Gemini, but outputs are uncalibrated text, accuracy trails specialist models, and results can shift between model versions.

  4. 04Behavioral Signals

    44.1% measured

    Call-center analytics suites

    A speech-analytics vendor whose emotion API measured strongest among the traditional SER vendors. Aimed at contact-center deployments rather than general developer self-serve.

  5. 05GPT-Audio 1.5

    43.3% measured

    OpenAI-stack teams

    Same pattern as Gemini: prompt-based emotion judgments from a general audio LLM. Convenient inside an existing OpenAI workflow; not calibrated, and mid-pack on accuracy.

  6. 06audEERING devAIce

    42.2% measured

    On-prem and embedded requirements

    From the makers of openSMILE, with SDK and on-premise options that pure cloud APIs lack. Measured accuracy trails the leaders, but deployment flexibility is the differentiator.

Hume’s current product is unranked here

Hume currently advertises Tagger for offline analysis across 600+ expression dimensions and Prosody for real-time signals, with Tagger access routed through Hume Research. Our 49.6% Hume row represents the legacy 48-dimension prosody endpoint, evaluated zero-shot on the 5,000-clip subset. It is retained as a dated historical snapshot, not evidence about current Tagger or Prosody. See Hume’s official product page.

FAQ

How was this list ranked?
By measured seven-class accuracy in the speech-emotion-bench July 2026 release. Open models were evaluated on the full 64,384-clip set; closed/API and audio-LLM systems used a fixed 5,000-clip stratified subset. All rows share the label mapping and scorer, but not the same number of clips. oruk builds the top measured entry and trained it in-distribution, while the compared systems were evaluated zero-shot. This is a vendor-published ranking; results and protocol are downloadable for checking.
What happened to Hume AI on this list?
Hume remains an emotion-measurement option. Its current page advertises offline Tagger and real-time Prosody, but neither current product was evaluated in this release. The 49.6% leaderboard row is a legacy 48-dimension Hume prosody snapshot on the 5,000-clip subset and must not be read as a score for current Tagger or Prosody.
Are LLM APIs like Gemini or GPT-Audio good enough for emotion detection?
For casual use, sometimes. On the fixed 5,000-clip API subset in this release, frontier multimodal APIs scored 40–46% seven-class accuracy. The highest-scoring open specialist measured 68.7% on the full 64,384 clips; the open-subset sensitivity check shifted open-model scores by less than two points. LLM outputs are free text rather than task-specific scores, so production thresholding and monitoring require extra work.
What should I check before picking an emotion API?
Four things: published, reproducible accuracy rather than demo videos; whether thresholds were validated on data like yours; deployment fit (file vs streaming, cloud vs on-prem, language coverage); and honest scope (outputs describe how speech sounds, not what someone feels inside).

oruk publishes this ranking and sells the top measured entry. Review the methodology and limitations before making a purchasing decision. Report a factual issue to access@oruk.ai.

Full 64-system leaderboard Hume AI alternatives oruk emotion API