Comparison
oruk vs Hume AI
oruk is not affiliated with, endorsed by, or sponsored by Hume AI, Inc. “Hume” and product names are trademarks of their owner, used here for identification and comparison only. Claims about Hume products are sourced below and dated; if anything is out of date, email access@oruk.ai and we will correct it.
Correction — July 20, 2026: an earlier version treated a legacy Hume Expression Measurement snapshot as evidence that Hume had left the category. Hume’s current official product page describes offline Tagger and real-time Prosody products. The comparison below now separates those products from the legacy API snapshot in our benchmark.
The honest verdict first: Hume and oruk overlap in voice emotion measurement, but they package it differently. Hume currently markets Tagger for offline, high-dimensional analysis and Prosody for real-time signals, alongside EVI and TTS. oruk offers a self-serve, synchronous API for transcription, emotion, and speaking style from recorded English audio. The 49.6% Hume result below belongs to Hume’s legacy 48-dimension prosody endpoint; it is not a score for current Tagger or Prosody.
Side by side
| oruk | Hume AI | |
|---|---|---|
| Product focus (July 2026) | Measuring speech: transcription, emotion, speaking style from audio files | Expression measurement (Tagger and Prosody), voice agents (EVI), and TTS |
| Offline emotion measurement | Yes — POST /v1/audio/emotions, synchronous | Tagger — batch analysis; Hume currently directs users to contact Research for access |
| Emotion output | 15 calibrated multilabel emotion scores + 16 speaking styles | Current Tagger advertises 600+ expression dimensions; the legacy prosody endpoint returned 48 |
| speech-emotion-bench accuracy | 77.6% (trained in-distribution; see methodology) | 49.6% for the legacy 48-dimension prosody snapshot on the 5,000-clip subset; current Tagger/Prosody not tested |
| Input | Prerecorded English audio files (WAV, FLAC, MP3, M4A, OGG, WebM) | Recorded audio through Tagger; live audio through Prosody/EVI; text for TTS |
| Streaming / real time | No — file-based API v1 | Yes — Prosody returns real-time signals; EVI is speech-to-speech |
| Pricing model | Per second of audio; emotion from $0.0060/min, analysis $0.0090/min | Current Tagger/Prosody access and pricing: contact Hume Research |
Benchmark numbers are same-harness measurements from speech-emotion-bench (77.6% oruk Spectra vs 49.6% Hume prosody); the oruk entry is trained in-distribution and the Hume legacy API row was scored on the fixed 5,000-clip subset. Open models use the full 64,384-clip evaluation. The caveats and protocol are on the methodology page. oruk pricing is version 2026-07-11 ($0.0060/min emotion, $0.0090/min unified analysis).
Choose Hume if
- You need offline Tagger analysis across Hume’s advertised 600+ expression dimensions and can use its contact-led access path.
- You need real-time emotion signals from the Prosody API.
- You are building a live voice agent and want speech-to-speech with expressive prosody (EVI 3).
- You need controllable, emotionally expressive text-to-speech (Octave).
Choose oruk if
- You need to measure emotion or speaking style in recorded calls, meetings, or interviews.
- You want a compact multilabel space with published thresholding and benchmark caveats rather than hundreds of dimensions.
- You want the transcript, emotion, and style from one request, priced per second of audio.
FAQ
- Is Hume AI still an option for measuring emotion in audio files?
- Yes. Hume’s current Expression Measurement page advertises Tagger for offline batch analysis and Prosody for real-time signals. The 49.6% benchmark row comes from a dated legacy 48-dimension prosody snapshot and should not be used to infer current product availability or performance. Hume currently asks prospective Tagger users to contact its Research team for access.
- When should I choose Hume over oruk?
- Choose Hume if you need its 600+ dimension Tagger taxonomy, real-time Prosody signals, or a speech-to-speech agent through EVI. Choose oruk if you need a self-serve synchronous REST API for recorded English audio that returns transcript, a smaller emotion/style label space, and per-second pricing.
- How were the benchmark numbers measured?
- On speech-emotion-bench: open models were scored on the full 64,384-clip evaluation, while closed/API and audio-LLM systems were scored on a fixed 5,000-clip stratified subset. All rows use the same seven-class mapping and scorer. The oruk entry is trained in-distribution while the compared systems are evaluated zero-shot. The Hume row is a legacy 48-dimension prosody snapshot, not current Tagger or Prosody.
- Does oruk detect the same expression dimensions as Hume?
- No. The legacy Hume prosody endpoint returned 48 dimensions, while current Tagger advertises 600+. oruk returns 15 emotion labels and 16 speaking-style labels using thresholds selected on held-out audio. Those thresholds still need validation on your speakers and domain; the migration guide maps only near-synonyms.
Sources
- Hume AI’s current Expression Measurement page, accessed July 20, 2026 — Tagger offline, Prosody real time, 600+ dimensions, 50+ languages, and contact-led access.
- Hume’s EVI FAQ, accessed July 20, 2026 — current EVI use of prosody measures and the distinction between expression labels and inner state.
- speech-emotion-bench protocol, sample counts, and downloadable per-system results: benchmark methodology. The Hume row is the legacy 48-dimension prosody endpoint, evaluated zero-shot on the 5,000-clip stratified subset; it is not a measurement of current Tagger or Prosody.
- oruk pricing: version 2026-07-11 on the pricing page.
Migration guide from Hume EM All Hume AI alternatives oruk emotion API
Measure emotion in your own audio
Create an account and run the oruk API on your recordings with $50 in trial credit — no card required, from $0.0060 per audio minute.
