Compare
Compare speech APIs
Compare published results, supported tasks, and pricing. Each comparison links to its sources and explains differences in evaluation conditions.
Explore the results in the benchmark leaderboard and its methodology.
Choose by the speech task
Speech recognition, emotion measurement, and voice generation solve different problems. Compare services that return the output you need before comparing a headline score or price.
- Transcripts, word timing, and speaker turns
- Start with the speech-to-text API and compare Deepgram or AssemblyAI. Check language coverage, streaming status, and transcription errors on your audio.
- Emotion and speaking style in recorded audio
- Compare the emotion API output with Hume expression measurement. Label sets, selected scores, and transcript sentiment are not interchangeable.
- Live transcription with phrase emotion
- Run the Realtime preview example and inspect its events. Oruk’s full speaking-style and unified-analysis outputs use separate file endpoints.
- Continuous valence and arousal
- Choose a dimensional model if you need continuous coordinates. Our valence and arousal guide explains how to evaluate those predictions and why Oruk’s categorical scores are a different output.
Roundups
Best speech emotion recognition APIs
Every system we measured, ranked, with the sample sizes and in-distribution disclosure that the ranking depends on.
Read →Hume AI alternatives
Managed APIs, open models, and audio LLMs that measure vocal emotion, against Hume’s current Tagger and Prosody products.
Read →
Head to head
Orukeet vs hosted Whisper
Transcription prices, tested speed, and accuracy evidence for Orukeet and low-cost hosted Whisper services.
Read →oruk vs Hume AI
How oruk compares with Hume Tagger, Prosody, EVI, and TTS, including which benchmark snapshot is legacy and why that matters.
Read →Oruk vs Valence AI
Pulse emotion analysis and its LiveKit integration compared with Oruk’s file and Realtime contracts, with a practical evaluation checklist.
Read →oruk vs Deepgram
Transcription against transcription, then where acoustic emotion and speaking style differ from text-derived sentiment.
Read →oruk vs AssemblyAI
Transcription, text sentiment versus acoustic emotion, speaking style, streaming, and what each charges per minute.
Read →