Speech understanding
What is speech understanding?
Speech understanding uses spoken audio to study words, delivery or meaning. This guide focuses on how the words were delivered: emotion labels, speaking style, prosody and changes over time. A transcript helps a reviewer read the conversation; acoustic annotations help them find passages to hear. Oruk’s English file API returns transcripts and selected vocal-expression labels. See the model-specific capabilities.
Speech understanding vs. speech recognition
AI transcription, also called automatic speech recognition (ASR) or speech-to-text, converts speech into written words. Accuracy depends on the language, recording conditions, vocabulary and model. Some systems also return word timings or speaker turns, but text alone does not preserve the original pitch, loudness and voice quality.
Two recordings of “great, thanks” can have the same transcript and different delivery. That makes the audio worth hearing; it does not make gratitude or sarcasm a fact a classifier can verify. A review application can combine a transcript with timed annotations and let a person replay each passage. Oruk’s analysis endpoint returns both layers for an English recording in one request.
The saved conversation walkthrough includes licensed audio and its actual JSON response. Compare the words, intervals and model labels without creating an account. If you only need text, start with the transcription API or the local Orukeet guide.
What acoustic speech models measure
- Emotion & affect
- Selected labels such as happy or frustrated describe learned vocal patterns. Their usefulness varies with the recording, speaker and evaluation setting.
- Speaking style
- A separate 16-label multilabel task describing how something was said. The labels characterize delivery, not intent or inner state.
- Sarcastic-sounding delivery
- A listener may interpret delivery differently from the literal words. A model label does not establish sarcasm, sincerity or intent.
- Prosody
- Pitch, pace, pauses and stress are measurable properties of a recording. Their interpretation depends on the speaker and context.
- Time-local change
- How emotion and speaking-style scores vary across segments of a recorded conversation.
- Frustration
- A model-selected label can help a reviewer locate a passage to hear in context. It is not a verified account of the speaker’s feelings.
Why it matters for conversation review
After a conversation ends, transcript-only review can miss frustrated or sarcastic-sounding delivery. Teams can combine transcripts with emotion and speaking-style scores to prioritize completed calls for human review. Voice-agent builders can apply the same analysis to prerecorded evaluation conversations to find responses that need revision. These are post-call review and evaluation workflows, not live routing or de-escalation.
Evaluate those review decisions on recordings your team can label. Our smile recording controls show why the input matters: adding leading silence changed model happy-tag decisions even though the speech samples were unchanged. The local review panel lets reviewers record which passages were useful, uncertain or unreviewed.
oruk models words and acoustic labels together
oruk API v1 transcribes English audio and returns selected multilabel emotion and speaking-style predictions. oruk also publishes measured results through dated evaluations and their replay artifacts. The production interface is a versioned REST API for prerecorded files.