Skip to content

Explore oruk

Speech-to-text, emotion, intent, tone of voice, meaning, sarcasm, and understanding.

Oruk is a research lab training speech foundation models to understand words, tone, and emotion directly from audio. Our API helps developers build AI that can recognize a hesitant answer or respond to a frustrated customer.

Backed by

a16z speedrunAlpha by a16z speedrunNVIDIA Inception Program

Research backgrounds

StanfordBerkeleyCambridge
try resonance-1

r u ok? understand how things are said

Hear the difference

Same words. Different delivery.

Loading the recording…

01 / Voice agents

Give your agent a better ear.

Let your agent pick up on hesitation, frustration, or excitement, then use that context to shape its response.

  • Phrase-level emotion
  • Conversation review
  • Response design
Explore live speech
Caller

“I’m not sure what to do next.”

worriedhesitant
Make room for a clearer explanation.

Illustrative workflow

Research highlights

All research

More audio. Less spend.

See all pricing
Billing period

Prices in USD. Included minutes reset monthly.

  • Hobby

    $9/ month

    Billed monthly

    Speech understanding
    250 min/month
    Transcription · Orukeet
    20,000 min/month

    7-day free trial

    Start 7-day free trial of Hobby

    $0 today. Then $9/month. Cancel before the trial ends.

  • Builder

    Recommended

    $49/ month

    Billed monthly

    Speech understanding
    2,500 min/month
    Transcription · Orukeet
    ≈ 109,000 min/month

    7-day free trial

    Start 7-day free trial of Builder

    $0 today. Then $49/month. Cancel before the trial ends.

  • Production

    $199/ month

    Billed monthly

    Speech understanding
    20,000 min/month
    Transcription · Orukeet
    ≈ 442,000 min/month

    7-day free trial

    Start 7-day free trial of Production

    $0 today. Then $199/month. Cancel before the trial ends.

  • Enterprise

    Contact us

    Tailored to your team

    Speech understanding
    Custom
    Transcription · Orukeet
    Custom

    7-day free trial

    Contact us about Enterprise

    Schedule a 30-minute call

How the two allowances work

Resonance, Fourier, Realtime, and Proficiency share the speech understanding allowance. Orukeet transcription has a separate allowance; optional tasks and request rounding reduce its estimated minutes. Allowances reset monthly. Trials include 25%; extra usage shares one spending cap.

Now serving: Resonance 2 — emotion and style, same pricing and existing plans.

Speech emotion

Accuracy ↑

+9.1 pts vs next best

oruk Resonance 177.8%
95% confidence interval 77.5 to 78.1
emotion2vec+ seed68.7%
95% confidence interval 68.3 to 69.1
EmotionThinker60.5%
95% confidence interval 59.1 to 61.8
SenseVoice Small55.7%
95% confidence interval 55.3 to 56.1
Hume AI legacy prosody49.6%
95% confidence interval 48.2 to 51.0
Gemini 3 Flash Preview46.0%
95% confidence interval 44.6 to 47.4
GPT-Audio 1.543.3%
95% confidence interval 41.9 to 44.7
Qwen2.5-Omni 7B31.3%
95% confidence interval 30.0 to 32.6

Oruk trained in-distribution; alternatives zero-shot. Open models: 64,384 clips; API/audio LLMs: 5,000. Hume is the legacy prosody endpoint.

July–August 2026 snapshot

Speech emotion: evaluation details

Sarcasm

Macro-F1 × 100 ↑

Vocal context vs transcript sentiment

Standard 5-fold

Oruk Resonance 156.2
95% confidence interval 52.4 to 60.1
Deepgram nova-351.7
95% confidence interval 47.8 to 55.3

Speaker-independent

Oruk Resonance 151.0
95% confidence interval 47.1 to 54.8
Deepgram nova-339.5
95% confidence interval 35.8 to 43.1

Cross-show

Oruk Resonance 146.2
95% confidence interval 40.9 to 51.5
Deepgram nova-329.9
95% confidence interval 27.3 to 32.2

Two identical research probes on MUStARD, not standalone API classifiers. Standard-split confidence intervals overlap.

Protocols & splits

Resonance 1 · July–August 2026 evaluation. Whiskers show 95% confidence intervals. Results describe the evaluated checkpoint.

Work with us

Talk to the people
behind the models.

Bring us your audio and your use case. We’ll help you find the right model and integration.

Partnerships

Build with Oruk

Tell us what you’re building and where speech fits.

Contact our team at access@oruk.ai.

Build with Oruk

Send audio. Get a transcript, speaker turns, and vocal context.

Common questions

What exactly does Oruk do?

Oruk turns speech into transcripts, speaker turns, and labels for emotion and delivery. Choose a model for recorded audio or live conversations.

How is it different from speech-to-text?

Speech-to-text tells you what was said. Oruk also analyzes how it sounded: frustrated, excited, hesitant, sarcastic, and more.

What can I build with it?

Voice agents with more context, searchable call reviews, expressive captions, research tools, and other products that work with speech.

Can I try my own audio?

Yes. Open the full demo to use your microphone or upload a short recording. You can try it without an account.

Try your own audio
Does it work in real time?

Use Resonance for a complete English recording: it returns a transcript, emotion and speaking-style labels, and timed segments. Use Realtime for live multilingual transcription with phrase-level emotion. Realtime is in preview and does not return the full file-analysis label set. Orukeet is an option for short English recordings or streamed utterances, with final text after commit.

Which languages and formats are supported?

File analysis supports English audio in WAV, FLAC, MP3, M4A, OGG, and WebM. Realtime supports 32 locales with automatic language detection. Language support for transcription does not imply that every emotion or style task is available in that language.

What happens to my recordings?

Oruk’s inference services discard audio and outputs after responding. Optional speaker diarization retains uploaded audio for up to 48 hours and speaker-label results for up to 24 hours. Request metadata is retained for billing, security, and support. Oruk does not use customer audio or outputs to train, fine-tune, or evaluate models without your explicit written agreement.

Data handling
How accurate are the results?

Performance varies with the model, language, recording, and task. The benchmark report includes dated results, confidence intervals, and evaluation conditions. Emotion scores describe vocal expression, not someone’s inner state.

Read the evaluations
How does pricing work?

Every plan includes all speech models. From $9/month with a 7-day free trial.

See pricing