Skip to content

Try Resonance 2 early.

Tell us a little about yourself and what you’re building.

A brief description of your use case.

We’ll use these details to review your application and contact you about access. Privacy policy

Explore oruk

Hear what the transcript misses.

Speech understanding models for transcription and vocal delivery. Choose a model for recorded audio or live speech, with language support and outputs that vary by model.

Backed by

a16z speedrunAlpha by a16z speedrunNVIDIA Inception Program

Research backgrounds

StanfordBerkeleyCambridge
try resonance-1

r u ok? understand how things are said

Hear the difference

Same words. Different delivery.

Loading the recording…

Give your agent a better ear.

Let your agent pick up on hesitation, frustration, or excitement, then use that context to shape its response.

  • Phrase-level emotion
  • Conversation review
  • Response design
Explore live speech
Caller

“I’m not sure what to do next.”

worriedhesitant
Make room for a clearer explanation.

Illustrative workflow

Research highlights

All research

More audio. Less spend.

See all pricing
Billing period

Prices in USD. Included minutes reset monthly.

  • Hobby

    $9/ month

    Billed monthly

    Speech understanding
    250 min/month
    Transcription · Orukeet
    20,000 min/month

    7-day free trial

    Start 7-day free trial of Hobby

    $0 today. Then $9/month. Cancel before the trial ends.

  • Builder

    Recommended

    $49/ month

    Billed monthly

    Speech understanding
    2,500 min/month
    Transcription · Orukeet
    ≈ 109,000 min/month

    7-day free trial

    Start 7-day free trial of Builder

    $0 today. Then $49/month. Cancel before the trial ends.

  • Production

    $199/ month

    Billed monthly

    Speech understanding
    20,000 min/month
    Transcription · Orukeet
    ≈ 442,000 min/month

    7-day free trial

    Start 7-day free trial of Production

    $0 today. Then $199/month. Cancel before the trial ends.

  • Enterprise

    Contact us

    Tailored to your team

    Speech understanding
    Custom
    Transcription · Orukeet
    Custom

    7-day free trial

    Contact us about Enterprise

    Schedule a 30-minute call

How the two allowances work

Resonance 1, Fourier, Realtime, and Proficiency share the speech understanding allowance. Orukeet transcription has a separate allowance; optional tasks and request rounding reduce its estimated minutes. Allowances reset monthly. Trials include 25%; extra usage shares one spending cap.

Now serving: Resonance 2 — emotion and style on existing plans; new organizations need separate approval.

Benchmarks

Oruk at the top of the evaluations below.

Full results & methodology
Independently benchmarked bySpeko

Human speech

Emotion accuracy ↑

91%

1stby score
  1. Oruk Resonance91%
  2. emotion2vec+ large88%
  3. qwen3-asr-flash65%

Synthetic speech

Emotion accuracy ↑

48%

1stby score
  1. Oruk Resonance48%
  2. MERaLiON-SER-v133%
  3. Oruk Fourier33%
Oruk evaluationsPublished methods & results

Speech emotion

7-class accuracy ↑

77.8%

1stby score
  1. Oruk Resonance 177.8%
    95% confidence interval 77.5 to 78.1
  2. emotion2vec+ seed68.7%
    95% confidence interval 68.3 to 69.1
  3. emotion2vec+ large68.6%
    95% confidence interval 68.2 to 69.0

July–August 2026 snapshot

Speech emotion: evaluation details

Sarcasm

Macro-F1 × 100 ↑

56.2

1stby score
  1. Oruk Resonance 156.2
    95% confidence interval 52.4 to 60.1
  2. Deepgram nova-351.7
    95% confidence interval 47.8 to 55.3

MUStARD · standard 5-fold

Oruk evaluation · MUStARD

Sarcasm across three evaluation splits

BeSimple leaderboard

Vocal Affect Bench

7-class accuracy ↑

49.3%

1stby score
  1. Oruk Resonance 249.3%
  2. Gemini 3.8 Flash45.7%
  3. Gemini 3.5 Flash44.3%

Oruk trained on all 280 clips; not held out.

Talk to the people
behind the models.

Bring us your audio and your use case. We’ll help you find the right model and integration.

Partnerships

Build with Oruk

Build with Oruk

Send audio. Get a transcript, speaker turns, and vocal context.

Common questions

What exactly does Oruk do?

Oruk turns speech into transcripts, speaker turns, and labels for emotion and delivery. Choose a model for recorded audio or live conversations.

How is it different from speech-to-text?

Speech-to-text tells you what was said. Oruk also analyzes how it sounded: frustrated, excited, hesitant, sarcastic, and more.

What can I build with it?

Voice agents with more context, searchable call reviews, expressive captions, research tools, and other products that work with speech.

Can I try my own audio?

Yes. Open the full demo to use your microphone or upload a short recording. You can try it without an account.

Try your own audio
Does it work in real time?

Use Resonance 1 for a complete English recording: it returns a transcript, emotion and speaking-style labels, and timed segments. Use Realtime for live multilingual transcription with phrase-level emotion. Realtime is in preview and does not return the full file-analysis label set. Orukeet is an option for short English recordings or streamed utterances, with final text after commit.

Which languages and formats are supported?

Spectra-2 transcribes files in 25 languages; Realtime supports 32 locales. Original Resonance 1 and Fourier support English in WAV, FLAC, MP3, M4A, OGG, or WebM. Formats and limits vary by model. Transcription coverage does not establish emotion accuracy.

What happens to my recordings?

Resonance 2 buffers audio in memory until reset or session close and pending work finishes. Stored results exclude audio; 24-hour authorized replay is not a deletion deadline. Optional diarization keeps audio up to 48h and results up to 24h. Metadata is kept for billing, security and support. No training, fine-tuning or evaluation without explicit written agreement.

Data handling
How accurate are the results?

Performance varies with the model, language, recording, and task. The benchmark report includes dated results, confidence intervals, and evaluation conditions. Emotion scores describe vocal expression, not someone’s inner state.

Read the evaluations
How does pricing work?

Plans start at $9/month with a 7-day free trial. New organizations need separate approval for Resonance 2.

See pricing
Cedartown Foods
OpenWhispr
Demosyne
Evitar
Intangibility
Lightberry
Ownkey for Windows
RentAHuman
TeamTalks
Cedartown Foods
OpenWhispr
Demosyne
Evitar
Intangibility
Lightberry
Ownkey for Windows
RentAHuman
TeamTalks
CandaceAI
AudioCpp-Bindings
Driftwood
Fulloch
JobAds
LocalAI
Phonely
Skribe
TranscriptionSuite
CandaceAI
AudioCpp-Bindings
Driftwood
Fulloch
JobAds
LocalAI
Phonely
Skribe
TranscriptionSuite
SalesEQ
AutoSubs
Edixir
Glaut
Kalmia
Minutes
Pipecat
Takeoff
UneeQ
SalesEQ
AutoSubs
Edixir
Glaut
Kalmia
Minutes
Pipecat
Takeoff
UneeQ
Cybernetic Physics
BabbleWise
ElliQ / Intuition Robotics
Hanson Robotics
Kiloforge
Native States
pyannoteAI
TapTalk
Upscale Cleaning Solutions
Cybernetic Physics
BabbleWise
ElliQ / Intuition Robotics
Hanson Robotics
Kiloforge
Native States
pyannoteAI
TapTalk
Upscale Cleaning Solutions
Wellspoken
Buzz
Ethisys
Hoid
LeadArray
Otter.ai
Remi / Reflect Technology
Taya
VoiceOps
Wellspoken
Buzz
Ethisys
Hoid
LeadArray
Otter.ai
Remi / Reflect Technology
Taya
VoiceOps