How much faster can the same speech model run?
Faster speech inference with the original acoustic weights.
Oruk is a research lab training speech foundation models to understand words, tone, and emotion directly from audio. Our API helps developers build AI that can recognize a hesitant answer or respond to a frustrated customer.
01 / Voice agents
Let your agent pick up on hesitation, frustration, or excitement, then use that context to shape its response.
“I’m not sure what to do next.”
Illustrative workflow
Faster speech inference with the original acoustic weights.
An open speech model for 25 languages, with fitted Gabor filters.
Past questions help a quantum memory choose which qubits to connect.
A speech model built from 499 fruit fly neurons.
Six speech systems learn the same filter shape, predicted by signal theory.
Speech models organize voices into shapes you can explore in 3-D.
Emotion structure transfers across languages, with a measurable gap.
Continuous emotion and speaking-style scores, alongside the words.
Mouth shape changes speech. Model happy tags need their own controls.
Shared color associations vary with language, climate, and context.
Prices in USD. Included minutes reset monthly.
$9/ month
Billed monthly
7-day free trial
Start 7-day free trial of Hobby$0 today. Then $9/month. Cancel before the trial ends.
$49/ month
Billed monthly
7-day free trial
Start 7-day free trial of Builder$0 today. Then $49/month. Cancel before the trial ends.
$199/ month
Billed monthly
7-day free trial
Start 7-day free trial of Production$0 today. Then $199/month. Cancel before the trial ends.
Contact us
Tailored to your team
Resonance, Fourier, Realtime, and Proficiency share the speech understanding allowance. Orukeet transcription has a separate allowance; optional tasks and request rounding reduce its estimated minutes. Allowances reset monthly. Trials include 25%; extra usage shares one spending cap.
Now serving: Resonance 2 — emotion and style, same pricing and existing plans.
+9.1 pts vs next best
Oruk trained in-distribution; alternatives zero-shot. Open models: 64,384 clips; API/audio LLMs: 5,000. Hume is the legacy prosody endpoint.
Vocal context vs transcript sentiment
Two identical research probes on MUStARD, not standalone API classifiers. Standard-split confidence intervals overlap.
Protocols & splitsResonance 1 · July–August 2026 evaluation. Whiskers show 95% confidence intervals. Results describe the evaluated checkpoint.
Work with us
Bring us your audio and your use case. We’ll help you find the right model and integration.
Send audio. Get a transcript, speaker turns, and vocal context.
Oruk turns speech into transcripts, speaker turns, and labels for emotion and delivery. Choose a model for recorded audio or live conversations.
Speech-to-text tells you what was said. Oruk also analyzes how it sounded: frustrated, excited, hesitant, sarcastic, and more.
Voice agents with more context, searchable call reviews, expressive captions, research tools, and other products that work with speech.
Yes. Open the full demo to use your microphone or upload a short recording. You can try it without an account.
Try your own audioUse Resonance for a complete English recording: it returns a transcript, emotion and speaking-style labels, and timed segments. Use Realtime for live multilingual transcription with phrase-level emotion. Realtime is in preview and does not return the full file-analysis label set. Orukeet is an option for short English recordings or streamed utterances, with final text after commit.
File analysis supports English audio in WAV, FLAC, MP3, M4A, OGG, and WebM. Realtime supports 32 locales with automatic language detection. Language support for transcription does not imply that every emotion or style task is available in that language.
Oruk’s inference services discard audio and outputs after responding. Optional speaker diarization retains uploaded audio for up to 48 hours and speaker-label results for up to 24 hours. Request metadata is retained for billing, security, and support. Oruk does not use customer audio or outputs to train, fine-tune, or evaluate models without your explicit written agreement.
Data handlingPerformance varies with the model, language, recording, and task. The benchmark report includes dated results, confidence intervals, and evaluation conditions. Emotion scores describe vocal expression, not someone’s inner state.
Read the evaluationsEvery plan includes all speech models. From $9/month with a 7-day free trial.
See pricing