Skip to content

Speech AI that actuallyunderstands you

Try it live — tap and speak

Built by researchers from

Stanford UniversityUC BerkeleyUC Santa BarbaraUniversity of Cambridge

Backed by

a16z speedrunalpha by a16z speedrun

Live demos

See it working before you build

The same API behind every demo: transcription, emotion, and speaking style from one call.

All demos

Evaluations · August 2026 report

Measured speech results

Full results & methodology

Speech emotion

77.8%+8.9 pts vs best open model
oruk Spectra 1Other systems
77.8
Spectra 1
68.7
e2v+ seed
68.6
e2v+ large
68.5
e2v+ base
63.6
e2v FT
60.5
EmoThinker
55.7
SenseVoice
55.1
e2v probe
52.4
Kimi-Audio
51.4
Qwen Omni

IEMOCAP

71.5%+7.9 pts vs next commercial API
emotion labelssentiment outputvalence accuracy % · 95% CIchance 33.3oruk Spectra 171.5Imentiv63.6Smallest.ai Pulse59.3Mistral Voxtral Small54.4Rev AI sentiment53.5Hume EVI prosody51.3Deepgram nova-3 sentiment46.3AssemblyAI sentiment36.0Gladia sentiment35.7

identical audio to every vendor · wider intervals mean the vendor stopped early at credit or rate limits · hover a row

Sarcasm

56.2%Above baseline on every split
orukaudio affectDeepgramtranscript sentimentmacro-F1 · 95% CIalways not sarcastic · 0.3330.562±0.0380.517±0.038Standard 5-foldspeaker-dependent0.510±0.0390.395±0.037Speaker-independentgrouped 5-fold0.462±0.0530.299±0.024Cross-showtrain BBT+GG → test Friends

Transcription

4.44%Lowest WER in the measured panel
4.44
Spectra 1
4.47
ASR A
5.50
Azure
5.70
Whisper L
6.10
Whisper M
6.53
ASR B
7.00
Whisper S
8.90
Google
9.70
Leopard
10.1
Cheetah

Per-emotion F1

Every emotion class, measured

oruk Spectra against Gemini 3 Flash Preview, the strongest frontier multimodal API in the August 2026 report, on all seven emotion classes.

oruk Spectra 1Gemini 3 Flash
0.91
0.18
Disgust
0.86
0.35
Fear
0.86
0.27
Surprise
0.81
0.46
Anger
0.76
0.53
Neutral
0.75
0.31
Sadness
0.74
0.50
Happiness

Models & pricing

A model for every conversation.

Flagship

oruk Resonance

Highest-accuracy transcription and unified analysis.

  • English transcription and acoustic context
  • 15 emotion labels and 16 speaking-style labels
  • Multilabel scores calibrated on held-out audio
PricingUSD / audio min
Transcription
$0.0080
Emotion + style
$0.0100
Unified analysis
$0.0120
oruk-resonanceJoin the waitlist
Efficient

oruk Spectra 1

The efficient choice for emotion and style.

  • Recommended for emotion and style endpoints
  • Returns calibrated multilabel predictions
  • Affect requests skip unnecessary transcript output
PricingUSD / audio min
Transcription
$0.0045
Emotion + style
$0.0075
Unified analysis
$0.0090
oruk-spectra-1Join the waitlist
Preview

oruk Spectra 2

A preview tier that is not yet serving traffic.

  • Listed in the versioned model catalog as preview
  • Not yet serving inference traffic
  • Will support transcription, emotion, style, and analysis
PricingUSD / audio min
Transcription
$0.0065
Emotion + style
$0.0100
Unified analysis
$0.0120
oruk-spectra-2Join the waitlist

What does an hour of audio cost?

Per audio hour
$0.72
Per audio minute
$0.0120
Audio per month
hours

$7.20 per month

The $50 trial credit covers about 69 hours at this rate.

50% less than buying transcript, emotion, and style separately.

Billed by the measured second. $50 trial credit, no card required. Volume rates and on-device licensing on the pricing page.

Scope

What oruk does

The API today

  • English transcription. Transcribes prerecorded English audio files via POST /v1/audio/transcriptions.

  • Multilabel emotion detection. 15 emotion labels with scores calibrated on held-out audio; a clip can carry several labels at once.

  • Speaking-style classification. 16 speaking-style labels (calibrated, multilabel) describing how something was said.

  • Combined affect. Emotion and speaking style in one call, skipping transcript output when it is not needed.

  • Unified analysis. One request returns transcript, calibrated labels, time-local segments, and tagged text.

API v1 · scope statement updated 2026-07-13. Full details on the capabilities page.

Company

A speech lab building machines that understand people

About the lab
  • 01

    Meaning over transcription

    A transcript preserves words but drops acoustic information. oruk models the transcript, emotion, and speaking style together so applications can use both.

  • 02

    Research in the open

    We benchmark against the field on its own terms and publish where it counts. Progress you can verify, not claims you have to take on faith.

  • 03

    Built for builders

    Understanding should be one API call away. We carry the hard parts so the teams shipping conversational AI to real people don’t have to.

Questions

Speech understanding, answered.

What does the oruk API do today?

oruk API v1 processes English audio files for transcription, emotion detection, speaking-style classification, or unified analysis. Unified responses include a transcript, calibrated labels, time-local segments, and tagged text. The current public contract is file based, not streaming.

What speech models does oruk offer?

The API serves two tiers today: Spectra 1 for efficient affect workloads and Resonance, the flagship, for the strongest transcription and unified analysis. Spectra 2 is listed in the catalog as a preview tier but is not yet serving traffic.

How does multilabel emotion detection work?

The API scores 15 emotion labels against thresholds calibrated on held-out audio. A clip may return several emotions. If no emotion crosses its threshold, the highest-scoring emotion is returned as a fallback. Long audio also includes segment-level labels so applications can follow changes over time.

How are speaking styles represented?

Speaking style is a separate 16-label multilabel output. A clip may have multiple styles or no selected style. Emotion and style can be requested independently, together through the affect endpoint, or alongside a transcript through unified analysis.

Is oruk a speech-to-text API?

Yes. Resonance is the recommended model for English transcription. The same API can also return acoustic emotion and speaking-style labels, which preserves useful information that a transcript alone does not contain.

How do developers access oruk?

Developers create an account, which activates immediately with $50 in trial credit, and generate a production API key in the developer portal. The v1 API uses bearer authentication and multipart file uploads. Keys are displayed only once.

Which languages does API v1 support?

API v1 supports English. Inputs may be mono or stereo and use WAV, FLAC, MP3, M4A, OGG, or WebM containers with common sample rates. oruk does not currently advertise multilingual production support for this API.

How is usage billed?

Inference is priced per audio minute and metered by measured second, with a one-second minimum. Each response includes the billable duration, rate, estimated cost, and pricing version. The developer portal shows the immutable credit ledger and request-level usage history.