Meet Spectra-2: transcription in 25 languages, emotion and speaking style in one request. Compare the models below for recordings, live speech and pronunciation scoring.
Orukeet is our open speech recognizer for 25 languages. Download its native, ONNX or NeMo weights for local transcription. The hosted Orukeet preview below currently serves English and offers optional speech tasks. Run Orukeet locally or read the model report.
API catalog and discovery
The base URL is https://speech-api.oruk.ai. Read the public machine-readable model catalog, or use an API key to query GET /v1/models. Subscription allowances and overage rates are on the pricing page.
Input
Realtime PCM16 or prerecorded audio files
Output
Transcripts, emotion, speaking style, or proficiency scores
Language
Spectra-2: 25 languages; Realtime: 32 locales
Limits
File size and duration vary by model
Billing
Monthly allowances for every model
Getting started
One Oruk key and subscription
oruk Spectra-2
New · Stable
oruk-spectra-2
Choose Spectra-2 for multilingual transcription with emotion and speaking-style scores in one request.
Mono 16 kHz WAV, finalized FLAC or raw PCM, 45 ms–60 seconds, up to 4 MiB per file. Scores describe the whole recording. Uses your existing Oruk key and shared speech understanding minutes.
Transcription supports these 25 languages. Language detection is automatic; send the recording without a language parameter. The codes below are for reference.
Bulgarian
bg
Croatian
hr
Czech
cs
Danish
da
Dutch
nl
English
en
Estonian
et
Finnish
fi
French
fr
German
de
Greek
el
Hungarian
hu
Italian
it
Latvian
lv
Lithuanian
lt
Maltese
mt
Polish
pl
Portuguese
pt
Romanian
ro
Russian
ru
Slovak
sk
Slovenian
sl
Spanish
es
Swedish
sv
Ukrainian
uk
Emotion and speaking-style scores use the same English label names in every response.
oruk Resonance 2
Preview
oruk-resonance-2
Choose Resonance 2 for emotion and speaking-style scores without transcription. Six signed axes keep opposite labels mutually exclusive, alongside 19 independent scores.
Accepts 0.1–120 seconds and 30 MiB per request. Clip-level acoustic scores; performance varies by language. Same plans and affect pricing as Resonance 1.
Choose Orukeet for fast English dictation and short recordings. Use native Oruk keys with a subscription, REST uploads, or audio streaming.
Preview: 60 seconds and 4 MiB per request, eight active requests or recordings per organization. Emotion detection and speaker diarization are optional tasks.
Start here for prerecorded English speech. Use one request for transcript, emotion, and speaking style; add speaker labels for calls with multiple people.
File input in English. Speaker labels are available on Resonance 1; they identify turns within a recording, not people across recordings.
Evaluate Fourier when your workflow needs transcription and its native emotion output together. It runs those tasks in parallel and returns the shared speaking-style output.
File input in English. Speaker diarization requires Resonance 1. Compare both models on representative recordings; no latency or accuracy advantage is implied.
Choose Proficiency for word-level pronunciation feedback, recording scores and legacy CEFR estimates from English speech.
Preview. Use audible English of at least five seconds; short reading prompts suit continuous pronunciation, while 30–60 seconds of spontaneous speech suit CEFR. Scores are estimates, not certification.
Stable identifies the supported file API; preview capabilities may change as they are evaluated. Test representative recordings before production use. See score interpretation, versioned evaluations, and data handling.
Plans and pricing
Resonance 2: Same price as Resonance 1. One audio minute uses one minute from your existing speech understanding allowance, with the same overage rate. No plan change or separate add-on.
Plans cover the speech model lineup. New organizations need separate approval for Resonance 2; an active trial or plan does not by itself grant access. Hobby is $9/month, Builder $49/month, and Production $199/month. Spectra-2, Resonance 1, Resonance 2, Fourier, Realtime, and Proficiency share 250, 2,500, or 20,000 monthly audio minutes, respectively, including supported diarization. Each plan also includes a separate Orukeet allowance, starting at up to 20,000 transcription minutes on Hobby. Orukeet uses $0.00045/minute ($0.027/hour); optional tasks draw from its allowance at their published rates.
Speech understanding subscription plans
Subscription
Monthly price
Speech understanding minutes
Additional minute
Hobby
$9
250
$0.020
Builder
$49
2,500
$0.015
Production
$199
20,000
$0.012
Standard self-serve plans start with a 7-day free trial: a card is required, $0 is charged today, and you can cancel before the trial ends. Promotional offers show their own terms at signup.
Both subscription allowances reset monthly, including on annual plans. Other-model audio uses a one-second minimum. Orukeet measures actual duration and rounds each request to one microdollar. Optional tasks use its allowance; excess usage is billed monthly within the shared spending cap. Usage cost fields show consumption, not the final invoice. Compare subscriptions and Enterprise options.
Endpoints
Choose the endpoint for your task and pass a supported model ID. Spectra-2, Resonance 1 and Fourier handle file analysis. Resonance 2 has a dedicated affect route. Realtime uses WebSocket, and Proficiency has a dedicated scoring endpoint.
POST /v1/audio/transcriptionsTranscript; language support varies by model
POST /v1/audio/emotionsMultilabel emotion scores
POST /v1/audio/stylesMultilabel speaking-style scores
POST /v1/audio/affectEmotion and style labels
POST /v1/audio/analysisTranscript, labels, segments, and tagged text
POST /v1/audio/proficiency0–1 word and recording scores, legacy CEFR, fluency, transcript
WS /v1/realtimeMultilingual transcription and phrase-level emotion