Skip to content

Try Resonance 2 early.

Tell us a little about yourself and what you’re building.

A brief description of your use case.

We’ll use these details to review your application and contact you about access. Privacy policy

Explore oruk

Model catalog

Oruk speech models

Meet Spectra-2: transcription in 25 languages, emotion and speaking style in one request. Compare the models below for recordings, live speech and pronunciation scoring.

Looking for a model to run locally?

Orukeet is our open speech recognizer for 25 languages. Download its native, ONNX or NeMo weights for local transcription. The hosted Orukeet preview below currently serves English and offers optional speech tasks. Run Orukeet locally or read the model report.

API catalog and discovery

The base URL is https://speech-api.oruk.ai. Read the public machine-readable model catalog, or use an API key to query GET /v1/models. Subscription allowances and overage rates are on the pricing page.

Input
Realtime PCM16 or prerecorded audio files
Output
Transcripts, emotion, speaking style, or proficiency scores
Language
Spectra-2: 25 languages; Realtime: 32 locales
Limits
File size and duration vary by model
Billing
Monthly allowances for every model
Getting started
One Oruk key and subscription

oruk Spectra-2

New · Stable
oruk-spectra-2

Choose Spectra-2 for multilingual transcription with emotion and speaking-style scores in one request.

Mono 16 kHz WAV, finalized FLAC or raw PCM, 45 ms–60 seconds, up to 4 MiB per file. Scores describe the whole recording. Uses your existing Oruk key and shared speech understanding minutes.

Read the Spectra-2 API referenceTry Spectra-2

Supported languages

Transcription supports these 25 languages. Language detection is automatic; send the recording without a language parameter. The codes below are for reference.

Bulgarian
bg
Croatian
hr
Czech
cs
Danish
da
Dutch
nl
English
en
Estonian
et
Finnish
fi
French
fr
German
de
Greek
el
Hungarian
hu
Italian
it
Latvian
lv
Lithuanian
lt
Maltese
mt
Polish
pl
Portuguese
pt
Romanian
ro
Russian
ru
Slovak
sk
Slovenian
sl
Spanish
es
Swedish
sv
Ukrainian
uk

Emotion and speaking-style scores use the same English label names in every response.

oruk Resonance 2

Preview
oruk-resonance-2

Choose Resonance 2 for emotion and speaking-style scores without transcription. Six signed axes keep opposite labels mutually exclusive, alongside 19 independent scores.

Accepts 0.1–120 seconds and 30 MiB per request. Clip-level acoustic scores; performance varies by language. Same plans and affect pricing as Resonance 1.

Read the Resonance 2 API reference

oruk Orukeet

Best value · Preview
oruk-orukeet

Choose Orukeet for fast English dictation and short recordings. Use native Oruk keys with a subscription, REST uploads, or audio streaming.

Preview: 60 seconds and 4 MiB per request, eight active requests or recordings per organization. Emotion detection and speaker diarization are optional tasks.

Explore Orukeet and its API

oruk Resonance 1

Recommended · Stable
oruk-resonance

Start here for prerecorded English speech. Use one request for transcript, emotion, and speaking style; add speaker labels for calls with multiple people.

File input in English. Speaker labels are available on Resonance 1; they identify turns within a recording, not people across recordings.

Run the Resonance 1 quickstart

oruk Fourier

Stable
oruk-fourier

Evaluate Fourier when your workflow needs transcription and its native emotion output together. It runs those tasks in parallel and returns the shared speaking-style output.

File input in English. Speaker diarization requires Resonance 1. Compare both models on representative recordings; no latency or accuracy advantage is implied.

Run the Fourier quickstart

Realtime

Preview
oruk-realtime

Choose Realtime for live PCM audio, transcript tokens in 32 locales, and phrase-level emotion events over WebSocket.

Preview. Phrase emotion is a separate event stream; speaking-style analysis remains in the English file API. Sessions are limited to 10 minutes.

Connect a live stream

Proficiency

Preview
oruk-proficiency-1

Choose Proficiency for word-level pronunciation feedback, recording scores and legacy CEFR estimates from English speech.

Preview. Use audible English of at least five seconds; short reading prompts suit continuous pronunciation, while 30–60 seconds of spontaneous speech suit CEFR. Scores are estimates, not certification.

Run the proficiency example

Stable identifies the supported file API; preview capabilities may change as they are evaluated. Test representative recordings before production use. See score interpretation, versioned evaluations, and data handling.

Plans and pricing

Resonance 2: Same price as Resonance 1. One audio minute uses one minute from your existing speech understanding allowance, with the same overage rate. No plan change or separate add-on.

Plans cover the speech model lineup. New organizations need separate approval for Resonance 2; an active trial or plan does not by itself grant access. Hobby is $9/month, Builder $49/month, and Production $199/month. Spectra-2, Resonance 1, Resonance 2, Fourier, Realtime, and Proficiency share 250, 2,500, or 20,000 monthly audio minutes, respectively, including supported diarization. Each plan also includes a separate Orukeet allowance, starting at up to 20,000 transcription minutes on Hobby. Orukeet uses $0.00045/minute ($0.027/hour); optional tasks draw from its allowance at their published rates.

Speech understanding subscription plans
SubscriptionMonthly priceSpeech understanding minutesAdditional minute
Hobby$9250$0.020
Builder$492,500$0.015
Production$19920,000$0.012

Standard self-serve plans start with a 7-day free trial: a card is required, $0 is charged today, and you can cancel before the trial ends. Promotional offers show their own terms at signup.

Both subscription allowances reset monthly, including on annual plans. Other-model audio uses a one-second minimum. Orukeet measures actual duration and rounds each request to one microdollar. Optional tasks use its allowance; excess usage is billed monthly within the shared spending cap. Usage cost fields show consumption, not the final invoice. Compare subscriptions and Enterprise options.

Endpoints

Choose the endpoint for your task and pass a supported model ID. Spectra-2, Resonance 1 and Fourier handle file analysis. Resonance 2 has a dedicated affect route. Realtime uses WebSocket, and Proficiency has a dedicated scoring endpoint.

  • POST /v1/audio/transcriptionsTranscript; language support varies by model
  • POST /v1/audio/emotionsMultilabel emotion scores
  • POST /v1/audio/stylesMultilabel speaking-style scores
  • POST /v1/audio/affectEmotion and style labels
  • POST /v1/audio/analysisTranscript, labels, segments, and tagged text
  • POST /v1/audio/proficiency0–1 word and recording scores, legacy CEFR, fluency, transcript
  • WS /v1/realtimeMultilingual transcription and phrase-level emotion

Full request and response documentation is in the API docs, rates are on the pricing page, and accuracy comparisons follow the benchmark methodology. Data handling is covered by the privacy policy and terms of service.