Skip to content
Docs Guide

SDK version 0.2.15

Python and TypeScript SDKs

Small typed clients for multipart upload, stable errors, request IDs, timeouts, and bounded retries. Use the versioned packages below for transcription, emotion, speaking style, speaker diarization, and proficiency. Python requires 3.10+; TypeScript requires Node.js 18+ or a compatible fetch runtime.

Use one Oruk API key across the model lineup. Spectra-2 transcribes 25 languages with emotion and speaking-style scores. Reuse an SDK client across recordings to reduce repeat-call latency. Start with a sample recording and your first request. Choose a recipe below for transcription, full speech analysis, or speaking proficiency; use Realtime for live audio.

Install
python -m pip install https://oruk.ai/sdk/oruk-0.2.15-py3-none-any.whl
View registry release 0.2.14 on PyPI

Set ORUK_API_KEY in your environment and use your own audio file. Download the complete Python program, or copy a standalone snippet below.

Terminal · subscription model workflows
curl --fail -O https://oruk.ai/examples/analyze-file.py
python analyze-file.py sample.wav --task analysis > result.json
# Other workflows; each command sends a separate request:
python analyze-file.py support-call.wav --diarize > speakers.json
python analyze-file.py speaking-sample.wav --task proficiency > proficiency.json
Spectra-2 · transcription and expression
import os
from oruk import Oruk

with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
    result = client.analyze("clip.flac", model="oruk-spectra-2")
    print(result["text"])
    print(result["emotions"], result["styles"])
Analyze audio
import os
from oruk import Oruk

with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
    result = client.analyze(
        "sample.wav",
        model="oruk-resonance",
    )

print(result["text"])
print(result["emotions"])
print(result["styles"])
Transcribe with Orukeet
import os
from oruk import Oruk

with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
    result = client.transcribe("sample.wav", model="oruk-orukeet")
print(result["text"])
# Optional task arguments: emotion_detection=True, diarize=True
Diarize speakers
import os
from oruk import Oruk

with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
    result = client.analyze("support-call.wav", model="oruk-resonance", diarize=True)
for seg in result["segments"]:
    print(seg["speaker"], seg.get("text"), seg.get("emotions", []))
Word pronunciation and legacy CEFR
import os
from oruk import Oruk

with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
    result = client.proficiency(
        "speaking-sample.wav", model="oruk-proficiency-1",
        transcript="The exact words spoken in your recording",
    )  # omit transcript for automatic transcription

p = result["proficiency"]
print(p["continuous_check"], p["score"])  # 0–1 or None
for word in p["words"]:
    print(word["text"], word["start_s"], word["end_s"], word["score"])
# CEFR has its own validity check and keeps the original 0–5 scale.
if result["check"]["status"] != "insufficient_audio":
    print(p["cefr"], p["cefr_score"], result["check"]["status"])

TypeScript

Package
Install
npm install @oruk-ai/sdk@0.2.15
View registry release 0.2.15 on npm

Set ORUK_API_KEY in your environment and use your own audio file. Download the complete TypeScript program. These terminal commands use Node.js 22+ with the tsx runner.

Terminal · subscription model workflows
npm install --save-dev tsx
curl --fail -O https://oruk.ai/examples/analyze-file.mts
npx tsx analyze-file.mts sample.wav --task analysis > result.json
# Other workflows; each command sends a separate request:
npx tsx analyze-file.mts support-call.wav --diarize > speakers.json
npx tsx analyze-file.mts speaking-sample.wav --task proficiency > proficiency.json
Spectra-2 · transcription and expression
import { readFile } from "node:fs/promises"
import { Oruk } from "@oruk-ai/sdk"

const apiKey = process.env.ORUK_API_KEY
if (!apiKey) throw new Error("Set ORUK_API_KEY in your environment.")
const client = new Oruk({ apiKey })
const result = await client.analyze({
  file: new Blob([await readFile("clip.flac")], { type: "audio/flac" }),
  filename: "clip.flac",
  model: "oruk-spectra-2",
})
console.log(result.text, result.emotions, result.styles)
Analyze audio
import { readFile } from "node:fs/promises"
import { Oruk } from "@oruk-ai/sdk"

const apiKey = process.env.ORUK_API_KEY
if (!apiKey) throw new Error("Set ORUK_API_KEY in your environment.")
const client = new Oruk({ apiKey })
const bytes = await readFile("sample.wav")
const result = await client.analyze({
  file: new Blob([new Uint8Array(bytes)], { type: "audio/wav" }),
  filename: "sample.wav",
  model: "oruk-resonance",
})

console.log(result.text, result.emotions, result.styles)
Transcribe with Orukeet
import { readFile } from "node:fs/promises"
import { Oruk } from "@oruk-ai/sdk"

const apiKey = process.env.ORUK_API_KEY
if (!apiKey) throw new Error("Set ORUK_API_KEY in your environment.")
const client = new Oruk({ apiKey })
const bytes = await readFile("sample.wav")
const result = await client.transcribe({
  file: new Blob([new Uint8Array(bytes)], { type: "audio/wav" }),
  filename: "sample.wav",
  model: "oruk-orukeet",
  // Optional task arguments: emotionDetection: true, diarize: true
})
console.log(result.text)

Node.js 22+ · Experimental · Batch transcription

Use Oruk with AI SDK 7.0.137

This Oruk-maintained experimental adapter supports TranscriptionModelV4 with @ai-sdk/provider 4.0.26. Download the experimental 0.0.0 preview directly from Oruk; it has no npm release or official Vercel listing. Strict compilation and synthetic transport tests passed for the pinned versions. A live provider request has not been qualified through this adapter.

Install · pinned experimental adapter
npm install --save-exact ai@7.0.137 https://oruk.ai/sdk/oruk-ai-sdk-provider-0.0.0.tgz
Download experimental package

Only oruk-spectra-2 batch transcription is exposed. Supply WAV or finalized FLAC bytes: mono 16 kHz, 45 ms–60 seconds, at most 4 MiB. The service validates the recording format and duration. This adapter does not decode or resample audio, stream, expose other models, or add timestamps, speaker labels or confidence.

Server-side Node.js · your recording sends one request
import { readFile } from "node:fs/promises"
import { randomUUID } from "node:crypto"
import { transcribe } from "ai"
import { createOruk, OrukAPICallError } from "@oruk/ai-sdk-provider"

// Server-side only. Use your own WAV or finalized FLAC recording.
const apiKey = process.env.ORUK_API_KEY
if (!apiKey) throw new Error("Set ORUK_API_KEY in your environment.")
const audio = await readFile("sample.wav")
const requestId = randomUUID() // retain with the file if the outcome is uncertain
const oruk = createOruk({ apiKey })
try {
  const result = await transcribe({
    model: oruk.transcription("oruk-spectra-2"),
    audio,
    maxRetries: 0,
    providerOptions: { oruk: { requestId } },
    abortSignal: AbortSignal.timeout(120_000),
  })
  console.log(result.text, result.providerMetadata.oruk)
} catch (error) {
  if (error instanceof OrukAPICallError) {
    console.error(error.code, error.requestId, error.responseRequestId)
  }
  throw error
}

The result preserves the transcript, duration, and all 15 emotion and 16 speaking-style scores in providerMetadata.oruk. Scores describe the clip's vocal expression, are independent rather than a probability distribution, and do not establish private feelings. Timed segments remain empty; no detected language is inferred from the model's language coverage. Native usage estimates are reference values, not invoices.

Every dispatched failure is non-retryable to AI SDK, including 429, 5xx, timeout and cancellation. A timeout or abort does not prove that inference or billing stopped. Retain the original request ID and exact file when the outcome is uncertain; a completed ID can return 409 without replaying the result. Review the Spectra-2 contract before a manual retry. Each new request uses your existing account allowance and billing terms.

Python · Pipecat integration · Release candidate

Stream transcription and phrase emotion with Pipecat

The separate pipecat-oruk adapter connects a Pipecat pipeline to Oruk Realtime. It returns interim transcripts, final text, and phrase-emotion events. Oruk maintains this community integration; it is not an upstream Pipecat release. The published 0.1.0rc3 hosted adapter source matches rc1, which was tested with Pipecat 1.8.1 on Python 3.11–3.14.

These macOS/Linux commands install the published package and fetch its matching example. Use a mono PCM16 WAV at 8–96 kHz, up to five minutes long. The example streams it at playback speed and exits with an error if the turn fails. On Windows, create the environment with py -3.12 -m venv .venv and activate it with .\.venv\Scripts\Activate.ps1 in PowerShell.

Terminal · Pipecat file-stream example
git clone https://github.com/Oruk-AI/pipecat-oruk.git
cd pipecat-oruk
git checkout fda0c0067bb6b21c1f5490720f3c1e521ea47d85
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install 'pipecat-oruk==0.1.0rc3' 'pipecat-ai==1.8.1'
# Set ORUK_API_KEY in the server environment, then use your WAV file:
python examples/pipecat_stream_file.py /path/to/recording.wav --wait-for-emotions

Import the service with from pipecat_oruk import OrukSTTService. Phrase estimates can arrive after the final text; wait_for_emotions=True waits for turn completion before delivering final text with all returned phrase events. Interim text remains immediate. Estimates describe vocal expression and may be missing or wrong; they are not facts about private feelings.

For microphone input, the repository includes a browser speech demo using WebRTC and Silero. The original rc1 production recording covers consecutive speech turns, cancellation, reconnection, and metering. It does not verify a complete STT/LLM/TTS conversation or assistant barge-in.

Choose the endpoint and interpret its output

Analysis uses Resonance 1 to return a transcript, emotion, and speaking style in one request. The emotion, style, and affect endpoints omit transcription. Proficiency now returns continuous word and recording scores on 0–1 alongside legacy CEFR. Use proficiency.continuous_check for pronunciation validity and check.status for CEFR validity and billing; null pronunciation scores are unavailable, not zero. Diarization labels are local to a recording, not identities or roles. The file SDKs do not open a realtime connection. See the endpoint reference or the realtime workflow.

Emotion and style arrays contain selected model scores, not necessarily every label. The highest-scoring emotion is returned when none clears its threshold; styles can be empty. These scores are not automatically calibrated probabilities of private feelings. See label interpretation.

Response usage may include reference rates or estimated_cost_usd. Those values are not your subscription invoice. Actual billing follows your plan's included minutes and overage terms. Each command processes and meters its recording separately; use plan pricing and your billing records for actual charges.

Continuous proficiency examples

Existing proficiency requests receive the new fields without changing authentication or the model ID. SDK 0.2.10 returns the full JSON response; its TypeScript declarations describe the legacy CEFR fields. For typed use of the new fields, use the current OpenAPI schema or the direct Node.js example below. The separate pronunciation check does not determine billing.

Python and TypeScript speech client behavior

The following defaults apply to the speech clients above. The separate experimental AI SDK adapter never automatically retries a dispatched request.

  • Unique X-Request-ID generated for each logical request and reused across retries
  • HTTP retries for 429, 500, 502, 503, and 504; other HTTP errors are not retried
  • Exponential backoff with jitter and two retries by default
  • 120-second default timeout, configurable per client
  • TypeScript also retries network/timeout failures; Python propagates httpx network errors
  • Structured status, code, and request ID on API errors
Continue to the production guide