Small typed clients for multipart upload, stable errors, request IDs, timeouts, and bounded retries. Use the versioned packages below for transcription, emotion, speaking style, speaker diarization, and proficiency. Python requires 3.10+; TypeScript requires Node.js 18+ or a compatible fetch runtime.
curl --fail -O https://oruk.ai/examples/analyze-file.py
python analyze-file.py sample.wav --task analysis > result.json
# Other workflows; each command sends a separate request:
python analyze-file.py support-call.wav --diarize > speakers.json
python analyze-file.py speaking-sample.wav --task proficiency > proficiency.json
Spectra-2 · transcription and expression
import os
from oruk import Oruk
with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
result = client.analyze("clip.flac", model="oruk-spectra-2")
print(result["text"])
print(result["emotions"], result["styles"])
Analyze audio
import os
from oruk import Oruk
with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
result = client.analyze(
"sample.wav",
model="oruk-resonance",
)
print(result["text"])
print(result["emotions"])
print(result["styles"])
Transcribe with Orukeet
import os
from oruk import Oruk
with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
result = client.transcribe("sample.wav", model="oruk-orukeet")
print(result["text"])
# Optional task arguments: emotion_detection=True, diarize=True
Diarize speakers
import os
from oruk import Oruk
with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
result = client.analyze("support-call.wav", model="oruk-resonance", diarize=True)
for seg in result["segments"]:
print(seg["speaker"], seg.get("text"), seg.get("emotions", []))
Word pronunciation and legacy CEFR
import os
from oruk import Oruk
with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
result = client.proficiency(
"speaking-sample.wav", model="oruk-proficiency-1",
transcript="The exact words spoken in your recording",
) # omit transcript for automatic transcription
p = result["proficiency"]
print(p["continuous_check"], p["score"]) # 0–1 or None
for word in p["words"]:
print(word["text"], word["start_s"], word["end_s"], word["score"])
# CEFR has its own validity check and keeps the original 0–5 scale.
if result["check"]["status"] != "insufficient_audio":
print(p["cefr"], p["cefr_score"], result["check"]["status"])
Set ORUK_API_KEY in your environment and use your own audio file. Download the complete TypeScript program. These terminal commands use Node.js 22+ with the tsx runner.
Terminal · subscription model workflows
npm install --save-dev tsx
curl --fail -O https://oruk.ai/examples/analyze-file.mts
npx tsx analyze-file.mts sample.wav --task analysis > result.json
# Other workflows; each command sends a separate request:
npx tsx analyze-file.mts support-call.wav --diarize > speakers.json
npx tsx analyze-file.mts speaking-sample.wav --task proficiency > proficiency.json
Spectra-2 · transcription and expression
import { readFile } from "node:fs/promises"
import { Oruk } from "@oruk-ai/sdk"
const apiKey = process.env.ORUK_API_KEY
if (!apiKey) throw new Error("Set ORUK_API_KEY in your environment.")
const client = new Oruk({ apiKey })
const result = await client.analyze({
file: new Blob([await readFile("clip.flac")], { type: "audio/flac" }),
filename: "clip.flac",
model: "oruk-spectra-2",
})
console.log(result.text, result.emotions, result.styles)
Analyze audio
import { readFile } from "node:fs/promises"
import { Oruk } from "@oruk-ai/sdk"
const apiKey = process.env.ORUK_API_KEY
if (!apiKey) throw new Error("Set ORUK_API_KEY in your environment.")
const client = new Oruk({ apiKey })
const bytes = await readFile("sample.wav")
const result = await client.analyze({
file: new Blob([new Uint8Array(bytes)], { type: "audio/wav" }),
filename: "sample.wav",
model: "oruk-resonance",
})
console.log(result.text, result.emotions, result.styles)
Transcribe with Orukeet
import { readFile } from "node:fs/promises"
import { Oruk } from "@oruk-ai/sdk"
const apiKey = process.env.ORUK_API_KEY
if (!apiKey) throw new Error("Set ORUK_API_KEY in your environment.")
const client = new Oruk({ apiKey })
const bytes = await readFile("sample.wav")
const result = await client.transcribe({
file: new Blob([new Uint8Array(bytes)], { type: "audio/wav" }),
filename: "sample.wav",
model: "oruk-orukeet",
// Optional task arguments: emotionDetection: true, diarize: true
})
console.log(result.text)
Node.js 22+ · Experimental · Batch transcription
Use Oruk with AI SDK 7.0.137
This Oruk-maintained experimental adapter supports TranscriptionModelV4 with @ai-sdk/provider 4.0.26. Download the experimental 0.0.0 preview directly from Oruk; it has no npm release or official Vercel listing. Strict compilation and synthetic transport tests passed for the pinned versions. A live provider request has not been qualified through this adapter.
Only oruk-spectra-2 batch transcription is exposed. Supply WAV or finalized FLAC bytes: mono 16 kHz, 45 ms–60 seconds, at most 4 MiB. The service validates the recording format and duration. This adapter does not decode or resample audio, stream, expose other models, or add timestamps, speaker labels or confidence.
Server-side Node.js · your recording sends one request
import { readFile } from "node:fs/promises"
import { randomUUID } from "node:crypto"
import { transcribe } from "ai"
import { createOruk, OrukAPICallError } from "@oruk/ai-sdk-provider"
// Server-side only. Use your own WAV or finalized FLAC recording.
const apiKey = process.env.ORUK_API_KEY
if (!apiKey) throw new Error("Set ORUK_API_KEY in your environment.")
const audio = await readFile("sample.wav")
const requestId = randomUUID() // retain with the file if the outcome is uncertain
const oruk = createOruk({ apiKey })
try {
const result = await transcribe({
model: oruk.transcription("oruk-spectra-2"),
audio,
maxRetries: 0,
providerOptions: { oruk: { requestId } },
abortSignal: AbortSignal.timeout(120_000),
})
console.log(result.text, result.providerMetadata.oruk)
} catch (error) {
if (error instanceof OrukAPICallError) {
console.error(error.code, error.requestId, error.responseRequestId)
}
throw error
}
The result preserves the transcript, duration, and all 15 emotion and 16 speaking-style scores in providerMetadata.oruk. Scores describe the clip's vocal expression, are independent rather than a probability distribution, and do not establish private feelings. Timed segments remain empty; no detected language is inferred from the model's language coverage. Native usage estimates are reference values, not invoices.
Every dispatched failure is non-retryable to AI SDK, including 429, 5xx, timeout and cancellation. A timeout or abort does not prove that inference or billing stopped. Retain the original request ID and exact file when the outcome is uncertain; a completed ID can return 409 without replaying the result. Review the Spectra-2 contract before a manual retry. Each new request uses your existing account allowance and billing terms.
Python · Pipecat integration · Release candidate
Stream transcription and phrase emotion with Pipecat
The separate pipecat-oruk adapter connects a Pipecat pipeline to Oruk Realtime. It returns interim transcripts, final text, and phrase-emotion events. Oruk maintains this community integration; it is not an upstream Pipecat release. The published 0.1.0rc3 hosted adapter source matches rc1, which was tested with Pipecat 1.8.1 on Python 3.11–3.14.
These macOS/Linux commands install the published package and fetch its matching example. Use a mono PCM16 WAV at 8–96 kHz, up to five minutes long. The example streams it at playback speed and exits with an error if the turn fails. On Windows, create the environment with py -3.12 -m venv .venv and activate it with .\.venv\Scripts\Activate.ps1 in PowerShell.
Terminal · Pipecat file-stream example
git clone https://github.com/Oruk-AI/pipecat-oruk.git
cd pipecat-oruk
git checkout fda0c0067bb6b21c1f5490720f3c1e521ea47d85
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install 'pipecat-oruk==0.1.0rc3' 'pipecat-ai==1.8.1'
# Set ORUK_API_KEY in the server environment, then use your WAV file:
python examples/pipecat_stream_file.py /path/to/recording.wav --wait-for-emotions
Import the service with from pipecat_oruk import OrukSTTService. Phrase estimates can arrive after the final text; wait_for_emotions=True waits for turn completion before delivering final text with all returned phrase events. Interim text remains immediate. Estimates describe vocal expression and may be missing or wrong; they are not facts about private feelings.
For microphone input, the repository includes a browser speech demo using WebRTC and Silero. The original rc1 production recording covers consecutive speech turns, cancellation, reconnection, and metering. It does not verify a complete STT/LLM/TTS conversation or assistant barge-in.
Analysis uses Resonance 1 to return a transcript, emotion, and speaking style in one request. The emotion, style, and affect endpoints omit transcription. Proficiency now returns continuous word and recording scores on 0–1 alongside legacy CEFR. Use proficiency.continuous_check for pronunciation validity and check.status for CEFR validity and billing; null pronunciation scores are unavailable, not zero. Diarization labels are local to a recording, not identities or roles. The file SDKs do not open a realtime connection. See the endpoint reference or the realtime workflow.
Emotion and style arrays contain selected model scores, not necessarily every label. The highest-scoring emotion is returned when none clears its threshold; styles can be empty. These scores are not automatically calibrated probabilities of private feelings. See label interpretation.
Response usage may include reference rates or estimated_cost_usd. Those values are not your subscription invoice. Actual billing follows your plan's included minutes and overage terms. Each command processes and meters its recording separately; use plan pricing and your billing records for actual charges.
Continuous proficiency examples
Existing proficiency requests receive the new fields without changing authentication or the model ID. SDK 0.2.10 returns the full JSON response; its TypeScript declarations describe the legacy CEFR fields. For typed use of the new fields, use the current OpenAPI schema or the direct Node.js example below. The separate pronunciation check does not determine billing.