Metadata-Version: 2.4
Name: oruk
Version: 0.2.15
Summary: Official Python client for the oruk Speech API: multilingual transcription, multilabel emotion and speaking-style scores, and unified audio analysis.
Project-URL: Homepage, https://oruk.ai
Project-URL: Documentation, https://oruk.ai/docs
Project-URL: Changelog, https://oruk.ai/changelog
Project-URL: Pricing, https://oruk.ai/pricing
Author-email: oruk labs <access@oruk.ai>
License: MIT
Keywords: audio analysis,emotion detection,oruk,paralinguistics,speech,speech emotion recognition,speech understanding,speech-to-text,transcription
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Multimedia :: Sound/Audio :: Analysis
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Requires-Dist: httpx<1,>=0.27
Provides-Extra: fast
Requires-Dist: soundfile<1,>=0.13; extra == 'fast'
Description-Content-Type: text/markdown

# oruk — Python client for the oruk Speech API

Official Python SDK for [oruk](https://oruk.ai), the speech lab building audio
models for transcription, multilabel emotion scores,
speaking-style classification, and unified audio analysis.

This SDK calls the file API. With Resonance, send a prerecorded English audio file (WAV, FLAC, MP3,
M4A, OGG, or WebM; up to 30 MB / 60 minutes), get structured results back.
Resonance is oruk’s flagship speech recognition model. Plans include audio minutes, measured by the second with a one-second minimum. The separate Realtime preview supports 32 locales and phrase-level emotion scores over WebSocket; see the [realtime reference](https://oruk.ai/docs#realtime).

## Install

```bash
python -m pip install https://oruk.ai/sdk/oruk-0.2.15-py3-none-any.whl
```

Requires Python 3.10 or newer. [Versioned wheel mirror](https://oruk.ai/sdk/oruk-0.2.15-py3-none-any.whl).

## Quickstart

Create an account at [oruk.ai](https://oruk.ai/auth/signup), choose a
[subscription plan](https://oruk.ai/pricing), and create an API key in the
developer portal. Standard self-serve plans begin with a 7-day trial
(card required, $0 today).

```python
import os
from oruk import Oruk

with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
    result = client.analyze("sample.wav", model="oruk-resonance")

print(result["text"])       # English transcript
print(result["emotions"])   # selected model scores; see interpretation below
print(result["styles"])     # selected speaking-style scores; can be empty
```

## Spectra-2

Spectra-2 transcribes 25 languages and returns 15 emotion scores and 16 speaking-style scores in one request. Use mono 16 kHz WAV, finalized FLAC, PCM16, or float32 audio from 45 ms to 60 seconds, up to 4 MiB. The [model reference](https://oruk.ai/docs/spectra-2) lists the supported languages and outputs.

```python
with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
    result = client.analyze("sample.flac", model="oruk-spectra-2")
    print(result["text"], result["emotions"], result["styles"])
```

The SDK prepares available speech sessions during ordinary requests to reduce repeat-call latency. For an application that needs to prepare before its first recording, `client.prepare_spectra2()` returns whether a session is ready. Ordinary requests remain available when it returns `False`. Explicit `request_id` values keep their existing behavior. Spectra-2 results cover the whole clip and do not include word timestamps or speaker segments.

For smaller uploads, install `soundfile>=0.13,<1` alongside this SDK and enable
`compress_audio=True` when creating the client. Eligible mono 16 kHz, 16-bit WAV
recordings are encoded as lossless FLAC only when the result is smaller. No audio
is resampled or reduced in precision. Reuse the same setting and original bytes
when retrying a request. Already encoded FLAC files pass through unchanged.

```python
with Oruk(api_key=os.environ["ORUK_API_KEY"], compress_audio=True) as client:
    client.prepare_spectra2()  # Before recording or a latency-sensitive upload.
    result = client.analyze("sample.wav", model="oruk-spectra-2")
```

Session capacity is tracked conservatively, including concurrent and uncertain
requests. A recording that exceeds the remaining capacity uses ordinary admission
directly, avoiding a predictably refused session request. The server still checks
authorization, capacity, expiry and billing on every call.

## Endpoints

| Method | API endpoint | Returns |
|---|---|---|
| `client.transcribe(file)` | `POST /v1/audio/transcriptions` | English transcript |
| `client.emotions(file)` | `POST /v1/audio/emotions` | Selected scores from 15 emotion labels, no transcription |
| `client.styles(file)` | `POST /v1/audio/styles` | Selected scores from 16 speaking-style labels, no transcription |
| `client.affect(file)` | `POST /v1/audio/affect` | emotion + style, no transcript |
| `client.analyze(file)` | `POST /v1/audio/analysis` | transcript, labels, segments, tagged text |
| `client.proficiency(file, transcript=None)` | `POST /v1/audio/proficiency` | Preview: CEFR band, 0–5 score, fluency, transcript |

Every method accepts a path, `Path`, or binary file object and an optional
`request_id=`. The first five endpoints use `oruk-resonance` by default;
`oruk-fourier` is another file model with its own output scope. Proficiency
uses `oruk-proficiency-1`, not Resonance or Fourier. A supplied proficiency
`transcript=` skips built-in transcription. See the [endpoint reference](https://oruk.ai/docs#reference)
before changing models; their capabilities are not interchangeable.
On the first five endpoints, with `model="oruk-resonance"`, pass `diarize=True`
(and optionally `num_speakers=` when you know the speaker count, 1–32) to label speakers: diarization locates the speaker turns,
then Resonance scores each speaker turn, so every segment carries a
`speaker` field. Other fields follow the endpoint: emotion-only output does
not gain a transcript. Speaker labels are local to the recording, not identities or roles. Diarization is included in plan minutes.

```python
with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
    result = client.analyze("support-call.wav", model="oruk-resonance", diarize=True)
for seg in result["segments"]:
    print(seg["speaker"], seg.get("text"), seg.get("emotions", []))
```

Emotion only: `client.emotions(...)` on Resonance runs the encoder and affect
head and never invokes the transcription decoder, so nothing is transcribed,
the result has no transcript. One audio minute uses one plan minute for either emotion-only or unified analysis; calling both separately processes the audio twice.

```python
with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
    result = client.emotions("support-call.wav", model="oruk-resonance")
print(result["emotions"])               # actual selected labels and scores
print(result.get("text"))               # None: no transcript is produced
```

The client preserves request IDs after ambiguous outcomes and across retries of HTTP 429, 500,
502, 503, and 504 with jittered backoff (two retries by default). Other HTTP
errors are not retried; network errors from httpx propagate to the caller. A first-attempt Spectra-2 session-capacity refusal can use an ordinary request within the configured retry limit, only after the server confirms that no inference was admitted.
A stable request ID helps tracing; it does not promise exactly-once processing. Errors raise `OrukAPIError`
with `status`, `code`, and `request_id` attributes.

## Runnable file workflows

Set `ORUK_API_KEY` in your environment, download the complete
[Python example](https://oruk.ai/examples/analyze-file.py), and use your own recording:

```bash
curl --fail -O https://oruk.ai/examples/analyze-file.py
python analyze-file.py sample.wav --task analysis > result.json
python analyze-file.py support-call.wav --task analysis --diarize > speakers.json
python analyze-file.py speaking-sample.wav --task proficiency > proficiency.json
```

Each command sends one logical request, with bounded HTTP retries if needed.
Running multiple commands processes the audio separately. `--help` describes
model selection, known speaker count, and an optional proficiency transcript
file. The program writes the complete API response to stdout and errors to
stderr. It does not fabricate missing labels or assume every proficiency
request was scored. For proficiency, use 30–60 seconds of spontaneous English
and inspect `check.status`; `insufficient_audio` does not establish a CEFR
level. A model estimate is not a language certificate.

## Interpreting scores and usage

Spectra-2 returns all 15 emotion and 16 speaking-style scores. For Resonance, these labels are model vocabularies,
not a guarantee that every response contains every label. Outputs are selected
by model thresholds. If no emotion clears its threshold, the highest-scoring
emotion is returned; styles can be empty. Several labels may be high and scores
need not sum to one. These thresholds do not establish calibrated probabilities
of a person's private feelings. Evaluate representative audio before choosing
application thresholds. See [label interpretation](https://oruk.ai/docs#labels)
and the [scope of the evaluations](https://oruk.ai/benchmarks/methodology).

`usage` includes measured and billable audio duration and may carry reference
fields such as `rate_per_minute_usd`, `estimated_cost_usd`, or `pricing_version`.
Those reference estimates are **not your subscription invoice**. Actual charges
follow the plan allowance, overage terms, and billing records. One minute of
audio uses one plan minute per request; separate calls process and meter the
file separately. Keep API keys in server-side code. This SDK's file methods do
not implement the separate [realtime WebSocket workflow](https://oruk.ai/guides/realtime-phrase-emotions).

## Links

- Documentation and API reference: <https://oruk.ai/docs>
- Capabilities and scope: <https://oruk.ai/capabilities>
- Pricing: <https://oruk.ai/pricing>
- Benchmarks: <https://oruk.ai/benchmarks/methodology>
- Service status: <https://oruk.ai/status>

## License

MIT


## Orukeet

Use model `oruk-orukeet` for English transcription up to 60 seconds / 4 MiB. Every subscription includes an Orukeet allowance; see [current plans and task rates](https://oruk.ai/pricing#orukeet). Optional emotion detection and speaker diarization draw from that same allowance. Extra usage shares the plan spending cap. See [the Orukeet contract](https://oruk.ai/docs#orukeet) for availability, outputs, streaming, and limits.
