Skip to content

Explore oruk

Transcription · Compared September 12, 2026

Orukeet and hosted Whisper

Orukeet matches the lowest public Whisper Turbo rate we verified: $0.006 per audio hour. It adds native Oruk keys, audio streaming before an utterance ends, and optional speech tasks.

The same base rate as NagaAI

Published transcription prices checked September 12, 2026
Service$/audio minute$/audio hour
Orukeet

Prepaid. Optional tasks priced separately.

0.00010.006
NagaAI · Whisper large-v3-turbo

Public model price, currently shown with a 50% discount.

0.00010.006
DeepInfra · Whisper large-v3-turbo

Public per-minute price.

0.00020.012

These are transcription-only rates before optional tasks, taxes, or provider-specific billing rules. Prices can change. Orukeet measures actual audio duration with no audio-duration minimum, then rounds the final request charge up to one microdollar. Subscription minutes do not apply.

Measured over the public API

Median final-audio-to-text latency
97 ms
Median transcription processing, streaming run
20 ms
Sustained direct REST requests, 8 concurrent
6.6/s

September 12, 2026. Native Oruk key, one client, one warm A100 in the US. Streaming: 60 real-time replays, p95 119 ms after the final audio frame; connection setup and recording time excluded. Throughput uses a one-minute direct REST run. REST includes upload, authorization, credit reservation, processing, and response. Base transcription only.

Streaming latency

Native-key streaming latency in milliseconds
Timing boundaryMedianp95p99
Final audio frame → final text97 ms119 ms198 ms
Server transcription processing20 ms38 ms42 ms

60 utterances across six clips (1.4411.72 seconds), sent as 20 ms mono 16 kHz PCM16 frames at real-time recording cadence. 0 errors. One connection was reused; its initial connection and ready event took 677 ms. Authorization and credit reservation run while audio arrives. These timings apply to final utterances; Orukeet does not emit partial word hypotheses.

REST throughput and latency

120 requests at each concurrency level on each public REST route
RouteConcurrentRequests/sAudio s/sMedianp95p99Errors
Standard11.211.1735 ms1,617 ms2,285 ms0/120
Standard22.220.2772 ms1,784 ms2,825 ms0/120
Standard44.036.5831 ms2,111 ms2,707 ms0/120
Standard85.147.01,220 ms3,082 ms4,822 ms0/120
Direct11.211.3592 ms1,893 ms2,370 ms0/120
Direct22.724.6554 ms1,806 ms2,458 ms0/120
Direct45.046.3623 ms1,581 ms1,858 ms0/120
Direct88.175.2695 ms2,210 ms2,921 ms0/120

Standard: speech-api.oruk.ai. Direct: orukeet-direct.oruk.ai. Both accept the same native key and task flags. Each row measures 120 requests across the same 12 clips (1.4423.02 seconds; 9.24 seconds on average). “Audio s/s” is successful audio seconds processed per wall-clock second.

Closed-loop load: each client starts its next request after the previous one finishes. Preloaded WAV files, HTTP/1.1 connection reuse, eight warmup requests per route excluded from these rows and retained in the download. No retries or outlier removal. Network distance, upload bandwidth, recording length, and other traffic affect results. This measures one client and one serving region.

One-minute sustained load

Direct REST

6.6 requests/s

61.2 audio seconds per second. 408 completed requests; 0 errors. Median 909 ms, p95 2,657 ms, p99 5,768 ms.

Server transcription processing: 26 ms median.

Standard REST

7.0 requests/s

64.6 audio seconds per second. 430 completed requests; 0 errors. Median 969 ms, p95 2,147 ms, p99 2,662 ms.

Server transcription processing: 26 ms median.

Eight concurrent clients per route issuing work for 60 seconds, then finishing in-flight requests. Throughput includes that final drain time. The same 12 clips cycle throughout; eight warmup calls per route are excluded. Streaming and optional tasks ran separately. Across all runs, including warmups: 1,905 requests, 0 errors.

Optional task latency

Five complete REST requests per optional task configuration
Tasks addedServer medianFull request medianFull request maxErrors
Emotion detection176 ms891 ms1,342 ms0/5
Speaker diarization5,330 ms6,479 ms7,336 ms0/5
Emotion detection + Speaker diarization5,000 ms5,750 ms6,416 ms0/5

Five sequential requests per configuration, standard REST route, one 14.61-second synthetic recording with two speakers. Server time includes transcription and the selected tasks. These small samples show observed ranges; optional tasks are excluded from the base-transcription measurements above.

Accuracy on the same English clips

On 2,741 Earnings22 clips containing 48,899 reference words, our optimized Orukeet serving build scored 10.20% word error rate. A local Whisper large-v3-turbo F16 run with beam size 5 scored 11.42%. Lower is better.

The counts are 4,990 and 5,584 word errors respectively, over 5.43 hours of audio. This is a local model comparison on one English domain, using the same reference normalization and evaluation clips. Hosted decoders may use different precision, settings, or preprocessing. We make no claim that either hosted Whisper service has this WER.

Download benchmark summary and provenance

Choosing an integration

Orukeet

For short English recordings and dictation: upload a file or stream audio before committing an utterance. Optional emotion detection and speaker diarization use task flags. Current limits are 60 seconds, 4 MiB, and eight active requests or recordings per organization.

Hosted Whisper Turbo

Consider the provider’s languages, file limits, timestamps, decoding controls, and regional behavior for your workload. NagaAI currently lists the same base rate; DeepInfra lists twice that rate. Compare end-to-end requests with your recordings before choosing on speed.

Orukeet currently serves from one US region. The measurements above use native Oruk keys, with optional task timings reported separately.