Transcription · Compared September 12, 2026
Orukeet and hosted Whisper
Orukeet matches the lowest public Whisper Turbo rate we verified: $0.006 per audio hour. It adds native Oruk keys, audio streaming before an utterance ends, and optional speech tasks.
The same base rate as NagaAI
| Service | $/audio minute | $/audio hour |
|---|---|---|
| Orukeet Prepaid. Optional tasks priced separately. | 0.0001 | 0.006 |
| NagaAI · Whisper large-v3-turbo Public model price, currently shown with a 50% discount. | 0.0001 | 0.006 |
| DeepInfra · Whisper large-v3-turbo Public per-minute price. | 0.0002 | 0.012 |
These are transcription-only rates before optional tasks, taxes, or provider-specific billing rules. Prices can change. Orukeet measures actual audio duration with no audio-duration minimum, then rounds the final request charge up to one microdollar. Subscription minutes do not apply.
Measured over the public API
- Median final-audio-to-text latency
- 97 ms
- Median transcription processing, streaming run
- 20 ms
- Sustained direct REST requests, 8 concurrent
- 6.6/s
September 12, 2026. Native Oruk key, one client, one warm A100 in the US. Streaming: 60 real-time replays, p95 119 ms after the final audio frame; connection setup and recording time excluded. Throughput uses a one-minute direct REST run. REST includes upload, authorization, credit reservation, processing, and response. Base transcription only.
Streaming latency
| Timing boundary | Median | p95 | p99 |
|---|---|---|---|
| Final audio frame → final text | 97 ms | 119 ms | 198 ms |
| Server transcription processing | 20 ms | 38 ms | 42 ms |
60 utterances across six clips (1.44–11.72 seconds), sent as 20 ms mono 16 kHz PCM16 frames at real-time recording cadence. 0 errors. One connection was reused; its initial connection and ready event took 677 ms. Authorization and credit reservation run while audio arrives. These timings apply to final utterances; Orukeet does not emit partial word hypotheses.
REST throughput and latency
| Route | Concurrent | Requests/s | Audio s/s | Median | p95 | p99 | Errors |
|---|---|---|---|---|---|---|---|
| Standard | 1 | 1.2 | 11.1 | 735 ms | 1,617 ms | 2,285 ms | 0/120 |
| Standard | 2 | 2.2 | 20.2 | 772 ms | 1,784 ms | 2,825 ms | 0/120 |
| Standard | 4 | 4.0 | 36.5 | 831 ms | 2,111 ms | 2,707 ms | 0/120 |
| Standard | 8 | 5.1 | 47.0 | 1,220 ms | 3,082 ms | 4,822 ms | 0/120 |
| Direct | 1 | 1.2 | 11.3 | 592 ms | 1,893 ms | 2,370 ms | 0/120 |
| Direct | 2 | 2.7 | 24.6 | 554 ms | 1,806 ms | 2,458 ms | 0/120 |
| Direct | 4 | 5.0 | 46.3 | 623 ms | 1,581 ms | 1,858 ms | 0/120 |
| Direct | 8 | 8.1 | 75.2 | 695 ms | 2,210 ms | 2,921 ms | 0/120 |
Standard: speech-api.oruk.ai. Direct: orukeet-direct.oruk.ai. Both accept the same native key and task flags. Each row measures 120 requests across the same 12 clips (1.44–23.02 seconds; 9.24 seconds on average). “Audio s/s” is successful audio seconds processed per wall-clock second.
Closed-loop load: each client starts its next request after the previous one finishes. Preloaded WAV files, HTTP/1.1 connection reuse, eight warmup requests per route excluded from these rows and retained in the download. No retries or outlier removal. Network distance, upload bandwidth, recording length, and other traffic affect results. This measures one client and one serving region.
One-minute sustained load
Direct REST
6.6 requests/s
61.2 audio seconds per second. 408 completed requests; 0 errors. Median 909 ms, p95 2,657 ms, p99 5,768 ms.
Server transcription processing: 26 ms median.
Standard REST
7.0 requests/s
64.6 audio seconds per second. 430 completed requests; 0 errors. Median 969 ms, p95 2,147 ms, p99 2,662 ms.
Server transcription processing: 26 ms median.
Eight concurrent clients per route issuing work for 60 seconds, then finishing in-flight requests. Throughput includes that final drain time. The same 12 clips cycle throughout; eight warmup calls per route are excluded. Streaming and optional tasks ran separately. Across all runs, including warmups: 1,905 requests, 0 errors.
Optional task latency
| Tasks added | Server median | Full request median | Full request max | Errors |
|---|---|---|---|---|
| Emotion detection | 176 ms | 891 ms | 1,342 ms | 0/5 |
| Speaker diarization | 5,330 ms | 6,479 ms | 7,336 ms | 0/5 |
| Emotion detection + Speaker diarization | 5,000 ms | 5,750 ms | 6,416 ms | 0/5 |
Five sequential requests per configuration, standard REST route, one 14.61-second synthetic recording with two speakers. Server time includes transcription and the selected tasks. These small samples show observed ranges; optional tasks are excluded from the base-transcription measurements above.
Accuracy on the same English clips
On 2,741 Earnings22 clips containing 48,899 reference words, our optimized Orukeet serving build scored 10.20% word error rate. A local Whisper large-v3-turbo F16 run with beam size 5 scored 11.42%. Lower is better.
The counts are 4,990 and 5,584 word errors respectively, over 5.43 hours of audio. This is a local model comparison on one English domain, using the same reference normalization and evaluation clips. Hosted decoders may use different precision, settings, or preprocessing. We make no claim that either hosted Whisper service has this WER.
Download benchmark summary and provenanceChoosing an integration
Orukeet
For short English recordings and dictation: upload a file or stream audio before committing an utterance. Optional emotion detection and speaker diarization use task flags. Current limits are 60 seconds, 4 MiB, and eight active requests or recordings per organization.
Hosted Whisper Turbo
Consider the provider’s languages, file limits, timestamps, decoding controls, and regional behavior for your workload. NagaAI currently lists the same base rate; DeepInfra lists twice that rate. Compare end-to-end requests with your recordings before choosing on speed.
Orukeet currently serves from one US region. The measurements above use native Oruk keys, with optional task timings reported separately.