Skip to content
oruk

Oruk research

OrukeetA new shape for
speech recognition.

An open speech model for 25 languages, built on Parakeet.

Meet Orukeet, our new open source speech recognition model. It comes with day-zero support in OpenWhispr and runtime optimization by Hoid for latency and throughput across user devices.

Inside Parakeet are thousands of small filters. Many settle into a familiar shape: a wave that fades at both ends. We built Orukeet by fitting that shape and keeping it fixed.

Orukeet builds on NVIDIA’s Parakeet TDT 0.6B v3. Half of its temporal filters are fitted Gabor functions. The rest of the model learns around them, using multilingual and multi-accent data.

Across 20,146 FLEURS recordings in 25 languages, pooled word error rate falls from Parakeet’s 11.01% to Orukeet’s 9.85%. That is a 10.6% relative reduction. Orukeet improves 23 of the 25 languages.

What changes inside the model

A Gabor combines a repeating wave with a smooth, bell-shaped envelope. The wave sets the rhythm; the envelope keeps it within a short stretch of time. Each filter gets its own fit.

Parakeet’s encoder has 24 FastConformer blocks. Attention gathers context across the recording, while small convolutions detect local patterns. We change the temporal filters inside those blocks. Each is just nine stored numbers, or taps.

Inside Orukeet627M parameters

Attention gathers context across the recording. Temporal convolutions detect local patterns. This is where Orukeet fixes 12,288 of the 24,576 nine-tap filters to fitted Gabor shapes.

One filter, before and afterMeasured weights · nine taps
−4Tap index · 0+4
Layer 3, channel 119. Relative RMS fit error: 2.74%. Both curves use the original kernel’s norm. The vertical scale includes the full fitted curve and stays fixed as the points move. These are the report’s four predetermined examples.

Starting with an adapted Parakeet checkpoint, we rank all 24,576 temporal depthwise filters by how closely a Gabor fits them. We replace the closest 12,288. Some layers get more replacements than others. Then we fine-tune the remaining weights while keeping every replacement fixed.

The released model still has 627 million parameters. It stores the fitted samples as ordinary convolution weights, so inference uses the same operations as Parakeet.

The opening animation illustrates fitting with three seven-point sets. The interactive plot above shows actual nine-tap kernels from the report.

Fewer word errors across languages

We decode the same recordings with both checkpoints under matched settings. Word error rate counts replaced, missing and extra words; lower is better. Pick a language to see the scores side by side.

Parakeet → Orukeet20,146 recordings

Word error rate · lower is better

Parakeet
11.01%
Orukeet
9.85%
0%15%

10.6% lower WER for Orukeet

Same recordings, matched NeMo decoding. Pooled WER adds word errors and reference words across recordings. Scoring and full results ↗
All 74 benchmark splits 61 lower WER with Orukeet
Download CSV ↓
Word error rate (%) · lower is better · 74 splits
SplitParakeetOrukeet
LibriSpeech test-clean1.531.46
LibriSpeech test-other3.142.86
FLEURS Bulgarian11.9210.37
FLEURS Croatian11.2910.20
FLEURS Czech11.128.97
FLEURS Danish17.1914.88
FLEURS Dutch6.405.60
FLEURS English4.283.82
FLEURS Estonian13.3210.44
FLEURS Finnish11.149.35
FLEURS French4.695.01
FLEURS German4.213.92
FLEURS Greek21.0730.81
FLEURS Hungarian13.6010.68
FLEURS Italian2.432.09
FLEURS Latvian21.7817.41
FLEURS Lithuanian20.9516.55
FLEURS Maltese19.2215.60
FLEURS Polish6.816.11
FLEURS Portuguese4.493.73
FLEURS Romanian11.449.34
FLEURS Russian4.894.72
FLEURS Slovak9.217.75
FLEURS Slovenian22.6222.11
FLEURS Spanish3.222.75
FLEURS Swedish13.3811.36
FLEURS Ukrainian6.005.39
EuroSpeech Bulgarian14.2213.04
EuroSpeech German13.4011.14
EuroSpeech Greek25.8326.35
EuroSpeech English24.4023.77
EuroSpeech Estonian34.6725.33
EuroSpeech Finnish16.6115.20
EuroSpeech French19.4214.28
EuroSpeech Croatian12.9312.56
EuroSpeech Italian10.9512.32
EuroSpeech Lithuanian38.4433.10
EuroSpeech Latvian57.1842.14
EuroSpeech Maltese36.8336.15
EuroSpeech Portuguese23.0823.81
EuroSpeech Slovak17.2914.91
EuroSpeech Slovenian48.4350.23
EuroSpeech Ukrainian13.6514.25
GSB AI8.717.98
GSB Chinese accent14.4913.56
GSB Filipino accent13.3012.79
GSB Indian accent6.505.59
GSB Japanese accent19.1517.78
GSB Scottish accent22.0820.35
GSB Singaporean accent13.8912.86
GSB agriculture6.205.84
GSB arts5.474.87
GSB biology3.673.31
GSB economics7.056.57
GSB engineering4.063.50
GSB entertainment10.408.87
GSB finance5.815.11
GSB humanities7.987.47
GSB law9.759.04
GSB medicine3.493.18
GSB military3.433.06
Golos crowd · Russian2.842.92
Golos far-field · Russian7.989.10
Lesbos Greek94.7893.55
Monsoon India4.123.78
NST Danish26.4911.59
NST Swedish16.5712.36
VoxPopuli Czech7.327.39
VoxPopuli Spanish6.076.20
VoxPopuli Hungarian12.0011.05
VoxPopuli Italian11.3711.82
VoxPopuli Dutch9.509.56
VoxPopuli Polish6.486.24
VoxPopuli Romanian11.4811.20

The report also covers read English, regional accents and specialist vocabulary. Orukeet has lower WER on 61 of 74 measured splits, including all 20 English accent and domain splits. On the 47-split accent and domain pool, WER falls from 16.72% to 15.25%.

How the comparison was run

The report uses NeMo greedy-batch TDT with identical 16 kHz audio, FP32 weights and BF16 CUDA autocast. Its pinned normalizer handles spelling, numbers and compound boundaries. Pooled scores add edit counts and normalized reference words; language averages weight each language equally.

These are comparisons of the released checkpoints on the report’s development benchmarks. Final adaptation and selection use LibriSpeech test-other; 6,118 recordings in the accent and domain suite also enter adaptation. The technical report records the procedure and every split.

Getting the words out faster

Hoid optimized Orukeet’s native runtime, including its Metal kernels and attention cache. The installer selects Metal on Apple silicon, CUDA on a detected NVIDIA GPU, or CPU.

Warm transcription · Apple M5 Max

260.5 ms39.6 ms

Parakeet · ONNX INT8 · CPUOrukeet · native Q8 · Metal
6.6×faster in this comparison

Median warm recognition time on 160 clips, three timed repeats per clip. Same machine and audio; Parakeet uses four CPU threads. Model loading, file decoding and app overhead are outside the timer. Measurement record ↗

OpenWhispr brought day-zero model support. Its merged integration downloads Orukeet from Hugging Face and runs it through the existing sherpa-onnx Parakeet path, with no additional runtime dependencies. OpenWhispr uses the ONNX version; the Metal timing above comes from the native runtime.

Run Orukeet

Use Orukeet for transcription jobs, media, server workers or local applications. The release includes NeMo source weights, ONNX INT8 and native Q8/F16 formats, all from the same checkpoint.

python -m pip install https://github.com/Oruk-AI/orukeet/releases/download/v0.1.1/orukeet-0.1.1-py3-none-any.whl
orukeet install --device auto --cache ./orukeet-cache --output installation.json

Then follow the usage guide to transcribe a recording. Code is MIT; weights and fitted kernels are CC BY-SA 4.0.