Oruk research
OrukeetA new shape for
speech recognition.
An open speech model for 25 languages, built on Parakeet.
Meet Orukeet, our new open source speech recognition model. It comes with day-zero support in OpenWhispr and runtime optimization by Hoid for latency and throughput across user devices.
Inside Parakeet are thousands of small filters. Many settle into a familiar shape: a wave that fades at both ends. We built Orukeet by fitting that shape and keeping it fixed.
Orukeet builds on NVIDIA’s Parakeet TDT 0.6B v3. Half of its temporal filters are fitted Gabor functions. The rest of the model learns around them, using multilingual and multi-accent data.
Across 20,146 FLEURS recordings in 25 languages, pooled word error rate falls from Parakeet’s 11.01% to Orukeet’s 9.85%. That is a 10.6% relative reduction. Orukeet improves 23 of the 25 languages.
What changes inside the model
A Gabor combines a repeating wave with a smooth, bell-shaped envelope. The wave sets the rhythm; the envelope keeps it within a short stretch of time. Each filter gets its own fit.
Parakeet’s encoder has 24 FastConformer blocks. Attention gathers context across the recording, while small convolutions detect local patterns. We change the temporal filters inside those blocks. Each is just nine stored numbers, or taps.
Attention gathers context across the recording. Temporal convolutions detect local patterns. This is where Orukeet fixes 12,288 of the 24,576 nine-tap filters to fitted Gabor shapes.
Starting with an adapted Parakeet checkpoint, we rank all 24,576 temporal depthwise filters by how closely a Gabor fits them. We replace the closest 12,288. Some layers get more replacements than others. Then we fine-tune the remaining weights while keeping every replacement fixed.
The released model still has 627 million parameters. It stores the fitted samples as ordinary convolution weights, so inference uses the same operations as Parakeet.
The opening animation illustrates fitting with three seven-point sets. The interactive plot above shows actual nine-tap kernels from the report.
Fewer word errors across languages
We decode the same recordings with both checkpoints under matched settings. Word error rate counts replaced, missing and extra words; lower is better. Pick a language to see the scores side by side.
Word error rate · lower is better
10.6% lower WER for Orukeet
All 74 benchmark splits 61 lower WER with Orukeet
| Split | Parakeet | Orukeet |
|---|---|---|
| LibriSpeech test-clean | 1.53 | 1.46 |
| LibriSpeech test-other | 3.14 | 2.86 |
| FLEURS Bulgarian | 11.92 | 10.37 |
| FLEURS Croatian | 11.29 | 10.20 |
| FLEURS Czech | 11.12 | 8.97 |
| FLEURS Danish | 17.19 | 14.88 |
| FLEURS Dutch | 6.40 | 5.60 |
| FLEURS English | 4.28 | 3.82 |
| FLEURS Estonian | 13.32 | 10.44 |
| FLEURS Finnish | 11.14 | 9.35 |
| FLEURS French | 4.69 | 5.01 |
| FLEURS German | 4.21 | 3.92 |
| FLEURS Greek | 21.07 | 30.81 |
| FLEURS Hungarian | 13.60 | 10.68 |
| FLEURS Italian | 2.43 | 2.09 |
| FLEURS Latvian | 21.78 | 17.41 |
| FLEURS Lithuanian | 20.95 | 16.55 |
| FLEURS Maltese | 19.22 | 15.60 |
| FLEURS Polish | 6.81 | 6.11 |
| FLEURS Portuguese | 4.49 | 3.73 |
| FLEURS Romanian | 11.44 | 9.34 |
| FLEURS Russian | 4.89 | 4.72 |
| FLEURS Slovak | 9.21 | 7.75 |
| FLEURS Slovenian | 22.62 | 22.11 |
| FLEURS Spanish | 3.22 | 2.75 |
| FLEURS Swedish | 13.38 | 11.36 |
| FLEURS Ukrainian | 6.00 | 5.39 |
| EuroSpeech Bulgarian | 14.22 | 13.04 |
| EuroSpeech German | 13.40 | 11.14 |
| EuroSpeech Greek | 25.83 | 26.35 |
| EuroSpeech English | 24.40 | 23.77 |
| EuroSpeech Estonian | 34.67 | 25.33 |
| EuroSpeech Finnish | 16.61 | 15.20 |
| EuroSpeech French | 19.42 | 14.28 |
| EuroSpeech Croatian | 12.93 | 12.56 |
| EuroSpeech Italian | 10.95 | 12.32 |
| EuroSpeech Lithuanian | 38.44 | 33.10 |
| EuroSpeech Latvian | 57.18 | 42.14 |
| EuroSpeech Maltese | 36.83 | 36.15 |
| EuroSpeech Portuguese | 23.08 | 23.81 |
| EuroSpeech Slovak | 17.29 | 14.91 |
| EuroSpeech Slovenian | 48.43 | 50.23 |
| EuroSpeech Ukrainian | 13.65 | 14.25 |
| GSB AI | 8.71 | 7.98 |
| GSB Chinese accent | 14.49 | 13.56 |
| GSB Filipino accent | 13.30 | 12.79 |
| GSB Indian accent | 6.50 | 5.59 |
| GSB Japanese accent | 19.15 | 17.78 |
| GSB Scottish accent | 22.08 | 20.35 |
| GSB Singaporean accent | 13.89 | 12.86 |
| GSB agriculture | 6.20 | 5.84 |
| GSB arts | 5.47 | 4.87 |
| GSB biology | 3.67 | 3.31 |
| GSB economics | 7.05 | 6.57 |
| GSB engineering | 4.06 | 3.50 |
| GSB entertainment | 10.40 | 8.87 |
| GSB finance | 5.81 | 5.11 |
| GSB humanities | 7.98 | 7.47 |
| GSB law | 9.75 | 9.04 |
| GSB medicine | 3.49 | 3.18 |
| GSB military | 3.43 | 3.06 |
| Golos crowd · Russian | 2.84 | 2.92 |
| Golos far-field · Russian | 7.98 | 9.10 |
| Lesbos Greek | 94.78 | 93.55 |
| Monsoon India | 4.12 | 3.78 |
| NST Danish | 26.49 | 11.59 |
| NST Swedish | 16.57 | 12.36 |
| VoxPopuli Czech | 7.32 | 7.39 |
| VoxPopuli Spanish | 6.07 | 6.20 |
| VoxPopuli Hungarian | 12.00 | 11.05 |
| VoxPopuli Italian | 11.37 | 11.82 |
| VoxPopuli Dutch | 9.50 | 9.56 |
| VoxPopuli Polish | 6.48 | 6.24 |
| VoxPopuli Romanian | 11.48 | 11.20 |
The report also covers read English, regional accents and specialist vocabulary. Orukeet has lower WER on 61 of 74 measured splits, including all 20 English accent and domain splits. On the 47-split accent and domain pool, WER falls from 16.72% to 15.25%.
How the comparison was run
The report uses NeMo greedy-batch TDT with identical 16 kHz audio, FP32 weights and BF16 CUDA autocast. Its pinned normalizer handles spelling, numbers and compound boundaries. Pooled scores add edit counts and normalized reference words; language averages weight each language equally.
These are comparisons of the released checkpoints on the report’s development benchmarks. Final adaptation and selection use LibriSpeech test-other; 6,118 recordings in the accent and domain suite also enter adaptation. The technical report records the procedure and every split.
Getting the words out faster
Hoid optimized Orukeet’s native runtime, including its Metal kernels and attention cache. The installer selects Metal on Apple silicon, CUDA on a detected NVIDIA GPU, or CPU.
260.5 ms39.6 ms
Median warm recognition time on 160 clips, three timed repeats per clip. Same machine and audio; Parakeet uses four CPU threads. Model loading, file decoding and app overhead are outside the timer. Measurement record ↗
OpenWhispr brought day-zero model support. Its merged integration downloads Orukeet from Hugging Face and runs it through the existing sherpa-onnx Parakeet path, with no additional runtime dependencies. OpenWhispr uses the ONNX version; the Metal timing above comes from the native runtime.
Run Orukeet
Use Orukeet for transcription jobs, media, server workers or local applications. The release includes NeMo source weights, ONNX INT8 and native Q8/F16 formats, all from the same checkpoint.
python -m pip install https://github.com/Oruk-AI/orukeet/releases/download/v0.1.1/orukeet-0.1.1-py3-none-any.whl
orukeet install --device auto --cache ./orukeet-cache --output installation.jsonThen follow the usage guide to transcribe a recording. Code is MIT; weights and fitted kernels are CC BY-SA 4.0.