Research
Notes from the speech lab
Original experiments, dated model evaluations, and explainers. Each article identifies what was tested, where the evidence ends, and how to inspect the work or try the relevant capability.
Speech models meet handwritten math
Oruk and Edixir are partnering to test speech models for spoken questions in Turkish AI math tutoring.
Read articleA voice. Many meanings.
How small differences in speech acquire social meaning. Follow /s/, released /t/, and pitch through the communities that give them meaning.
Read articleMaking Orukeet twice as fast
A speech model can run faster without losing layers. We trace a 2.18× warm-latency gain, then count the word errors on a separate seven-language test.
Read articleResonance-2: hear how it was said
Meet our next speech emotion model. Listen to six deliveries of the same sentence, inspect real API scores, and see what changed in training and evaluation.
Read articleHow much faster can the same speech model run?
An experimental Resonance 2 runtime preserves 500 greedy transcripts while reducing inference time. We compare decoder settings, cache policies and affect readouts.
Read articleWhat a smile does to a voice
You can hear a smile in a whisper. Explore how mouth shape changes speech, then follow happy labels through Fourier and Orukeet across 546 clips.
Read articleOrukeet: a new shape for speech recognition
Meet our open source speech model for 25 languages, with day-zero OpenWhispr support and runtime optimization by Hoid. Explore the filters and results.
Read articleHow a quantum memory learns what to connect
A quantum memory can use past questions to learn which qubits to connect. Some kinds of questions take more practice than others.
Read articleWe taught a fruit fly to hear human emotion
499 reconstructed fly neurons, human voices, and a trained readout. Listen through the circuit, compare six emotions, and switch neurons off yourself.
Read articleColor is shared. Meaning has an accent.
Across 132 studies, 42,266 people, and 64 countries, humans keep finding a shared color grammar. Language, climate, context, and culture give it an accent.
Read articleEverything else a voice carries
Part three of How machines hear. The words got solved. Who was speaking, how activated they were, and whether they meant it get discarded at three separate stages of a modern pipeline — each for a defensible local reason, and mostly by accident.
Read articleFour ways to hear a sentence
Part two of How machines hear. Encoder-only, encoder-decoder, transducer, decoder-only. The cleanest way to tell them apart is not their block diagrams but the clock: feed all four the same sentence and ask when each one is allowed to speak.
Read articleNobody knows where the words are
Part one of How machines hear. A second of speech is sixteen thousand numbers and nobody tells you which of them are the words. That single missing piece of information is the reason speech architectures look the way they do.
Read articleUniversal, with an accent
Fifty years of cross-cultural emotion research says vocal emotion is universal with an accent. We re-ran the experiment inside frozen speech encoders across seven languages: the same four-emotion arrangement appears in every language, probes lose accuracy at every border, and one subtraction buys up to a third of it back.
Read articleWhen the words lie: measuring sarcasm in speech
Sarcasm inverts the words. On MUStARD, transcript sentiment drops below a constant classifier on unseen shows, while acoustic scores keep most of their ground.
Read articleA score you can check
The full accounting of how our speech emotion models score: 64 systems on our own benchmark with the conflicts stated, nine commercial APIs on IEMOCAP with identical audio, and sarcasm on MUStARD — with the methodology and every system’s failure modes, including ours.
Read articleThe shape of a voice
Run a sentence through a speech model and it becomes a shape. We traced thousands of real recordings through four models in rotatable 3-D: emotions separate into threads nobody trained them to draw, four architectures agree on the layout, and Whisper rebuilds the vowel chart phoneticians drew a century ago.
Read articlePitch contours and perceived emotion
A pitch contour can shift the odds of several interpretations without belonging to any one emotion. In a pilot study, six unsupervised contour families coincided with different distributions of listener labels, but no family pointed to a single answer.
Read articleSpeech models learn physics
We examined convolution filters across six speech systems and found that training repeatedly produces Gabor-like shapes predicted by signal theory.
Read articleWhat speech encoders hear
We mapped every layer of four speech encoders and found that emotion, speaker identity, and sentence content settle at different depths.
Read articleHow to read oruk research
A guide to our public speech research: how model snapshots, datasets, metrics, and interactive experiments relate to the API you can use today.
Read article




















