Live demo
Captions that show how it was said
Pick 30 seconds of a sample clip, paste a link to your own audio or video, or point it at a YouTube video, and the API transcribes it for real, right now. Every caption carries the full label set for that moment — each emotion scored, plus how it was delivered — so you can watch delivery shift across a scene instead of getting one label for the whole thing.
A direct link to an audio or video file — .mp4, .mov, .webm, .mp3, .m4a, .wav. Not a page that contains a video.
Loading clip…
Pick a clip or paste a link, choose a 30-second window, then run it. The audio goes to the live API in short slices, so each caption carries every label measured for that moment rather than one reading averaged across the scene.
Labels it can return
- happy
- excited
- hopeful
- proud
- relieved
- surprised
- neutral
- sad
- worried
- disappointed
- scared
- embarrassed
- angry
- frustrated
- disgusted
Plus a speaking-style score for delivery. Every label is a calibrated reading of how the speech sounds, not a claim about what the speaker meant.
His Girl Friday (1940), Columbia Pictures — public domain. Source · Rights
Or caption a YouTube video
Same API, different plumbing. The player below is YouTube’s own and stays untouched — nothing is downloaded and nothing is drawn over it. Instead your browser shares the tab’s audio with this page, the way a screen recorder does, and the captions appear in the column beside the video.
Paste a link, press play, then capture 30 seconds. Your browser shares the tab’s audio with this page — nothing is downloaded from YouTube and the player is left exactly as it is.
How this works
Real calls, not a recording
Every run hits the same public endpoint you would call, with the same models and the same pricing. Nothing here is pre-computed — re-run the same window and you are billing real inference again.
Cut into 5-second slices
The window is analysed in short slices rather than one request. That is what gives each caption its own timing and its own emotion reading; a single call over 30 seconds averages the affect into one label and loses the arc.
Nothing is downloaded
Sample clips are public domain and hosted by us; a link you paste is fetched as a plain file. For YouTube, no server of ours contacts YouTube at all — the embed plays normally and your browser shares the audio it is already producing.
Worth being precise about one limit: captions are timed to the slice, not to the word. The API returns segment-level timing, so a cue turns over every few seconds rather than highlighting each word as it lands. Emotion and speaking-style labels describe how speech sounds — they are acoustic measurements, not claims about what a speaker meant or felt.
