Skip to content

Live demo

Captions that show how it was said

Pick 30 seconds of a clip and the API transcribes it for real, right now. The captions carry an emotion reading for each moment, so you can watch delivery shift across a scene instead of getting one label for the whole thing.

Loading clip…

0:060:36
0:001:15

Pick a clip and a 30-second window, then run it. The audio goes to the live API in short slices, so each caption carries the emotion measured for that moment rather than one label averaged across the scene.

Labels it can return

  • happy
  • excited
  • hopeful
  • proud
  • relieved
  • surprised
  • neutral
  • sad
  • worried
  • disappointed
  • scared
  • embarrassed
  • angry
  • frustrated
  • disgusted

Plus a speaking-style score for delivery. Every label is a calibrated reading of how the speech sounds, not a claim about what the speaker meant.

His Girl Friday (1940), Columbia Pictures — public domain. Source · Rights

How this works

Real calls, not a recording

Every run hits the same public endpoint you would call, with the same models and the same pricing. Nothing here is pre-computed — re-run the same window and you are billing real inference again.

Cut into 5-second slices

The window is analysed in short slices rather than one request. That is what gives each caption its own timing and its own emotion reading; a single call over 30 seconds averages the affect into one label and loses the arc.

Clips we can legally use

Everything in the library is public domain and hosted by us, which is also why the captions can be drawn over the picture. Embedded players from video platforms forbid overlays.

Worth being precise about one limit: captions are timed to the slice, not to the word. The API returns segment-level timing, so a cue turns over every few seconds rather than highlighting each word as it lands. Emotion and speaking-style labels describe how speech sounds — they are acoustic measurements, not claims about what a speaker meant or felt.