IEMOCAP dataset
Also called Interactive Emotional Dyadic Motion Capture database.
IEMOCAP is an audiovisual corpus of emotional speech recorded at the University of Southern California SAIL lab and released by Busso and colleagues in 2008. It contains roughly twelve hours of dyadic conversation between ten actors, arranged as five sessions of two speakers each. Half the material follows scripts chosen to elicit particular emotions and half is improvised from hypothetical scenarios, which gives the corpus a mix of controlled and comparatively natural delivery. Markers on the face, head and hands were motion-captured alongside the audio and video, which is why the database carries motion capture in its name.
Utterances carry both categorical labels — anger, happiness, sadness, neutrality, frustration, excitement, fear, surprise, disgust and other — and dimensional ratings for valence, activation and dominance. Multiple annotators labelled each utterance, and agreement is uneven across categories, so most published work keeps only the utterances where annotators agreed.
Nearly every paper reporting IEMOCAP numbers uses the same four-class convention: anger, sadness, neutral, and happiness with excitement merged into it. The merge exists because happy and excited are acoustically similar and annotators separated them inconsistently. This convention matters when reading results, because four-class IEMOCAP accuracy is not comparable to numbers from any other class set, and papers occasionally report a different mapping without saying so.
The corpus has well-known limitations. The speech is acted rather than spontaneous, so emotional displays are clearer than they are in real recordings. Ten speakers is a small pool, which makes speaker-independent evaluation fragile and rewards models that memorise voices — reporting a random split rather than a leave-one-session-out split inflates results substantially. And the recordings are studio-clean, so performance on IEMOCAP says little about behaviour on telephone audio.
What makes IEMOCAP useful to us is precisely that we had no hand in building it. We sent identical audio from the test split — 2,008 clips — through nine commercial speech APIs and scored every response with one mapping to negative, neutral and positive, against a chance rate of 33 percent. oruk Spectra reached 71.5 percent on valence polarity; the next best commercial system reached 63.6. The products that read only the transcript clustered near chance, which is the expected result when the signal being measured is in the delivery.
That run also showed how uneven vendor evaluation is in practice. Four vendors could not complete the run on their free tiers, one returned no sentiment for 33 clips, and one has no neutral class at all — which is why our published table carries per-system sample sizes and confidence intervals rather than bare point estimates, including one four-class result where a competitor sits above us on the 92 clips its tier allowed.
Related terms
- CREMA-D dataset
- The Crowd-sourced Emotional Multimodal Actors Dataset — 7,442 acted clips from 91 actors, notable for its demographic diversity and its crowd-sourced perceptual ratings.
- RAVDESS dataset
- The Ryerson Audio-Visual Database of Emotional Speech and Song — 24 actors, eight emotions, two fixed sentences. Widely used, and narrow enough that models tuned on it generalize poorly.
- MELD dataset
- The Multimodal EmotionLines Dataset — roughly 13,000 utterances from Friends dialogues, built so that emotion has to be read in conversational context.
Nine commercial APIs on IEMOCAP Speech emotion APIs, measured All terms Documentation
