CREMA-D dataset
Also called Crowd-sourced Emotional Multimodal Actors Dataset.
CREMA-D, published by Cao and colleagues in 2014, is an audiovisual corpus of acted emotional speech containing 7,442 clips from 91 actors. Each actor speaks a fixed set of twelve sentences, chosen to be emotionally neutral in wording, in six emotional states: anger, disgust, fear, happiness, sadness and neutral. Non-neutral readings were also recorded at four intensity levels, which gives the corpus a graded structure most emotion datasets lack.
Its most distinctive feature is how the labels were produced. Rather than treating the emotion the actor was asked to portray as ground truth, the creators collected ratings from more than 2,400 crowd workers, who judged clips in three conditions: audio only, video only, and audiovisual. That design yields perceptual labels — what listeners actually heard — alongside the intended category, and it makes explicit that the two often disagree. For work on how speech sounds rather than what a speaker felt, a perceptual label is the more defensible target.
The actor pool is unusually diverse for a corpus of this age, spanning a wide age range and several self-identified racial and ethnic groups. That matters because emotion models trained on narrow speaker pools tend to transfer badly, and CREMA-D at least allows the question to be asked.
We used 3,276 CREMA-D recordings to look inside four frozen speech encoders, passing the audio through each and mapping every layer to see where emotional structure emerges. Reducing those representations to three dimensions showed emotion becoming linearly separable at intermediate layers rather than at the output, which is consistent with the general finding that the most transferable speech representations sit in the middle of these networks rather than at the top.
The corpus has the limitations that come with acted speech and a fixed script. Twelve sentences repeated across 91 actors means lexical content is held constant, which is useful for isolating delivery but also means models can learn the sentences. Portrayed emotion is more exaggerated than spontaneous emotion, so scores on CREMA-D overstate performance on real recordings. And the recordings are clean studio audio. It is a good instrument for controlled comparisons and a poor proxy for deployment conditions.
Related terms
- IEMOCAP dataset
- The Interactive Emotional Dyadic Motion Capture database — twelve hours of acted dyadic conversation from ten actors, and the most widely used academic benchmark for speech emotion recognition.
- RAVDESS dataset
- The Ryerson Audio-Visual Database of Emotional Speech and Song — 24 actors, eight emotions, two fixed sentences. Widely used, and narrow enough that models tuned on it generalize poorly.
- MELD dataset
- The Multimodal EmotionLines Dataset — roughly 13,000 utterances from Friends dialogues, built so that emotion has to be read in conversational context.
What four encoders did with CREMA-D Score your own audio All terms Documentation
