Speech emotion
Seven-class emotion, 64 systems
- Sample
- 64,384 held-out clips; closed/API and audio-LLM systems on a fixed 5,000-clip stratified subsample
- Metric
- Accuracy across the shared seven-class mapping; macro F1 in the downloadable report
- Protocol
- Open models use the full 64,384 clips; closed/API and audio-LLM systems use the fixed 5,000-clip subset. Every row uses the same seven-class mapping and scorer within its assigned sample; refusals and API errors count as neutral predictions.
- Exclusions and scope
- Audio Flamingo 3 is excluded: it returned unparseable output on every clip in two environment configurations.
- Confidence intervals
- Not reported in this release because per-example bootstrap artifacts are not available for every compared system.
