# Prospective independent speech-emotion evaluation

Protocol version: September 9, 2026. **Status: design only; no independent audio
set has been registered and no result is claimed.** The downloadable September 7
CREMA-D run remains a same-source diagnostic. This protocol does not change its
inputs or promote its results to held-out evidence.

## Question and claim

Measure the current `oruk-fourier` English file emotion endpoint on licensed,
human-recorded audio excluded from the evaluated model's development. The primary
question is whether its highest returned emotion agrees with a listener's dominant
perceived vocal emotion in a fixed six-class task. This is not a test of the
speaker's inner state, continuous valence/arousal, the full 15-label multilabel
contract, transcription, or live conversation latency.

A corpus supplied by a third party is not necessarily independent of training.
An operator assertion is not independent verification. Describe each separately.

## Freeze before inference

1. Obtain human recordings with documented permission for this API evaluation.
   Record the original source, immutable revision, license/permission evidence,
   collection period, speaker/session identifiers, and permission to redistribute
   audio and annotations separately. Use original mono PCM16 16 kHz WAV, 3–12
   seconds, without clipping or denoising. This first protocol deliberately uses
   the existing runner's exact input contract; other formats need a separately
   frozen preprocessing protocol. Exclude synthetic voices and private customer
   recordings without the required permissions. Record exclusions before scoring.
2. Use at least three independent adult English-proficient listeners, blinded to
   model outputs and vendor identity. Each chooses the dominant perceived vocal
   label from anger, disgust, fear, happiness, neutral, sadness, other, or
   uncertain. Keep every anonymized vote. Two agreeing votes among the six target
   labels establish the reference; otherwise mark the clip ambiguous/out of scope.
   Report agreement and all excluded counts. These are task-specific reference
   judgments, not calibrated clinical or psychological measurements.
3. Target 600 eligible clips: 100 per class, at least 40 speakers, no more than 15
   selected clips per speaker. Iterate eligible clips in ascending SHA-256 order,
   accepting a clip only while its class and speaker caps permit. If the rule
   cannot fill all six classes or the speaker minimum, stop and revise the protocol
   before any API inference; do not substitute based on model outputs. This is a
   deliberately balanced diagnostic, not an estimate of real-world prevalence.
4. Bind the served emotion artifact to complete provenance for all model
   components and previous checkpoints. Audit exact original bytes, canonical
   decoded audio, near duplicates, source, speakers and sessions against training,
   pretraining, validation, calibration and model-selection inventories. Merely
   checking the final training split is insufficient. Record unavailable
   inventories explicitly. Reject overlap or label the narrower scope honestly.
   A newly collected corpus helps establish temporal separation, but the component
   freeze dates and collection history must support that claim.
5. Freeze the input manifest, annotations, exclusions, taxonomy, selection code,
   model attestation, scorer version and execution plan with SHA-256 hashes and an
   independently timestamped record. Complete the accompanying readiness template.
   The human reviewer must inspect supporting evidence: populated fields alone do
   not certify independence. No execution is registered until actual files exist.

## Execute once and preserve failures

Use the existing `runner.py` with a dedicated internal evaluation identity excluded
from acquisition metrics, original WAV bytes, and `oruk-fourier`. Register the
actual summed conservative duration bound and exactly one request per selected
clip. The outer limit for this proposed protocol is 600 requests / 7,200 rounded
audio seconds; this document does not authorize a purchase or a run on its own.
The runner performs no retries and preserves every error and unattempted row in
its append-only journal. Do not rerun difficult clips or tune thresholds on results.

Capture the public alias and actual revision if returned. A missing revision must
remain null. Independently record and hash the loaded emotion component before and
after the run. Preserve complete task-response JSON privately, excluding headers
and credentials; publish response fields only after checking the audio license,
privacy permissions and response contents. Keep all safe per-recording receipts,
input hashes, usage totals, API failures, timestamps, code/runtime versions and the
exact permission-qualified download instructions. Verify the usage ledger against
the request IDs; reference cost estimates are not actual invoices.

## Score and publish

Use the existing fixed top-returned-label projection and spelling map. Empty or
out-of-taxonomy winners abstain; do not discard unsupported labels before taking
the maximum. Errors and abstentions stay in the denominator. Publish accuracy,
coverage, fixed-six-label macro F1 and UAR, per-class support/precision/recall/F1,
and the complete confusion matrix. Report whole-speaker 95% bootstrap intervals
with 2,000 replicates and seed 20260817. Report the number of independent groups,
missing speaker identifiers and collection limits beside the interval.

Provide a working offline replay from licensed original audio and saved receipts.
Keep intended-acted, listener-labeled and natural conversational results separate.
Include failed runs and weak classes. Any post-hoc error analysis must be labeled
as such. Do not claim superiority over competitors without a separately frozen,
matched-input comparison of their current products under compatible contracts.

## Why the next run is still pending

The available public CREMA-D recordings have a usable published data license, but
training exclusion for the served model is not established. Other common corpora
are also not automatically new to the model. A fresh, authorized audio collection
or a verified excluded corpus, its listener labels, and model-bound development
provenance are the dependencies for an independent result. The existing replay
is usable now; this future protocol is not a substitute for that missing evidence.

Source context: [CREMA-D's original repository and license](https://github.com/CheyneyComputerScience/CREMA-D)
and [Google's guidance on useful original evidence for AI search](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide).
