# Replay the Resonance-2 acted-emotion diagnostic

This package contains all **546 saved Resonance-2 response receipts** from the
September 17, 2026 public API run, including every incorrect and out-of-taxonomy
answer. It reproduces **337/546 correct (61.7%)**, **170 scoring abstentions**,
and the published **58.1–65.9%** actor-bootstrap interval. All 546 requests
returned HTTP 200. A scoring abstention is not a transport failure or necessarily
the API's own `abstained` flag.

## Run offline

Download and extract [the complete archive](https://oruk.ai/research/resonance-2/reproduction-2026-09-19.zip),
then run inside the extracted directory:

```sh
python3 replay.py --check
```

Python 3.12 or newer; standard library only. No installation, API key, audio
download, model inference or network access is needed. The script verifies every
archived file's byte count and SHA-256 before importing the frozen scorer. It then
recomputes every prediction, all published metrics and all 2,000 bootstrap
replicates. A changed/missing/extra receipt, changed prediction, or changed result
fails rather than being silently omitted. Redirect stdout outside this directory
if saving the result. Hashes detect accidental changes; they are not an independent
attestation of Oruk's source data.

## What the rule actually does

`protocol.json` was recorded at 2026-09-17 21:30:22 UTC, before the first request.
Its primary adapter takes the API's threshold-selected `labels` under the `f1`
regime. The package also includes the exact 31 selection thresholds extracted
from the retained calibration artifact whose SHA-256 matches every response.
Replay independently reconstructs the saved selection: `score > 0` and
`score >= threshold`, ordered by descending score (stable calibration-label
order for exact ties). It then removes speaking-style labels using the pre-existing **15
emotion** inventory below, not the benchmark's six classes:

`happy excited hopeful sad worried angry frustrated disappointed scared disgusted surprised embarrassed proud relieved neutral`

Among those selected emotions, rank by descending score and break an exact tie by
lexical label order. Only then map `angry→anger`, `happy→happiness`, `sad→sadness`,
`fearful/scared→fear`, `disgusted→disgust`, and `surprised→surprise`. Canonical
spellings stay unchanged. If the winner is outside `anger, disgust, fear,
happiness, neutral, sadness`, it abstains. An empty selected-emotion list, API
abstention or terminal request failure also abstains. All abstentions count as
incorrect and remain in the denominator. Scores are never renormalized.

The secondary analysis was also registered before inference. It ranks **all 15
continuous emotion scores** without using threshold-selected labels, retaining
the same alias and abstention rule. It gives **346/546 (63.4%)**, with 153
abstentions. It is a sensitivity analysis, not a replacement for the primary
61.7% result. The scorer checks both against their original saved results.

Accuracy uses all 546 recordings. Fixed-taxonomy macro F1 and UAR give all six
classes equal weight; absent classes contribute zero. Confidence intervals draw
91 whole actors with replacement, preserving all six recordings per drawn actor,
for 2,000 iterations with Python `random.Random(20260817).choices`. Percentiles
use linear interpolation. The unchanged `reference_runner.py` contains the exact
metric, alias, tie-breaking and bootstrap implementation used by the original run.

## Input and model provenance

- Input manifest SHA-256:
  `960149e3b798955f4db4d2bea42f373e30d923371abb4e2dfb89e055687fd4d7`.
- CREMA-D source revision: `1658cd342dff90010aa843eaeebd53610a08b1dc`.
  All 91 actors, the IEO sentence (“It's eleven o'clock”), six intended acted
  emotions each; medium intensity except neutral XX. The manifest supplies exact
  immutable URLs, audio SHA-256, byte/sample counts, actor IDs and intended labels.
- Endpoint: `POST https://speech-api.oruk.ai/v1/audio/resonance-2?regime=f1`.
  Uploads used `clip.wav` and the model ID, without gold labels, source filenames,
  actor metadata or transcripts. Two concurrent requests, at most one retry for
  transport or 429/5xx failures; the actual retained run needed 546 HTTP 200
  requests and no retries.
- Response-declared model revision:
  `b34570175184c22066f9d3b595d52ef62281adb9a1d51c332d3874b8ff66832b`.
- Response-declared calibration revision:
  `ccd5948dc3a3cbfd628b158ccf9f95ac582da38c2e8f242128c1e465615e1c07`.

The archive includes complete saved parsed response JSON: all 31 scores, six
axes, 19 unipolar scores, selected labels, revision fields, duration, usage,
timestamps and transport metadata. Each receipt is copied byte-for-byte from the
local saved evidence. The recorded transport `response_sha256` describes the
original HTTP response bytes; those exact wire bytes were **not retained**.
Verifying the archived receipt is not a rehash of the original wire body.
No authorization headers, API keys, account records or private recordings are
included. Usage cost fields are historical estimates, not invoices or current
subscription pricing.

This offline replay does not download or rehash the WAV files. The original
`input-verification.json` records their pre-run verification; the manifest makes
separate audio verification possible. Model/calibration identities are fields
declared by the recorded response, not an independent attestation of the loaded
worker. The model weights, pre-calibration logits, calibration slopes and biases
are not in this package. Replay verifies selection from the stored continuous
scores using the extracted numeric thresholds; it does not refit them, reconstruct
scores from logits, or rerun model inference. A new call
to today's preview endpoint is a new measurement and might produce different data.

## Scope and limitations

This is a **single-sentence acted-emotion diagnostic**, not an independent held-out
leaderboard or evidence of general conversational accuracy. Training overlap has
not been audited. Targets are intended emotions, not listener consensus. The
interval is descriptive and conditional on these predictions; it does not correct
training exposure or model-selection bias. There is no matched human comparison.

`published-benchmark-results.json` is an unchanged snapshot of the article's
published data, included to check its Resonance-2 row and 546 predictions. This
package **does not replay the September 12 comparator runs**: their output
adapters differ (Hume's fixed EVI six-axis projection, Google's descriptive best
of two registered models, and transcription-plus-GPT-5.5 pipelines). Those rows
remain historical context, not freshly verified competitor inference.

The separate 31-category human-rating F1 evaluation in the article is not replayed
here. Its development holdouts were reused; prior pretraining exposure is not
certified. Private human recordings/ratings are not published in this archive.
Likewise, this package does not make the runtime article's aggregate-only latency
results independently replayable.

## Files and licenses

- `receipts/*.json`: all 546 original response receipts, no omitted failures.
- `manifest.json`, `protocol.json`, `input-verification.json`: original inputs,
  registered rules and pre-run audio verification.
- `scored-rows.json`, `secondary-scored-rows.json`, `scores.json`: unchanged
  expected predictions, metrics and intervals.
- `reference_runner.py`: unchanged original scoring implementation; replay calls
  only its offline scoring functions, not its separate download/inference commands.
- `f1-thresholds.json`: exact selection cuts, label order, source artifact hash
  and extraction pointer; no training examples or fitted classifier weights.
- `replay.py`: validation and offline adapter wrapper added September 19.
- `files.sha256.json`: byte counts and SHA-256 for all files above and this README.

CREMA-D: Cao, Cooper, Keutmann, Gur, Nenkova and Verma (2014),
[paper](https://doi.org/10.1109/TAFFC.2014.2336244),
[source](https://github.com/CheyneyComputerScience/CREMA-D/tree/1658cd342dff90010aa843eaeebd53610a08b1dc).
As with the original public diagnostic, the downloadable data are made available
under [ODbL 1.0](https://opendatacommons.org/licenses/odbl/1.0/); CREMA-D individual
contents are under [DbCL 1.0](https://opendatacommons.org/licenses/dbcl/1.0/).
No additional audio is bundled. Oruk generated the model outputs; they are not
additional human annotations. The two Python files are provided under MIT; see
`LICENSE-code.txt`.
