Current analysis index
# Current SongBench: source-bound machine estimates

Check source hashes and versions first. Each result identifies the exact audio
and UTF-8 lyric bytes it analyzed. A title or filename is not a source identity.
Do not transfer timings between alternate recordings, edited audio, or revised
lyrics. The current index maps every active catalog placement to its result;
shared results are allowed only for identical audio and lyric hashes.
Use source.lyrics.document_url for exact UTF-8 source bytes and Unicode codepoint
spans. Older album lyric URLs may normalize whitespace or line endings; do not
apply these source spans to those rendered/normalized documents.

## Two different analysis layers

The existing GM SongBench v0.4 packets remain legacy anatomy/provenance reports.
Their RMS, waveform and spectrogram data are measured audio proxies. Their tempo
is rough autocorrelation, and their lyric-section times are proportional
line-count scaffolding, NOT measured lyric/audio alignment. Human/canon context
in those packets is context, not a new listening test or machine measurement.

The current layer contains newly executed, source-bound tempo candidates and
gated word/line placement estimates. Read each component's status and limitations:
coverage of every song does not mean every analysis succeeded or every word was
located. Failed, unsupported and abstained results remain explicitly represented.
Top-level packet status and the index's analysis_status_counts describe LYRIC
ALIGNMENT ONLY, not combined analysis success. Read tempo.status separately;
even COMPLETE_ESTIMATE lyric alignment does not imply successful tempo analysis.

Tempo execution records retain decoder_qualification=UNVERIFIED and
execution_mode=PINNED_DECODER_UNQUALIFIED. Pinned implementation hashes and
successful execution on a particular input do not establish general decoder
qualification or musical accuracy. Synthetic checks do not promote these labels.
Failed or unsupported tempo may have null input/execution provenance when no
such identity was established; nulls do not imply a decoder was executed.

## Reading estimates safely

- Tempo candidates are heuristic alternatives, not guaranteed beats or canonical
  BPM. Ranking ratios are not calibrated probabilities; half/double-time ambiguity
  can remain. A null canonical tempo must remain unassigned.
- Lyric timings are machine estimates, never human-verified by publication.
  Forced alignment and a separate unprompted transcription use the same model;
  agreement is not independent ground truth.
- Null primary start/end means unresolved. Candidate times are diagnostic only.
  Do not silently substitute candidates, interpolate missing words, or turn nulls
  into zero. Partial line bounds cover only resolved words, not the whole line.
- Raw model probabilities are uncalibrated scores, not confidence in correctness.
  COMPLETE_ESTIMATE is still an estimate; PARTIAL_ESTIMATE and ABSTAINED must stay
  visible. Neither status makes a result edit-ready.
- No automatic cuts, lyric rewrites, rankings, musical-quality verdicts or canon
  changes are authorized. Human listening and explicit human approval control edits.

JSON is the machine-readable source of truth; Markdown and this guide are mirrors.
Use the current index and its hashed per-result files. Build freshness does not
change the original analysis generation time. Historical results for another
audio hash are not current results for an edited recording.