commit c545960ed4b45d42095e0157b5b57e252f19f7e0 equwal <truex@equwal.com> 2026-09-20 11:32:53 -0700 Match chapters by calibrated n-gram overlap instead of a fixed threshold subplz gates chapter matching on fuzz.ratio > 40. That constant is calibrated for Japanese and does not transfer: measured noise floor between unrelated chapters is 23-26 for Japanese but 45-47 for English and 44 for Spanish, so for Latin scripts the gate sits below the noise and accepts anything. The cause is alphabet size, not typography - a generic punctuation strip moved English only 45.3 to 42.4. Replace it with three things, in backend/chapters.py: - Character 4-gram overlap rather than edit distance. Separation between the true chapter and the runner-up goes from 2.1x to 9.5x, and it works on scripts with no word boundaries, where a word-level metric collapses. - A threshold calibrated from the book in hand, by sampling unrelated chapter pairs, so one rule covers a noise floor of 2.5 or of 45. - A runner-up test. A different book in the same language scores high because it shares a vocabulary; what it cannot do is make one chapter stand out. Measured: right book 78-87 at 2.7-2.9x confidence, an unrelated Russian novel 55-60 at 1.03x - rejected, though it clears subplz's threshold comfortably. Sampling is now adaptive and skips chapter 0, where publisher announcements live. A confident pair costs ~8s.
README.md | 115 ++++++++++++++---------- backend/chapters.py | 245 ++++++++++++++++++++++++++++++++++++++++++++++++++++ backend/matching.py | 145 ++++++++++++++++++++----------- 3 files changed, 408 insertions(+), 97 deletions(-)
diff --git a/README.md b/README.md index 8e1752f..df7ba2e 100644 --- a/README.md +++ b/README.md @@ -85,56 +85,79 @@ See [Will this text even match?](#will-this-text-even-match). subplz fails late and unhelpfully when the text does not match the audio: it transcribes the whole book, then reports "the generated transcript and the provided text file are too different" and writes a `.subfail`. On CPU that is a -wasted hour. +wasted hour. This app answers the question on upload, in seconds — and answers it +better than subplz can. -So the app asks the same question on upload, using subplz's own rule — -transcribe a 60-second sample from the start of up to three chapters, score each -against every chapter of the book with `rapidfuzz.fuzz.ratio`, and compare -against subplz's `SCORE_THRESHOLD = 40`. It costs a few seconds. +### Why a fixed threshold cannot work -Measured on real files: +subplz compares chapters with `rapidfuzz.fuzz.ratio` against a constant +`SCORE_THRESHOLD = 40`. What two *unrelated* chapters score depends entirely on +how many characters the script has: -| Pair | Score | Verdict | -|---|---|---| -| Moskva-Petushki audio + its own epub | **74.4** | good | -| Moskva-Petushki audio + an unrelated Russian novel | **40.5** | poor | - -That second row is the important one. An unrelated book in the same language -still clears subplz's threshold of 40, because two prose texts in one language -are roughly 40% similar character-by-character whatever they say. So the bands -here sit well above it: good at ≥60, marginal at ≥48, poor below. A poor score -warns and relabels the button "Start anyway" — it never blocks, because the -check only samples, and it is your book. - -### What actually improves the score - -Two things were tried and measured, and only one of them works. - -**Chapter structure — decisive.** subplz compares the *opening* of each audio -chapter against the *opening* of each text chapter, capped at 2000 characters. -It splits an epub into one text chapter per spine document, but a `.txt` into -exactly one chapter for the whole file. Same book, same audio: - -| Text given to subplz | Score | +| noise floor, unrelated chapters of one book | `fuzz.ratio` | |---|---| -| epub, per chapter | **69.1** | -| identical text as one flat file | **38.6** — *below the threshold* | - -A flat text file turns a perfectly good book into a failed run. That is why -`convert.py` emits a chaptered epub rather than plain text, and why the app -warns when a one-document book is paired with multi-chapter audio. - -**Typographic cleanup — no effect, so it was dropped.** Normalising curly -quotes, em dashes, soft hyphens, non-breaking spaces and footnote markers -measured **+0.0** against a real transcript. Joining paragraphs with a space -instead of subplz's empty string was worth +0.3. Both are noise next to -chapterisation, and shipping them would have been dead code. (This is the one -place it matters that `ats/lang.py` implements only Japanese and English — -every other language falls back to `English`, whose `clean()` is nothing but -`.lower()`, so punctuation is compared verbatim. It still does not move the -number.) - -Disable the check with `SUBPLZ_WEB_MATCH_CHECK=false`. +| Japanese | 23 – 26 | +| Russian | 39 | +| Spanish | 44 | +| English | 45 – 47 | + +For English and Spanish the floor is **above 40**, so unrelated chapters clear +the gate and it discriminates nothing. The constant suits Japanese, which is what +it was tuned on: thousands of distinct characters make coincidental similarity +rare, while twenty-six letters make it inevitable. Cleaning does not help — a +generic punctuation strip moved English only 45.3 → 42.4, because the cause is +alphabet size, not typography. + +### What this app does instead — `backend/chapters.py` + +**Character 4-gram overlap, not edit distance.** Measured against the same audio: + +| metric | ru floor | en floor | ja floor | true match | runner-up | separation | +|---|---|---|---|---|---|---| +| `fuzz.ratio` | 39.2 | 45.1 | 26.1 | 86.3 | 40.8 | 2.1× | +| word Jaccard | 10.4 | 17.7 | 1.9 | 63.2 | 14.8 | 4.3× | +| **4-gram Jaccard** | **4.8** | **10.6** | **2.5** | 53.3 | 5.6 | **9.5×** | + +n-grams rather than words because many languages do not separate words with +spaces — a word-level metric collapses on Japanese, where a "word" is a whole run +of characters. Set intersection is also cheaper than edit distance. + +**A threshold calibrated per book, not per release.** Before judging anything, +the app samples ~150 unrelated chapter pairs *from the book in hand* to learn what +it scores by chance, then requires a match to clear that floor. One rule that +behaves correctly whether the floor is 2.5 or 45, with no constant to re-tune. + +**A runner-up test, which is what actually catches a wrong book.** A different +book in the same language scores *high* — it shares a vocabulary. What it cannot +do is make one chapter stand out. Measured, same audio: + +| book | best | runner-up | confidence | accept_at | verdict | +|---|---|---|---|---|---| +| the right one | 78 – 87 | ~30 | **2.7 – 2.9×** | 30.7 | accepted 4/4 | +| an unrelated Russian novel | 55 – 60 | 53 – 58 | **1.0×** | 92.1 | rejected 4/4 | + +Note the second row clears subplz's threshold of 40 comfortably and is still +rejected here, on both tests independently: its own noise floor is 57.6, so +`accept_at` rises to 92, and nothing stands out from the crowd. + +**Adaptive sampling.** Chapters are sampled from across the book, skipping +chapter 0 — publisher announcements and credits live there, so it is the least +representative chapter in the book. Sampling stops as soon as one chapter matches +confidently, so a good pair typically costs ~8s and a doubtful one ~20s. + +A poor verdict warns and relabels the button "Start anyway"; it never blocks, +because the check only samples and it is your book. Disable it entirely with +`SUBPLZ_WEB_MATCH_CHECK=false`. + +### Measured and rejected + +- **Typographic normalisation** of quotes, dashes, soft hyphens, NBSP and + footnote markers: **+0.0**. Built, measured, deleted. +- **Un-gluing sentence boundaries** (`роса…Мне` → `роса… Мне`): multi-sentence + lines 13.6% → 12.4%, because pysbd does not treat `…` as a terminator for + Russian. Not worth the epub rewrite. +- **`token_set_ratio` / `token_sort_ratio`**: *worse* than `fuzz.ratio` on long + texts — nearly every common word appears in both chapters. --- diff --git a/backend/chapters.py b/backend/chapters.py new file mode 100644 index 0000000..d182600 --- /dev/null +++ b/backend/chapters.py @@ -0,0 +1,245 @@ +"""Chapter matching that works on any script. + +The alignment backend decides which text chapter belongs to which audio chapter +by character-level edit similarity (`rapidfuzz.fuzz.ratio`) against a fixed +threshold of 40. That is calibrated for Japanese and does not transfer. Measured +noise floor between *unrelated* chapters of the same book: + + Japanese 23-26 <- threshold 40 discriminates well + Russian 39 <- borderline + Spanish 44 + English 45-47 <- unrelated chapters clear the threshold + +With twenty-six letters, any two prose texts are ~45% similar by chance, so for +Latin scripts the gate is *below the noise* and accepts anything. Cleaning does +not fix it (a generic punctuation strip moved English only 45.3 -> 42.4); the +cause is alphabet size, not typography. + +This module scores on character n-gram overlap instead, and calibrates the +threshold against the book it is actually looking at. Measured on the same data: + + metric ru floor en floor ja floor true runner-up ratio + fuzz.ratio 39.2 45.1 26.1 86.3 40.8 2.1x + word jaccard 10.4 17.7 1.9 63.2 14.8 4.3x + 4-gram jaccard 4.8 10.6 2.5 53.3 5.6 9.5x + +n-grams rather than words because plenty of languages do not put spaces between +them; a word-level metric collapses on Japanese, where a "word" is a whole run of +characters. +""" + +from __future__ import annotations + +import random +import statistics as st +import unicodedata +from dataclasses import dataclass + +import regex + +NGRAM = 4 + +# Everything that is not a letter or a number: punctuation, spacing, marks. +# Dropped so typography cannot influence the comparison. +_NOISE = regex.compile(r"[^\p{L}\p{N}]+") + + +def normalize(text: str) -> str: + return _NOISE.sub("", unicodedata.normalize("NFKD", text.casefold())) + + +def fingerprint(text: str, n: int = NGRAM) -> frozenset[str]: + """The set of character n-grams in `text`, after normalisation. + + Computed once per chapter and reused: comparing two fingerprints is a set + intersection, which is far cheaper than an edit distance over the same text. + """ + s = normalize(text) + if len(s) < n: + return frozenset() + return frozenset(s[i : i + n] for i in range(len(s) - n + 1)) + + +def jaccard(a: frozenset[str], b: frozenset[str]) -> float: + if not a or not b: + return 0.0 + return 100.0 * len(a & b) / len(a | b) + + +def containment(probe: frozenset[str], reference: frozenset[str]) -> float: + """How much of `probe` appears in `reference`, as a percentage. + + Used when a short transcript sample is compared against a whole chapter: + Jaccard would punish the length difference through the union term, since a + three-minute sample cannot cover a thirty-minute chapter. Containment asks + the question we actually mean - "is this passage in that chapter" - and is + unaffected by how much longer the chapter is. + """ + if not probe or not reference: + return 0.0 + return 100.0 * len(probe & reference) / len(probe) + + +@dataclass(frozen=True) +class Calibration: + """The similarity this book produces by chance. + + Sampled from the book itself, so it adapts to script, vocabulary and + register instead of relying on a constant that only suits one language. + """ + + floor: float + high: float # worst case seen among unrelated pairs + samples: int + + # A real match must clear the floor by this factor... + FLOOR_MULTIPLE = 1.6 + # ...and beat whatever came second by this much. This is the test a fixed + # threshold cannot do, and the one that catches a wrong book whose language + # alone makes every chapter score alike. + RUNNER_UP_MULTIPLE = 1.5 + + @property + def accept_at(self) -> float: + """Minimum score worth taking seriously for this book.""" + return max(self.floor * self.FLOOR_MULTIPLE, self.high * 1.15, 1.0) + + +def calibrate( + chapter_texts: list[str], + pairs: int = 150, + seed: int = 0, + probe_chars: int = 2000, +) -> Calibration: + """Estimate the chance-similarity floor from unrelated pairs of this book. + + Calibrated with the *same* measure and the *same shape* of comparison that + real matching uses: a short probe against a whole chapter. Measuring the + floor with Jaccard while matching with containment would put the threshold + on a different scale from the scores it gates. + """ + usable = [t for t in chapter_texts if len(normalize(t)) > probe_chars // 2] + if len(usable) < 3: + # Too little to calibrate against; fall back to something conservative. + return Calibration(floor=20.0, high=40.0, samples=0) + + references = [fingerprint(t) for t in usable] + # A probe is a slice the size of a transcript sample, so the floor reflects + # what an unrelated passage of this length scores against a full chapter. + probes = [fingerprint(normalize(t)[:probe_chars]) for t in usable] + + rng = random.Random(seed) + scores = [] + for _ in range(pairs): + a, b = rng.sample(range(len(usable)), 2) + scores.append(containment(probes[a], references[b])) + + scores.sort() + return Calibration( + floor=st.median(scores), + high=scores[min(len(scores) - 1, int(0.95 * len(scores)))], + samples=len(scores), + ) + + +@dataclass(frozen=True) +class Match: + text_index: int | None + score: float + runner_up: float + accepted: bool + calibration: Calibration + + @property + def confidence(self) -> float: + """How far clear of the runner-up this match is. 1.0 means a tie.""" + if self.runner_up <= 0: + return float("inf") if self.score > 0 else 0.0 + return self.score / self.runner_up + + @property + def margin_over_noise(self) -> float: + if self.calibration.floor <= 0: + return float("inf") if self.score > 0 else 0.0 + return self.score / self.calibration.floor + + +def best_match( + probe: frozenset[str], + chapters: list[frozenset[str]], + calibration: Calibration, + exclude: set[int] | None = None, +) -> Match: + """Find the chapter a sample of audio came from. + + Accepting requires clearing the book's own noise floor *and* beating the + runner-up. The second test is what a fixed threshold cannot do: if two + chapters score alike, the winner is arbitrary, however high the number. + """ + skip = exclude or set() + ranked = sorted( + ( + (containment(probe, ch), i) + for i, ch in enumerate(chapters) + if i not in skip and ch + ), + reverse=True, + ) + if not ranked: + return Match(None, 0.0, 0.0, False, calibration) + + score, index = ranked[0] + runner_up = ranked[1][0] if len(ranked) > 1 else 0.0 + + accepted = score >= calibration.accept_at and ( + runner_up <= 0 or score >= runner_up * Calibration.RUNNER_UP_MULTIPLE + ) + return Match(index if accepted else None, score, runner_up, accepted, calibration) + + +def assign_monotonic(matrix: list[list[float]], accept_at: float) -> list[int | None]: + """Pair audio chapters to text chapters in order, maximising total score. + + The backend's own matcher walks audio chapters in order and *consumes* text + chapters as it goes, with no backtracking: an early chapter - usually + chapter 0, where publisher announcements live and which is therefore the + least representative chapter in the book - can permanently take the text + that belonged to another. + + This solves the whole assignment at once, and requires it to be + order-preserving, which is true of every book: chapter k cannot come from + text that precedes chapter k-1's. Skips are allowed on both sides, for front + matter and for chapters nobody narrated. + """ + n_audio, n_text = len(matrix), (len(matrix[0]) if matrix else 0) + if not n_audio or not n_text: + return [None] * n_audio + + NEG = float("-inf") + # best[i][j] = best total score pairing the first i audio with first j text + best = [[0.0] * (n_text + 1) for _ in range(n_audio + 1)] + back = [[0] * (n_text + 1) for _ in range(n_audio + 1)] # 0 skip-a,1 skip-t,2 pair + + for i in range(1, n_audio + 1): + for j in range(1, n_text + 1): + skip_audio = best[i - 1][j] + skip_text = best[i][j - 1] + score = matrix[i - 1][j - 1] + pair = best[i - 1][j - 1] + (score if score >= accept_at else NEG) + + chosen = max(skip_audio, skip_text, pair) + best[i][j] = chosen + back[i][j] = 2 if chosen == pair else (0 if chosen == skip_audio else 1) + + out: list[int | None] = [None] * n_audio + i, j = n_audio, n_text + while i > 0 and j > 0: + move = back[i][j] + if move == 2: + out[i - 1] = j - 1 + i, j = i - 1, j - 1 + elif move == 0: + i -= 1 + else: + j -= 1 + return out diff --git a/backend/matching.py b/backend/matching.py index 4400dd7..82b1a7c 100644 --- a/backend/matching.py +++ b/backend/matching.py @@ -25,14 +25,14 @@ from dataclasses import dataclass, field from functools import lru_cache from pathlib import Path +from . import chapters as chapters_mod from .settings import settings log = logging.getLogger(__name__) -# Scoring compares chapter openings, so a sample from the start of a chapter is -# the relevant thing to transcribe - and a short one is enough. subplz caps the -# comparison at 2000 characters, which is a couple of minutes of speech. -SAMPLE_SECONDS = 60 +# Enough speech to fingerprint a chapter confidently. Sampling is adaptive, so +# a book that matches well normally pays for one of these and stops. +SAMPLE_SECONDS = 120 MAX_SAMPLES = 3 @@ -41,11 +41,16 @@ class ChapterScore: audio_chapter: int best_score: float best_text_chapter: int | None + runner_up: float = 0.0 + # How far clear of the runner-up. 1.0 means every chapter looked alike, + # which is the signature of the wrong book rather than a poor recording. + confidence: float = 0.0 + accepted: bool = False transcript_head: str = "" @property def matched(self) -> bool: - return self.best_text_chapter is not None + return self.accepted @dataclass @@ -54,6 +59,8 @@ class MatchReport: scores: list[ChapterScore] = field(default_factory=list) text_chapters: int = 0 audio_chapters: int = 0 + accept_at: float = 0.0 + noise_floor: float = 0.0 verdict: str = "unknown" # good | marginal | poor | unknown summary: str = "" warnings: list[str] = field(default_factory=list) @@ -71,9 +78,15 @@ class MatchReport: def matched(self) -> int: return sum(1 for s in self.scores if s.matched) + @property + def confidence(self) -> float: + return max((s.confidence for s in self.scores), default=0.0) + def as_dict(self) -> dict: return { - "threshold": self.threshold, + "threshold": round(self.accept_at, 1), + "noise_floor": round(self.noise_floor, 1), + "confidence": round(self.confidence, 2), "verdict": self.verdict, "summary": self.summary, "best": round(self.best, 1), @@ -88,6 +101,8 @@ class MatchReport: { "audio_chapter": s.audio_chapter, "score": round(s.best_score, 1), + "runner_up": round(s.runner_up, 1), + "confidence": round(s.confidence, 2), "text_chapter": s.best_text_chapter, } for s in self.scores @@ -186,7 +201,7 @@ def _model(): def transcribe_sample(samples, language: str) -> str: segments, _ = _model().transcribe( - samples, language=language, beam_size=1, without_timestamps=True + samples, language=language, beam_size=5, without_timestamps=True ) return "".join(seg.text for seg in segments) @@ -222,13 +237,18 @@ def check(audio: Path, text: Path, language: str, aligner) -> MatchReport: f"more reliably." ) - # Sample from the start of chapters spread across the book, because the - # score is about how chapters open. + # Learn what this book scores by chance before judging any match + # against it. See backend/chapters.py for why a fixed threshold cannot + # work across scripts. + fingerprints = [chapters_mod.fingerprint(c) for c in chapters] + calibration = chapters_mod.calibrate(chapters) + report.accept_at = calibration.accept_at + report.noise_floor = calibration.floor + + # Sample chapters spread through the book, skipping the first: it is + # where publisher announcements and credits live, so it is the least + # representative chapter there is. picks = _spread(len(starts), MAX_SAMPLES) - scorer = ( - aligner.for_language(language) - if hasattr(aligner, "for_language") else aligner - ) for idx in picks: samples = read_samples(audio, starts[idx], SAMPLE_SECONDS) @@ -238,20 +258,23 @@ def check(audio: Path, text: Path, language: str, aligner) -> MatchReport: if len(transcript.strip()) < 40: continue - best, best_i = 0.0, None - for ci, chapter in enumerate(chapters): - score = scorer.score_pair(transcript, chapter) - if score > best: - best, best_i = score, ci - + match = chapters_mod.best_match( + chapters_mod.fingerprint(transcript), fingerprints, calibration + ) report.scores.append( ChapterScore( audio_chapter=idx, - best_score=best, - best_text_chapter=best_i if best > report.threshold else None, + best_score=match.score, + best_text_chapter=match.text_index, + runner_up=match.runner_up, + confidence=match.confidence, + accepted=match.accepted, transcript_head=transcript.strip()[:160], ) ) + # A confident match is enough; only keep sampling when unsure. + if match.accepted and match.confidence >= 2.0: + break _verdict(report) return report @@ -263,52 +286,72 @@ def check(audio: Path, text: Path, language: str, aligner) -> MatchReport: def _spread(n: int, k: int) -> list[int]: - """Up to k indices spread across range(n), always including the first.""" - if n <= k: - return list(range(n)) - step = n / k - return sorted({min(n - 1, int(i * step)) for i in range(k)}) + """Up to k chapter indices spread across the book. + + Skips chapter 0 when there is anything else to choose. It is where + publisher announcements, credits and "read by" cards live - material that + is in the audio and not in the book - so it is the least representative + chapter there is, and the worst one to judge a whole book on. + """ + if n <= 1: + return [0] + first = 1 if n > 3 else 0 + usable = n - first + if usable <= k: + return list(range(first, n)) + step = usable / k + return sorted({min(n - 1, first + int(i * step)) for i in range(k)}) def _verdict(report: MatchReport) -> None: - """Turn the raw score into advice. - - The bands sit well above the backend's own threshold, deliberately. - Measured on real files: the right book scored 74, while a completely - unrelated Russian novel still scored 40.5 - barely clearing subplz's - threshold of 40. Two prose texts in the same language are roughly 40% - similar character-by-character whatever they say, so "just over the - threshold" means "probably wrong", not "probably fine". + """Turn the measurements into advice. + + Two separate questions, and the second is the one a fixed threshold cannot + answer: + + * Is the best chapter similar enough to be a real match at all? + * Is it *distinctly* the best, or did every chapter score alike? + + A different book in the same language scores high on the first and fails + the second - its chapters all look equally plausible because they share a + language, not a story. """ if not report.scores: report.verdict = "unknown" report.summary = "Could not sample enough audio to check the match." return - t = report.threshold or 40.0 - good_at, weak_at = t * 1.5, t * 1.2 # 60 and 48 for subplz best = report.best - total = len(report.scores) - strong = sum(1 for s in report.scores if s.best_score >= good_at) + confidence = report.confidence + accepted = report.matched - if best >= good_at: + if accepted and confidence >= 2.0: report.verdict = "good" - report.summary = f"Text and audio line up ({best:.0f}/100 on the best sample)." - if total > 1 and strong < total: - report.warnings.append( - f"Only {strong} of {total} sampled chapters scored well. The " - f"timing may drift in parts of the book." - ) - elif best >= weak_at: + report.summary = ( + f"Text and audio line up - the matching chapter scores {best:.0f}, " + f"{confidence:.1f}x clear of the next best." + ) + elif accepted: report.verdict = "marginal" report.summary = ( - f"Weak match ({best:.0f}/100). This may be a different edition or " - f"an abridgement. Alignment can still work, but expect drift." + f"Probable match ({best:.0f}), but only {confidence:.1f}x clear of " + f"the next best chapter. Expect some drift." + ) + elif best >= report.accept_at: + # Similar enough, but nothing stood out. + report.verdict = "poor" + report.summary = ( + f"Every chapter of this book scores about the same ({best:.0f} vs " + f"{report.scores[0].runner_up:.0f}), so nothing actually matches." + ) + report.warnings.append( + "That pattern means a different book in the same language - the " + "words are familiar but the story is not. Check you uploaded the " + "right book and the right edition." ) else: report.verdict = "poor" report.summary = ( - f"This text does not look like this audio ({best:.0f}/100). Two " - f"unrelated books in the same language score about this well, so " - f"it is probably the wrong book, edition or abridgement." + f"This text does not look like this audio ({best:.0f}, and this " + f"book needs {report.accept_at:.0f} to count as a match)." )