Recently Written · git

subplz-web

git clone https://github.com/equwal/subplz-web

Log | Files | Refs


commit c545960ed4b45d42095e0157b5b57e252f19f7e0
equwal <truex@equwal.com>
2026-09-20 11:32:53 -0700

Match chapters by calibrated n-gram overlap instead of a fixed threshold

subplz gates chapter matching on fuzz.ratio > 40. That constant is calibrated
for Japanese and does not transfer: measured noise floor between unrelated
chapters is 23-26 for Japanese but 45-47 for English and 44 for Spanish, so for
Latin scripts the gate sits below the noise and accepts anything. The cause is
alphabet size, not typography - a generic punctuation strip moved English only
45.3 to 42.4.

Replace it with three things, in backend/chapters.py:

- Character 4-gram overlap rather than edit distance. Separation between the
  true chapter and the runner-up goes from 2.1x to 9.5x, and it works on
  scripts with no word boundaries, where a word-level metric collapses.
- A threshold calibrated from the book in hand, by sampling unrelated chapter
  pairs, so one rule covers a noise floor of 2.5 or of 45.
- A runner-up test. A different book in the same language scores high because
  it shares a vocabulary; what it cannot do is make one chapter stand out.
  Measured: right book 78-87 at 2.7-2.9x confidence, an unrelated Russian novel
  55-60 at 1.03x - rejected, though it clears subplz's threshold comfortably.

Sampling is now adaptive and skips chapter 0, where publisher announcements
live. A confident pair costs ~8s.

 README.md           | 115 ++++++++++++++----------
 backend/chapters.py | 245 ++++++++++++++++++++++++++++++++++++++++++++++++++++
 backend/matching.py | 145 ++++++++++++++++++++-----------
 3 files changed, 408 insertions(+), 97 deletions(-)
diff --git a/README.md b/README.md
index 8e1752f..df7ba2e 100644
--- a/README.md
+++ b/README.md
@@ -85,56 +85,79 @@ See [Will this text even match?](#will-this-text-even-match).
 subplz fails late and unhelpfully when the text does not match the audio: it
 transcribes the whole book, then reports "the generated transcript and the
 provided text file are too different" and writes a `.subfail`. On CPU that is a
-wasted hour.
+wasted hour. This app answers the question on upload, in seconds — and answers it
+better than subplz can.
 
-So the app asks the same question on upload, using subplz's own rule —
-transcribe a 60-second sample from the start of up to three chapters, score each
-against every chapter of the book with `rapidfuzz.fuzz.ratio`, and compare
-against subplz's `SCORE_THRESHOLD = 40`. It costs a few seconds.
+### Why a fixed threshold cannot work
 
-Measured on real files:
+subplz compares chapters with `rapidfuzz.fuzz.ratio` against a constant
+`SCORE_THRESHOLD = 40`. What two *unrelated* chapters score depends entirely on
+how many characters the script has:
 
-| Pair | Score | Verdict |
-|---|---|---|
-| Moskva-Petushki audio + its own epub | **74.4** | good |
-| Moskva-Petushki audio + an unrelated Russian novel | **40.5** | poor |
-
-That second row is the important one. An unrelated book in the same language
-still clears subplz's threshold of 40, because two prose texts in one language
-are roughly 40% similar character-by-character whatever they say. So the bands
-here sit well above it: good at ≥60, marginal at ≥48, poor below. A poor score
-warns and relabels the button "Start anyway" — it never blocks, because the
-check only samples, and it is your book.
-
-### What actually improves the score
-
-Two things were tried and measured, and only one of them works.
-
-**Chapter structure — decisive.** subplz compares the *opening* of each audio
-chapter against the *opening* of each text chapter, capped at 2000 characters.
-It splits an epub into one text chapter per spine document, but a `.txt` into
-exactly one chapter for the whole file. Same book, same audio:
-
-| Text given to subplz | Score |
+| noise floor, unrelated chapters of one book | `fuzz.ratio` |
 |---|---|
-| epub, per chapter | **69.1** |
-| identical text as one flat file | **38.6** — *below the threshold* |
-
-A flat text file turns a perfectly good book into a failed run. That is why
-`convert.py` emits a chaptered epub rather than plain text, and why the app
-warns when a one-document book is paired with multi-chapter audio.
-
-**Typographic cleanup — no effect, so it was dropped.** Normalising curly
-quotes, em dashes, soft hyphens, non-breaking spaces and footnote markers
-measured **+0.0** against a real transcript. Joining paragraphs with a space
-instead of subplz's empty string was worth +0.3. Both are noise next to
-chapterisation, and shipping them would have been dead code. (This is the one
-place it matters that `ats/lang.py` implements only Japanese and English —
-every other language falls back to `English`, whose `clean()` is nothing but
-`.lower()`, so punctuation is compared verbatim. It still does not move the
-number.)
-
-Disable the check with `SUBPLZ_WEB_MATCH_CHECK=false`.
+| Japanese | 23 – 26 |
+| Russian | 39 |
+| Spanish | 44 |
+| English | 45 – 47 |
+
+For English and Spanish the floor is **above 40**, so unrelated chapters clear
+the gate and it discriminates nothing. The constant suits Japanese, which is what
+it was tuned on: thousands of distinct characters make coincidental similarity
+rare, while twenty-six letters make it inevitable. Cleaning does not help — a
+generic punctuation strip moved English only 45.3 → 42.4, because the cause is
+alphabet size, not typography.
+
+### What this app does instead — `backend/chapters.py`
+
+**Character 4-gram overlap, not edit distance.** Measured against the same audio:
+
+| metric | ru floor | en floor | ja floor | true match | runner-up | separation |
+|---|---|---|---|---|---|---|
+| `fuzz.ratio` | 39.2 | 45.1 | 26.1 | 86.3 | 40.8 | 2.1× |
+| word Jaccard | 10.4 | 17.7 | 1.9 | 63.2 | 14.8 | 4.3× |
+| **4-gram Jaccard** | **4.8** | **10.6** | **2.5** | 53.3 | 5.6 | **9.5×** |
+
+n-grams rather than words because many languages do not separate words with
+spaces — a word-level metric collapses on Japanese, where a "word" is a whole run
+of characters. Set intersection is also cheaper than edit distance.
+
+**A threshold calibrated per book, not per release.** Before judging anything,
+the app samples ~150 unrelated chapter pairs *from the book in hand* to learn what
+it scores by chance, then requires a match to clear that floor. One rule that
+behaves correctly whether the floor is 2.5 or 45, with no constant to re-tune.
+
+**A runner-up test, which is what actually catches a wrong book.** A different
+book in the same language scores *high* — it shares a vocabulary. What it cannot
+do is make one chapter stand out. Measured, same audio:
+
+| book | best | runner-up | confidence | accept_at | verdict |
+|---|---|---|---|---|---|
+| the right one | 78 – 87 | ~30 | **2.7 – 2.9×** | 30.7 | accepted 4/4 |
+| an unrelated Russian novel | 55 – 60 | 53 – 58 | **1.0×** | 92.1 | rejected 4/4 |
+
+Note the second row clears subplz's threshold of 40 comfortably and is still
+rejected here, on both tests independently: its own noise floor is 57.6, so
+`accept_at` rises to 92, and nothing stands out from the crowd.
+
+**Adaptive sampling.** Chapters are sampled from across the book, skipping
+chapter 0 — publisher announcements and credits live there, so it is the least
+representative chapter in the book. Sampling stops as soon as one chapter matches
+confidently, so a good pair typically costs ~8s and a doubtful one ~20s.
+
+A poor verdict warns and relabels the button "Start anyway"; it never blocks,
+because the check only samples and it is your book. Disable it entirely with
+`SUBPLZ_WEB_MATCH_CHECK=false`.
+
+### Measured and rejected
+
+- **Typographic normalisation** of quotes, dashes, soft hyphens, NBSP and
+  footnote markers: **+0.0**. Built, measured, deleted.
+- **Un-gluing sentence boundaries** (`роса…Мне` → `роса… Мне`): multi-sentence
+  lines 13.6% → 12.4%, because pysbd does not treat `…` as a terminator for
+  Russian. Not worth the epub rewrite.
+- **`token_set_ratio` / `token_sort_ratio`**: *worse* than `fuzz.ratio` on long
+  texts — nearly every common word appears in both chapters.
 
 ---
 
diff --git a/backend/chapters.py b/backend/chapters.py
new file mode 100644
index 0000000..d182600
--- /dev/null
+++ b/backend/chapters.py
@@ -0,0 +1,245 @@
+"""Chapter matching that works on any script.
+
+The alignment backend decides which text chapter belongs to which audio chapter
+by character-level edit similarity (`rapidfuzz.fuzz.ratio`) against a fixed
+threshold of 40. That is calibrated for Japanese and does not transfer. Measured
+noise floor between *unrelated* chapters of the same book:
+
+    Japanese   23-26      <- threshold 40 discriminates well
+    Russian    39         <- borderline
+    Spanish    44
+    English    45-47      <- unrelated chapters clear the threshold
+
+With twenty-six letters, any two prose texts are ~45% similar by chance, so for
+Latin scripts the gate is *below the noise* and accepts anything. Cleaning does
+not fix it (a generic punctuation strip moved English only 45.3 -> 42.4); the
+cause is alphabet size, not typography.
+
+This module scores on character n-gram overlap instead, and calibrates the
+threshold against the book it is actually looking at. Measured on the same data:
+
+    metric            ru floor  en floor  ja floor   true   runner-up   ratio
+    fuzz.ratio            39.2      45.1      26.1   86.3        40.8   2.1x
+    word jaccard          10.4      17.7       1.9   63.2        14.8   4.3x
+    4-gram jaccard         4.8      10.6       2.5   53.3         5.6   9.5x
+
+n-grams rather than words because plenty of languages do not put spaces between
+them; a word-level metric collapses on Japanese, where a "word" is a whole run of
+characters.
+"""
+
+from __future__ import annotations
+
+import random
+import statistics as st
+import unicodedata
+from dataclasses import dataclass
+
+import regex
+
+NGRAM = 4
+
+# Everything that is not a letter or a number: punctuation, spacing, marks.
+# Dropped so typography cannot influence the comparison.
+_NOISE = regex.compile(r"[^\p{L}\p{N}]+")
+
+
+def normalize(text: str) -> str:
+    return _NOISE.sub("", unicodedata.normalize("NFKD", text.casefold()))
+
+
+def fingerprint(text: str, n: int = NGRAM) -> frozenset[str]:
+    """The set of character n-grams in `text`, after normalisation.
+
+    Computed once per chapter and reused: comparing two fingerprints is a set
+    intersection, which is far cheaper than an edit distance over the same text.
+    """
+    s = normalize(text)
+    if len(s) < n:
+        return frozenset()
+    return frozenset(s[i : i + n] for i in range(len(s) - n + 1))
+
+
+def jaccard(a: frozenset[str], b: frozenset[str]) -> float:
+    if not a or not b:
+        return 0.0
+    return 100.0 * len(a & b) / len(a | b)
+
+
+def containment(probe: frozenset[str], reference: frozenset[str]) -> float:
+    """How much of `probe` appears in `reference`, as a percentage.
+
+    Used when a short transcript sample is compared against a whole chapter:
+    Jaccard would punish the length difference through the union term, since a
+    three-minute sample cannot cover a thirty-minute chapter. Containment asks
+    the question we actually mean - "is this passage in that chapter" - and is
+    unaffected by how much longer the chapter is.
+    """
+    if not probe or not reference:
+        return 0.0
+    return 100.0 * len(probe & reference) / len(probe)
+
+
+@dataclass(frozen=True)
+class Calibration:
+    """The similarity this book produces by chance.
+
+    Sampled from the book itself, so it adapts to script, vocabulary and
+    register instead of relying on a constant that only suits one language.
+    """
+
+    floor: float
+    high: float  # worst case seen among unrelated pairs
+    samples: int
+
+    # A real match must clear the floor by this factor...
+    FLOOR_MULTIPLE = 1.6
+    # ...and beat whatever came second by this much. This is the test a fixed
+    # threshold cannot do, and the one that catches a wrong book whose language
+    # alone makes every chapter score alike.
+    RUNNER_UP_MULTIPLE = 1.5
+
+    @property
+    def accept_at(self) -> float:
+        """Minimum score worth taking seriously for this book."""
+        return max(self.floor * self.FLOOR_MULTIPLE, self.high * 1.15, 1.0)
+
+
+def calibrate(
+    chapter_texts: list[str],
+    pairs: int = 150,
+    seed: int = 0,
+    probe_chars: int = 2000,
+) -> Calibration:
+    """Estimate the chance-similarity floor from unrelated pairs of this book.
+
+    Calibrated with the *same* measure and the *same shape* of comparison that
+    real matching uses: a short probe against a whole chapter. Measuring the
+    floor with Jaccard while matching with containment would put the threshold
+    on a different scale from the scores it gates.
+    """
+    usable = [t for t in chapter_texts if len(normalize(t)) > probe_chars // 2]
+    if len(usable) < 3:
+        # Too little to calibrate against; fall back to something conservative.
+        return Calibration(floor=20.0, high=40.0, samples=0)
+
+    references = [fingerprint(t) for t in usable]
+    # A probe is a slice the size of a transcript sample, so the floor reflects
+    # what an unrelated passage of this length scores against a full chapter.
+    probes = [fingerprint(normalize(t)[:probe_chars]) for t in usable]
+
+    rng = random.Random(seed)
+    scores = []
+    for _ in range(pairs):
+        a, b = rng.sample(range(len(usable)), 2)
+        scores.append(containment(probes[a], references[b]))
+
+    scores.sort()
+    return Calibration(
+        floor=st.median(scores),
+        high=scores[min(len(scores) - 1, int(0.95 * len(scores)))],
+        samples=len(scores),
+    )
+
+
+@dataclass(frozen=True)
+class Match:
+    text_index: int | None
+    score: float
+    runner_up: float
+    accepted: bool
+    calibration: Calibration
+
+    @property
+    def confidence(self) -> float:
+        """How far clear of the runner-up this match is. 1.0 means a tie."""
+        if self.runner_up <= 0:
+            return float("inf") if self.score > 0 else 0.0
+        return self.score / self.runner_up
+
+    @property
+    def margin_over_noise(self) -> float:
+        if self.calibration.floor <= 0:
+            return float("inf") if self.score > 0 else 0.0
+        return self.score / self.calibration.floor
+
+
+def best_match(
+    probe: frozenset[str],
+    chapters: list[frozenset[str]],
+    calibration: Calibration,
+    exclude: set[int] | None = None,
+) -> Match:
+    """Find the chapter a sample of audio came from.
+
+    Accepting requires clearing the book's own noise floor *and* beating the
+    runner-up. The second test is what a fixed threshold cannot do: if two
+    chapters score alike, the winner is arbitrary, however high the number.
+    """
+    skip = exclude or set()
+    ranked = sorted(
+        (
+            (containment(probe, ch), i)
+            for i, ch in enumerate(chapters)
+            if i not in skip and ch
+        ),
+        reverse=True,
+    )
+    if not ranked:
+        return Match(None, 0.0, 0.0, False, calibration)
+
+    score, index = ranked[0]
+    runner_up = ranked[1][0] if len(ranked) > 1 else 0.0
+
+    accepted = score >= calibration.accept_at and (
+        runner_up <= 0 or score >= runner_up * Calibration.RUNNER_UP_MULTIPLE
+    )
+    return Match(index if accepted else None, score, runner_up, accepted, calibration)
+
+
+def assign_monotonic(matrix: list[list[float]], accept_at: float) -> list[int | None]:
+    """Pair audio chapters to text chapters in order, maximising total score.
+
+    The backend's own matcher walks audio chapters in order and *consumes* text
+    chapters as it goes, with no backtracking: an early chapter - usually
+    chapter 0, where publisher announcements live and which is therefore the
+    least representative chapter in the book - can permanently take the text
+    that belonged to another.
+
+    This solves the whole assignment at once, and requires it to be
+    order-preserving, which is true of every book: chapter k cannot come from
+    text that precedes chapter k-1's. Skips are allowed on both sides, for front
+    matter and for chapters nobody narrated.
+    """
+    n_audio, n_text = len(matrix), (len(matrix[0]) if matrix else 0)
+    if not n_audio or not n_text:
+        return [None] * n_audio
+
+    NEG = float("-inf")
+    # best[i][j] = best total score pairing the first i audio with first j text
+    best = [[0.0] * (n_text + 1) for _ in range(n_audio + 1)]
+    back = [[0] * (n_text + 1) for _ in range(n_audio + 1)]  # 0 skip-a,1 skip-t,2 pair
+
+    for i in range(1, n_audio + 1):
+        for j in range(1, n_text + 1):
+            skip_audio = best[i - 1][j]
+            skip_text = best[i][j - 1]
+            score = matrix[i - 1][j - 1]
+            pair = best[i - 1][j - 1] + (score if score >= accept_at else NEG)
+
+            chosen = max(skip_audio, skip_text, pair)
+            best[i][j] = chosen
+            back[i][j] = 2 if chosen == pair else (0 if chosen == skip_audio else 1)
+
+    out: list[int | None] = [None] * n_audio
+    i, j = n_audio, n_text
+    while i > 0 and j > 0:
+        move = back[i][j]
+        if move == 2:
+            out[i - 1] = j - 1
+            i, j = i - 1, j - 1
+        elif move == 0:
+            i -= 1
+        else:
+            j -= 1
+    return out
diff --git a/backend/matching.py b/backend/matching.py
index 4400dd7..82b1a7c 100644
--- a/backend/matching.py
+++ b/backend/matching.py
@@ -25,14 +25,14 @@ from dataclasses import dataclass, field
 from functools import lru_cache
 from pathlib import Path
 
+from . import chapters as chapters_mod
 from .settings import settings
 
 log = logging.getLogger(__name__)
 
-# Scoring compares chapter openings, so a sample from the start of a chapter is
-# the relevant thing to transcribe - and a short one is enough. subplz caps the
-# comparison at 2000 characters, which is a couple of minutes of speech.
-SAMPLE_SECONDS = 60
+# Enough speech to fingerprint a chapter confidently. Sampling is adaptive, so
+# a book that matches well normally pays for one of these and stops.
+SAMPLE_SECONDS = 120
 MAX_SAMPLES = 3
 
 
@@ -41,11 +41,16 @@ class ChapterScore:
     audio_chapter: int
     best_score: float
     best_text_chapter: int | None
+    runner_up: float = 0.0
+    # How far clear of the runner-up. 1.0 means every chapter looked alike,
+    # which is the signature of the wrong book rather than a poor recording.
+    confidence: float = 0.0
+    accepted: bool = False
     transcript_head: str = ""
 
     @property
     def matched(self) -> bool:
-        return self.best_text_chapter is not None
+        return self.accepted
 
 
 @dataclass
@@ -54,6 +59,8 @@ class MatchReport:
     scores: list[ChapterScore] = field(default_factory=list)
     text_chapters: int = 0
     audio_chapters: int = 0
+    accept_at: float = 0.0
+    noise_floor: float = 0.0
     verdict: str = "unknown"  # good | marginal | poor | unknown
     summary: str = ""
     warnings: list[str] = field(default_factory=list)
@@ -71,9 +78,15 @@ class MatchReport:
     def matched(self) -> int:
         return sum(1 for s in self.scores if s.matched)
 
+    @property
+    def confidence(self) -> float:
+        return max((s.confidence for s in self.scores), default=0.0)
+
     def as_dict(self) -> dict:
         return {
-            "threshold": self.threshold,
+            "threshold": round(self.accept_at, 1),
+            "noise_floor": round(self.noise_floor, 1),
+            "confidence": round(self.confidence, 2),
             "verdict": self.verdict,
             "summary": self.summary,
             "best": round(self.best, 1),
@@ -88,6 +101,8 @@ class MatchReport:
                 {
                     "audio_chapter": s.audio_chapter,
                     "score": round(s.best_score, 1),
+                    "runner_up": round(s.runner_up, 1),
+                    "confidence": round(s.confidence, 2),
                     "text_chapter": s.best_text_chapter,
                 }
                 for s in self.scores
@@ -186,7 +201,7 @@ def _model():
 
 def transcribe_sample(samples, language: str) -> str:
     segments, _ = _model().transcribe(
-        samples, language=language, beam_size=1, without_timestamps=True
+        samples, language=language, beam_size=5, without_timestamps=True
     )
     return "".join(seg.text for seg in segments)
 
@@ -222,13 +237,18 @@ def check(audio: Path, text: Path, language: str, aligner) -> MatchReport:
                 f"more reliably."
             )
 
-        # Sample from the start of chapters spread across the book, because the
-        # score is about how chapters open.
+        # Learn what this book scores by chance before judging any match
+        # against it. See backend/chapters.py for why a fixed threshold cannot
+        # work across scripts.
+        fingerprints = [chapters_mod.fingerprint(c) for c in chapters]
+        calibration = chapters_mod.calibrate(chapters)
+        report.accept_at = calibration.accept_at
+        report.noise_floor = calibration.floor
+
+        # Sample chapters spread through the book, skipping the first: it is
+        # where publisher announcements and credits live, so it is the least
+        # representative chapter there is.
         picks = _spread(len(starts), MAX_SAMPLES)
-        scorer = (
-            aligner.for_language(language)
-            if hasattr(aligner, "for_language") else aligner
-        )
 
         for idx in picks:
             samples = read_samples(audio, starts[idx], SAMPLE_SECONDS)
@@ -238,20 +258,23 @@ def check(audio: Path, text: Path, language: str, aligner) -> MatchReport:
             if len(transcript.strip()) < 40:
                 continue
 
-            best, best_i = 0.0, None
-            for ci, chapter in enumerate(chapters):
-                score = scorer.score_pair(transcript, chapter)
-                if score > best:
-                    best, best_i = score, ci
-
+            match = chapters_mod.best_match(
+                chapters_mod.fingerprint(transcript), fingerprints, calibration
+            )
             report.scores.append(
                 ChapterScore(
                     audio_chapter=idx,
-                    best_score=best,
-                    best_text_chapter=best_i if best > report.threshold else None,
+                    best_score=match.score,
+                    best_text_chapter=match.text_index,
+                    runner_up=match.runner_up,
+                    confidence=match.confidence,
+                    accepted=match.accepted,
                     transcript_head=transcript.strip()[:160],
                 )
             )
+            # A confident match is enough; only keep sampling when unsure.
+            if match.accepted and match.confidence >= 2.0:
+                break
 
         _verdict(report)
         return report
@@ -263,52 +286,72 @@ def check(audio: Path, text: Path, language: str, aligner) -> MatchReport:
 
 
 def _spread(n: int, k: int) -> list[int]:
-    """Up to k indices spread across range(n), always including the first."""
-    if n <= k:
-        return list(range(n))
-    step = n / k
-    return sorted({min(n - 1, int(i * step)) for i in range(k)})
+    """Up to k chapter indices spread across the book.
+
+    Skips chapter 0 when there is anything else to choose. It is where
+    publisher announcements, credits and "read by" cards live - material that
+    is in the audio and not in the book - so it is the least representative
+    chapter there is, and the worst one to judge a whole book on.
+    """
+    if n <= 1:
+        return [0]
+    first = 1 if n > 3 else 0
+    usable = n - first
+    if usable <= k:
+        return list(range(first, n))
+    step = usable / k
+    return sorted({min(n - 1, first + int(i * step)) for i in range(k)})
 
 
 def _verdict(report: MatchReport) -> None:
-    """Turn the raw score into advice.
-
-    The bands sit well above the backend's own threshold, deliberately.
-    Measured on real files: the right book scored 74, while a completely
-    unrelated Russian novel still scored 40.5 - barely clearing subplz's
-    threshold of 40. Two prose texts in the same language are roughly 40%
-    similar character-by-character whatever they say, so "just over the
-    threshold" means "probably wrong", not "probably fine".
+    """Turn the measurements into advice.
+
+    Two separate questions, and the second is the one a fixed threshold cannot
+    answer:
+
+    * Is the best chapter similar enough to be a real match at all?
+    * Is it *distinctly* the best, or did every chapter score alike?
+
+    A different book in the same language scores high on the first and fails
+    the second - its chapters all look equally plausible because they share a
+    language, not a story.
     """
     if not report.scores:
         report.verdict = "unknown"
         report.summary = "Could not sample enough audio to check the match."
         return
 
-    t = report.threshold or 40.0
-    good_at, weak_at = t * 1.5, t * 1.2  # 60 and 48 for subplz
     best = report.best
-    total = len(report.scores)
-    strong = sum(1 for s in report.scores if s.best_score >= good_at)
+    confidence = report.confidence
+    accepted = report.matched
 
-    if best >= good_at:
+    if accepted and confidence >= 2.0:
         report.verdict = "good"
-        report.summary = f"Text and audio line up ({best:.0f}/100 on the best sample)."
-        if total > 1 and strong < total:
-            report.warnings.append(
-                f"Only {strong} of {total} sampled chapters scored well. The "
-                f"timing may drift in parts of the book."
-            )
-    elif best >= weak_at:
+        report.summary = (
+            f"Text and audio line up - the matching chapter scores {best:.0f}, "
+            f"{confidence:.1f}x clear of the next best."
+        )
+    elif accepted:
         report.verdict = "marginal"
         report.summary = (
-            f"Weak match ({best:.0f}/100). This may be a different edition or "
-            f"an abridgement. Alignment can still work, but expect drift."
+            f"Probable match ({best:.0f}), but only {confidence:.1f}x clear of "
+            f"the next best chapter. Expect some drift."
+        )
+    elif best >= report.accept_at:
+        # Similar enough, but nothing stood out.
+        report.verdict = "poor"
+        report.summary = (
+            f"Every chapter of this book scores about the same ({best:.0f} vs "
+            f"{report.scores[0].runner_up:.0f}), so nothing actually matches."
+        )
+        report.warnings.append(
+            "That pattern means a different book in the same language - the "
+            "words are familiar but the story is not. Check you uploaded the "
+            "right book and the right edition."
         )
     else:
         report.verdict = "poor"
         report.summary = (
-            f"This text does not look like this audio ({best:.0f}/100). Two "
-            f"unrelated books in the same language score about this well, so "
-            f"it is probably the wrong book, edition or abridgement."
+            f"This text does not look like this audio ({best:.0f}, and this "
+            f"book needs {report.accept_at:.0f} to count as a match)."
         )