README.md (19397 bytes)
1 # SubPlz Web 2 3 Line an audiobook up with its ebook, sentence by sentence. 4 5  6 7 Drop in an audiobook and the ebook it was read from. You get back subtitles 8 timed to the narration, with the wording taken from your own book rather than 9 from a machine's guess at what it heard. Two things to do with that: 10 11 - **HoshiReader whispersync** — the `.srt` is the timing file. 12 - **Subtitled video** — a YouTube-ready MP4 with a selectable caption track. 13 14 Free in the browser, without limit; see [Pricing](#pricing). 15 16 --- 17 18 ## Quick start 19 20 ```powershell 21 .\run.ps1 22 ``` 23 24 Opens <http://127.0.0.1:8420>. First run creates `.venv` and installs 25 everything, including the alignment backend — several minutes, because it 26 pulls torch. 27 28 Already set up: 29 30 ```powershell 31 .venv\Scripts\python.exe -m uvicorn backend.main:app --port 8420 32 ``` 33 34 Requires **ffmpeg on PATH** and **Python 3.11** — subplz pins `>=3.10,<3.12`, 35 so 3.12+ will not work. 36 37 --- 38 39 ## About "no Whisper" 40 41 This app never runs `subplz gen`, the mode that transcribes a book from scratch. 42 Only `subplz sync` is reachable, and `gen` is not exposed anywhere in the API. 43 44 That said, `subplz sync` is itself built on Whisper, and there is no way around 45 that. It works like this: 46 47 1. A **tiny** Whisper model makes a rough transcript of the audio. 48 2. That transcript is aligned to *your* ebook text with Needleman–Wunsch. 49 3. The subtitle text that gets written is **your book's text**, not Whisper's. 50 51 So Whisper is used as a timing device, not as a source of words. Nothing it 52 mis-hears reaches the `.srt`. If "no Whisper" meant "no model downloads and no 53 transcription step at all", subplz cannot do that — you would need a different 54 aligner (aeneas, Montreal Forced Aligner, WhisperX) and a different tool. 55 56 --- 57 58 ## Input formats 59 60 **Audio** — `m4b`, `mp3`, `m4a`, `opus`, `flac`, `wav`, `mkv` and friends. 61 A single file, or **a folder of per-chapter files**: drop all 44 mp3s and they 62 are merged into one chaptered file, in natural order (`9.mp3` before `10.mp3`). 63 Files can arrive one drop at a time; the upload starts once both halves are in. 64 65 **Book** — `epub`, `txt`, `srt`, `vtt`, `ass`, `fb2`, `fb2.zip`, and the 66 Kindle formats `mobi`, `azw3`, `azw` and `prc`. In the browser every one of 67 them is read in the tab (`frontend/engine/book.js`; the Kindle formats through 68 [foliate-js](https://github.com/johnfactotum/foliate-js), fetched by 69 `tools/fetch_vendor.py`). A server job converts the same formats on upload. 70 71 Conversion on the server delegates rather than parsing ebook formats by hand 72 (`backend/convert.py`): 73 74 | From | How | 75 |------|-----| 76 | `azw3` / KF8 | the `mobi` package unpacks it straight to epub | 77 | `mobi` (older) | `mobi` unpacks to HTML, ebooklib rebuilds the epub | 78 | `fb2`, `fb2.zip` | it is XML; lxml reads it, ebooklib writes the epub | 79 80 `mobi` is the only added dependency; lxml, BeautifulSoup and ebooklib already 81 ship with subplz. 82 83 **Conversion must preserve chapters**, and that is not a stylistic preference. 84 See [Will this text even match?](#will-this-text-even-match). 85 86 --- 87 88 ## Will this text even match? 89 90 subplz fails late and unhelpfully when the text does not match the audio: it 91 transcribes the whole book, then reports "the generated transcript and the 92 provided text file are too different" and writes a `.subfail`. On CPU that is a 93 wasted hour. This app answers the question on upload, in seconds — and answers it 94 better than subplz can. 95 96 ### Why a fixed threshold cannot work 97 98 subplz compares chapters with `rapidfuzz.fuzz.ratio` against a constant 99 `SCORE_THRESHOLD = 40`. What two *unrelated* chapters score depends entirely on 100 how many characters the script has: 101 102 | noise floor, unrelated chapters of one book | `fuzz.ratio` | 103 |---|---| 104 | Japanese | 23 – 26 | 105 | Russian | 39 | 106 | Spanish | 44 | 107 | English | 45 – 47 | 108 109 For English and Spanish the floor is **above 40**, so unrelated chapters clear 110 the gate and it discriminates nothing. The constant suits Japanese, which is what 111 it was tuned on: thousands of distinct characters make coincidental similarity 112 rare, while twenty-six letters make it inevitable. Cleaning does not help — a 113 generic punctuation strip moved English only 45.3 → 42.4, because the cause is 114 alphabet size, not typography. 115 116 ### What this app does instead — `backend/chapters.py` 117 118 **Character 4-gram overlap, not edit distance.** Measured against the same audio: 119 120 | metric | ru floor | en floor | ja floor | true match | runner-up | separation | 121 |---|---|---|---|---|---|---| 122 | `fuzz.ratio` | 39.2 | 45.1 | 26.1 | 86.3 | 40.8 | 2.1× | 123 | word Jaccard | 10.4 | 17.7 | 1.9 | 63.2 | 14.8 | 4.3× | 124 | **4-gram Jaccard** | **4.8** | **10.6** | **2.5** | 53.3 | 5.6 | **9.5×** | 125 126 n-grams rather than words because many languages do not separate words with 127 spaces — a word-level metric collapses on Japanese, where a "word" is a whole run 128 of characters. Set intersection is also cheaper than edit distance. 129 130 **A threshold calibrated per book, not per release.** Before judging anything, 131 the app samples ~150 unrelated chapter pairs *from the book in hand* to learn what 132 it scores by chance, then requires a match to clear that floor. One rule that 133 behaves correctly whether the floor is 2.5 or 45, with no constant to re-tune. 134 135 **A runner-up test, which is what actually catches a wrong book.** A different 136 book in the same language scores *high* — it shares a vocabulary. What it cannot 137 do is make one chapter stand out. Measured, same audio: 138 139 | book | best | runner-up | confidence | accept_at | verdict | 140 |---|---|---|---|---|---| 141 | the right one | 78 – 87 | ~30 | **2.7 – 2.9×** | 30.7 | accepted 4/4 | 142 | an unrelated Russian novel | 55 – 60 | 53 – 58 | **1.0×** | 92.1 | rejected 4/4 | 143 144 Note the second row clears subplz's threshold of 40 comfortably and is still 145 rejected here, on both tests independently: its own noise floor is 57.6, so 146 `accept_at` rises to 92, and nothing stands out from the crowd. 147 148 **Adaptive sampling.** Chapters are sampled from across the book, skipping 149 chapter 0 — publisher announcements and credits live there, so it is the least 150 representative chapter in the book. Sampling stops as soon as one chapter matches 151 confidently, so a good pair typically costs ~8s and a doubtful one ~20s. 152 153 A poor verdict warns and relabels the button "Start anyway"; it never blocks, 154 because the check only samples and it is your book. Disable it entirely with 155 `SUBPLZ_WEB_MATCH_CHECK=false`. 156 157 ### Measured and rejected 158 159 - **Typographic normalisation** of quotes, dashes, soft hyphens, NBSP and 160 footnote markers: **+0.0**. Built, measured, deleted. 161 - **Un-gluing sentence boundaries** (`роса…Мне` → `роса… Мне`): multi-sentence 162 lines 13.6% → 12.4%, because pysbd does not treat `…` as a terminator for 163 Russian. Not worth the epub rewrite. 164 - **`token_set_ratio` / `token_sort_ratio`**: *worse* than `fuzz.ratio` on long 165 texts — nearly every common word appears in both chapters. 166 167 --- 168 169 ## Languages 170 171 97 languages. The catch upstream does not document: subplz splits sentences with 172 `pysbd`, which supports **23** languages and raises `ValueError` on anything 173 else — **Portuguese and Finnish included**. subplz can use `stanza` instead when 174 given `--nlp`, which covers 74 more. 175 176 This app resolves that per request. `backend/languages.json` records which 177 splitter each language needs, and the aligner adds `--nlp` automatically. You 178 never see the failure. 179 180 | Language | Splitter | Notes | 181 |---|---|---| 182 | Spanish, Russian, Japanese | pysbd | fast path | 183 | Portuguese, Finnish | stanza | one-off model download on first use | 184 185 subplz also defaults `--language` **and** `--lang` to Japanese. The aligner 186 always sets both explicitly, so a non-Japanese book cannot silently run as 187 Japanese. 188 189 Regenerate the registry after upgrading pysbd or stanza: 190 191 ```powershell 192 .venv\Scripts\python.exe -m backend.gen_languages 193 ``` 194 195 --- 196 197 ## How it works 198 199 ``` 200 drop files ──► POST /api/uploads stage, pair, convert, detect language 201 │ 202 ▼ (draft — correct the language here) 203 POST /api/jobs/{id}/start entitlement check, enqueue 204 │ 205 ▼ 206 queue ──► runner ──► aligner ──► .srt ──► .mp4 207 │ 208 ▼ 209 storage + metadata.json 210 ``` 211 212 Upload and start are separate calls on purpose: a wrong language guess costs a 213 click instead of a re-upload and a wasted multi-hour run. 214 215 ### Audio is prepared twice, deliberately 216 217 The aligner is fed a **16 kHz mono** copy; the video keeps the original quality. 218 That is not tidiness — handing subplz 44.1 kHz stereo crashed ctranslate2 here 219 (integer divide by zero, part-way through a chapter), reproducibly. Doing the 220 conversion ourselves also repairs damaged input: the error-tolerant ffmpeg flags 221 drop corrupt frames instead of letting a single bad chapter abort the whole run. 222 223 ### Files 224 225 | Path | Role | 226 |---|---| 227 | `backend/api.py` | HTTP routes | 228 | `backend/aligner.py` | the alignment backend, behind an interface | 229 | `backend/runner.py` | stages inputs, drives the aligner, collects artifacts | 230 | `backend/convert.py` | fb2/mobi/azw3 → a chaptered epub | 231 | `backend/matching.py` | the preflight match score | 232 | `backend/render.py` | the YouTube MP4 | 233 | `backend/detect.py` | pairs the dropped files, detects the language | 234 | `backend/languages.py` | language registry and splitter routing | 235 | `backend/queue.py` | in-process or Redis dispatch | 236 | `backend/storage.py` | local disk or S3 | 237 | `backend/billing.py` | who may do what: free in the browser, a credit for a server job | 238 | `backend/payments.py` | Stripe: checkout, idempotent fulfilment, webhooks | 239 | `backend/accounts.py` | one email, one account; folding anonymous work in | 240 | `backend/auth.py` | emailed one-time sign-in links | 241 | `backend/pricing.py` | the plan catalogue, with the market it was set against | 242 | `frontend/` | vanilla HTML/CSS/JS, no build step | 243 | `tools/client.py` | CLI client, and a worked example of the API | 244 | `worker.py` | standalone worker for the Redis backend | 245 246 ### Outputs 247 248 | Artifact | What | 249 |---|---| 250 | `<name>.<lang>.srt` | the subtitles | 251 | `<name>.<lang>.mp4` | cover + audio + soft caption track (`mov_text`) | 252 | `metadata.json` | language, model, splitter, cue count, timing span | 253 | `subplz.log` | the full run log — the only way to debug a bad alignment | 254 255 A job that runs in the browser tab can also make `<name>.read-along.epub`: the 256 epub with the narration inside it, as EPUB 3 Media Overlays 257 (`frontend/engine/epub.js`). Thorium, Storyteller and other EPUB 3 readers play 258 it and highlight each line. The cues do not say where in the pages their words 259 are, so the engine finds each cue's words again in the text of the pages, in 260 order, and puts a `<span id>` around them. An EPUB 2 book becomes EPUB 3 (a 261 navigation document is made from the NCX). The audio must be MP3 or AAC in 262 m4a/m4b, which are the types an EPUB 3 reader must play. `tests/engine/epub.test.mjs` 263 follows each overlay as a reader does, and gives the result to the W3C 264 `epubcheck` when `EPUBCHECK` points to its jar. 265 266 Uploaded media is deleted once a job succeeds. A **failed** job keeps its inputs 267 so you can fix the language and retry without re-uploading. 268 269 --- 270 271 ## Swapping the alignment backend 272 273 subplz is one implementation of `aligner.Aligner`, not a hard dependency. 274 Everything subplz-specific — its argument names, its Japanese defaults, the 275 shape of its progress output, where it writes the result — lives in 276 `SubPlzAligner`. To replace it: 277 278 1. subclass `Aligner` (`build_command`, `progress_reader`, `locate_output`) 279 2. register it in `ALIGNERS` 280 3. set `SUBPLZ_WEB_ALIGNER` to its name 281 282 The API, queue, storage and job runner do not change. 283 284 --- 285 286 ## Video 287 288 A still cover image at 1 fps, the audio, and the subtitles as a **selectable 289 track** rather than burned in — so the file stays small, the encode stays fast, 290 and the viewer can turn captions off. The cover is the largest image in the 291 epub, or a plain dark card when there is not one. 292 293 The H.264 encoder is **probed, not assumed**: `ffmpeg -encoders` lists encoders 294 that were compiled in, including hardware ones on machines with no such 295 hardware, so the app encodes one test frame with each candidate and takes the 296 first that actually works. On this machine that is `h264_amf`; a build with 297 `libx264` will prefer that. Set `SUBPLZ_WEB_VIDEO_ENCODER` to force one, or 298 `SUBPLZ_WEB_RENDER_VIDEO=false` to skip video entirely. 299 300 --- 301 302 ## CLI 303 304 ```powershell 305 .venv\Scripts\python.exe tools\client.py submit book.m4b book.epub --wait 306 .venv\Scripts\python.exe tools\client.py submit .\chapters\ book.epub --wait 307 .venv\Scripts\python.exe tools\client.py list 308 .venv\Scripts\python.exe tools\client.py download <job_id> --dir out\ 309 ``` 310 311 Use this rather than `curl` for non-ASCII filenames: curl on a non-UTF-8 console 312 mangles multipart filenames, and this client does not. 313 314 --- 315 316 ## Pricing 317 318 All of the code is public, and anyone may host it. What this site sells is the 319 use of its operator's machines, and nothing else. 320 321 **Free and paid are split by where the work is done, not by what comes out.** 322 323 | Tier | Where it runs | You get | Costs | 324 |---|---|---|---| 325 | free | the visitor's browser (or the Android app) | each output: `.srt`, `.mkv`, `.mp4`, the read-along `.epub` | nothing, without limit | 326 | cloud | this server's hardware: a large speech model on a GPU | the same, in minutes and not hours, from any device | one credit | 327 328 | Plan | Price | Per book | 329 |---|---|---| 330 | One book | $4.99 | $4.99 | 331 | 5 books | $16.99 | $3.40 | 332 | 20 books | $39.00 | $1.95 | 333 334 A book costs about $0.64 to convert on a rented GPU. Above about $5 a technical 335 buyer wraps a raw alignment API (ElevenLabs: $2.20 for a 10-hour book). Below 336 $3 the fixed card fee takes too much. There is no unlimited plan: use comes in 337 bursts, and one heavy user of such a plan costs more than the plan brings in. 338 An operator who wants a recurring plan adds one with `SUBPLZ_WEB_PLANS_JSON`. 339 Every number is an env var — see `backend/pricing.py`. 340 341 A job in the browser costs the server nothing, so it is never counted and never 342 refused. A server job takes its credit at the start. A job that fails or is 343 cancelled gets the credit back. 344 345 The page offers the server job on the confirm screen ("Convert on our server"): 346 the files are uploaded, the job runs on the server, and the result waits in the 347 list of conversions, so the visitor can close the tab. The page asks for the 348 credit before the upload, not after it. 349 350 The server does not keep a visitor's files past their use 351 (`backend/retention.py`, once an hour): uploads go when the job succeeds, or 352 after 24 hours when it did not; results go after 7 days. `/terms.html` says the 353 same to the visitor, with the numbers read from the server, and shows 354 `SUBPLZ_WEB_CONTACT_EMAIL`. 355 356 `SUBPLZ_WEB_CLOUD_ENABLED` says that fast conversion is on offer. While it is 357 false the page shows no way to buy credits, whatever `SUBPLZ_WEB_BILLING_ENABLED` 358 says: credits that buy nothing must not be for sale. 359 360 **An account is an email.** It gets attached either by following an emailed 361 one-time link (no passwords anywhere) or by paying, since Stripe collects one. 362 Whatever the device converted while anonymous is folded into the account, and 363 the cookie is re-pointed at it - including when the payment lands by webhook 364 with no browser attached (`Account.merged_into`). 365 366 **Payments are Stripe Checkout**, with prices sent inline from 367 `backend/pricing.py`, so there is nothing to configure in the Stripe dashboard 368 beyond a webhook. Two rules: the browser returning from Stripe is never treated 369 as proof of payment (the session is re-read from Stripe), and fulfilment is 370 idempotent on the checkout session id, because Stripe reports one payment 371 several times. 372 373 ### Turning payments on 374 375 ```bash 376 # on the server, as root - each prompts for the value with input hidden 377 tools/set-secret.sh SUBPLZ_WEB_STRIPE_SECRET_KEY # sk_live_... (or sk_test_...) 378 tools/set-secret.sh SUBPLZ_WEB_STRIPE_WEBHOOK_SECRET # whsec_... 379 tools/set-secret.sh SUBPLZ_WEB_SMTP_PASSWORD # for sign-in emails 380 ``` 381 382 In Stripe: Developers -> Webhooks -> add `https://<host>/api/billing/webhook` 383 with `checkout.session.completed`, `checkout.session.async_payment_succeeded`, 384 `customer.subscription.created`, `.updated` and `.deleted`. Then set 385 `SUBPLZ_WEB_BILLING_ENABLED=true`, `SUBPLZ_WEB_PUBLIC_BASE_URL`, 386 `SUBPLZ_WEB_COOKIE_SECURE=true` and the `SMTP_*` values in `.env` and restart. 387 388 Not handled: refunds and disputes (do them in the Stripe dashboard and adjust 389 `purchased_credits` by hand), and tax (Stripe Tax is one checkout parameter away 390 once you are registered somewhere). 391 392 --- 393 394 ## Releases 395 396 Every deploy is a tag (`vMAJOR.MINOR.PATCH`) with a GitHub release; `/healthz` 397 reports which one is running. Tag, push, then deploy that tag: 398 399 ```bash 400 git tag -a v2.0.1 -m "what changed" && git push origin main --tags 401 gh release create v2.0.1 --notes "what changed" 402 ``` 403 404 --- 405 406 ## Scaling out 407 408 Everything environment-specific is an env var with a localhost default. See 409 `.env.example`. 410 411 | Concern | localhost | public | 412 |---|---|---| 413 | Queue | thread pool in the API process | Redis + `worker.py` | 414 | Storage | `./data/artifacts` | S3, presigned download URLs | 415 | Database | SQLite | Postgres | 416 | Identity | cookie | cookie + emailed sign-in links (`SMTP_*`) | 417 | Billing | off, nothing locked | `SUBPLZ_WEB_BILLING_ENABLED=true` + Stripe keys | 418 419 ```bash 420 SUBPLZ_WEB_QUEUE_BACKEND=redis \ 421 SUBPLZ_WEB_REDIS_URL=redis://redis:6379/0 \ 422 SUBPLZ_WEB_DATABASE_URL=postgresql+psycopg://user:pass@host/subplz \ 423 SUBPLZ_WEB_STORAGE_BACKEND=s3 SUBPLZ_WEB_S3_BUCKET=subplz-artifacts \ 424 SUBPLZ_WEB_DEVICE=cuda \ 425 python worker.py 426 ``` 427 428 With an external queue the API copies staged uploads into shared storage before 429 enqueuing, and the worker pulls them down, so the API and workers do not need a 430 shared filesystem. 431 432 ### Before going public 433 434 - Set the Stripe keys, the webhook and SMTP (see *Turning payments on*). 435 - Put a reverse proxy in front for TLS and upload limits. 436 - Add a retention job — audiobooks are large and artifacts are kept forever. 437 - Run workers on hardware that can take it. Alignment needs roughly 2–3 GB of 438 RAM; a 1 GB VPS will OOM. A 4h40m Russian audiobook took ~45 minutes on 439 `tiny`/CPU with 15 threads here, and would take many hours on one vCPU. 440 441 --- 442 443 ## Performance 444 445 Alignment is dominated by the Whisper pass. `tiny` on CPU is the slow path; a 446 GPU with `--device cuda` is far faster. Chaptered files are processed chapter by 447 chapter, which is also what makes the progress bar meaningful — the runner 448 counts chapter completions rather than trusting the per-chapter bar, which 449 restarts at 0% for every chapter. 450 451 Single files longer than about four hours can exhaust RAM, per upstream. Prefer 452 chaptered `m4b`, or a folder of per-chapter files. 453 454 ## Licence 455 456 AGPL-3.0: see `LICENSE`. You may use, change and host this, and you must give 457 your users the source of what you host. `NOTICE` has the licences of the work 458 this is built on (SubPlz, whisper.cpp, Whisper).