Recently Written · git

subplz-web

git clone https://github.com/equwal/subplz-web

Log | Files | Refs


README.md (19397 bytes)

1 # SubPlz Web
2 
3 Line an audiobook up with its ebook, sentence by sentence.
4 
5 ![subread.space: drop an audiobook and its ebook, get subtitles, videos and a read-along book](docs/screenshots/home.png)
6 
7 Drop in an audiobook and the ebook it was read from. You get back subtitles
8 timed to the narration, with the wording taken from your own book rather than
9 from a machine's guess at what it heard. Two things to do with that:
10 
11 - **HoshiReader whispersync** — the `.srt` is the timing file.
12 - **Subtitled video** — a YouTube-ready MP4 with a selectable caption track.
13 
14 Free in the browser, without limit; see [Pricing](#pricing).
15 
16 ---
17 
18 ## Quick start
19 
20 ```powershell
21 .\run.ps1
22 ```
23 
24 Opens <http://127.0.0.1:8420>. First run creates `.venv` and installs
25 everything, including the alignment backend — several minutes, because it
26 pulls torch.
27 
28 Already set up:
29 
30 ```powershell
31 .venv\Scripts\python.exe -m uvicorn backend.main:app --port 8420
32 ```
33 
34 Requires **ffmpeg on PATH** and **Python 3.11** — subplz pins `>=3.10,<3.12`,
35 so 3.12+ will not work.
36 
37 ---
38 
39 ## About "no Whisper"
40 
41 This app never runs `subplz gen`, the mode that transcribes a book from scratch.
42 Only `subplz sync` is reachable, and `gen` is not exposed anywhere in the API.
43 
44 That said, `subplz sync` is itself built on Whisper, and there is no way around
45 that. It works like this:
46 
47 1. A **tiny** Whisper model makes a rough transcript of the audio.
48 2. That transcript is aligned to *your* ebook text with Needleman–Wunsch.
49 3. The subtitle text that gets written is **your book's text**, not Whisper's.
50 
51 So Whisper is used as a timing device, not as a source of words. Nothing it
52 mis-hears reaches the `.srt`. If "no Whisper" meant "no model downloads and no
53 transcription step at all", subplz cannot do that — you would need a different
54 aligner (aeneas, Montreal Forced Aligner, WhisperX) and a different tool.
55 
56 ---
57 
58 ## Input formats
59 
60 **Audio** — `m4b`, `mp3`, `m4a`, `opus`, `flac`, `wav`, `mkv` and friends.
61 A single file, or **a folder of per-chapter files**: drop all 44 mp3s and they
62 are merged into one chaptered file, in natural order (`9.mp3` before `10.mp3`).
63 Files can arrive one drop at a time; the upload starts once both halves are in.
64 
65 **Book** — `epub`, `txt`, `srt`, `vtt`, `ass`, `fb2`, `fb2.zip`, and the
66 Kindle formats `mobi`, `azw3`, `azw` and `prc`. In the browser every one of
67 them is read in the tab (`frontend/engine/book.js`; the Kindle formats through
68 [foliate-js](https://github.com/johnfactotum/foliate-js), fetched by
69 `tools/fetch_vendor.py`). A server job converts the same formats on upload.
70 
71 Conversion on the server delegates rather than parsing ebook formats by hand
72 (`backend/convert.py`):
73 
74 | From | How |
75 |------|-----|
76 | `azw3` / KF8 | the `mobi` package unpacks it straight to epub |
77 | `mobi` (older) | `mobi` unpacks to HTML, ebooklib rebuilds the epub |
78 | `fb2`, `fb2.zip` | it is XML; lxml reads it, ebooklib writes the epub |
79 
80 `mobi` is the only added dependency; lxml, BeautifulSoup and ebooklib already
81 ship with subplz.
82 
83 **Conversion must preserve chapters**, and that is not a stylistic preference.
84 See [Will this text even match?](#will-this-text-even-match).
85 
86 ---
87 
88 ## Will this text even match?
89 
90 subplz fails late and unhelpfully when the text does not match the audio: it
91 transcribes the whole book, then reports "the generated transcript and the
92 provided text file are too different" and writes a `.subfail`. On CPU that is a
93 wasted hour. This app answers the question on upload, in seconds — and answers it
94 better than subplz can.
95 
96 ### Why a fixed threshold cannot work
97 
98 subplz compares chapters with `rapidfuzz.fuzz.ratio` against a constant
99 `SCORE_THRESHOLD = 40`. What two *unrelated* chapters score depends entirely on
100 how many characters the script has:
101 
102 | noise floor, unrelated chapters of one book | `fuzz.ratio` |
103 |---|---|
104 | Japanese | 23 – 26 |
105 | Russian | 39 |
106 | Spanish | 44 |
107 | English | 45 – 47 |
108 
109 For English and Spanish the floor is **above 40**, so unrelated chapters clear
110 the gate and it discriminates nothing. The constant suits Japanese, which is what
111 it was tuned on: thousands of distinct characters make coincidental similarity
112 rare, while twenty-six letters make it inevitable. Cleaning does not help — a
113 generic punctuation strip moved English only 45.3 → 42.4, because the cause is
114 alphabet size, not typography.
115 
116 ### What this app does instead — `backend/chapters.py`
117 
118 **Character 4-gram overlap, not edit distance.** Measured against the same audio:
119 
120 | metric | ru floor | en floor | ja floor | true match | runner-up | separation |
121 |---|---|---|---|---|---|---|
122 | `fuzz.ratio` | 39.2 | 45.1 | 26.1 | 86.3 | 40.8 | 2.1× |
123 | word Jaccard | 10.4 | 17.7 | 1.9 | 63.2 | 14.8 | 4.3× |
124 | **4-gram Jaccard** | **4.8** | **10.6** | **2.5** | 53.3 | 5.6 | **9.5×** |
125 
126 n-grams rather than words because many languages do not separate words with
127 spaces — a word-level metric collapses on Japanese, where a "word" is a whole run
128 of characters. Set intersection is also cheaper than edit distance.
129 
130 **A threshold calibrated per book, not per release.** Before judging anything,
131 the app samples ~150 unrelated chapter pairs *from the book in hand* to learn what
132 it scores by chance, then requires a match to clear that floor. One rule that
133 behaves correctly whether the floor is 2.5 or 45, with no constant to re-tune.
134 
135 **A runner-up test, which is what actually catches a wrong book.** A different
136 book in the same language scores *high* — it shares a vocabulary. What it cannot
137 do is make one chapter stand out. Measured, same audio:
138 
139 | book | best | runner-up | confidence | accept_at | verdict |
140 |---|---|---|---|---|---|
141 | the right one | 78 – 87 | ~30 | **2.7 – 2.9×** | 30.7 | accepted 4/4 |
142 | an unrelated Russian novel | 55 – 60 | 53 – 58 | **1.0×** | 92.1 | rejected 4/4 |
143 
144 Note the second row clears subplz's threshold of 40 comfortably and is still
145 rejected here, on both tests independently: its own noise floor is 57.6, so
146 `accept_at` rises to 92, and nothing stands out from the crowd.
147 
148 **Adaptive sampling.** Chapters are sampled from across the book, skipping
149 chapter 0 — publisher announcements and credits live there, so it is the least
150 representative chapter in the book. Sampling stops as soon as one chapter matches
151 confidently, so a good pair typically costs ~8s and a doubtful one ~20s.
152 
153 A poor verdict warns and relabels the button "Start anyway"; it never blocks,
154 because the check only samples and it is your book. Disable it entirely with
155 `SUBPLZ_WEB_MATCH_CHECK=false`.
156 
157 ### Measured and rejected
158 
159 - **Typographic normalisation** of quotes, dashes, soft hyphens, NBSP and
160   footnote markers: **+0.0**. Built, measured, deleted.
161 - **Un-gluing sentence boundaries** (`роса…Мне` → `роса… Мне`): multi-sentence
162   lines 13.6% → 12.4%, because pysbd does not treat `…` as a terminator for
163   Russian. Not worth the epub rewrite.
164 - **`token_set_ratio` / `token_sort_ratio`**: *worse* than `fuzz.ratio` on long
165   texts — nearly every common word appears in both chapters.
166 
167 ---
168 
169 ## Languages
170 
171 97 languages. The catch upstream does not document: subplz splits sentences with
172 `pysbd`, which supports **23** languages and raises `ValueError` on anything
173 else — **Portuguese and Finnish included**. subplz can use `stanza` instead when
174 given `--nlp`, which covers 74 more.
175 
176 This app resolves that per request. `backend/languages.json` records which
177 splitter each language needs, and the aligner adds `--nlp` automatically. You
178 never see the failure.
179 
180 | Language | Splitter | Notes |
181 |---|---|---|
182 | Spanish, Russian, Japanese | pysbd | fast path |
183 | Portuguese, Finnish | stanza | one-off model download on first use |
184 
185 subplz also defaults `--language` **and** `--lang` to Japanese. The aligner
186 always sets both explicitly, so a non-Japanese book cannot silently run as
187 Japanese.
188 
189 Regenerate the registry after upgrading pysbd or stanza:
190 
191 ```powershell
192 .venv\Scripts\python.exe -m backend.gen_languages
193 ```
194 
195 ---
196 
197 ## How it works
198 
199 ```
200 drop files ──► POST /api/uploads        stage, pair, convert, detect language
201                      │
202                      ▼                  (draft — correct the language here)
203                POST /api/jobs/{id}/start    entitlement check, enqueue
204                      │
205                      ▼
206                queue ──► runner ──► aligner ──► .srt ──► .mp4
207                                         │
208                                         ▼
209                                    storage + metadata.json
210 ```
211 
212 Upload and start are separate calls on purpose: a wrong language guess costs a
213 click instead of a re-upload and a wasted multi-hour run.
214 
215 ### Audio is prepared twice, deliberately
216 
217 The aligner is fed a **16 kHz mono** copy; the video keeps the original quality.
218 That is not tidiness — handing subplz 44.1 kHz stereo crashed ctranslate2 here
219 (integer divide by zero, part-way through a chapter), reproducibly. Doing the
220 conversion ourselves also repairs damaged input: the error-tolerant ffmpeg flags
221 drop corrupt frames instead of letting a single bad chapter abort the whole run.
222 
223 ### Files
224 
225 | Path | Role |
226 |---|---|
227 | `backend/api.py` | HTTP routes |
228 | `backend/aligner.py` | the alignment backend, behind an interface |
229 | `backend/runner.py` | stages inputs, drives the aligner, collects artifacts |
230 | `backend/convert.py` | fb2/mobi/azw3 → a chaptered epub |
231 | `backend/matching.py` | the preflight match score |
232 | `backend/render.py` | the YouTube MP4 |
233 | `backend/detect.py` | pairs the dropped files, detects the language |
234 | `backend/languages.py` | language registry and splitter routing |
235 | `backend/queue.py` | in-process or Redis dispatch |
236 | `backend/storage.py` | local disk or S3 |
237 | `backend/billing.py` | who may do what: free in the browser, a credit for a server job |
238 | `backend/payments.py` | Stripe: checkout, idempotent fulfilment, webhooks |
239 | `backend/accounts.py` | one email, one account; folding anonymous work in |
240 | `backend/auth.py` | emailed one-time sign-in links |
241 | `backend/pricing.py` | the plan catalogue, with the market it was set against |
242 | `frontend/` | vanilla HTML/CSS/JS, no build step |
243 | `tools/client.py` | CLI client, and a worked example of the API |
244 | `worker.py` | standalone worker for the Redis backend |
245 
246 ### Outputs
247 
248 | Artifact | What |
249 |---|---|
250 | `<name>.<lang>.srt` | the subtitles |
251 | `<name>.<lang>.mp4` | cover + audio + soft caption track (`mov_text`) |
252 | `metadata.json` | language, model, splitter, cue count, timing span |
253 | `subplz.log` | the full run log — the only way to debug a bad alignment |
254 
255 A job that runs in the browser tab can also make `<name>.read-along.epub`: the
256 epub with the narration inside it, as EPUB 3 Media Overlays
257 (`frontend/engine/epub.js`). Thorium, Storyteller and other EPUB 3 readers play
258 it and highlight each line. The cues do not say where in the pages their words
259 are, so the engine finds each cue's words again in the text of the pages, in
260 order, and puts a `<span id>` around them. An EPUB 2 book becomes EPUB 3 (a
261 navigation document is made from the NCX). The audio must be MP3 or AAC in
262 m4a/m4b, which are the types an EPUB 3 reader must play. `tests/engine/epub.test.mjs`
263 follows each overlay as a reader does, and gives the result to the W3C
264 `epubcheck` when `EPUBCHECK` points to its jar.
265 
266 Uploaded media is deleted once a job succeeds. A **failed** job keeps its inputs
267 so you can fix the language and retry without re-uploading.
268 
269 ---
270 
271 ## Swapping the alignment backend
272 
273 subplz is one implementation of `aligner.Aligner`, not a hard dependency.
274 Everything subplz-specific — its argument names, its Japanese defaults, the
275 shape of its progress output, where it writes the result — lives in
276 `SubPlzAligner`. To replace it:
277 
278 1. subclass `Aligner` (`build_command`, `progress_reader`, `locate_output`)
279 2. register it in `ALIGNERS`
280 3. set `SUBPLZ_WEB_ALIGNER` to its name
281 
282 The API, queue, storage and job runner do not change.
283 
284 ---
285 
286 ## Video
287 
288 A still cover image at 1 fps, the audio, and the subtitles as a **selectable
289 track** rather than burned in — so the file stays small, the encode stays fast,
290 and the viewer can turn captions off. The cover is the largest image in the
291 epub, or a plain dark card when there is not one.
292 
293 The H.264 encoder is **probed, not assumed**: `ffmpeg -encoders` lists encoders
294 that were compiled in, including hardware ones on machines with no such
295 hardware, so the app encodes one test frame with each candidate and takes the
296 first that actually works. On this machine that is `h264_amf`; a build with
297 `libx264` will prefer that. Set `SUBPLZ_WEB_VIDEO_ENCODER` to force one, or
298 `SUBPLZ_WEB_RENDER_VIDEO=false` to skip video entirely.
299 
300 ---
301 
302 ## CLI
303 
304 ```powershell
305 .venv\Scripts\python.exe tools\client.py submit book.m4b book.epub --wait
306 .venv\Scripts\python.exe tools\client.py submit .\chapters\ book.epub --wait
307 .venv\Scripts\python.exe tools\client.py list
308 .venv\Scripts\python.exe tools\client.py download <job_id> --dir out\
309 ```
310 
311 Use this rather than `curl` for non-ASCII filenames: curl on a non-UTF-8 console
312 mangles multipart filenames, and this client does not.
313 
314 ---
315 
316 ## Pricing
317 
318 All of the code is public, and anyone may host it. What this site sells is the
319 use of its operator's machines, and nothing else.
320 
321 **Free and paid are split by where the work is done, not by what comes out.**
322 
323 | Tier | Where it runs | You get | Costs |
324 |---|---|---|---|
325 | free | the visitor's browser (or the Android app) | each output: `.srt`, `.mkv`, `.mp4`, the read-along `.epub` | nothing, without limit |
326 | cloud | this server's hardware: a large speech model on a GPU | the same, in minutes and not hours, from any device | one credit |
327 
328 | Plan | Price | Per book |
329 |---|---|---|
330 | One book | $4.99 | $4.99 |
331 | 5 books | $16.99 | $3.40 |
332 | 20 books | $39.00 | $1.95 |
333 
334 A book costs about $0.64 to convert on a rented GPU. Above about $5 a technical
335 buyer wraps a raw alignment API (ElevenLabs: $2.20 for a 10-hour book). Below
336 $3 the fixed card fee takes too much. There is no unlimited plan: use comes in
337 bursts, and one heavy user of such a plan costs more than the plan brings in.
338 An operator who wants a recurring plan adds one with `SUBPLZ_WEB_PLANS_JSON`.
339 Every number is an env var — see `backend/pricing.py`.
340 
341 A job in the browser costs the server nothing, so it is never counted and never
342 refused. A server job takes its credit at the start. A job that fails or is
343 cancelled gets the credit back.
344 
345 The page offers the server job on the confirm screen ("Convert on our server"):
346 the files are uploaded, the job runs on the server, and the result waits in the
347 list of conversions, so the visitor can close the tab. The page asks for the
348 credit before the upload, not after it.
349 
350 The server does not keep a visitor's files past their use
351 (`backend/retention.py`, once an hour): uploads go when the job succeeds, or
352 after 24 hours when it did not; results go after 7 days. `/terms.html` says the
353 same to the visitor, with the numbers read from the server, and shows
354 `SUBPLZ_WEB_CONTACT_EMAIL`.
355 
356 `SUBPLZ_WEB_CLOUD_ENABLED` says that fast conversion is on offer. While it is
357 false the page shows no way to buy credits, whatever `SUBPLZ_WEB_BILLING_ENABLED`
358 says: credits that buy nothing must not be for sale.
359 
360 **An account is an email.** It gets attached either by following an emailed
361 one-time link (no passwords anywhere) or by paying, since Stripe collects one.
362 Whatever the device converted while anonymous is folded into the account, and
363 the cookie is re-pointed at it - including when the payment lands by webhook
364 with no browser attached (`Account.merged_into`).
365 
366 **Payments are Stripe Checkout**, with prices sent inline from
367 `backend/pricing.py`, so there is nothing to configure in the Stripe dashboard
368 beyond a webhook. Two rules: the browser returning from Stripe is never treated
369 as proof of payment (the session is re-read from Stripe), and fulfilment is
370 idempotent on the checkout session id, because Stripe reports one payment
371 several times.
372 
373 ### Turning payments on
374 
375 ```bash
376 # on the server, as root - each prompts for the value with input hidden
377 tools/set-secret.sh SUBPLZ_WEB_STRIPE_SECRET_KEY       # sk_live_... (or sk_test_...)
378 tools/set-secret.sh SUBPLZ_WEB_STRIPE_WEBHOOK_SECRET   # whsec_...
379 tools/set-secret.sh SUBPLZ_WEB_SMTP_PASSWORD           # for sign-in emails
380 ```
381 
382 In Stripe: Developers -> Webhooks -> add `https://<host>/api/billing/webhook`
383 with `checkout.session.completed`, `checkout.session.async_payment_succeeded`,
384 `customer.subscription.created`, `.updated` and `.deleted`. Then set
385 `SUBPLZ_WEB_BILLING_ENABLED=true`, `SUBPLZ_WEB_PUBLIC_BASE_URL`,
386 `SUBPLZ_WEB_COOKIE_SECURE=true` and the `SMTP_*` values in `.env` and restart.
387 
388 Not handled: refunds and disputes (do them in the Stripe dashboard and adjust
389 `purchased_credits` by hand), and tax (Stripe Tax is one checkout parameter away
390 once you are registered somewhere).
391 
392 ---
393 
394 ## Releases
395 
396 Every deploy is a tag (`vMAJOR.MINOR.PATCH`) with a GitHub release; `/healthz`
397 reports which one is running. Tag, push, then deploy that tag:
398 
399 ```bash
400 git tag -a v2.0.1 -m "what changed" && git push origin main --tags
401 gh release create v2.0.1 --notes "what changed"
402 ```
403 
404 ---
405 
406 ## Scaling out
407 
408 Everything environment-specific is an env var with a localhost default. See
409 `.env.example`.
410 
411 | Concern | localhost | public |
412 |---|---|---|
413 | Queue | thread pool in the API process | Redis + `worker.py` |
414 | Storage | `./data/artifacts` | S3, presigned download URLs |
415 | Database | SQLite | Postgres |
416 | Identity | cookie | cookie + emailed sign-in links (`SMTP_*`) |
417 | Billing | off, nothing locked | `SUBPLZ_WEB_BILLING_ENABLED=true` + Stripe keys |
418 
419 ```bash
420 SUBPLZ_WEB_QUEUE_BACKEND=redis \
421 SUBPLZ_WEB_REDIS_URL=redis://redis:6379/0 \
422 SUBPLZ_WEB_DATABASE_URL=postgresql+psycopg://user:pass@host/subplz \
423 SUBPLZ_WEB_STORAGE_BACKEND=s3 SUBPLZ_WEB_S3_BUCKET=subplz-artifacts \
424 SUBPLZ_WEB_DEVICE=cuda \
425 python worker.py
426 ```
427 
428 With an external queue the API copies staged uploads into shared storage before
429 enqueuing, and the worker pulls them down, so the API and workers do not need a
430 shared filesystem.
431 
432 ### Before going public
433 
434 - Set the Stripe keys, the webhook and SMTP (see *Turning payments on*).
435 - Put a reverse proxy in front for TLS and upload limits.
436 - Add a retention job — audiobooks are large and artifacts are kept forever.
437 - Run workers on hardware that can take it. Alignment needs roughly 2–3 GB of
438   RAM; a 1 GB VPS will OOM. A 4h40m Russian audiobook took ~45 minutes on
439   `tiny`/CPU with 15 threads here, and would take many hours on one vCPU.
440 
441 ---
442 
443 ## Performance
444 
445 Alignment is dominated by the Whisper pass. `tiny` on CPU is the slow path; a
446 GPU with `--device cuda` is far faster. Chaptered files are processed chapter by
447 chapter, which is also what makes the progress bar meaningful — the runner
448 counts chapter completions rather than trusting the per-chapter bar, which
449 restarts at 0% for every chapter.
450 
451 Single files longer than about four hours can exhaust RAM, per upstream. Prefer
452 chaptered `m4b`, or a folder of per-chapter files.
453 
454 ## Licence
455 
456 AGPL-3.0: see `LICENSE`. You may use, change and host this, and you must give
457 your users the source of what you host. `NOTICE` has the licences of the work
458 this is built on (SubPlz, whisper.cpp, Whisper).