docs/reference.md (14810 bytes)
1 # desktop-subtitle-replay 2 3 Live subtitles for anything playing on your desktop, plus a mining workflow for 4 the clips you save afterwards. Built for language learning: understand it now, 5 turn it into Anki cards later. 6 7 Runs entirely on your machine. No API keys, no cloud, no internet after the 8 model downloads. 9 10 ``` 11 ┌─► OBS overlay (captions on the stream) 12 desktop audio ──────┼─► reader.html (selectable text — Yomitan mines it) 13 (live) └─► captions.txt/log (plain text) 14 15 replay clip ────────┬─► clip.srt (subtitles) 16 (offline) ├─► clip.html (video + hoverable synced subs) 17 └─► clip.anki.tsv (one card per sentence + audio) 18 ``` 19 20 Two halves, because they want opposite things. Live needs speed and accepts 21 mistakes. Replay needs accuracy and does not care about time. 22 23 ## Requirements 24 25 - Windows 10/11 26 - Python 3.9–3.12 (3.11 recommended) 27 - OBS Studio 28+ (only for the live overlay) 28 - ~2 GB disk for the model cache 29 - A GPU is *not* required, but it changes what is possible — see below 30 31 ## Install 32 33 ```bash 34 .\setup.ps1 35 ``` 36 37 ## Start everything 38 39 ```bash 40 .\start.ps1 41 ``` 42 43 Runs both halves: live captions, plus a watcher that subtitles every replay clip 44 OBS saves and serves it as a mining page. `-Lang ru`, `-LiveOnly`, `-ReplayOnly` 45 and `-Model` adjust it. The two halves are also usable separately, below. 46 47 ## Live captions 48 49 ```bash 50 .\run.ps1 51 ``` 52 53 Prints three URLs: 54 55 | page | where it goes | 56 |---|---| 57 | `overlay.html` | OBS Browser Source — captions burned into the stream | 58 | `reader.html` | **your real browser** — selectable text for Yomitan | 59 | `control.html` | language and model switching while running | 60 61 For OBS: **+ → Browser**, paste the overlay URL, **1920 × 1080**, untick 62 **Shutdown source when not visible**. 63 64 For mining: open `reader.html` in the browser where Yomitan is installed. Each 65 sentence is a plain DOM text node, so Yomitan's popup and its sentence field 66 work normally. The timestamp is drawn with CSS rather than text, so it never 67 gets absorbed into the sentence you mine. 68 69 The reader keeps the whole session, and restores it from `captions.log` if you 70 reload, so you can scroll back to something said minutes ago. `A+`/`A−` resize, 71 `Follow` toggles auto-scroll (turn it off while working through a line), 72 double-click pins a line, `Copy all` grabs the transcript. 73 74 ### Changing language and model mid-session 75 76 Open `control.html`. Language switches on one click or keys `1`–`9` and applies 77 to the next sentence — no reload. Model switching reloads and pauses captions 78 for a few seconds. Defaults cover `fi,ru,ja,es,pt,en,auto`: 79 80 ```bash 81 .\run.ps1 --langs fi,ru,ja,es,pt,en,auto --lang fi 82 ``` 83 84 **Whisper has no regional variants.** Argentine Spanish is `es`; there is no 85 `es-AR`. It handles Rioplatense pronunciation and *voseo*, but normalises 86 toward standard orthography and will not reliably reproduce regional slang. 87 Brazilian and European Portuguese are both `pt`. 88 89 ### How `auto` decides 90 91 Whisper's own detection answers with any of 99 languages, decides afresh on 92 every call, and has little to go on in a two-word utterance. `auto` here is 93 built on top of it and constrained to the languages you actually configured: 94 95 1. **Restricted to `--langs`.** A Finnish clip cannot come back as Estonian. 96 2. **Close relatives are folded in.** Whisper splits mass between neighbours, 97 so Estonian counts toward Finnish, Ukrainian and Bulgarian toward Russian, 98 Galician toward Portuguese, Catalan toward Spanish. Otherwise a clip can 99 lose because `fi` and `et` split the vote and a third language wins. 100 3. **Reliability gate.** Restricting the set deletes the competitors and so 101 inflates confidence — `en 0.38, ko 0.25, nn 0.10` looks like a commanding 102 `en` once `ko` and `nn` are dropped. If too little mass lands inside your 103 set, it declines to guess and lets Whisper handle that one segment. 104 4. **Evidence accumulates over time**, weighted by segment length, instead of 105 each segment deciding alone. 106 5. **Hysteresis.** A new language must lead by a margin for several 107 consecutive detections before it takes over, so one odd segment cannot 108 flip the caption language mid-conversation. 109 110 Detection runs on a schedule (`--detect-every`, default 6 s) rather than every 111 segment. It costs an encoder pass, but passing an explicit language into 112 Whisper skips its *internal* detection, so the steady-state cost is roughly 113 neutral. 114 115 Measured on a Finnish clip in short windows, plain Whisper committed to `en` 116 on an ambiguous leading window; the detector declined to guess there and then 117 held `fi` across every remaining window. 118 119 | flag | default | effect | 120 |---|---|---| 121 | `--detect-every` | `6.0` | seconds between detection passes | 122 | `--detect-min-audio` | `1.6` | skip detection on shorter segments | 123 | `--detect-margin` | `1.3` | how far a challenger must lead to switch | 124 | `--detect-hold` | `2` | consecutive detections before switching | 125 126 The control panel shows `auto → ja` once it settles. Pinning with `1`–`9` is 127 still the most reliable option when you already know the language. 128 129 ### Choosing what gets captioned 130 131 Default loopback captures everything your speakers play. Your own microphone is 132 *not* included unless OBS monitors it, and music gets transcribed too. 133 134 With [VB-Audio Virtual Cable](https://vb-audio.com/Cable/), send only what you 135 want captioned: 136 137 1. OBS → **Settings → Audio → Advanced → Monitoring Device** = `CABLE Input` 138 2. Audio Mixer → gear on each source → **Advanced Audio Properties** → 139 **Audio Monitoring** = **Monitor and Output** 140 3. Leave music and alerts on **Monitor Off** 141 142 ```bash 143 .\run.ps1 --mic --audio-device CABLE 144 ``` 145 146 ## Subtitling replay clips 147 148 ```bash 149 .\.venv\Scripts\python.exe subtitle.py clip.mp4 150 ``` 151 152 Writes `clip.srt` and `clip.html` next to the clip. The HTML page is the mining 153 surface: video on the left, every sentence listed as selectable text, click to 154 seek, `Loop cue` to repeat a line while you work it out. 155 156 Watch your replay folder and subtitle clips automatically as OBS saves them. 157 With no folder given, it reads OBS's own config to find where recordings go 158 (honouring Simple vs Advanced output mode): 159 160 ```bash 161 .\.venv\Scripts\python.exe subtitle.py --watch --serve 162 ``` 163 164 `--serve` matters more than it looks. Opening `clip.html` from `file://` 165 requires enabling Yomitan's *Allow access to file URLs*, and serving it through 166 `python -m http.server` **silently breaks video seeking**, because that server 167 ignores HTTP Range requests. `--serve` runs a range-capable server, so scrubbing 168 works and Yomitan needs no extra permission. 169 170 ### Anki cards 171 172 ```bash 173 .\.venv\Scripts\python.exe subtitle.py clip.mp4 --anki 174 ``` 175 176 Produces `clip.anki.tsv` (one row per sentence) plus a media folder of 177 per-sentence audio clips, with columns: sentence, translation, `[sound:…]`, 178 `<img>`, source file, timestamp. Copy the media into your Anki 179 `collection.media` and import the TSV. 180 181 This is the batch path. If you mine word-by-word with Yomitan, use `clip.html` 182 instead and let Yomitan build the cards — it captures the sentence context on 183 its own, which is usually what you want. 184 185 Screenshots need Pillow (`pip install pillow`); without it the image column is 186 left empty and everything else still works. 187 188 **Two things about `--anki-translate`.** First, `large-v3-turbo` is a 189 transcription-only fine-tune and *cannot translate* — asked to, it silently 190 returns the source language. Translation is therefore routed to `small` unless 191 you set `--translate-model`. Second, Whisper's Finnish→English is genuinely 192 weak: it rendered *"sen verran syrjäisillä seuduilla"* as "the lake of Sennvera". 193 For learning, Yomitan's dictionary lookups are far more trustworthy than a 194 machine-translated card back. 195 196 ## Speed, honestly 197 198 Measured on a 16-core CPU, no CUDA, `int8`, on real Finnish speech: 199 200 | model | RTF | verdict | 201 |---|---|---| 202 | `tiny` | 0.07 | fast, poor Finnish | 203 | `base` | 0.10 | fast, weak Finnish | 204 | **`small`** (live default) | **0.57** | the practical ceiling on CPU | 205 | `large-v3-turbo` (replay default) | 1.77 | too slow live, ideal offline | 206 207 RTF is measured over a whole file. **Per segment during a live stream it is 208 worse** — short utterances pay a fixed encoder cost, so real sessions show 0.55 209 to 1.4, occasionally decoding slower than real time. Expect captions **3–8 s 210 behind the speaker**, not 1 s. That is inherent to the approach, not a bug. 211 212 Benchmark your own machine: 213 214 ```bash 215 .\.venv\Scripts\python.exe bench.py --models small,medium 216 ``` 217 218 Synthetic audio flatters models because there are fewer tokens to decode. Point 219 it at a real recording for a number you can trust: 220 221 ```bash 222 .\.venv\Scripts\python.exe bench.py --wav selftest.wav 223 ``` 224 225 ### Why it is not truly real-time 226 227 Whisper is an offline encoder-decoder over fixed 30-second windows. It cannot 228 emit a word until it has a chunk to process, so this waits for a pause and 229 transcribes the finished utterance. That is chunked pseudo-streaming, not 230 streaming ASR. 231 232 Engines that genuinely stream — Kaldi/Vosk online decoding, sherpa-onnx 233 Zipformer transducers — emit ~200–500 ms after the sound. Vosk covers Russian, 234 Japanese, Spanish and Portuguese, **but has no Finnish model**. Nothing 235 genuinely streaming covers this language set; Whisper covers all of it and is 236 not streaming. 237 238 `--stream` implements LocalAgreement streaming (Macháček et al.): re-decode a 239 growing buffer, commit only the prefix two consecutive decodes agree on. It is 240 **off by default because it measured worse here on both axes**, on the same 241 17.6 s Finnish clip: 242 243 | mode | speed | output | 244 |---|---|---| 245 | chunked `small` | 0.57× | *"Nyt ollaan taas sen verran syrjäisillä seuduilla ja harvakseltaan kuljetuilla seuduilla."* | 246 | `--stream small` | 3.58× | too slow to run | 247 | `--stream base` | 0.96× | *"Tolaan taas sen verran Syrjää Näissä ei sillä seudulla…"* | 248 | `--stream tiny` | 0.34× | *"Kösitäästä. Ja tolaa on… parvaksiautaa"* | 249 250 The models fast enough to stream are too weak for Finnish; the model good 251 enough for Finnish is 3.6× too slow. **With a CUDA GPU this inverts** — run 252 `--stream --compute-device cuda --compute float16 --model large-v3` and 253 LocalAgreement becomes the better mode. 254 255 To reduce latency without it, shorten the silence needed to close a caption and 256 free up CPU: 257 258 ```bash 259 .\run.ps1 --pause 0.4 --no-partials 260 ``` 261 262 ### Better Finnish at the same speed 263 264 Stock `small` is a generalist. A Finnish-fine-tuned `small` is the same size, so 265 the same speed, but markedly better at Finnish: 266 267 ```bash 268 .\get-finnish-model.ps1 269 ``` 270 271 Once built, `run.ps1` picks it up automatically whenever `--lang fi` and no 272 `--model` is given. `--model small` overrides it; other languages ignore it. 273 274 Measured on the same 17.6 s Finnish clip: 275 276 | model | transcription | RTF | 277 |---|---|---| 278 | `small` stock | "harv**o**kseltaan kuljetuilla" ✗ | 0.20 | 279 | **`fi-small-ct2`** | "harv**a**kseltaan kuljetuilla" ✓ | **0.21** | 280 | `large-v3-turbo` | "harvakseltaan kuljetuilla" ✓ | 0.57 | 281 282 Turbo-level accuracy at stock-`small` speed — 2.7× faster than turbo for the 283 same correct word. That is why it is the live default when present. 284 285 **Keep turbo for replay clips, though.** The fine-tune punctuates less 286 reliably, running sentences together, and the offline path splits cues on 287 sentence boundaries — so turbo yields cleaner subtitle cues and better Anki 288 card boundaries. Live, the VAD does the segmenting, so this costs nothing. 289 290 The conversion needs ~8 GB free and pulls torch, used only for that one step. 291 `-BuildRoot D:\somewhere` builds elsewhere; delete `.venv-convert` under the 292 build root afterwards to reclaim about 5 GB. 293 294 ## Syncing captions to the picture 295 296 The overlay is composited before any stream delay or replay buffer, so captions 297 ride along with the video automatically. 298 299 To line captions up with lips, delay the picture by the caption latency: 300 301 - **Render Delay** filter of `2500`–`4000` ms on video sources 302 - matching **Sync Offset** on audio sources (Advanced Audio Properties) 303 304 If you already run a replay buffer, this costs you nothing. 305 306 ## Options 307 308 `livecap.py`: 309 310 | flag | default | effect | 311 |---|---|---| 312 | `--lang` / `--langs` | `fi` / `fi,ru,ja,es,pt,en,auto` | active language, and the panel's buttons | 313 | `--model` | `small`, or the Finnish fine-tune if built | model name or local CTranslate2 directory | 314 | `--mic` / `--audio-device` | loopback / auto | capture an input device instead | 315 | `--pause` | `0.65` | silence that closes a caption; lower is snappier | 316 | `--min-speech` | `0.45` | ignore bursts shorter than this | 317 | `--vad-floor` / `--vad-ratio` | `0.004` / `3.0` | speech gate, absolute and relative | 318 | `--no-partials` | off | finished sentences only; roughly halves CPU | 319 | `--stream` | off | LocalAgreement streaming (see above) | 320 | `--translate` | off | add an English line | 321 | `--compute-device` | `cpu` | `cuda` with an NVIDIA GPU | 322 323 `subtitle.py`: 324 325 | flag | default | effect | 326 |---|---|---| 327 | `--model` | `large-v3-turbo` | accuracy over speed | 328 | `--watch FOLDER` | — | subtitle clips as they appear | 329 | `--serve [PORT]` | — | range-capable server so seeking works | 330 | `--anki` / `--anki-translate` | off | card export, with English backs | 331 | `--format` | `srt` | `srt`, `vtt`, `txt` | 332 | `--width` / `--max-chars` | `42` / `84` | line and cue length | 333 334 ## Diagnostics 335 336 ```bash 337 .\run.ps1 --selftest 12 338 ``` 339 340 Records 12 s, writes `selftest.wav`, transcribes it and reports the real-time 341 factor. Run this first. 342 343 ```bash 344 .\run.ps1 --meter 345 ``` 346 347 Level meter with the VAD gate, for tuning `--vad-floor`. `--verbose` logs every 348 segment the VAD sends. 349 350 | symptom | cause | 351 |---|---| 352 | `no speech recognised in N.Ns segment` | music/noise, or wrong `--lang` | 353 | nothing in the log at all | VAD never fires — check `--meter` and device | 354 | `backlog full` | model too slow; use a smaller one | 355 | video will not seek | server ignoring Range requests — use `--serve` | 356 | red dot in overlay/reader | lost the WebSocket; it retries automatically | 357 358 Whisper hallucinates stock phrases over silence (`Tekstitys: YLE`, 359 `Kiitos kun katsoit!`, `Thanks for watching`). Short results matching those are 360 filtered, alongside a no-speech-probability threshold. 361 362 ## Tests 363 364 ```bash 365 .\.venv\Scripts\python.exe tests\test_smoke.py 366 ``` 367 368 ```bash 369 .\.venv\Scripts\python.exe tests\test_subtitle.py 370 ``` 371 372 ## License 373 374 MIT — see [LICENSE](LICENSE). Uses 375 [faster-whisper](https://github.com/SYSTRAN/faster-whisper) (MIT) and OpenAI's 376 Whisper models. Streaming mode follows the LocalAgreement policy from 377 Macháček, Dabre & Bojar, *Turning Whisper into Real-Time Transcription System* 378 (2023).