Recently Written · git

desktop-subtitle-replay

git clone https://github.com/equwal/desktop-subtitle-replay

Log | Files | Refs


docs/reference.md (14810 bytes)

1 # desktop-subtitle-replay
2 
3 Live subtitles for anything playing on your desktop, plus a mining workflow for
4 the clips you save afterwards. Built for language learning: understand it now,
5 turn it into Anki cards later.
6 
7 Runs entirely on your machine. No API keys, no cloud, no internet after the
8 model downloads.
9 
10 ```
11                     ┌─► OBS overlay        (captions on the stream)
12 desktop audio ──────┼─► reader.html        (selectable text — Yomitan mines it)
13    (live)           └─► captions.txt/log   (plain text)
14 
15 replay clip ────────┬─► clip.srt           (subtitles)
16    (offline)        ├─► clip.html          (video + hoverable synced subs)
17                     └─► clip.anki.tsv      (one card per sentence + audio)
18 ```
19 
20 Two halves, because they want opposite things. Live needs speed and accepts
21 mistakes. Replay needs accuracy and does not care about time.
22 
23 ## Requirements
24 
25 - Windows 10/11
26 - Python 3.9–3.12 (3.11 recommended)
27 - OBS Studio 28+ (only for the live overlay)
28 - ~2 GB disk for the model cache
29 - A GPU is *not* required, but it changes what is possible — see below
30 
31 ## Install
32 
33 ```bash
34 .\setup.ps1
35 ```
36 
37 ## Start everything
38 
39 ```bash
40 .\start.ps1
41 ```
42 
43 Runs both halves: live captions, plus a watcher that subtitles every replay clip
44 OBS saves and serves it as a mining page. `-Lang ru`, `-LiveOnly`, `-ReplayOnly`
45 and `-Model` adjust it. The two halves are also usable separately, below.
46 
47 ## Live captions
48 
49 ```bash
50 .\run.ps1
51 ```
52 
53 Prints three URLs:
54 
55 | page | where it goes |
56 |---|---|
57 | `overlay.html` | OBS Browser Source — captions burned into the stream |
58 | `reader.html` | **your real browser** — selectable text for Yomitan |
59 | `control.html` | language and model switching while running |
60 
61 For OBS: **+ → Browser**, paste the overlay URL, **1920 × 1080**, untick
62 **Shutdown source when not visible**.
63 
64 For mining: open `reader.html` in the browser where Yomitan is installed. Each
65 sentence is a plain DOM text node, so Yomitan's popup and its sentence field
66 work normally. The timestamp is drawn with CSS rather than text, so it never
67 gets absorbed into the sentence you mine.
68 
69 The reader keeps the whole session, and restores it from `captions.log` if you
70 reload, so you can scroll back to something said minutes ago. `A+`/`A−` resize,
71 `Follow` toggles auto-scroll (turn it off while working through a line),
72 double-click pins a line, `Copy all` grabs the transcript.
73 
74 ### Changing language and model mid-session
75 
76 Open `control.html`. Language switches on one click or keys `1`–`9` and applies
77 to the next sentence — no reload. Model switching reloads and pauses captions
78 for a few seconds. Defaults cover `fi,ru,ja,es,pt,en,auto`:
79 
80 ```bash
81 .\run.ps1 --langs fi,ru,ja,es,pt,en,auto --lang fi
82 ```
83 
84 **Whisper has no regional variants.** Argentine Spanish is `es`; there is no
85 `es-AR`. It handles Rioplatense pronunciation and *voseo*, but normalises
86 toward standard orthography and will not reliably reproduce regional slang.
87 Brazilian and European Portuguese are both `pt`.
88 
89 ### How `auto` decides
90 
91 Whisper's own detection answers with any of 99 languages, decides afresh on
92 every call, and has little to go on in a two-word utterance. `auto` here is
93 built on top of it and constrained to the languages you actually configured:
94 
95 1. **Restricted to `--langs`.** A Finnish clip cannot come back as Estonian.
96 2. **Close relatives are folded in.** Whisper splits mass between neighbours,
97    so Estonian counts toward Finnish, Ukrainian and Bulgarian toward Russian,
98    Galician toward Portuguese, Catalan toward Spanish. Otherwise a clip can
99    lose because `fi` and `et` split the vote and a third language wins.
100 3. **Reliability gate.** Restricting the set deletes the competitors and so
101    inflates confidence — `en 0.38, ko 0.25, nn 0.10` looks like a commanding
102    `en` once `ko` and `nn` are dropped. If too little mass lands inside your
103    set, it declines to guess and lets Whisper handle that one segment.
104 4. **Evidence accumulates over time**, weighted by segment length, instead of
105    each segment deciding alone.
106 5. **Hysteresis.** A new language must lead by a margin for several
107    consecutive detections before it takes over, so one odd segment cannot
108    flip the caption language mid-conversation.
109 
110 Detection runs on a schedule (`--detect-every`, default 6 s) rather than every
111 segment. It costs an encoder pass, but passing an explicit language into
112 Whisper skips its *internal* detection, so the steady-state cost is roughly
113 neutral.
114 
115 Measured on a Finnish clip in short windows, plain Whisper committed to `en`
116 on an ambiguous leading window; the detector declined to guess there and then
117 held `fi` across every remaining window.
118 
119 | flag | default | effect |
120 |---|---|---|
121 | `--detect-every` | `6.0` | seconds between detection passes |
122 | `--detect-min-audio` | `1.6` | skip detection on shorter segments |
123 | `--detect-margin` | `1.3` | how far a challenger must lead to switch |
124 | `--detect-hold` | `2` | consecutive detections before switching |
125 
126 The control panel shows `auto → ja` once it settles. Pinning with `1`–`9` is
127 still the most reliable option when you already know the language.
128 
129 ### Choosing what gets captioned
130 
131 Default loopback captures everything your speakers play. Your own microphone is
132 *not* included unless OBS monitors it, and music gets transcribed too.
133 
134 With [VB-Audio Virtual Cable](https://vb-audio.com/Cable/), send only what you
135 want captioned:
136 
137 1. OBS → **Settings → Audio → Advanced → Monitoring Device** = `CABLE Input`
138 2. Audio Mixer → gear on each source → **Advanced Audio Properties** →
139    **Audio Monitoring** = **Monitor and Output**
140 3. Leave music and alerts on **Monitor Off**
141 
142 ```bash
143 .\run.ps1 --mic --audio-device CABLE
144 ```
145 
146 ## Subtitling replay clips
147 
148 ```bash
149 .\.venv\Scripts\python.exe subtitle.py clip.mp4
150 ```
151 
152 Writes `clip.srt` and `clip.html` next to the clip. The HTML page is the mining
153 surface: video on the left, every sentence listed as selectable text, click to
154 seek, `Loop cue` to repeat a line while you work it out.
155 
156 Watch your replay folder and subtitle clips automatically as OBS saves them.
157 With no folder given, it reads OBS's own config to find where recordings go
158 (honouring Simple vs Advanced output mode):
159 
160 ```bash
161 .\.venv\Scripts\python.exe subtitle.py --watch --serve
162 ```
163 
164 `--serve` matters more than it looks. Opening `clip.html` from `file://`
165 requires enabling Yomitan's *Allow access to file URLs*, and serving it through
166 `python -m http.server` **silently breaks video seeking**, because that server
167 ignores HTTP Range requests. `--serve` runs a range-capable server, so scrubbing
168 works and Yomitan needs no extra permission.
169 
170 ### Anki cards
171 
172 ```bash
173 .\.venv\Scripts\python.exe subtitle.py clip.mp4 --anki
174 ```
175 
176 Produces `clip.anki.tsv` (one row per sentence) plus a media folder of
177 per-sentence audio clips, with columns: sentence, translation, `[sound:…]`,
178 `<img>`, source file, timestamp. Copy the media into your Anki
179 `collection.media` and import the TSV.
180 
181 This is the batch path. If you mine word-by-word with Yomitan, use `clip.html`
182 instead and let Yomitan build the cards — it captures the sentence context on
183 its own, which is usually what you want.
184 
185 Screenshots need Pillow (`pip install pillow`); without it the image column is
186 left empty and everything else still works.
187 
188 **Two things about `--anki-translate`.** First, `large-v3-turbo` is a
189 transcription-only fine-tune and *cannot translate* — asked to, it silently
190 returns the source language. Translation is therefore routed to `small` unless
191 you set `--translate-model`. Second, Whisper's Finnish→English is genuinely
192 weak: it rendered *"sen verran syrjäisillä seuduilla"* as "the lake of Sennvera".
193 For learning, Yomitan's dictionary lookups are far more trustworthy than a
194 machine-translated card back.
195 
196 ## Speed, honestly
197 
198 Measured on a 16-core CPU, no CUDA, `int8`, on real Finnish speech:
199 
200 | model | RTF | verdict |
201 |---|---|---|
202 | `tiny` | 0.07 | fast, poor Finnish |
203 | `base` | 0.10 | fast, weak Finnish |
204 | **`small`** (live default) | **0.57** | the practical ceiling on CPU |
205 | `large-v3-turbo` (replay default) | 1.77 | too slow live, ideal offline |
206 
207 RTF is measured over a whole file. **Per segment during a live stream it is
208 worse** — short utterances pay a fixed encoder cost, so real sessions show 0.55
209 to 1.4, occasionally decoding slower than real time. Expect captions **3–8 s
210 behind the speaker**, not 1 s. That is inherent to the approach, not a bug.
211 
212 Benchmark your own machine:
213 
214 ```bash
215 .\.venv\Scripts\python.exe bench.py --models small,medium
216 ```
217 
218 Synthetic audio flatters models because there are fewer tokens to decode. Point
219 it at a real recording for a number you can trust:
220 
221 ```bash
222 .\.venv\Scripts\python.exe bench.py --wav selftest.wav
223 ```
224 
225 ### Why it is not truly real-time
226 
227 Whisper is an offline encoder-decoder over fixed 30-second windows. It cannot
228 emit a word until it has a chunk to process, so this waits for a pause and
229 transcribes the finished utterance. That is chunked pseudo-streaming, not
230 streaming ASR.
231 
232 Engines that genuinely stream — Kaldi/Vosk online decoding, sherpa-onnx
233 Zipformer transducers — emit ~200–500 ms after the sound. Vosk covers Russian,
234 Japanese, Spanish and Portuguese, **but has no Finnish model**. Nothing
235 genuinely streaming covers this language set; Whisper covers all of it and is
236 not streaming.
237 
238 `--stream` implements LocalAgreement streaming (Macháček et al.): re-decode a
239 growing buffer, commit only the prefix two consecutive decodes agree on. It is
240 **off by default because it measured worse here on both axes**, on the same
241 17.6 s Finnish clip:
242 
243 | mode | speed | output |
244 |---|---|---|
245 | chunked `small` | 0.57× | *"Nyt ollaan taas sen verran syrjäisillä seuduilla ja harvakseltaan kuljetuilla seuduilla."* |
246 | `--stream small` | 3.58× | too slow to run |
247 | `--stream base` | 0.96× | *"Tolaan taas sen verran Syrjää Näissä ei sillä seudulla…"* |
248 | `--stream tiny` | 0.34× | *"Kösitäästä. Ja tolaa on… parvaksiautaa"* |
249 
250 The models fast enough to stream are too weak for Finnish; the model good
251 enough for Finnish is 3.6× too slow. **With a CUDA GPU this inverts** — run
252 `--stream --compute-device cuda --compute float16 --model large-v3` and
253 LocalAgreement becomes the better mode.
254 
255 To reduce latency without it, shorten the silence needed to close a caption and
256 free up CPU:
257 
258 ```bash
259 .\run.ps1 --pause 0.4 --no-partials
260 ```
261 
262 ### Better Finnish at the same speed
263 
264 Stock `small` is a generalist. A Finnish-fine-tuned `small` is the same size, so
265 the same speed, but markedly better at Finnish:
266 
267 ```bash
268 .\get-finnish-model.ps1
269 ```
270 
271 Once built, `run.ps1` picks it up automatically whenever `--lang fi` and no
272 `--model` is given. `--model small` overrides it; other languages ignore it.
273 
274 Measured on the same 17.6 s Finnish clip:
275 
276 | model | transcription | RTF |
277 |---|---|---|
278 | `small` stock | "harv**o**kseltaan kuljetuilla" ✗ | 0.20 |
279 | **`fi-small-ct2`** | "harv**a**kseltaan kuljetuilla" ✓ | **0.21** |
280 | `large-v3-turbo` | "harvakseltaan kuljetuilla" ✓ | 0.57 |
281 
282 Turbo-level accuracy at stock-`small` speed — 2.7× faster than turbo for the
283 same correct word. That is why it is the live default when present.
284 
285 **Keep turbo for replay clips, though.** The fine-tune punctuates less
286 reliably, running sentences together, and the offline path splits cues on
287 sentence boundaries — so turbo yields cleaner subtitle cues and better Anki
288 card boundaries. Live, the VAD does the segmenting, so this costs nothing.
289 
290 The conversion needs ~8 GB free and pulls torch, used only for that one step.
291 `-BuildRoot D:\somewhere` builds elsewhere; delete `.venv-convert` under the
292 build root afterwards to reclaim about 5 GB.
293 
294 ## Syncing captions to the picture
295 
296 The overlay is composited before any stream delay or replay buffer, so captions
297 ride along with the video automatically.
298 
299 To line captions up with lips, delay the picture by the caption latency:
300 
301 - **Render Delay** filter of `2500`–`4000` ms on video sources
302 - matching **Sync Offset** on audio sources (Advanced Audio Properties)
303 
304 If you already run a replay buffer, this costs you nothing.
305 
306 ## Options
307 
308 `livecap.py`:
309 
310 | flag | default | effect |
311 |---|---|---|
312 | `--lang` / `--langs` | `fi` / `fi,ru,ja,es,pt,en,auto` | active language, and the panel's buttons |
313 | `--model` | `small`, or the Finnish fine-tune if built | model name or local CTranslate2 directory |
314 | `--mic` / `--audio-device` | loopback / auto | capture an input device instead |
315 | `--pause` | `0.65` | silence that closes a caption; lower is snappier |
316 | `--min-speech` | `0.45` | ignore bursts shorter than this |
317 | `--vad-floor` / `--vad-ratio` | `0.004` / `3.0` | speech gate, absolute and relative |
318 | `--no-partials` | off | finished sentences only; roughly halves CPU |
319 | `--stream` | off | LocalAgreement streaming (see above) |
320 | `--translate` | off | add an English line |
321 | `--compute-device` | `cpu` | `cuda` with an NVIDIA GPU |
322 
323 `subtitle.py`:
324 
325 | flag | default | effect |
326 |---|---|---|
327 | `--model` | `large-v3-turbo` | accuracy over speed |
328 | `--watch FOLDER` | — | subtitle clips as they appear |
329 | `--serve [PORT]` | — | range-capable server so seeking works |
330 | `--anki` / `--anki-translate` | off | card export, with English backs |
331 | `--format` | `srt` | `srt`, `vtt`, `txt` |
332 | `--width` / `--max-chars` | `42` / `84` | line and cue length |
333 
334 ## Diagnostics
335 
336 ```bash
337 .\run.ps1 --selftest 12
338 ```
339 
340 Records 12 s, writes `selftest.wav`, transcribes it and reports the real-time
341 factor. Run this first.
342 
343 ```bash
344 .\run.ps1 --meter
345 ```
346 
347 Level meter with the VAD gate, for tuning `--vad-floor`. `--verbose` logs every
348 segment the VAD sends.
349 
350 | symptom | cause |
351 |---|---|
352 | `no speech recognised in N.Ns segment` | music/noise, or wrong `--lang` |
353 | nothing in the log at all | VAD never fires — check `--meter` and device |
354 | `backlog full` | model too slow; use a smaller one |
355 | video will not seek | server ignoring Range requests — use `--serve` |
356 | red dot in overlay/reader | lost the WebSocket; it retries automatically |
357 
358 Whisper hallucinates stock phrases over silence (`Tekstitys: YLE`,
359 `Kiitos kun katsoit!`, `Thanks for watching`). Short results matching those are
360 filtered, alongside a no-speech-probability threshold.
361 
362 ## Tests
363 
364 ```bash
365 .\.venv\Scripts\python.exe tests\test_smoke.py
366 ```
367 
368 ```bash
369 .\.venv\Scripts\python.exe tests\test_subtitle.py
370 ```
371 
372 ## License
373 
374 MIT — see [LICENSE](LICENSE). Uses
375 [faster-whisper](https://github.com/SYSTRAN/faster-whisper) (MIT) and OpenAI's
376 Whisper models. Streaming mode follows the LocalAgreement policy from
377 Macháček, Dabre & Bojar, *Turning Whisper into Real-Time Transcription System*
378 (2023).