README.md (2966 bytes)
1 # desktop-subtitle-replay 2 3 **Live subtitles for anything playing on your screen — then turn what you just heard into Anki cards.** 4 5 Watch a stream in Finnish. Read along as it happens. Hit your replay buffer on 6 the sentence you didn't catch, and get it back as a clip with hoverable 7 subtitles you can mine with [Yomitan](https://yomitan.wiki/). 8 9 Everything runs on your machine. No API keys, no cloud, no upload. 10 11 --- 12 13 ## Why this exists 14 15 Subtitle tools caption *your* speech for *your* viewers. This does the 16 opposite: it captions what you're listening to, so you can follow along in a 17 language you're still learning — and keeps the audio so you can study it later. 18 19 **Live and replay want opposite things**, so they get different engines: 20 21 | | live | replay | 22 |---|---|---| 23 | needs | speed | accuracy | 24 | model | `small` (~0.2 RTF) | `large-v3-turbo` | 25 | output | OBS overlay + reader page | `.srt`, mining page, Anki cards | 26 27 The live model transcribed *"harvokseltaan"*. The replay model gets 28 *"harvakseltaan"* — the actual word. You want both. 29 30 ## Mining is the point 31 32 Yomitan reads **DOM text**, not pixels. So subtitles are rendered as real, 33 selectable text — each sentence its own text node, timestamps drawn in CSS so 34 they never contaminate the sentence you mine. Hover a word, get a definition, 35 make a card, with the sentence context captured automatically. 36 37 That works on the live reader *and* on every replay clip. 38 39 ## Start 40 41 ```bash 42 .\setup.ps1 43 ``` 44 45 ```bash 46 .\start.ps1 47 ``` 48 49 That's it. Three URLs get printed: 50 51 - **reader** — open in the browser where Yomitan lives 52 - **overlay** — paste into an OBS Browser Source 53 - **control** — switch language with `1`–`9`, swap models, no restart 54 55 Save a replay in OBS and its mining page appears automatically. 56 57 ## Languages 58 59 Any Whisper language. Automatic detection is constrained to the ones you 60 actually use, so a Finnish clip can't come back as Estonian: 61 62 ```bash 63 .\start.ps1 -Lang ja 64 ``` 65 66 ```bash 67 .\run.ps1 --langs fi,ru,ja,es,pt,en --lang auto 68 ``` 69 70 ## What it won't do 71 72 **It runs about 3–8 seconds behind.** Whisper is not a streaming model — it 73 can't emit a word until it has a chunk to process, so this waits for a pause 74 and transcribes the finished sentence. Genuinely streaming engines exist, but 75 none of them cover this language set. If you need lip-sync, delay your video 76 by the same amount; against a replay buffer that costs nothing. 77 78 **It wants CPU.** No GPU required, and `small` holds real time on 16 cores — 79 but Japanese is tighter than European languages. With a CUDA GPU, everything 80 here gets better and `--stream` becomes viable. 81 82 **It will occasionally invent a sentence.** Whisper hallucinates over silence. 83 The common stock phrases are filtered, in every language it supports. 84 85 --- 86 87 Full options, tuning, benchmarks and the reasoning behind them: 88 **[docs/reference.md](docs/reference.md)** 89 90 MIT. Built on [faster-whisper](https://github.com/SYSTRAN/faster-whisper).