Recently Written · git

desktop-subtitle-replay

git clone https://github.com/equwal/desktop-subtitle-replay

Log | Files | Refs


README.md (2966 bytes)

1 # desktop-subtitle-replay
2 
3 **Live subtitles for anything playing on your screen — then turn what you just heard into Anki cards.**
4 
5 Watch a stream in Finnish. Read along as it happens. Hit your replay buffer on
6 the sentence you didn't catch, and get it back as a clip with hoverable
7 subtitles you can mine with [Yomitan](https://yomitan.wiki/).
8 
9 Everything runs on your machine. No API keys, no cloud, no upload.
10 
11 ---
12 
13 ## Why this exists
14 
15 Subtitle tools caption *your* speech for *your* viewers. This does the
16 opposite: it captions what you're listening to, so you can follow along in a
17 language you're still learning — and keeps the audio so you can study it later.
18 
19 **Live and replay want opposite things**, so they get different engines:
20 
21 |  | live | replay |
22 |---|---|---|
23 | needs | speed | accuracy |
24 | model | `small` (~0.2 RTF) | `large-v3-turbo` |
25 | output | OBS overlay + reader page | `.srt`, mining page, Anki cards |
26 
27 The live model transcribed *"harvokseltaan"*. The replay model gets
28 *"harvakseltaan"* — the actual word. You want both.
29 
30 ## Mining is the point
31 
32 Yomitan reads **DOM text**, not pixels. So subtitles are rendered as real,
33 selectable text — each sentence its own text node, timestamps drawn in CSS so
34 they never contaminate the sentence you mine. Hover a word, get a definition,
35 make a card, with the sentence context captured automatically.
36 
37 That works on the live reader *and* on every replay clip.
38 
39 ## Start
40 
41 ```bash
42 .\setup.ps1
43 ```
44 
45 ```bash
46 .\start.ps1
47 ```
48 
49 That's it. Three URLs get printed:
50 
51 - **reader** — open in the browser where Yomitan lives
52 - **overlay** — paste into an OBS Browser Source
53 - **control** — switch language with `1`–`9`, swap models, no restart
54 
55 Save a replay in OBS and its mining page appears automatically.
56 
57 ## Languages
58 
59 Any Whisper language. Automatic detection is constrained to the ones you
60 actually use, so a Finnish clip can't come back as Estonian:
61 
62 ```bash
63 .\start.ps1 -Lang ja
64 ```
65 
66 ```bash
67 .\run.ps1 --langs fi,ru,ja,es,pt,en --lang auto
68 ```
69 
70 ## What it won't do
71 
72 **It runs about 3–8 seconds behind.** Whisper is not a streaming model — it
73 can't emit a word until it has a chunk to process, so this waits for a pause
74 and transcribes the finished sentence. Genuinely streaming engines exist, but
75 none of them cover this language set. If you need lip-sync, delay your video
76 by the same amount; against a replay buffer that costs nothing.
77 
78 **It wants CPU.** No GPU required, and `small` holds real time on 16 cores —
79 but Japanese is tighter than European languages. With a CUDA GPU, everything
80 here gets better and `--stream` becomes viable.
81 
82 **It will occasionally invent a sentence.** Whisper hallucinates over silence.
83 The common stock phrases are filtered, in every language it supports.
84 
85 ---
86 
87 Full options, tuning, benchmarks and the reasoning behind them:
88 **[docs/reference.md](docs/reference.md)**
89 
90 MIT. Built on [faster-whisper](https://github.com/SYSTRAN/faster-whisper).