commit 9f8898667f6a95057bf21f130d51f73b0d11d420 equwal <truex@equwal.com> 2026-08-26 03:45:31 -0700 Add replay mining workflow and offline subtitler The live captioner only served half the need: subtitles are read in the moment to follow along, but the saved replay clips are where Anki cards get made, and those two jobs want opposite trade-offs. Live needs speed and tolerates errors; replay needs accuracy and ignores latency. subtitle.py transcribes clips offline with large-v3-turbo, which is far too slow live but markedly better - it produces "harvakseltaan" where small produces "harvokseltaan". It writes SRT/VTT, a per-clip mining page, and optional Anki rows with per-sentence audio and screenshots. Mining happens through Yomitan, which reads DOM text, so the deliverable is hoverable HTML rather than pixels: reader.html for live and clip.html for replays, both keeping each sentence as a bare text node. Timestamps render from CSS attr() so Yomitan cannot absorb them into the mined sentence. Cues are grouped by sentence first and only then split to fit a line, so a card never begins mid-sentence; short trailing runts fold back into the previous cue. Fixes found while testing: - Video seeking was silently impossible over http, because Python's stock handler ignores Range requests. Added a 206-capable server; scrubbing is most of what mining involves. - Subtitle sync only followed timeupdate, which does not fire while paused, so scrubbing never updated the text. Now also bound to seeked. - wrap() gave up when no split fit the width and emitted one long line. --stream implements LocalAgreement streaming but stays off: measured on a 17.6s Finnish clip it was worse on both axes, at 3.58x real time with small, and garbled with the smaller models fast enough to keep up. It is the right mode on a GPU and is documented as such. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
.github/workflows/ci.yml | 7 +- README.md | 357 +++++++++++---------- control.html | 2 + livecap.py | 211 +++++++++++- reader.html | 232 ++++++++++++++ requirements.txt | 1 + subtitle.py | 816 +++++++++++++++++++++++++++++++++++++++++++++++ tests/test_subtitle.py | 202 ++++++++++++ 8 files changed, 1660 insertions(+), 168 deletions(-)
diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index f032299..c91b793 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -26,12 +26,15 @@ jobs: pip install -r requirements.txt - name: Byte-compile - run: python -m compileall -q livecap.py bench.py + run: python -m compileall -q livecap.py bench.py subtitle.py - name: CLI smoke check run: | python livecap.py --help python bench.py --help + python subtitle.py --help - name: Tests - run: python tests/test_smoke.py + run: | + python tests/test_smoke.py + python tests/test_subtitle.py diff --git a/README.md b/README.md index 8f78e18..13d5e15 100644 --- a/README.md +++ b/README.md @@ -1,25 +1,32 @@ -# livecap — live captions for OBS +# desktop-subtitle-replay -Real-time speech captions for OBS Studio on Windows. Captures audio, detects -speech, transcribes it with Whisper, and renders it into a transparent Browser -Source overlay. +Live subtitles for anything playing on your desktop, plus a mining workflow for +the clips you save afterwards. Built for language learning: understand it now, +turn it into Anki cards later. Runs entirely on your machine. No API keys, no cloud, no internet after the -model is downloaded. Built for Finnish, works with any language Whisper -supports. +model downloads. ``` -audio device ──► VAD segmenter ──► faster-whisper ──► WebSocket ──► overlay.html - └─► captions.txt (GDI+ fallback) + ┌─► OBS overlay (captions on the stream) +desktop audio ──────┼─► reader.html (selectable text — Yomitan mines it) + (live) └─► captions.txt/log (plain text) + +replay clip ────────┬─► clip.srt (subtitles) + (offline) ├─► clip.html (video + hoverable synced subs) + └─► clip.anki.tsv (one card per sentence + audio) ``` +Two halves, because they want opposite things. Live needs speed and accepts +mistakes. Replay needs accuracy and does not care about time. + ## Requirements - Windows 10/11 - Python 3.9–3.12 (3.11 recommended) -- OBS Studio 28+ +- OBS Studio 28+ (only for the live overlay) - ~2 GB disk for the model cache -- A GPU is *not* required — this is tuned to run on CPU +- A GPU is *not* required, but it changes what is possible — see below ## Install @@ -27,132 +34,174 @@ audio device ──► VAD segmenter ──► faster-whisper ──► WebSocke .\setup.ps1 ``` -Creates a `.venv` and installs dependencies. One-time. - -## Run +## Live captions ```bash .\run.ps1 ``` -Prints a Browser Source URL. In OBS: **+ → Browser**, paste it, set -**1920 × 1080**, and untick **Shutdown source when not visible**. +Prints three URLs: -``` -http://127.0.0.1:8777/overlay.html?ws=8765&lines=2&size=42&hide=8 -``` +| page | where it goes | +|---|---| +| `overlay.html` | OBS Browser Source — captions burned into the stream | +| `reader.html` | **your real browser** — selectable text for Yomitan | +| `control.html` | language and model switching while running | -The overlay is transparent and places captions in the lower third. Keep the -console window open while streaming; Ctrl+C stops it. +For OBS: **+ → Browser**, paste the overlay URL, **1920 × 1080**, untick +**Shutdown source when not visible**. -Check your devices first if needed: +For mining: open `reader.html` in the browser where Yomitan is installed. Each +sentence is a plain DOM text node, so Yomitan's popup and its sentence field +work normally. The timestamp is drawn with CSS rather than text, so it never +gets absorbed into the sentence you mine. + +### Changing language and model mid-session + +Open `control.html`. Language switches on one click or keys `1`–`9` and applies +to the next sentence — no reload. Model switching reloads and pauses captions +for a few seconds. Defaults cover `fi,ru,ja,es,pt,en,auto`: ```bash -.\run.ps1 --list-devices +.\run.ps1 --langs fi,ru,ja,es,pt,en,auto --lang fi ``` -## Changing language and model while streaming +**Whisper has no regional variants.** Argentine Spanish is `es`; there is no +`es-AR`. It handles Rioplatense pronunciation and *voseo*, but normalises +toward standard orthography and will not reliably reproduce regional slang. +Brazilian and European Portuguese are both `pt`. -`run.ps1` also prints a control panel URL. Open it in any browser (a second -monitor, a phone on the same machine, or an OBS dock via -**Docks → Custom Browser Docks**): +**`auto` re-detects per segment**, which flips on short or noisy audio and can +mislabel mid-sentence. For deliberate language changes the `1`–`9` keys are far +more reliable. -``` -http://127.0.0.1:8777/control.html?ws=8765 -``` +### Choosing what gets captioned + +Default loopback captures everything your speakers play. Your own microphone is +*not* included unless OBS monitors it, and music gets transcribed too. -- **Language** — one click, or press `1`–`9`. Applies to the next sentence. - No restart and no model reload, so it is effectively instant. -- **Model** — a dropdown. Swapping reloads the model, which pauses captions for - a few seconds; the panel shows a *loading* state while it happens. -- **Toggles** — English translation, live partial text, and clear captions. -- A live feed of what is being recognised, so you can sanity-check without - looking at the stream. +With [VB-Audio Virtual Cable](https://vb-audio.com/Cable/), send only what you +want captioned: -Choose which languages get buttons: +1. OBS → **Settings → Audio → Advanced → Monitoring Device** = `CABLE Input` +2. Audio Mixer → gear on each source → **Advanced Audio Properties** → + **Audio Monitoring** = **Monitor and Output** +3. Leave music and alerts on **Monitor Off** ```bash -.\run.ps1 --langs fi,ru,ja,es,pt,en,auto --lang fi +.\run.ps1 --mic --audio-device CABLE ``` -That default set covers Finnish, Russian, Japanese, Spanish, Portuguese, -English and auto-detect. Any [Whisper language code](https://github.com/openai/whisper#available-models-and-languages) -works. - -Two things worth knowing: +## Subtitling replay clips -- **Whisper has no regional variants.** Argentine Spanish is `es` — there is no - `es-AR`. The model handles Rioplatense pronunciation and *voseo* as part of - `es`, but it normalises toward standard orthography and will not reliably - reproduce regional slang. The same applies to Brazilian vs. European - Portuguese: both are `pt`. -- **`auto` re-detects per segment**, which sounds ideal for multilingual streams - but flips on short or noisy utterances and can mislabel mid-sentence. For a - stream that switches language deliberately, the `1`–`9` keys are far more - reliable than `auto`. +```bash +.\.venv\Scripts\python.exe subtitle.py clip.mp4 +``` -Wider models are better at non-English audio. If you mostly caption Russian or -Japanese, `medium` is worth the latency if your CPU can take it — check with -`bench.py` first. +Writes `clip.srt` and `clip.html` next to the clip. The HTML page is the mining +surface: video on the left, every sentence listed as selectable text, click to +seek, `Loop cue` to repeat a line while you work it out. -## Choosing what gets captioned +Watch your replay folder and subtitle clips automatically as OBS saves them: -**Default (`--loopback`)** captures everything your speakers play — guests, -video, game audio, music. Zero setup. The catch: your own microphone is *not* -included unless OBS monitors it, and music gets fed to Whisper too. +```bash +.\.venv\Scripts\python.exe subtitle.py --watch "C:\Users\you\Videos" --serve +``` -**Recommended: route a dedicated mix through a virtual cable.** With -[VB-Audio Virtual Cable](https://vb-audio.com/Cable/) installed: +`--serve` matters more than it looks. Opening `clip.html` from `file://` +requires enabling Yomitan's *Allow access to file URLs*, and serving it through +`python -m http.server` **silently breaks video seeking**, because that server +ignores HTTP Range requests. `--serve` runs a range-capable server, so scrubbing +works and Yomitan needs no extra permission. -1. OBS → **Settings → Audio → Advanced → Monitoring Device** = - `CABLE Input (VB-Audio Virtual Cable)` -2. Audio Mixer → gear icon on each source you want captioned → - **Advanced Audio Properties** → **Audio Monitoring** = **Monitor and Output** - (this keeps the source audible to viewers; *Monitor Only* would mute it) -3. Leave music, game and alert sources on **Monitor Off** -4. Run: +### Anki cards ```bash -.\run.ps1 --mic --audio-device CABLE +.\.venv\Scripts\python.exe subtitle.py clip.mp4 --anki ``` -Accuracy improves noticeably once Whisper stops trying to transcribe your -background music. +Produces `clip.anki.tsv` (one row per sentence) plus a media folder of +per-sentence audio clips, with columns: sentence, translation, `[sound:…]`, +`<img>`, source file, timestamp. Copy the media into your Anki +`collection.media` and import the TSV. + +This is the batch path. If you mine word-by-word with Yomitan, use `clip.html` +instead and let Yomitan build the cards — it captures the sentence context on +its own, which is usually what you want. -## Model selection +Screenshots need Pillow (`pip install pillow`); without it the image column is +left empty and everything else still works. -Measured on a 16-core CPU with no CUDA GPU, `int8` quantisation: +## Speed, honestly -| model | real speech RTF | latency per segment | Finnish quality | -|---|---|---|---| -| `tiny` | 0.07 | ~0.4 s | poor | -| `base` | 0.10 | ~0.6 s | weak | -| **`small`** (default) | **0.57** | **~2.5 s** | good | -| `large-v3-turbo` | 1.77 | ~10.6 s | unusable — falls behind | +Measured on a 16-core CPU, no CUDA, `int8`, on real Finnish speech: + +| model | RTF | verdict | +|---|---|---| +| `tiny` | 0.07 | fast, poor Finnish | +| `base` | 0.10 | fast, weak Finnish | +| **`small`** (live default) | **0.57** | the practical ceiling on CPU | +| `large-v3-turbo` (replay default) | 1.77 | too slow live, ideal offline | -RTF (real-time factor) below ~0.6 keeps up with continuous speech. Without a -CUDA GPU, `small` is the largest model that stays real-time, which is why it is -the default. With an NVIDIA GPU, add `--compute-device cuda --compute float16` -and `large-v3` becomes viable. +RTF is measured over a whole file. **Per segment during a live stream it is +worse** — short utterances pay a fixed encoder cost, so real sessions show 0.55 +to 1.4, occasionally decoding slower than real time. Expect captions **3–8 s +behind the speaker**, not 1 s. That is inherent to the approach, not a bug. Benchmark your own machine: ```bash -.\.venv\Scripts\python.exe .\bench.py --models small,medium +.\.venv\Scripts\python.exe bench.py --models small,medium ``` -Synthetic audio makes models look faster than they are, because there are fewer -tokens to decode. For a realistic number, point it at a recording: +Synthetic audio flatters models because there are fewer tokens to decode. Point +it at a real recording for a number you can trust: ```bash -.\.venv\Scripts\python.exe .\bench.py --wav selftest.wav +.\.venv\Scripts\python.exe bench.py --wav selftest.wav +``` + +### Why it is not truly real-time + +Whisper is an offline encoder-decoder over fixed 30-second windows. It cannot +emit a word until it has a chunk to process, so this waits for a pause and +transcribes the finished utterance. That is chunked pseudo-streaming, not +streaming ASR. + +Engines that genuinely stream — Kaldi/Vosk online decoding, sherpa-onnx +Zipformer transducers — emit ~200–500 ms after the sound. Vosk covers Russian, +Japanese, Spanish and Portuguese, **but has no Finnish model**. Nothing +genuinely streaming covers this language set; Whisper covers all of it and is +not streaming. + +`--stream` implements LocalAgreement streaming (Macháček et al.): re-decode a +growing buffer, commit only the prefix two consecutive decodes agree on. It is +**off by default because it measured worse here on both axes**, on the same +17.6 s Finnish clip: + +| mode | speed | output | +|---|---|---| +| chunked `small` | 0.57× | *"Nyt ollaan taas sen verran syrjäisillä seuduilla ja harvakseltaan kuljetuilla seuduilla."* | +| `--stream small` | 3.58× | too slow to run | +| `--stream base` | 0.96× | *"Tolaan taas sen verran Syrjää Näissä ei sillä seudulla…"* | +| `--stream tiny` | 0.34× | *"Kösitäästä. Ja tolaa on… parvaksiautaa"* | + +The models fast enough to stream are too weak for Finnish; the model good +enough for Finnish is 3.6× too slow. **With a CUDA GPU this inverts** — run +`--stream --compute-device cuda --compute float16 --model large-v3` and +LocalAgreement becomes the better mode. + +To reduce latency without it, shorten the silence needed to close a caption and +free up CPU: + +```bash +.\run.ps1 --pause 0.4 --no-partials ``` ### Better Finnish at the same speed -Stock `small` is a generalist. A Finnish-fine-tuned `small` is the same size — -so the same speed — but markedly better at Finnish. -`get-finnish-model.ps1` fetches one and converts it to CTranslate2: +Stock `small` is a generalist. A Finnish-fine-tuned `small` is the same size, so +the same speed, but markedly better at Finnish: ```bash .\get-finnish-model.ps1 @@ -162,45 +211,48 @@ so the same speed — but markedly better at Finnish. .\run.ps1 --model .\build\models\fi-small-ct2 ``` -The conversion needs ~8 GB free and pulls in torch, which is used only for that -one-off step. Use `-BuildRoot D:\somewhere` to build on another drive, and -delete `.venv-convert` under the build root afterwards to reclaim the space. +Needs ~8 GB free and pulls torch for the one-off conversion; use +`-BuildRoot D:\somewhere` to build elsewhere, then delete `.venv-convert`. -## Caption latency and stream delay +## Syncing captions to the picture -The overlay is composited into the program feed *before* any stream delay or -replay buffer, so captions ride along with the video automatically — a buffer -needs no special handling. +The overlay is composited before any stream delay or replay buffer, so captions +ride along with the video automatically. -A caption appears roughly **2.5 s after the sentence ends**: inference time plus -the 0.65 s of silence used to detect the end of the sentence. To lock captions -to lips, delay the picture to match: +To line captions up with lips, delay the picture by the caption latency: -- **Render Delay** filter of `2500` ms on your video sources -- **Sync Offset** of `2500` ms on the audio sources - (Advanced Audio Properties) +- **Render Delay** filter of `2500`–`4000` ms on video sources +- matching **Sync Offset** on audio sources (Advanced Audio Properties) -If you already run a replay buffer or stream delay, the extra 2.5 s costs you -nothing. +If you already run a replay buffer, this costs you nothing. ## Options +`livecap.py`: + | flag | default | effect | |---|---|---| -| `--lang` | `fi` | `fi`, `en`, `sv`, … or `auto` | -| `--model` | `small` | model name or a local CTranslate2 directory | -| `--mic` | off | capture an input device instead of desktop output | -| `--audio-device` | auto | index from `--list-devices`, or part of the name | -| `--pause` | `0.65` | silence (s) that ends a caption; lower = snappier, more fragments | +| `--lang` / `--langs` | `fi` / `fi,ru,ja,es,pt,en,auto` | active language, and the panel's buttons | +| `--model` | `small` | model name or local CTranslate2 directory | +| `--mic` / `--audio-device` | loopback / auto | capture an input device instead | +| `--pause` | `0.65` | silence that closes a caption; lower is snappier | | `--min-speech` | `0.45` | ignore bursts shorter than this | -| `--max-seg` | `11.0` | force a cut during non-stop speech | -| `--vad-floor` | `0.004` | absolute level gate; raise if noise triggers captions | -| `--vad-ratio` | `3.0` | gate relative to the running noise floor | -| `--no-partials` | off | only show finished sentences; roughly halves CPU | -| `--translate` | off | add an English line beneath the original | -| `--lines` / `--size` | `2` / `42` | overlay line count and font size | -| `--hide` | `8` | fade the overlay out after N idle seconds | -| `--compute-device` | `cpu` | `cuda` if you have an NVIDIA GPU | +| `--vad-floor` / `--vad-ratio` | `0.004` / `3.0` | speech gate, absolute and relative | +| `--no-partials` | off | finished sentences only; roughly halves CPU | +| `--stream` | off | LocalAgreement streaming (see above) | +| `--translate` | off | add an English line | +| `--compute-device` | `cpu` | `cuda` with an NVIDIA GPU | + +`subtitle.py`: + +| flag | default | effect | +|---|---|---| +| `--model` | `large-v3-turbo` | accuracy over speed | +| `--watch FOLDER` | — | subtitle clips as they appear | +| `--serve [PORT]` | — | range-capable server so seeking works | +| `--anki` / `--anki-translate` | off | card export, with English backs | +| `--format` | `srt` | `srt`, `vtt`, `txt` | +| `--width` / `--max-chars` | `42` / `84` | line and cue length | ## Diagnostics @@ -208,65 +260,42 @@ nothing. .\run.ps1 --selftest 12 ``` -Records 12 s, writes `selftest.wav`, transcribes it once and reports the -real-time factor. Run this first — speak or play audio during those 12 seconds -and check what comes back. +Records 12 s, writes `selftest.wav`, transcribes it and reports the real-time +factor. Run this first. ```bash .\run.ps1 --meter ``` -Live level meter showing the VAD gate, for tuning `--vad-floor`. Add -`--verbose` to log every segment the VAD decides to send. +Level meter with the VAD gate, for tuning `--vad-floor`. `--verbose` logs every +segment the VAD sends. | symptom | cause | |---|---| -| `no speech recognised in N.Ns segment` | music/noise reaching Whisper, or wrong `--lang` | -| nothing at all in the log | VAD never fires — check `--meter`, wrong device | -| `backlog full` | model too slow for real time; use a smaller one | -| red dot in the overlay | overlay lost the WebSocket; it retries automatically | +| `no speech recognised in N.Ns segment` | music/noise, or wrong `--lang` | +| nothing in the log at all | VAD never fires — check `--meter` and device | +| `backlog full` | model too slow; use a smaller one | +| video will not seek | server ignoring Range requests — use `--serve` | +| red dot in overlay/reader | lost the WebSocket; it retries automatically | -Whisper likes to hallucinate stock phrases over silence (`Tekstitys: YLE`, +Whisper hallucinates stock phrases over silence (`Tekstitys: YLE`, `Kiitos kun katsoit!`, `Thanks for watching`). Short results matching those are -filtered out, alongside a no-speech-probability threshold. +filtered, alongside a no-speech-probability threshold. -## Overlay styling +## Tests -Append query parameters to the Browser Source URL: +```bash +.\.venv\Scripts\python.exe tests\test_smoke.py +``` -| param | example | effect | -|---|---|---| -| `size` | `size=52` | font size in px | -| `lines` | `lines=3` | how many lines stay on screen | -| `align` | `align=top` | `top`, `center`, default bottom | -| `fg` | `fg=ffe066` | text colour (hex, no `#`) | -| `accent` | `accent=7dd3fc` | translation line colour | -| `box` | `box=0` | remove the dark background pill | -| `hide` | `hide=0` | never auto-hide | -| `tr` | `tr=1` | show translation lines (with `--translate`) | - -## Text output - -`captions.txt` holds the last few lines for an OBS **Text (GDI+)** source with -*Read from file*, if you would rather avoid browser sources. `captions.log` -keeps the full timestamped transcript of the session, which doubles as stream -notes. - -## How it works - -`livecap.py` pulls 32 ms blocks from a WASAPI device via `soundcard`, tracks a -running noise floor, and marks blocks as speech when they exceed -`max(noise × ratio, floor)`. Speech accumulates into a segment; a segment closes -on `--pause` of trailing silence or at `--max-seg`. Closed segments go to -faster-whisper on a worker thread and are published as `final`. While a segment -is still open, idle worker time is spent transcribing the partial buffer and -publishing lower-confidence `partial` text, so captions appear before the -speaker finishes. Results are broadcast over WebSocket to `overlay.html` and -mirrored to disk. +```bash +.\.venv\Scripts\python.exe tests\test_subtitle.py +``` ## License -MIT — see [LICENSE](LICENSE). - -Uses [faster-whisper](https://github.com/SYSTRAN/faster-whisper) (MIT) and -OpenAI's Whisper models. +MIT — see [LICENSE](LICENSE). Uses +[faster-whisper](https://github.com/SYSTRAN/faster-whisper) (MIT) and OpenAI's +Whisper models. Streaming mode follows the LocalAgreement policy from +Macháček, Dabre & Bojar, *Turning Whisper into Real-Time Transcription System* +(2023). diff --git a/control.html b/control.html index bf7f083..4e391e4 100644 --- a/control.html +++ b/control.html @@ -81,6 +81,7 @@ <div class="row"> <button id="translate">English translation</button> <button id="partials">Live partial text</button> + <button id="test">Send test caption</button> <button id="clear">Clear captions</button> </div> </div> @@ -217,6 +218,7 @@ el("translate").onclick = () => state && send("set_translate", !state.translate); el("partials").onclick = () => state && send("set_partials", !state.partials); el("clear").onclick = () => send("clear"); + el("test").onclick = () => send("test"); document.addEventListener("keydown", (e) => { if (e.target.tagName === "SELECT") return; diff --git a/livecap.py b/livecap.py index 0693dce..2b4201b 100644 --- a/livecap.py +++ b/livecap.py @@ -283,6 +283,8 @@ def http_server(args): log(" " + url) log("Control panel (open in any browser):") log(" http://127.0.0.1:%d/control.html?ws=%d" % (args.http_port, args.ws_port)) + log("Reader for Yomitan mining (open in your real browser):") + log(" http://127.0.0.1:%d/reader.html?ws=%d" % (args.http_port, args.ws_port)) return srv @@ -351,6 +353,13 @@ class Controller: elif cmd == "clear": self.bus.publish({"type": "clear"}) + elif cmd == "test": + # Lets you position the OBS overlay without having to talk. + text = val if isinstance(val, str) and val.strip() else \ + "Testiteksti — проверка — テスト — caption preview" + self.bus.publish({"type": "final", "text": text, "tr": "", + "ts": time.time()}) + elif cmd == "status": self.push_status() @@ -504,6 +513,190 @@ class Transcriber: # ---------------------------------------------------------- segmentation ---- +SENT_END = (".", "!", "?", "…", "。", "!", "?") + + +def norm_word(w): + return re.sub(r"[^\w]", "", w.strip().lower(), flags=re.UNICODE) + + +class StreamDecoder: + """LocalAgreement streaming (Macháček et al.). + + Whisper cannot decode incrementally, so instead we re-decode a growing + buffer and only commit the prefix that two consecutive hypotheses agree + on. Agreement is a good proxy for stability: text that survives another + decode with more audio behind it rarely changes again. This trades CPU + for latency - words appear while someone is still talking, rather than + a whole sentence landing after they stop. + """ + + def __init__(self, tr, bus, args): + self.tr, self.bus, self.args = tr, bus, args + self.native = [] # unconsumed audio at capture rate + self.prev = [] # previous hypothesis, uncommitted part + self.sentence = [] # committed words of the sentence in progress + self.speech = 0.0 # seconds of speech currently buffered + + # ---- audio ------------------------------------------------------- + def add(self, block, voiced, dur): + self.native.append(block) + if voiced: + self.speech += dur + + def buffered_seconds(self, sr): + return sum(len(b) for b in self.native) / sr + + def _audio16(self, sr): + return to_whisper(np.concatenate(self.native), sr) + + def _trim(self, cut_s, sr): + """Drop audio up to cut_s seconds, keeping block boundaries simple.""" + drop = int(cut_s * sr) + merged = np.concatenate(self.native) + merged = merged[min(drop, len(merged)):] + self.native = [merged] if len(merged) else [] + self.speech = max(0.0, self.speech - cut_s) + + # ---- decoding ---------------------------------------------------- + def _hypothesis(self, sr): + a = self.args + audio = self._audio16(sr) + # Whisper invents text when handed a very short buffer, and in + # streaming that invention gets committed before real audio arrives. + if len(audio) < int(a.stream_min_audio * TARGET_SR): + return [], audio + segs, _ = self.tr.model.transcribe( + audio, + language=None if a.lang == "auto" else a.lang, + beam_size=1, + temperature=0.0, + condition_on_previous_text=False, + initial_prompt=(self.tr.context or None) if not a.no_context else None, + vad_filter=False, + word_timestamps=True, + no_speech_threshold=0.6, + log_prob_threshold=-1.0, + ) + words = [] + for s in segs: + words.extend(getattr(s, "words", None) or []) + return words, audio + + def step(self, sr): + words, _ = self._hypothesis(sr) + if not words: + return + + k = 0 + while (k < len(words) and k < len(self.prev) + and norm_word(words[k].word) == norm_word(self.prev[k].word) + and norm_word(words[k].word)): + k += 1 + + confirmed, rest = words[:k], words[k:] + if confirmed: + cut = confirmed[-1].end + self.sentence.extend(confirmed) + self._trim(cut, sr) + # Remaining words are now measured against a shorter buffer. + for w in rest: + w.start = max(0.0, w.start - cut) + w.end = max(0.0, w.end - cut) + self.prev = rest + + text = self._text(self.sentence) + tail = self._text(rest) + if confirmed and text and text.strip().endswith(SENT_END): + self.flush() + elif text or tail: + self.bus.publish({"type": "partial", + "text": (text + " " + tail).strip(), + "ts": time.time()}) + + @staticmethod + def _text(words): + return re.sub(r"\s+", " ", "".join(w.word for w in words)).strip() + + def flush(self, drop_audio=False): + """Emit the sentence built so far as a final caption.""" + text = self._text(self.sentence) + self.sentence = [] + if drop_audio: + self.native, self.prev, self.speech = [], [], 0.0 + if not text or looks_hallucinated(text): + if text: + log("dropped: %r" % text) + return + if not self.args.no_context: + self.tr.context = (self.tr.context + " " + text)[-220:] + log("FINAL %s" % text) + self.bus.publish({"type": "final", "text": text, "tr": "", "ts": time.time()}) + + +def stream_segmenter(cap, tr, args, stop, bus): + """Low-latency path: continuous re-decode with LocalAgreement commits.""" + dur = cap.blocksize / cap.sr + noise = 1e-4 + dec = StreamDecoder(tr, bus, args) + silence = 0.0 + last_step = 0.0 + log("streaming mode: committing on agreement every %.1fs" % args.stream_interval) + + while not stop.is_set(): + drained = 0 + while True: + try: + blk = cap.q.get_nowait() + except queue.Empty: + break + drained += 1 + rms = float(np.sqrt(np.mean(blk * blk)) + 1e-12) + if rms < noise: + noise = 0.90 * noise + 0.10 * rms + else: + noise = 0.995 * noise + 0.005 * rms + voiced = rms > max(noise * args.vad_ratio, args.vad_floor) + silence = 0.0 if voiced else silence + dur + if voiced or dec.native: + dec.add(blk, voiced, dur) + + if not drained: + time.sleep(0.02) + + now = time.time() + buffered = dec.buffered_seconds(cap.sr) + + # A long pause ends the sentence: commit whatever is left. + if dec.native and silence >= args.pause: + if dec.speech >= args.min_speech: + dec.step(cap.sr) + if dec.prev: + dec.sentence.extend(dec.prev) + dec.prev = [] + dec.flush(drop_audio=True) + silence = 0.0 + last_step = now + continue + + if buffered >= args.max_seg: + dec.step(cap.sr) + if dec.prev: + dec.sentence.extend(dec.prev) + dec.prev = [] + dec.flush(drop_audio=True) + last_step = now + continue + + if (dec.speech >= args.min_speech + and now - last_step >= args.stream_interval): + last_step = now + try: + dec.step(cap.sr) + except Exception as e: + log("stream decode error:", repr(e)) + + def segmenter(cap, tr, jobs, args, stop): dur = cap.blocksize / cap.sr noise = 1e-4 @@ -663,6 +856,15 @@ def build_parser(): g.add_argument("--no-partials", dest="partials", action="store_false", help="only show finished sentences (lower CPU)") g.add_argument("--partial-every", type=float, default=0.9) + g.add_argument("--stream", action="store_true", + help="LocalAgreement streaming: words appear while someone is " + "still talking instead of after they stop. Costs a lot " + "more CPU - pair it with a smaller --model") + g.add_argument("--stream-interval", type=float, default=0.8, + help="how often to re-decode the buffer in --stream mode") + g.add_argument("--stream-min-audio", type=float, default=1.5, + help="do not decode until this much audio is buffered; " + "shorter buffers make Whisper hallucinate") g = p.add_argument_group("output") g.add_argument("--ws-port", type=int, default=8765) @@ -712,12 +914,17 @@ def main(): loop.run_until_complete(ws_server(bus, args, ctl, ws_stop)) threading.Thread(target=run_loop, daemon=True).start() - threading.Thread(target=tr.worker, args=(jobs, stop_ev, ctl), daemon=True).start() + if not args.stream: + threading.Thread(target=tr.worker, args=(jobs, stop_ev, ctl), + daemon=True).start() try: with Capture(dev, args.loopback) as cap: log("lang=%s model=%s -- Ctrl+C to stop" % (args.lang, args.model)) - segmenter(cap, tr, jobs, args, stop_ev) + if args.stream: + stream_segmenter(cap, tr, args, stop_ev, bus) + else: + segmenter(cap, tr, jobs, args, stop_ev) except KeyboardInterrupt: print() log("stopping") diff --git a/reader.html b/reader.html new file mode 100644 index 0000000..77289f5 --- /dev/null +++ b/reader.html @@ -0,0 +1,232 @@ +<!doctype html> +<html lang="en"> +<head> +<meta charset="utf-8"> +<meta name="viewport" content="width=device-width, initial-scale=1"> +<title>livecap reader</title> +<style> + :root { + --bg: #0e1116; --panel: #161a21; --line: #252b36; + --fg: #edf1f7; --dim: #8b95a7; --accent: #7dd3fc; --pin: #fbbf24; + --size: 30px; + } + * { box-sizing: border-box; } + html, body { height: 100%; margin: 0; } + body { + background: var(--bg); color: var(--fg); + display: flex; flex-direction: column; + font-family: "Inter", "Segoe UI", "Yu Gothic UI", "Meiryo", + "Noto Sans CJK JP", "Noto Sans", system-ui, sans-serif; + } + + header { + display: flex; gap: 10px; align-items: center; flex-wrap: wrap; + padding: 10px 16px; background: var(--panel); + border-bottom: 1px solid var(--line); font-size: 13px; + } + header .grow { flex: 1; } + button { + font: inherit; font-size: 13px; color: var(--fg); cursor: pointer; + background: #1f2531; border: 1px solid var(--line); + border-radius: 6px; padding: 6px 11px; + } + button:hover { background: #29313f; } + button.on { background: var(--accent); border-color: var(--accent); color: #06202c; } + .dot { width: 8px; height: 8px; border-radius: 50%; display: inline-block; } + .dot.up { background: #4ade80; box-shadow: 0 0 7px #4ade80; } + .dot.down { background: #ef4444; box-shadow: 0 0 7px #ef4444; } + .lang { color: var(--dim); } + + /* The mining surface. Everything here must stay real, selectable DOM text + so Yomitan can scan it and lift the surrounding sentence. */ + #log { + flex: 1; overflow-y: auto; padding: 24px 28px 40vh; + display: flex; flex-direction: column; gap: 14px; + } + .line { + font-size: var(--size); line-height: 1.75; + max-width: 34em; padding: 6px 10px; border-radius: 6px; + border-left: 3px solid transparent; + user-select: text; -webkit-user-select: text; + cursor: text; scroll-margin-bottom: 40vh; + } + .line:hover { background: #151a22; border-left-color: var(--line); } + .line.partial { color: var(--dim); border-left-color: var(--pin); } + .line.pinned { background: #1a1f29; border-left-color: var(--pin); } + /* Rendered from a data attribute so the timestamp is not DOM text - + otherwise Yomitan can absorb it into the sentence it mines. */ + .line::before { + content: attr(data-time); + display: block; font-size: 12px; color: var(--dim); + margin-bottom: 2px; user-select: none; font-variant-numeric: tabular-nums; + } + .line .tr { + display: block; font-size: calc(var(--size) * .52); + color: var(--accent); margin-top: 4px; line-height: 1.5; + } + .empty { color: var(--dim); font-size: 15px; padding: 8px 10px; } + + #jump { + position: fixed; right: 26px; bottom: 26px; display: none; + background: var(--accent); color: #06202c; border-color: var(--accent); + font-weight: 600; box-shadow: 0 4px 16px rgba(0,0,0,.5); + } + #jump.show { display: block; } +</style> +</head> +<body> + +<header> + <span id="conn"><span class="dot down"></span></span> + <span class="lang" id="lang">—</span> + <span class="grow"></span> + <button id="smaller" title="Smaller text">A−</button> + <button id="bigger" title="Larger text">A+</button> + <button id="follow" class="on" title="Auto-scroll to newest">Follow</button> + <button id="copy" title="Copy the whole transcript">Copy all</button> + <button id="clear">Clear</button> +</header> + +<div id="log"><div class="empty">Waiting for speech… hover any word with Yomitan to mine it.</div></div> +<button id="jump">↓ New captions</button> + +<script> +(function () { + const q = new URLSearchParams(location.search); + const WS_PORT = parseInt(q.get("ws") || "8765", 10); + + const log = document.getElementById("log"); + const jump = document.getElementById("jump"); + let follow = true; + let partialEl = null; + let size = parseInt(localStorage.getItem("livecap.size") || q.get("size") || "30", 10); + + function applySize() { + document.documentElement.style.setProperty("--size", size + "px"); + try { localStorage.setItem("livecap.size", String(size)); } catch (e) {} + } + applySize(); + + function atBottom() { + return log.scrollHeight - log.scrollTop - log.clientHeight < 80; + } + + function scrollDown() { + log.scrollTop = log.scrollHeight; + jump.classList.remove("show"); + } + + log.addEventListener("scroll", () => { + if (!follow) return; + if (!atBottom()) jump.classList.add("show"); + else jump.classList.remove("show"); + }); + + function stamp(ts) { + const d = ts ? new Date(ts * 1000) : new Date(); + return d.toTimeString().slice(0, 8); + } + + function clearEmpty() { + const e = log.querySelector(".empty"); + if (e) e.remove(); + } + + function makeLine(m, isPartial) { + const div = document.createElement("div"); + div.className = "line" + (isPartial ? " partial" : ""); + + div.dataset.time = stamp(m.ts) + (isPartial ? " · …" : ""); + + // Sentence text as a bare text node: Yomitan reads this directly. + div.appendChild(document.createTextNode(m.text)); + + if (m.tr) { + const tr = document.createElement("span"); + tr.className = "tr"; + tr.textContent = m.tr; + div.appendChild(tr); + } + + div.addEventListener("dblclick", () => { + div.classList.toggle("pinned"); + }); + return div; + } + + function onFinal(m) { + clearEmpty(); + if (partialEl) { partialEl.remove(); partialEl = null; } + const wasBottom = atBottom(); + log.appendChild(makeLine(m, false)); + while (log.children.length > 500) log.removeChild(log.firstChild); + if (follow && wasBottom) scrollDown(); + else if (follow) jump.classList.add("show"); + } + + function onPartial(m) { + clearEmpty(); + const wasBottom = atBottom(); + if (partialEl) partialEl.remove(); + partialEl = makeLine(m, true); + log.appendChild(partialEl); + if (follow && wasBottom) scrollDown(); + } + + document.getElementById("bigger").onclick = () => { size = Math.min(72, size + 3); applySize(); }; + document.getElementById("smaller").onclick = () => { size = Math.max(14, size - 3); applySize(); }; + document.getElementById("follow").onclick = (e) => { + follow = !follow; + e.target.className = follow ? "on" : ""; + if (follow) scrollDown(); + }; + document.getElementById("clear").onclick = () => { + log.innerHTML = '<div class="empty">Cleared.</div>'; + partialEl = null; + }; + document.getElementById("copy").onclick = async () => { + const text = [...log.querySelectorAll(".line")] + .filter((n) => !n.classList.contains("partial")) + .map((n) => [...n.childNodes].filter((c) => c.nodeType === 3) + .map((c) => c.textContent).join("")) + .join("\n"); + try { + await navigator.clipboard.writeText(text); + const b = document.getElementById("copy"); + b.textContent = "Copied"; + setTimeout(() => (b.textContent = "Copy all"), 1200); + } catch (e) {} + }; + jump.onclick = scrollDown; + + function connect() { + const ws = new WebSocket("ws://127.0.0.1:" + WS_PORT); + const conn = document.getElementById("conn"); + + ws.onopen = () => { + conn.innerHTML = '<span class="dot up"></span>'; + ws.send(JSON.stringify({ cmd: "status" })); + }; + ws.onclose = () => { + conn.innerHTML = '<span class="dot down"></span>'; + setTimeout(connect, 1500); + }; + ws.onerror = () => ws.close(); + ws.onmessage = (ev) => { + let m; + try { m = JSON.parse(ev.data); } catch (e) { return; } + if (m.type === "final") onFinal(m); + else if (m.type === "partial") onPartial(m); + else if (m.type === "status") { + document.getElementById("lang").textContent = m.lang + " · " + m.model; + } else if (m.type === "clear") { + if (partialEl) { partialEl.remove(); partialEl = null; } + } + }; + } + + connect(); +})(); +</script> +</body> +</html> diff --git a/requirements.txt b/requirements.txt index 9965d4b..9f5c16e 100644 --- a/requirements.txt +++ b/requirements.txt @@ -2,4 +2,5 @@ faster-whisper>=1.1.0 soundcard>=0.4.3 numpy>=1.26 scipy>=1.11 +pillow>=10.0 websockets>=12.0 diff --git a/subtitle.py b/subtitle.py new file mode 100644 index 0000000..91aedb7 --- /dev/null +++ b/subtitle.py @@ -0,0 +1,816 @@ +#!/usr/bin/env python3 +""" +subtitle.py - subtitle recordings and replay-buffer clips. + +Latency does not matter here, so this uses a much larger model than the live +captioner can afford and produces far better text, especially for Finnish, +Russian and Japanese. + + python subtitle.py clip.mp4 + python subtitle.py *.mkv --lang ru + python subtitle.py --watch "C:\\Users\\me\\Videos" # auto-subtitle new clips + python subtitle.py clip.mp4 --format vtt --translate + +Writes clip.srt next to clip.mp4. OBS, VLC, YouTube and Premiere all read it. +No ffmpeg needed - audio is decoded through PyAV, which ships with +faster-whisper. +""" + +import argparse +import os +import re +import sys +import threading +import time +from http.server import SimpleHTTPRequestHandler, ThreadingHTTPServer +from pathlib import Path + +VIDEO_EXT = {".mp4", ".mkv", ".mov", ".flv", ".webm", ".avi", ".ts", + ".m4a", ".mp3", ".wav", ".opus", ".flac", ".ogg"} + +for _s in (sys.stdout, sys.stderr): + try: + _s.reconfigure(encoding="utf-8", errors="replace") + except (AttributeError, ValueError): + pass + + +def log(*a): + print("[" + time.strftime("%H:%M:%S") + "]", *a, flush=True) + + +# ------------------------------------------------------------------ cues ---- + +def ts(seconds, sep=","): + if seconds < 0: + seconds = 0.0 + ms = int(round(seconds * 1000)) + h, ms = divmod(ms, 3600000) + m, ms = divmod(ms, 60000) + s, ms = divmod(ms, 1000) + return "%02d:%02d:%02d%s%03d" % (h, m, s, sep, ms) + + +def wrap(text, width): + """Split into at most two balanced lines, the way subtitles are normally set.""" + text = " ".join(text.split()) + if len(text) <= width: + return text + words = text.split() + if len(words) < 2: + return text + + fits, over = None, None + for i in range(1, len(words)): + a, b = " ".join(words[:i]), " ".join(words[i:]) + longest, balance = max(len(a), len(b)), abs(len(a) - len(b)) + if longest <= width: + if fits is None or balance < fits[0]: + fits = (balance, a, b) + # Fallback for text with no split that fits: overflow as little as + # possible rather than emitting one very long line. + if over is None or (longest, balance) < (over[0], over[1]): + over = (longest, balance, a, b) + + if fits: + return fits[1] + "\n" + fits[2] + return over[2] + "\n" + over[3] + + +SENT_END = (".", "!", "?", "…", "。", "!", "?", "؟") + + +class _Word: + """Stand-in when Whisper returns a segment without word timings.""" + + def __init__(self, start, end, word): + self.start, self.end, self.word = start, end, word + + +def collect_words(segments): + words = [] + for seg in segments: + ws = list(getattr(seg, "words", None) or []) + if ws: + words.extend(ws) + elif seg.text.strip(): + words.append(_Word(seg.start, seg.end, seg.text.strip() + " ")) + return words + + +def group_sentences(words, max_dur, max_gap): + """Split on sentence endings first - this is the boundary that matters. + + Long pauses and a hard duration cap act only as fallbacks, so a cue never + starts mid-sentence just because a character budget ran out. + """ + out, cur = [], [] + for w in words: + if cur: + gap = w.start - cur[-1].end + dur = w.end - cur[0].start + if gap > max_gap or dur > max_dur * 2.5: + out.append(cur) + cur = [] + cur.append(w) + if w.word.strip().endswith(SENT_END): + out.append(cur) + cur = [] + if cur: + out.append(cur) + return [g for g in out if "".join(x.word for x in g).strip()] + + +def text_of(group): + return " ".join("".join(w.word for w in group).split()) + + +def build_cues(segments, max_chars, max_dur, max_gap): + """Sentences first, then subdivide any that are too long to display.""" + cues = [] + for group in group_sentences(collect_words(segments), max_dur, max_gap): + chunks, chunk = [], [] + for w in group: + if chunk: + chars = sum(len(x.word) for x in chunk) + len(w.word) + dur = w.end - chunk[0].start + if chars > max_chars or dur > max_dur: + chunks.append(chunk) + chunk = [] + chunk.append(w) + if chunk: + chunks.append(chunk) + + # Greedy filling can strand a word or two on the last line; fold a + # runt back into its predecessor rather than flashing it alone. + if len(chunks) > 1 and len("".join(w.word for w in chunks[-1]).strip()) < 16: + chunks[-2].extend(chunks.pop()) + + for c in chunks: + cues.append((c[0].start, c[-1].end, text_of(c))) + return cues + + +def write_srt(cues, path, width): + with open(path, "w", encoding="utf-8") as fh: + for i, (start, end, text) in enumerate(cues, 1): + fh.write("%d\n%s --> %s\n%s\n\n" + % (i, ts(start), ts(end), wrap(text, width))) + + +def write_vtt(cues, path, width): + with open(path, "w", encoding="utf-8") as fh: + fh.write("WEBVTT\n\n") + for start, end, text in cues: + fh.write("%s --> %s\n%s\n\n" + % (ts(start, "."), ts(end, "."), wrap(text, width))) + + +def write_txt(cues, path, width): + with open(path, "w", encoding="utf-8") as fh: + for _, _, text in cues: + fh.write(text + "\n") + + +WRITERS = {"srt": write_srt, "vtt": write_vtt, "txt": write_txt} + + +# ------------------------------------------------------------------ anki ---- + +def anki_media_dir(): + """Anki's collection.media for the default profile, if we can find it.""" + base = Path(os.environ.get("APPDATA", "")) / "Anki2" + if not base.is_dir(): + return None + for profile in sorted(base.iterdir()): + media = profile / "collection.media" + if media.is_dir(): + return media + return None + + +def slice_audio(audio, sr, start, end, pad): + lo = max(0, int((start - pad) * sr)) + hi = min(len(audio), int((end + pad) * sr)) + return audio[lo:hi] + + +def write_audio_clip(samples, sr, path): + """mp3 when PyAV has an encoder for it, otherwise wav.""" + import numpy as np + + pcm = np.clip(samples, -1.0, 1.0) + if path.suffix == ".mp3": + try: + import av + + with av.open(str(path), "w") as container: + stream = container.add_stream("mp3", rate=sr) + stream.layout = "mono" + frame = av.AudioFrame.from_ndarray( + (pcm * 32767).astype(np.int16).reshape(1, -1), + format="s16", layout="mono") + frame.rate = sr + for packet in stream.encode(frame): + container.mux(packet) + for packet in stream.encode(None): + container.mux(packet) + return path + except Exception: + path = path.with_suffix(".wav") + + import wave + + with wave.open(str(path), "wb") as w: + w.setnchannels(1) + w.setsampwidth(2) + w.setframerate(sr) + w.writeframes((pcm * 32767).astype(np.int16).tobytes()) + return path + + +def grab_frames(video, times, out_paths, width): + """Best-effort screenshots. Returns the paths actually written.""" + try: + import av + except ImportError: + return {} + written = {} + try: + with av.open(str(video)) as container: + if not container.streams.video: + return {} + vs = container.streams.video[0] + vs.thread_type = "AUTO" + for i, (t, dest) in enumerate(zip(times, out_paths)): + try: + container.seek(int(t / vs.time_base), stream=vs) + for frame in container.decode(vs): + if frame.time is None or frame.time + 0.001 < t: + continue + img = frame.to_image() + if width and img.width > width: + img = img.resize((width, max(1, round(img.height * width / img.width)))) + img.save(str(dest), quality=82) + written[i] = dest + break + except Exception: + continue + except Exception: + return written + return written + + +def esc(text): + return text.replace("\t", " ").replace("\n", " ").strip() + + +# ----------------------------------------------------------- mining page ---- + +MINE_PAGE = """<!doctype html> +<html lang="__LANG__"> +<head> +<meta charset="utf-8"> +<title>__TITLE__</title> +<style> + :root { --bg:#0e1116; --panel:#161a21; --line:#252b36; --fg:#edf1f7; + --dim:#8b95a7; --accent:#7dd3fc; --size:26px; } + * { box-sizing:border-box; } + html,body { height:100%; margin:0; } + body { background:var(--bg); color:var(--fg); display:flex; flex-direction:column; + font-family:"Inter","Segoe UI","Yu Gothic UI","Meiryo","Noto Sans CJK JP", + "Noto Sans",system-ui,sans-serif; } + header { display:flex; gap:10px; align-items:center; flex-wrap:wrap; padding:10px 16px; + background:var(--panel); border-bottom:1px solid var(--line); font-size:13px; } + header .grow { flex:1; } + button { font:inherit; font-size:13px; color:var(--fg); cursor:pointer; background:#1f2531; + border:1px solid var(--line); border-radius:6px; padding:6px 11px; } + button:hover { background:#29313f; } + button.on { background:var(--accent); border-color:var(--accent); color:#06202c; } + main { flex:1; display:flex; min-height:0; flex-wrap:wrap; } + #left { flex:1 1 460px; min-width:320px; display:flex; flex-direction:column; + border-right:1px solid var(--line); } + video { width:100%; background:#000; max-height:52vh; } + #cur { padding:16px 20px; font-size:var(--size); line-height:1.7; min-height:3em; + user-select:text; border-top:1px solid var(--line); } + #list { flex:1 1 380px; min-width:300px; overflow-y:auto; padding:14px 18px 30vh; } + .cue { padding:8px 11px; border-radius:6px; border-left:3px solid transparent; + font-size:calc(var(--size)*.82); line-height:1.7; user-select:text; margin-bottom:8px; } + .cue::before { content:attr(data-time); display:block; font-size:11px; color:var(--dim); + user-select:none; margin-bottom:2px; font-variant-numeric:tabular-nums; } + .cue:hover { background:#151a22; } + .cue.active { background:#1a222e; border-left-color:var(--accent); } + .jump { float:right; font-size:11px; color:var(--dim); cursor:pointer; user-select:none; } +</style> +</head> +<body> +<header> + <b>__TITLE__</b> + <span class="grow"></span> + <button id="smaller">A-</button> + <button id="bigger">A+</button> + <button id="follow" class="on">Follow</button> + <button id="loop">Loop cue</button> + <button id="copy">Copy all</button> +</header> +<main> + <div id="left"> + <video id="v" src="__VIDEO__" controls preload="metadata"></video> + <div id="cur"></div> + </div> + <div id="list"></div> +</main> +<script> +const CUES = __CUES__; +const v = document.getElementById("v"); +const list = document.getElementById("list"); +const cur = document.getElementById("cur"); +let follow = true, looping = false, active = -1; + +function fmt(t) { + const m = Math.floor(t / 60), s = Math.floor(t % 60); + return (m < 10 ? "0" : "") + m + ":" + (s < 10 ? "0" : "") + s; +} + +CUES.forEach((c, i) => { + const d = document.createElement("div"); + d.className = "cue"; + d.dataset.time = fmt(c.start); + const j = document.createElement("span"); + j.className = "jump"; + j.textContent = "play"; + j.onclick = (e) => { e.stopPropagation(); v.currentTime = c.start; v.play(); }; + d.appendChild(j); + d.appendChild(document.createTextNode(c.text)); + d.onclick = () => { v.currentTime = c.start; }; + list.appendChild(d); +}); + +function setActive(i) { + if (i === active) return; + const prev = list.children[active]; + if (prev) prev.classList.remove("active"); + active = i; + const el = list.children[i]; + cur.textContent = ""; + if (!el) return; + el.classList.add("active"); + cur.appendChild(document.createTextNode(CUES[i].text)); + if (follow) el.scrollIntoView({ block: "center", behavior: "smooth" }); +} + +function syncNow() { + const t = v.currentTime; + if (looping && active >= 0 && !v.paused) { + const c = CUES[active]; + if (t > c.end + 0.05) { v.currentTime = c.start; return; } + } + let i = -1; + for (let k = 0; k < CUES.length; k++) { + if (t >= CUES[k].start - 0.05 && t <= CUES[k].end + 0.35) { i = k; break; } + } + // While scrubbing between cues, keep showing the one just passed. + if (i < 0) { + for (let k = CUES.length - 1; k >= 0; k--) { + if (t >= CUES[k].start) { i = k; break; } + } + } + if (i >= 0) setActive(i); +} + +// timeupdate only fires during playback; seeked covers scrubbing while paused, +// which is most of what mining actually involves. +["timeupdate", "seeked", "loadedmetadata", "play"].forEach( + (ev) => v.addEventListener(ev, syncNow)); + +document.getElementById("bigger").onclick = () => bump(3); +document.getElementById("smaller").onclick = () => bump(-3); +function bump(d) { + const s = parseInt(getComputedStyle(document.documentElement) + .getPropertyValue("--size"), 10) + d; + document.documentElement.style.setProperty("--size", + Math.max(14, Math.min(60, s)) + "px"); +} +document.getElementById("follow").onclick = (e) => { + follow = !follow; e.target.className = follow ? "on" : ""; +}; +document.getElementById("loop").onclick = (e) => { + looping = !looping; e.target.className = looping ? "on" : ""; +}; +document.getElementById("copy").onclick = async () => { + try { + await navigator.clipboard.writeText(CUES.map((c) => c.text).join("\\n")); + const b = document.getElementById("copy"); + b.textContent = "Copied"; setTimeout(() => b.textContent = "Copy all", 1200); + } catch (e) {} +}; +document.addEventListener("keydown", (e) => { + if (e.key === " ") { e.preventDefault(); v.paused ? v.play() : v.pause(); } + if (e.key === "ArrowLeft" && active > 0) v.currentTime = CUES[active - 1].start; + if (e.key === "ArrowRight" && active < CUES.length - 1) v.currentTime = CUES[active + 1].start; +}); +</script> +</body> +</html> +""" + + +def write_mine_page(path, cues, lang): + """A self-contained page: the clip plus selectable, synced subtitles. + + Yomitan mines from real DOM text, so the value is in the subtitles being + hoverable HTML rather than pixels burned into the video. + """ + import json + + data = [{"start": round(s, 3), "end": round(e, 3), "text": t} for s, e, t in cues] + html = (MINE_PAGE + .replace("__CUES__", json.dumps(data, ensure_ascii=False)) + .replace("__VIDEO__", path.name.replace('"', "%22")) + .replace("__TITLE__", path.stem.replace("<", "<")) + .replace("__LANG__", "ja" if lang == "ja" else (lang or "en"))) + out = path.with_suffix(".html") + out.write_text(html, encoding="utf-8") + return out + + +def export_anki(path, cues, translations, args, log=log): + """One row per sentence: text, translation, audio, screenshot.""" + from faster_whisper.audio import decode_audio + + media = Path(args.anki_media) if args.anki_media else path.parent / (path.stem + "_media") + media.mkdir(parents=True, exist_ok=True) + + sr = 24000 + try: + audio = decode_audio(str(path), sampling_rate=sr) + except Exception as e: + log("cannot re-decode audio for cards: %s" % e) + return None + + stem = "".join(c if (c.isalnum() or c in "-_") else "_" for c in path.stem)[:48] + audio_names, image_names = [], [] + + for i, (start, end, _) in enumerate(cues, 1): + clip = slice_audio(audio, sr, start, end, args.anki_pad) + dest = write_audio_clip(clip, sr, media / ("%s_%04d.mp3" % (stem, i))) + audio_names.append(dest.name) + + if args.anki_images: + mids = [(s + e) / 2 for s, e, _ in cues] + dests = [media / ("%s_%04d.jpg" % (stem, i)) for i in range(1, len(cues) + 1)] + got = grab_frames(path, mids, dests, args.anki_image_width) + image_names = [dests[i].name if i in got else "" for i in range(len(cues))] + if not got: + log("no screenshots (PyAV/Pillow could not decode video frames)") + else: + image_names = [""] * len(cues) + + tsv = path.with_suffix(".anki.tsv") + with open(tsv, "w", encoding="utf-8", newline="") as fh: + for i, (start, end, text) in enumerate(cues): + row = [ + esc(text), + esc(translations[i]) if translations else "", + "[sound:%s]" % audio_names[i], + '<img src="%s">' % image_names[i] if image_names[i] else "", + esc(path.name), + ts(start), + ] + fh.write("\t".join(row) + "\n") + + log("%d cards -> %s" % (len(cues), tsv.name)) + log(" media -> %s" % media) + if not args.anki_media: + target = anki_media_dir() + if target: + log(" copy the media files into: %s" % target) + else: + log(" copy the media files into your Anki collection.media folder") + return tsv + + +# ------------------------------------------------------------ processing ---- + +class Subtitler: + def __init__(self, args): + from faster_whisper import WhisperModel + + self.args = args + log("loading %r on %s/%s ..." % (args.model, args.device, args.compute)) + t0 = time.time() + self.model = WhisperModel(args.model, device=args.device, + compute_type=args.compute, + cpu_threads=args.threads, num_workers=1) + log("model ready in %.1fs" % (time.time() - t0)) + + def run(self, path): + from faster_whisper.audio import decode_audio + + a = self.args + out = path.with_suffix("." + a.format) + if out.exists() and not a.overwrite: + log("skip (exists): %s" % out.name) + return out + + try: + audio = decode_audio(str(path), sampling_rate=16000) + except Exception as e: + log("cannot decode %s: %s" % (path.name, e)) + return None + dur = len(audio) / 16000 + if dur < 0.2: + log("skip (no audio): %s" % path.name) + return None + + t0 = time.time() + segments, info = self.model.transcribe( + audio, + language=None if a.lang == "auto" else a.lang, + task="translate" if a.translate else "transcribe", + beam_size=a.beam, + temperature=[0.0, 0.2, 0.4, 0.6, 0.8, 1.0], + condition_on_previous_text=True, + vad_filter=True, + word_timestamps=True, + no_speech_threshold=0.6, + log_prob_threshold=-1.0, + ) + segments = list(segments) + took = time.time() - t0 + + if a.anki: + # Cards want whole sentences; subtitles want display-sized chunks. + groups = group_sentences(collect_words(segments), a.max_dur, a.max_gap) + card_cues = [(g[0].start, g[-1].end, text_of(g)) for g in groups] + else: + card_cues = [] + + cues = build_cues(segments, a.max_chars, a.max_dur, a.max_gap) + if not cues: + log("no speech found in %s" % path.name) + return None + + WRITERS[a.format](cues, out, a.width) + log("%s -> %s (%d cues, %s, %.0fs audio in %.0fs)" + % (path.name, out.name, len(cues), + getattr(info, "language", a.lang), dur, took)) + + if a.mine: + page = write_mine_page(path, cues, getattr(info, "language", a.lang)) + log("mining page -> %s" % page.name) + + if a.anki: + translations = self.translate_cues(audio, card_cues) if a.anki_translate else None + export_anki(path, card_cues, translations, a) + return out + + def translate_cues(self, audio, cues): + """English for the back of each card, translated per sentence.""" + out = [] + log("translating %d sentences for card backs..." % len(cues)) + for start, end, _ in cues: + lo, hi = max(0, int(start * 16000)), min(len(audio), int(end * 16000)) + clip = audio[lo:hi] + if len(clip) < 1600: + out.append("") + continue + try: + segs, _ = self.model.transcribe( + clip, language=None if self.args.lang == "auto" else self.args.lang, + task="translate", beam_size=self.args.beam, temperature=0.0, + condition_on_previous_text=False, vad_filter=False, + without_timestamps=True) + out.append(" ".join(s.text.strip() for s in segs).strip()) + except Exception: + out.append("") + return out + + +class _Limited: + """Feeds copyfile only the bytes belonging to the requested range.""" + + def __init__(self, fh, remaining): + self.fh, self.remaining = fh, remaining + + def read(self, size=-1): + if self.remaining <= 0: + return b"" + if size is None or size < 0: + size = self.remaining + data = self.fh.read(min(size, self.remaining)) + self.remaining -= len(data) + return data + + def close(self): + self.fh.close() + + +class RangeHandler(SimpleHTTPRequestHandler): + """Static files with HTTP Range support. + + Browsers refuse to seek in a <video> unless the server answers range + requests, and Python's stock handler does not - so scrubbing a clip, which + is most of what mining involves, would silently not work. + """ + + def log_message(self, *a): + pass + + def end_headers(self): + self.send_header("Accept-Ranges", "bytes") + SimpleHTTPRequestHandler.end_headers(self) + + def send_head(self): + rng = self.headers.get("Range") + if not rng: + return SimpleHTTPRequestHandler.send_head(self) + + path = self.translate_path(self.path) + if os.path.isdir(path): + return SimpleHTTPRequestHandler.send_head(self) + try: + fh = open(path, "rb") + except OSError: + self.send_error(404, "File not found") + return None + + size = os.fstat(fh.fileno()).st_size + m = re.match(r"bytes=(\d*)-(\d*)\s*$", rng.strip()) + if not m or (not m.group(1) and not m.group(2)): + fh.close() + self.send_error(400, "Malformed Range") + return None + + if not m.group(1): # bytes=-N (last N bytes) + start, end = max(0, size - int(m.group(2))), size - 1 + else: + start = int(m.group(1)) + end = int(m.group(2)) if m.group(2) else size - 1 + end = min(end, size - 1) + + if start >= size or start > end: + fh.close() + self.send_response(416) + self.send_header("Content-Range", "bytes */%d" % size) + self.send_header("Content-Length", "0") + self.end_headers() + return None + + self.send_response(206) + self.send_header("Content-Type", self.guess_type(path)) + self.send_header("Content-Range", "bytes %d-%d/%d" % (start, end, size)) + self.send_header("Content-Length", str(end - start + 1)) + self.end_headers() + fh.seek(start) + return _Limited(fh, end - start + 1) + + +def serve(folder, port): + from functools import partial + + handler = partial(RangeHandler, directory=str(folder)) + srv = ThreadingHTTPServer(("127.0.0.1", port), handler) + threading.Thread(target=srv.serve_forever, daemon=True).start() + log("serving %s at http://127.0.0.1:%d/ (video seeking enabled)" % (folder, port)) + pages = sorted(p.name for p in Path(folder).glob("*.html")) + for name in pages[-10:]: + log(" http://127.0.0.1:%d/%s" % (port, name)) + return srv + + +def stable(path, checks=3, delay=1.0): + """Wait until a file stops growing, so we do not read a half-written clip.""" + last = -1 + for _ in range(120): + try: + size = path.stat().st_size + except OSError: + return False + if size == last: + checks -= 1 + if checks <= 0: + return size > 0 + else: + checks = 3 + last = size + time.sleep(delay) + return False + + +def watch(sub, folder, args): + folder = Path(folder) + if not folder.is_dir(): + raise SystemExit("not a folder: %s" % folder) + seen = {p.resolve() for p in folder.iterdir() + if p.suffix.lower() in VIDEO_EXT} + log("watching %s for new clips (%d already present) - Ctrl+C to stop" + % (folder, len(seen))) + while True: + try: + for p in sorted(folder.iterdir()): + if p.suffix.lower() not in VIDEO_EXT: + continue + r = p.resolve() + if r in seen: + continue + seen.add(r) + log("new clip: %s" % p.name) + if stable(p): + sub.run(p) + else: + log("gave up waiting for %s to finish writing" % p.name) + time.sleep(args.poll) + except KeyboardInterrupt: + print() + log("stopped") + return + + +def main(): + p = argparse.ArgumentParser( + description="Subtitle recordings and replay clips", + formatter_class=argparse.ArgumentDefaultsHelpFormatter) + p.add_argument("files", nargs="*", help="video/audio files to subtitle") + p.add_argument("--watch", metavar="FOLDER", + help="watch a folder and subtitle clips as they appear") + p.add_argument("--poll", type=float, default=3.0, help="watch interval (s)") + + p.add_argument("--model", default="large-v3-turbo", + help="quality matters more than speed here") + p.add_argument("--lang", default="fi", help="language code, or 'auto'") + p.add_argument("--device", default="cpu", choices=["cpu", "cuda"]) + p.add_argument("--compute", default="int8") + p.add_argument("--threads", type=int, default=max(2, (os.cpu_count() or 8) - 2)) + p.add_argument("--beam", type=int, default=5) + p.add_argument("--translate", action="store_true", + help="write English subtitles instead of the original language") + + p.add_argument("--format", default="srt", choices=sorted(WRITERS)) + p.add_argument("--width", type=int, default=42, help="max characters per line") + p.add_argument("--max-chars", type=int, default=84, help="max characters per cue") + p.add_argument("--max-dur", type=float, default=6.0, help="max seconds per cue") + p.add_argument("--max-gap", type=float, default=0.8, + help="silence (s) that forces a new cue") + p.add_argument("--overwrite", action="store_true", + help="re-subtitle files that already have output") + + g = p.add_argument_group("mining") + g.add_argument("--mine", dest="mine", action="store_true", default=True, + help="write a browser page with the clip and hoverable subtitles") + g.add_argument("--no-mine", dest="mine", action="store_false") + g.add_argument("--serve", nargs="?", type=int, const=8778, default=None, + metavar="PORT", + help="serve the mining pages over http with video seeking, " + "so Yomitan works without file-URL permissions") + + g = p.add_argument_group("anki cards") + g.add_argument("--anki", action="store_true", + help="also export one card per sentence, with clipped audio") + g.add_argument("--anki-translate", action="store_true", + help="add an English translation to each card") + g.add_argument("--anki-images", dest="anki_images", action="store_true", default=True, + help="grab a screenshot per card") + g.add_argument("--no-anki-images", dest="anki_images", action="store_false") + g.add_argument("--anki-image-width", type=int, default=640) + g.add_argument("--anki-pad", type=float, default=0.25, + help="seconds of padding around each audio clip") + g.add_argument("--anki-media", default=None, + help="write media straight into Anki's collection.media") + + args = p.parse_args() + if not args.files and not args.watch: + p.error("give some files, or --watch a folder") + + sub = Subtitler(args) + + done = [] + for pattern in args.files: + path = Path(pattern) + matches = [path] if path.exists() else sorted(Path().glob(pattern)) + if not matches: + log("no such file: %s" % pattern) + for m in matches: + sub.run(m) + done.append(m) + + srv = None + if args.serve: + folder = Path(args.watch) if args.watch else ( + done[0].parent if done else Path(".")) + srv = serve(folder.resolve(), args.serve) + + if args.watch: + watch(sub, args.watch, args) + elif srv is not None: + log("Ctrl+C to stop serving") + try: + while True: + time.sleep(3600) + except KeyboardInterrupt: + print() + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/tests/test_subtitle.py b/tests/test_subtitle.py new file mode 100644 index 0000000..9e5d1e3 --- /dev/null +++ b/tests/test_subtitle.py @@ -0,0 +1,202 @@ +"""Tests for the offline subtitler. Run with: python tests\\test_subtitle.py""" + +import sys +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) + +import subtitle + + +class W: + """Minimal stand-in for a faster-whisper word.""" + + def __init__(self, start, end, word): + self.start, self.end, self.word = start, end, word + + +class Seg: + def __init__(self, words, text=None, start=0.0, end=0.0): + self.words = words + self.text = text if text is not None else "".join(w.word for w in words) + self.start, self.end = start, end + + +def words(spec, t0=0.0, step=0.4): + out, t = [], t0 + for token in spec.split(" "): + out.append(W(t, t + step, token + " ")) + t += step + return out + + +def test_timestamp_format(): + assert subtitle.ts(0) == "00:00:00,000" + assert subtitle.ts(1.5) == "00:00:01,500" + assert subtitle.ts(3661.25) == "01:01:01,250" + assert subtitle.ts(-5) == "00:00:00,000" + assert subtitle.ts(1.5, ".") == "00:00:01.500" + + +def test_wrap_balances_two_lines(): + out = subtitle.wrap("aaa bbb ccc ddd", 8) + assert out == "aaa bbb\nccc ddd", repr(out) + + +def test_wrap_leaves_short_text_alone(): + assert subtitle.wrap("short", 42) == "short" + + +def test_wrap_falls_back_when_nothing_fits(): + """A long sentence with no split under the limit still gets two lines.""" + text = ("Ollaan taas sen verran syrjaisilla seuduilla ja " + "harvakseltaan kuljetuilla seuduilla.") + out = subtitle.wrap(text, 42) + lines = out.split("\n") + assert len(lines) == 2, out + assert max(len(x) for x in lines) < len(text), out + + +def test_wrap_single_word_cannot_split(): + assert "\n" not in subtitle.wrap("Rindfleischetikettierungsgesetz", 5) + + +def test_sentences_split_on_terminators(): + ws = words("Yksi kaksi.") + words("Kolme nelja.", t0=2.0) + groups = subtitle.group_sentences(ws, max_dur=6.0, max_gap=0.8) + assert len(groups) == 2, [subtitle.text_of(g) for g in groups] + assert subtitle.text_of(groups[0]) == "Yksi kaksi." + + +def test_sentences_split_on_long_gap(): + ws = words("yksi kaksi") + words("kolme nelja", t0=9.0) + groups = subtitle.group_sentences(ws, max_dur=6.0, max_gap=0.8) + assert len(groups) == 2 + + +def test_cue_never_starts_mid_sentence_orphan(): + """A trailing runt is folded back rather than shown alone.""" + ws = words("aaaa bbbb cccc dddd eeee ffff gggg hhhh iiii jjjj kkkk") + cues = subtitle.build_cues([Seg(ws)], max_chars=30, max_dur=99, max_gap=9) + assert cues + for _, _, text in cues: + assert len(text) >= 16 or len(cues) == 1, cues + + +def test_segment_without_word_timings_still_produces_a_cue(): + seg = Seg([], text="Ei sanatason aikaleimoja.", start=1.0, end=3.0) + cues = subtitle.build_cues([seg], 84, 6.0, 0.8) + assert len(cues) == 1 + assert cues[0][2] == "Ei sanatason aikaleimoja." + assert cues[0][0] == 1.0 and cues[0][1] == 3.0 + + +def test_srt_roundtrip(tmp=None): + import tempfile + + cues = [(0.0, 1.5, "Ensimmainen."), (2.0, 3.25, "Toinen rivi tassa.")] + with tempfile.TemporaryDirectory() as d: + out = Path(d) / "x.srt" + subtitle.write_srt(cues, out, 42) + text = out.read_text(encoding="utf-8") + assert "1\n00:00:00,000 --> 00:00:01,500\nEnsimmainen." in text, text + assert "2\n00:00:02,000 --> 00:00:03,250" in text, text + assert text.endswith("\n\n") + + +def test_vtt_has_header_and_dot_timestamps(): + import tempfile + + with tempfile.TemporaryDirectory() as d: + out = Path(d) / "x.vtt" + subtitle.write_vtt([(0.0, 1.0, "Moi")], out, 42) + text = out.read_text(encoding="utf-8") + assert text.startswith("WEBVTT") + assert "00:00:00.000 --> 00:00:01.000" in text + + +def test_mine_page_embeds_cues_and_escapes_nothing_odd(): + import json + import tempfile + + cues = [(0.0, 1.0, "今日はいい天気ですね。"), (1.0, 2.0, "Kyllä se tästä.")] + with tempfile.TemporaryDirectory() as d: + video = Path(d) / "clip.mp4" + video.write_bytes(b"") + page = subtitle.write_mine_page(video, cues, "ja") + html = page.read_text(encoding="utf-8") + assert page.name == "clip.html" + assert 'src="clip.mp4"' in html + # Non-Latin text must survive verbatim for Yomitan to scan it. + assert "今日はいい天気ですね。" in html + assert "Kyllä se tästä." in html + assert "__CUES__" not in html and "__VIDEO__" not in html + start = html.index("const CUES = ") + len("const CUES = ") + data = json.loads(html[start:html.index("\n", start)].rstrip(";")) + assert len(data) == 2 and data[0]["text"] == "今日はいい天気ですね。" + + +def test_esc_strips_tabs_and_newlines(): + assert subtitle.esc("a\tb\nc ") == "a b c" + + +def test_range_requests_are_served(): + """Without 206 support a browser will not seek in a <video> at all.""" + import http.client + import tempfile + + payload = bytes(range(256)) * 8 # 2048 bytes + with tempfile.TemporaryDirectory() as d: + (Path(d) / "clip.mp4").write_bytes(payload) + srv = subtitle.serve(Path(d), 0) # port 0 = pick a free one + port = srv.server_address[1] + try: + def req(headers): + c = http.client.HTTPConnection("127.0.0.1", port, timeout=5) + c.request("GET", "/clip.mp4", headers=headers) + r = c.getresponse() + body = r.read() + c.close() + return r, body + + r, body = req({}) + assert r.status == 200, r.status + assert r.getheader("Accept-Ranges") == "bytes" + assert body == payload + + r, body = req({"Range": "bytes=10-19"}) + assert r.status == 206, r.status + assert r.getheader("Content-Range") == "bytes 10-19/2048" + assert body == payload[10:20], body + + r, body = req({"Range": "bytes=2040-"}) + assert r.status == 206 + assert body == payload[2040:] + + r, body = req({"Range": "bytes=-8"}) + assert r.status == 206 + assert body == payload[-8:] + + r, _ = req({"Range": "bytes=99999-"}) + assert r.status == 416, r.status + finally: + srv.shutdown() + srv.server_close() + + +def main(): + tests = [v for k, v in sorted(globals().items()) if k.startswith("test_")] + failed = 0 + for fn in tests: + try: + fn() + print(" PASS %s" % fn.__name__) + except Exception as e: + failed += 1 + print(" FAIL %s: %r" % (fn.__name__, e)) + print("\n%d passed, %d failed" % (len(tests) - failed, failed)) + return 1 if failed else 0 + + +if __name__ == "__main__": + sys.exit(main())