User Tools

Site Tools


whisperx_transcription


Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
whisperx_transcription [2026/10/10 17:18] – created millerjswhisperx_transcription [2026/10/10 17:20] (current) – [Local Speech-to-Text with WhisperX] millerjs
Line 1: Line 1:
-hold+====== Local Speech-to-Text with WhisperX ====== 
 +This is running on a SFF Optiplex 7070 with an i5-9500 (6 cores / 6 threads, AVX2, 32 GB RAM, Debian 13). No GPU, no cloud, just the CPU. A 50-minute recording takes about 48 minutes end-to-end with the large-v3 model in int8 mode. Not fast, but the machine is usually idle so I just kick it off and walk away. At some point I'd like to put a low profile GPU in here to help with accelerating these kinds of tasks, but this works fine for now. 
 + 
 +===== System Prerequisites ===== 
 +Debian 13 (Trixie) ships everything you need in the main repos. FFmpeg handles decoding the audio files (m4a, mp3, wav, whatever), and the rest is standard Python tooling. 
 + 
 +<code> 
 +sudo apt update 
 +sudo apt install -y ffmpeg 
 +sudo apt install -y python3 python3-venv python3-dev python3-pip build-essential pkg-config 
 +</code> 
 + 
 +Verify: 
 + 
 +<code> 
 +$ ffmpeg -version 
 +ffmpeg version 7.1.5-0+deb13u1 
 + 
 +$ python3 --version 
 +Python 3.13.5 
 +</code> 
 + 
 +===== Python Virtual Environment ===== 
 +WhisperX pulls in PyTorch, pyannote, faster-whisper, and a pile of other dependencies — about 2 GB worth. Keep it all isolated in a venv so it doesn't trash your system Python. 
 + 
 +I staged everything in ''~/dev/whisperx'' but you can put it wherever you want. 
 + 
 +<code> 
 +mkdir -p ~/dev/whisperx 
 +cd ~/dev/whisperx 
 +python3 -m venv venv 
 +source venv/bin/activate 
 +pip install --upgrade pip 
 +</code> 
 + 
 +===== Install WhisperX ===== 
 +Still in the activated venv: 
 + 
 +<code> 
 +pip install whisperx 
 +</code> 
 + 
 +This takes a few minutes. When it's done, verify the CLI works: 
 + 
 +<code> 
 +$ whisperx --help 
 +usage: whisperx [-h] [--model MODEL] [--model_cache_only MODEL_CACHE_ONLY] 
 +                [--model_dir MODEL_DIR] [--device DEVICE] 
 +... 
 +</code> 
 + 
 +If ''whisperx'' isn't found on your PATH, use ''python -m whisperx --help'' instead. 
 + 
 +===== Hugging Face Setup (required for diarization) ===== 
 +This is the step people most often miss, and skipping it produces a confusing ''401 Unauthorized'' or ''403 Forbidden'' error later. The pyannote diarization models are gated on Hugging Face — you have to create an account and explicitly accept the licenses before you can download them. 
 + 
 +==== Create an account ==== 
 +Go to [[https://huggingface.co/join|huggingface.co/join]] and sign up. It's free. 
 + 
 +==== Accept the model licenses ==== 
 +Visit each of these two pages and click "Agree and access repository": 
 + 
 +  - [[https://huggingface.co/pyannote/speaker-diarization-3.1|pyannote/speaker-diarization-3.1]] — the main diarization pipeline 
 +  - [[https://huggingface.co/pyannote/segmentation-3.0|pyannote/segmentation-3.0]] — a sub-model used by the above 
 + 
 +You have to accept **both**. If you only accept the first one, you'll get a 403 later. 
 + 
 +==== Generate an access token ==== 
 +  - Go to [[https://huggingface.co/settings/tokens|huggingface.co/settings/tokens]] 
 +  - Click "New token" 
 +  - Name it something like ''whisperx-local'' 
 +  - Type: **Read** (sufficient — you only need to download models) 
 +  - Copy the token — it starts with ''hf_...'' 
 + 
 +==== Make the token available ==== 
 +Add it to your ''~/.bashrc'' so it persists across sessions. Make sure to use ''export'' — without it, you're just setting a shell variable that never makes it into the environment, and you'll spend way too long figuring out why your token "isn't working." 
 + 
 +<code> 
 +echo 'export HF_TOKEN=hf_your_token_here' >> ~/.bashrc 
 +source ~/.bashrc 
 +</code> 
 + 
 +Verify it's actually exported: 
 + 
 +<code> 
 +$ echo "${HF_TOKEN:+set (${HF_TOKEN:0:6}...)}" 
 +set (hf_ZQj...) 
 +</code> 
 + 
 +==== Verify token access ==== 
 +Before you kick off a long transcription, take 5 seconds to verify your token can actually access both models. WhisperX installs ''huggingface_hub'' as a dependency, so you don't need to install anything extra. I wrote a small script for this: 
 + 
 +<code python test_hf_token.py> 
 +import sys, os 
 +from huggingface_hub import HfApi 
 + 
 +token = os.environ["HF_TOKEN"] 
 +api = HfApi(token=token) 
 + 
 +models = [ 
 +    "pyannote/speaker-diarization-3.1", 
 +    "pyannote/segmentation-3.0", 
 +] 
 + 
 +all_ok = True 
 +for m in models: 
 +    try: 
 +        info = api.model_info(m) 
 +        print(f"PASS: {m} (sha={info.sha[:8]})") 
 +    except Exception as e: 
 +        all_ok = False 
 +        print(f"FAIL: {m} -> {e}") 
 + 
 +sys.exit(0 if all_ok else 1) 
 +</code> 
 + 
 +<code> 
 +$ source ~/dev/whisperx/venv/bin/activate 
 +$ python3 test_hf_token.py 
 +PASS: pyannote/speaker-diarization-3.1 (sha=84fd2591) 
 +PASS: pyannote/segmentation-3.0 (sha=e66f3d3b) 
 +</code> 
 + 
 +Do this before running a transcription. It catches token and license issues in seconds instead of after a 20-minute transcription+alignment run that dies at the very end. 
 + 
 +===== The Diarization Problem (and how to fix it) ===== 
 +This is where things got interesting. WhisperX 3.8.6 (with pyannote.audio 4.0.7) has **two separate code paths** that both try to access a model called ''pyannote/speaker-diarization-community-1'' — a newer gated model that I hadn't accepted the license for, and that the setup guides I was following didn't even mention. Both will fail with ''GatedRepoError: 403'' and waste your time. 
 + 
 +==== Layer 1: the --diarize_model flag ==== 
 +WhisperX 3.8.6 defaults to ''pyannote/speaker-diarization-community-1''. The model you actually accepted the license for is ''pyannote/speaker-diarization-3.1''. You have to explicitly tell WhisperX to use the right one: 
 + 
 +<code> 
 +--diarize_model pyannote/speaker-diarization-3.1 
 +</code> 
 + 
 +==== Layer 2: the PLDA patch ==== 
 +Even with the correct ''--diarize_model'' flag, I hit a second 403 error. The speaker-diarization-3.1 pipeline uses ''AgglomerativeClustering'', which doesn't need PLDA (a speaker verification sub-model). But pyannote.audio 4.0.7's ''SpeakerDiarization.__init__'' **unconditionally loads PLDA** from ''pyannote/speaker-diarization-community-1'' anyway, regardless of which clustering method is actually being used. So it tries to download ''plda/xvec_transform.npz'' from a gated repo you don't have access to, and dies. 
 + 
 +The fix is to patch the pipeline to only load PLDA when ''VBxClustering'' is actually used. The file you need to edit is: 
 + 
 +''venv/lib/python3.13/site-packages/pyannote/audio/pipelines/speaker_diarization.py'' 
 + 
 +Find this in ''SpeakerDiarization.__init__'' (around line 230): 
 + 
 +<code python> 
 +        self.plda = plda 
 +        self._plda = get_plda(plda, token=token, cache_dir=cache_dir) 
 + 
 +        self.klustering = clustering 
 +</code> 
 + 
 +Replace it with: 
 + 
 +<code python> 
 +        self.plda = plda 
 +        self.klustering = clustering 
 + 
 +        # Only load PLDA when VBxClustering is used; AgglomerativeClustering doesn't need it. 
 +        # Loading it unconditionally triggers a download from the gated community-1 repo. 
 +        if self.klustering == "VBxClustering": 
 +            self._plda = get_plda(plda, token=token, cache_dir=cache_dir) 
 +        else: 
 +            self._plda = None 
 +</code> 
 + 
 +Then verify the pipeline loads cleanly: 
 + 
 +<code> 
 +$ source venv/bin/activate 
 +$ export HF_TOKEN=$(grep '^export HF_TOKEN=' ~/.bashrc | head -1 | sed 's/^export HF_TOKEN=//') 
 +$ python3 -c "from pyannote.audio import Pipeline; p = Pipeline.from_pretrained('pyannote/speaker-diarization-3.1', token='$HF_TOKEN'); print('OK:', p.klustering, p._plda)" 
 +OK: AgglomerativeClustering None 
 +</code> 
 + 
 +If you see ''AgglomerativeClustering None'', you're good. If you see a 403 error, the patch didn't apply correctly. 
 + 
 +**Note:** This patch lives inside the venv. If you ever recreate the venv or force-reinstall pyannote-audio, you'll need to re-apply it. 
 + 
 +===== Running a Transcription ===== 
 +Once everything is set up and the patch is applied, you're ready to transcribe. Drop your audio file somewhere and run: 
 + 
 +<code> 
 +source ~/dev/whisperx/venv/bin/activate 
 +export HF_TOKEN=$(grep '^export HF_TOKEN=' ~/.bashrc | head -1 | sed 's/^export HF_TOKEN=//') 
 + 
 +whisperx recording.m4a \ 
 +  --model large-v3 \ 
 +  --language en \ 
 +  --diarize \ 
 +  --diarize_model pyannote/speaker-diarization-3.1 \ 
 +  --hf_token "$HF_TOKEN" \ 
 +  --min_speakers 2 \ 
 +  --max_speakers 3 \ 
 +  --compute_type int8 \ 
 +  --output_dir ./transcripts 
 +</code> 
 + 
 +What each flag does: 
 + 
 +^ Flag ^ Purpose ^ 
 +| ''--model large-v3'' | Most accurate Whisper model. Options: tiny, base, small, medium, large-v3. Bigger = more accurate but slower. | 
 +| ''--language en'' | Skip language auto-detection (saves time). Remove if your recordings aren't in English. | 
 +| ''--diarize'' | Enable speaker diarization (the whole point). | 
 +| ''--diarize_model'' | Must be ''pyannote/speaker-diarization-3.1'' — see the section above. | 
 +| ''--hf_token'' | Your Hugging Face token. | 
 +| ''--min_speakers'' / ''--max_speakers'' | Hint the expected speaker count. Helps diarization accuracy for small meetings. | 
 +| ''--compute_type int8'' | Quantize to 8-bit. Significantly reduces memory and speeds up CPU inference. Highly recommended for CPU-only. | 
 +| ''--output_dir'' | Where to write output files. | 
 + 
 +The first time you run with a given model size, WhisperX downloads the model (~3 GB for large-v3). The pyannote models (~100 MB) and the wav2vec2 alignment model (~360 MB) also download on first run. Subsequent runs use cached models from ''~/.cache/huggingface/''. 
 + 
 +===== Output Files ===== 
 +WhisperX creates several files in the output directory: 
 + 
 +^ File ^ Description ^ 
 +| ''recording.srt'' | Subtitle format with speaker labels (most readable) | 
 +| ''recording.vtt'' | WebVTT format | 
 +| ''recording.tsv'' | Tab-separated, one segment per line | 
 +| ''recording.json'' | Full structured JSON with word-level timestamps | 
 +| ''recording.txt'' | Plain text, no timestamps | 
 + 
 +The SRT file looks like this: 
 + 
 +<code> 
 +1 
 +00:03:06,340 --> 00:03:08,121 
 +[SPEAKER_00]: Hey, good morning. 
 + 
 +2 
 +00:03:09,241 --> 00:03:09,882 
 +[SPEAKER_01]: Good morning. 
 + 
 +3 
 +00:03:12,403 --> 00:03:13,443 
 +Early part of the day. 
 +</code> 
 + 
 +Speaker labels appear at the start of each speaker's turn. Segments without a label continue from the previous speaker. 
 + 
 +===== Performance on the i5-9500 ===== 
 +I ran a 50.8-minute meeting recording through the full pipeline. Here's how the time broke down: 
 + 
 +^ Stage ^ Time ^ 
 +| Voice activity detection | ~2 sec | 
 +| Transcription (large-v3, int8) | ~18 min | 
 +| Forced alignment (wav2vec2) | ~3 min | 
 +| Diarization (pyannote 3.1) | ~27 min | 
 +| **Total** | **~48 min** | 
 + 
 +So roughly realtime for a 50-minute file. The diarization step is the surprise — it takes longer than the transcription itself on CPU. The Whisper model is well-optimized via CTranslate2's int8 path, but pyannote's speaker embedding extraction is pure PyTorch with no quantization. 
 + 
 +Memory usage peaked around 7 GB (well within the 32 GB available). CPU was pinned at 390% (about 4 of the 6 cores fully utilized). 
 + 
 +===== Convenience Wrapper ===== 
 +I got tired of typing all those flags, so I made a wrapper script. It activates the venv, checks that ''HF_TOKEN'' is set, and runs WhisperX with my usual defaults. 
 + 
 +<code bash transcribe.sh> 
 +#!/bin/bash 
 +# Usage: transcribe.sh <audio_file> [model_size] [min_speakers] [max_speakers] 
 + 
 +set -euo pipefail 
 + 
 +source "$HOME/dev/whisperx/venv/bin/activate" 
 + 
 +AUDIO_FILE="${1:?Usage: transcribe.sh <audio_file> [model] [min_spk] [max_spk]}" 
 +MODEL="${2:-large-v3}" 
 +MIN_SPK="${3:-2}" 
 +MAX_SPK="${4:-3}" 
 + 
 +if [[ -z "${HF_TOKEN:-}" ]]; then 
 +  echo "ERROR: HF_TOKEN is not set. Export it first:" >&2 
 +  echo "  export HF_TOKEN=hf_your_token_here" >&2 
 +  exit 1 
 +fi 
 + 
 +whisperx "$AUDIO_FILE" \ 
 +  --model "$MODEL" \ 
 +  --language en \ 
 +  --diarize \ 
 +  --diarize_model pyannote/speaker-diarization-3.1 \ 
 +  --hf_token "$HF_TOKEN" \ 
 +  --min_speakers "$MIN_SPK" \ 
 +  --max_speakers "$MAX_SPK" \ 
 +  --compute_type int8 \ 
 +  --output_dir ./transcripts 
 +</code> 
 + 
 +Put it somewhere in your ''$PATH'' and make it executable. Then transcribing is just: 
 + 
 +<code> 
 +$ transcribe.sh meeting.m4a 
 +$ transcribe.sh meeting.m4a medium 2 2   # quick test: medium model, exactly 2 speakers 
 +</code> 
 + 
 +===== Batch Transcription ===== 
 +For running through a pile of recordings, kick it off in the background and walk away: 
 + 
 +<code> 
 +nohup bash -c ' 
 +  for f in ~/Meetings/*.m4a; do 
 +    echo "=== $(date) — Transcribing: $f ===" 
 +    transcribe.sh "$f" 
 +  done 
 +  echo "=== $(date) — All done ===" 
 +' > ~/transcribe.log 2>&1 & 
 + 
 +# Check progress: 
 +tail -f ~/transcribe.log 
 +</code> 
 + 
 +Or use ''tmux'' / ''screen'' if you prefer to keep a live session. 
 + 
 +===== Troubleshooting ===== 
 +^ Problem ^ Fix ^ 
 +| ''401 Unauthorized'' or ''403 Forbidden'' | You didn't accept both HF model licenses, your token is wrong, or you're hitting the community-1 model — see the diarization section above. | 
 +| ''GatedRepoError'' for ''speaker-diarization-community-1'' | Two fixes needed: (1) pass ''--diarize_model pyannote/speaker-diarization-3.1'', and (2) apply the PLDA patch. Both are described above. | 
 +| ''ffmpeg not found'' | ''sudo apt install ffmpeg'' | 
 +| Out of memory / killed | Use ''--compute_type int8'' and/or a smaller model (''medium'' instead of ''large-v3''). With 32 GB RAM this shouldn't happen. | 
 +| Diarization puts everything on one speaker | Try narrowing ''--min_speakers'' / ''--max_speakers'', or check if the recording has a lot of crosstalk. | 
 +| Wrong language detected | Always specify ''--language en'' (or the appropriate code). | 
 +| Slow on CPU | Expected. Use ''--compute_type int8'' and ''medium'' model for speed; reserve ''large-v3'' for when accuracy matters most. | 
 +| ''ModuleNotFoundError'' after opening new terminal | You forgot to activate the venv: ''source ~/dev/whisperx/venv/bin/activate'' (or use the wrapper script which does it for you). | 
 +| Models re-downloading every run | Ensure ''HF_TOKEN'' is exported. Models cache in ''~/.cache/huggingface/'' — check write permissions if issues persist. | 
 +| ''HF_TOKEN'' not available after ''source ~/.bashrc'' | You used ''HF_TOKEN=...'' instead of ''export HF_TOKEN=...'' in ''~/.bashrc''. Without ''export'', it's a shell variable, not an environment variable. | 
whisperx_transcription.1791652690.txt.gz · Last modified: by millerjs

Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki