====== Local Speech-to-Text with WhisperX ====== This is running on a SFF Optiplex 7070 with an i5-9500 (6 cores / 6 threads, AVX2, 32 GB RAM, Debian 13). No GPU, no cloud, just the CPU. A 50-minute recording takes about 48 minutes end-to-end with the large-v3 model in int8 mode. Not fast, but the machine is usually idle so I just kick it off and walk away. At some point I'd like to put a low profile GPU in here to help with accelerating these kinds of tasks, but this works fine for now. ===== System Prerequisites ===== Debian 13 (Trixie) ships everything you need in the main repos. FFmpeg handles decoding the audio files (m4a, mp3, wav, whatever), and the rest is standard Python tooling. sudo apt update sudo apt install -y ffmpeg sudo apt install -y python3 python3-venv python3-dev python3-pip build-essential pkg-config Verify: $ ffmpeg -version ffmpeg version 7.1.5-0+deb13u1 $ python3 --version Python 3.13.5 ===== Python Virtual Environment ===== WhisperX pulls in PyTorch, pyannote, faster-whisper, and a pile of other dependencies — about 2 GB worth. Keep it all isolated in a venv so it doesn't trash your system Python. I staged everything in ''~/dev/whisperx'' but you can put it wherever you want. mkdir -p ~/dev/whisperx cd ~/dev/whisperx python3 -m venv venv source venv/bin/activate pip install --upgrade pip ===== Install WhisperX ===== Still in the activated venv: pip install whisperx This takes a few minutes. When it's done, verify the CLI works: $ whisperx --help usage: whisperx [-h] [--model MODEL] [--model_cache_only MODEL_CACHE_ONLY] [--model_dir MODEL_DIR] [--device DEVICE] ... If ''whisperx'' isn't found on your PATH, use ''python -m whisperx --help'' instead. ===== Hugging Face Setup (required for diarization) ===== This is the step people most often miss, and skipping it produces a confusing ''401 Unauthorized'' or ''403 Forbidden'' error later. The pyannote diarization models are gated on Hugging Face — you have to create an account and explicitly accept the licenses before you can download them. ==== Create an account ==== Go to [[https://huggingface.co/join|huggingface.co/join]] and sign up. It's free. ==== Accept the model licenses ==== Visit each of these two pages and click "Agree and access repository": - [[https://huggingface.co/pyannote/speaker-diarization-3.1|pyannote/speaker-diarization-3.1]] — the main diarization pipeline - [[https://huggingface.co/pyannote/segmentation-3.0|pyannote/segmentation-3.0]] — a sub-model used by the above You have to accept **both**. If you only accept the first one, you'll get a 403 later. ==== Generate an access token ==== - Go to [[https://huggingface.co/settings/tokens|huggingface.co/settings/tokens]] - Click "New token" - Name it something like ''whisperx-local'' - Type: **Read** (sufficient — you only need to download models) - Copy the token — it starts with ''hf_...'' ==== Make the token available ==== Add it to your ''~/.bashrc'' so it persists across sessions. Make sure to use ''export'' — without it, you're just setting a shell variable that never makes it into the environment, and you'll spend way too long figuring out why your token "isn't working." echo 'export HF_TOKEN=hf_your_token_here' >> ~/.bashrc source ~/.bashrc Verify it's actually exported: $ echo "${HF_TOKEN:+set (${HF_TOKEN:0:6}...)}" set (hf_ZQj...) ==== Verify token access ==== Before you kick off a long transcription, take 5 seconds to verify your token can actually access both models. WhisperX installs ''huggingface_hub'' as a dependency, so you don't need to install anything extra. I wrote a small script for this: import sys, os from huggingface_hub import HfApi token = os.environ["HF_TOKEN"] api = HfApi(token=token) models = [ "pyannote/speaker-diarization-3.1", "pyannote/segmentation-3.0", ] all_ok = True for m in models: try: info = api.model_info(m) print(f"PASS: {m} (sha={info.sha[:8]})") except Exception as e: all_ok = False print(f"FAIL: {m} -> {e}") sys.exit(0 if all_ok else 1) $ source ~/dev/whisperx/venv/bin/activate $ python3 test_hf_token.py PASS: pyannote/speaker-diarization-3.1 (sha=84fd2591) PASS: pyannote/segmentation-3.0 (sha=e66f3d3b) Do this before running a transcription. It catches token and license issues in seconds instead of after a 20-minute transcription+alignment run that dies at the very end. ===== The Diarization Problem (and how to fix it) ===== This is where things got interesting. WhisperX 3.8.6 (with pyannote.audio 4.0.7) has **two separate code paths** that both try to access a model called ''pyannote/speaker-diarization-community-1'' — a newer gated model that I hadn't accepted the license for, and that the setup guides I was following didn't even mention. Both will fail with ''GatedRepoError: 403'' and waste your time. ==== Layer 1: the --diarize_model flag ==== WhisperX 3.8.6 defaults to ''pyannote/speaker-diarization-community-1''. The model you actually accepted the license for is ''pyannote/speaker-diarization-3.1''. You have to explicitly tell WhisperX to use the right one: --diarize_model pyannote/speaker-diarization-3.1 ==== Layer 2: the PLDA patch ==== Even with the correct ''--diarize_model'' flag, I hit a second 403 error. The speaker-diarization-3.1 pipeline uses ''AgglomerativeClustering'', which doesn't need PLDA (a speaker verification sub-model). But pyannote.audio 4.0.7's ''SpeakerDiarization.__init__'' **unconditionally loads PLDA** from ''pyannote/speaker-diarization-community-1'' anyway, regardless of which clustering method is actually being used. So it tries to download ''plda/xvec_transform.npz'' from a gated repo you don't have access to, and dies. The fix is to patch the pipeline to only load PLDA when ''VBxClustering'' is actually used. The file you need to edit is: ''venv/lib/python3.13/site-packages/pyannote/audio/pipelines/speaker_diarization.py'' Find this in ''SpeakerDiarization.__init__'' (around line 230): self.plda = plda self._plda = get_plda(plda, token=token, cache_dir=cache_dir) self.klustering = clustering Replace it with: self.plda = plda self.klustering = clustering # Only load PLDA when VBxClustering is used; AgglomerativeClustering doesn't need it. # Loading it unconditionally triggers a download from the gated community-1 repo. if self.klustering == "VBxClustering": self._plda = get_plda(plda, token=token, cache_dir=cache_dir) else: self._plda = None Then verify the pipeline loads cleanly: $ source venv/bin/activate $ export HF_TOKEN=$(grep '^export HF_TOKEN=' ~/.bashrc | head -1 | sed 's/^export HF_TOKEN=//') $ python3 -c "from pyannote.audio import Pipeline; p = Pipeline.from_pretrained('pyannote/speaker-diarization-3.1', token='$HF_TOKEN'); print('OK:', p.klustering, p._plda)" OK: AgglomerativeClustering None If you see ''AgglomerativeClustering None'', you're good. If you see a 403 error, the patch didn't apply correctly. **Note:** This patch lives inside the venv. If you ever recreate the venv or force-reinstall pyannote-audio, you'll need to re-apply it. ===== Running a Transcription ===== Once everything is set up and the patch is applied, you're ready to transcribe. Drop your audio file somewhere and run: source ~/dev/whisperx/venv/bin/activate export HF_TOKEN=$(grep '^export HF_TOKEN=' ~/.bashrc | head -1 | sed 's/^export HF_TOKEN=//') whisperx recording.m4a \ --model large-v3 \ --language en \ --diarize \ --diarize_model pyannote/speaker-diarization-3.1 \ --hf_token "$HF_TOKEN" \ --min_speakers 2 \ --max_speakers 3 \ --compute_type int8 \ --output_dir ./transcripts What each flag does: ^ Flag ^ Purpose ^ | ''--model large-v3'' | Most accurate Whisper model. Options: tiny, base, small, medium, large-v3. Bigger = more accurate but slower. | | ''--language en'' | Skip language auto-detection (saves time). Remove if your recordings aren't in English. | | ''--diarize'' | Enable speaker diarization (the whole point). | | ''--diarize_model'' | Must be ''pyannote/speaker-diarization-3.1'' — see the section above. | | ''--hf_token'' | Your Hugging Face token. | | ''--min_speakers'' / ''--max_speakers'' | Hint the expected speaker count. Helps diarization accuracy for small meetings. | | ''--compute_type int8'' | Quantize to 8-bit. Significantly reduces memory and speeds up CPU inference. Highly recommended for CPU-only. | | ''--output_dir'' | Where to write output files. | The first time you run with a given model size, WhisperX downloads the model (~3 GB for large-v3). The pyannote models (~100 MB) and the wav2vec2 alignment model (~360 MB) also download on first run. Subsequent runs use cached models from ''~/.cache/huggingface/''. ===== Output Files ===== WhisperX creates several files in the output directory: ^ File ^ Description ^ | ''recording.srt'' | Subtitle format with speaker labels (most readable) | | ''recording.vtt'' | WebVTT format | | ''recording.tsv'' | Tab-separated, one segment per line | | ''recording.json'' | Full structured JSON with word-level timestamps | | ''recording.txt'' | Plain text, no timestamps | The SRT file looks like this: 1 00:03:06,340 --> 00:03:08,121 [SPEAKER_00]: Hey, good morning. 2 00:03:09,241 --> 00:03:09,882 [SPEAKER_01]: Good morning. 3 00:03:12,403 --> 00:03:13,443 Early part of the day. Speaker labels appear at the start of each speaker's turn. Segments without a label continue from the previous speaker. ===== Performance on the i5-9500 ===== I ran a 50.8-minute meeting recording through the full pipeline. Here's how the time broke down: ^ Stage ^ Time ^ | Voice activity detection | ~2 sec | | Transcription (large-v3, int8) | ~18 min | | Forced alignment (wav2vec2) | ~3 min | | Diarization (pyannote 3.1) | ~27 min | | **Total** | **~48 min** | So roughly realtime for a 50-minute file. The diarization step is the surprise — it takes longer than the transcription itself on CPU. The Whisper model is well-optimized via CTranslate2's int8 path, but pyannote's speaker embedding extraction is pure PyTorch with no quantization. Memory usage peaked around 7 GB (well within the 32 GB available). CPU was pinned at 390% (about 4 of the 6 cores fully utilized). ===== Convenience Wrapper ===== I got tired of typing all those flags, so I made a wrapper script. It activates the venv, checks that ''HF_TOKEN'' is set, and runs WhisperX with my usual defaults. #!/bin/bash # Usage: transcribe.sh [model_size] [min_speakers] [max_speakers] set -euo pipefail source "$HOME/dev/whisperx/venv/bin/activate" AUDIO_FILE="${1:?Usage: transcribe.sh [model] [min_spk] [max_spk]}" MODEL="${2:-large-v3}" MIN_SPK="${3:-2}" MAX_SPK="${4:-3}" if [[ -z "${HF_TOKEN:-}" ]]; then echo "ERROR: HF_TOKEN is not set. Export it first:" >&2 echo " export HF_TOKEN=hf_your_token_here" >&2 exit 1 fi whisperx "$AUDIO_FILE" \ --model "$MODEL" \ --language en \ --diarize \ --diarize_model pyannote/speaker-diarization-3.1 \ --hf_token "$HF_TOKEN" \ --min_speakers "$MIN_SPK" \ --max_speakers "$MAX_SPK" \ --compute_type int8 \ --output_dir ./transcripts Put it somewhere in your ''$PATH'' and make it executable. Then transcribing is just: $ transcribe.sh meeting.m4a $ transcribe.sh meeting.m4a medium 2 2 # quick test: medium model, exactly 2 speakers ===== Batch Transcription ===== For running through a pile of recordings, kick it off in the background and walk away: nohup bash -c ' for f in ~/Meetings/*.m4a; do echo "=== $(date) — Transcribing: $f ===" transcribe.sh "$f" done echo "=== $(date) — All done ===" ' > ~/transcribe.log 2>&1 & # Check progress: tail -f ~/transcribe.log Or use ''tmux'' / ''screen'' if you prefer to keep a live session. ===== Troubleshooting ===== ^ Problem ^ Fix ^ | ''401 Unauthorized'' or ''403 Forbidden'' | You didn't accept both HF model licenses, your token is wrong, or you're hitting the community-1 model — see the diarization section above. | | ''GatedRepoError'' for ''speaker-diarization-community-1'' | Two fixes needed: (1) pass ''--diarize_model pyannote/speaker-diarization-3.1'', and (2) apply the PLDA patch. Both are described above. | | ''ffmpeg not found'' | ''sudo apt install ffmpeg'' | | Out of memory / killed | Use ''--compute_type int8'' and/or a smaller model (''medium'' instead of ''large-v3''). With 32 GB RAM this shouldn't happen. | | Diarization puts everything on one speaker | Try narrowing ''--min_speakers'' / ''--max_speakers'', or check if the recording has a lot of crosstalk. | | Wrong language detected | Always specify ''--language en'' (or the appropriate code). | | Slow on CPU | Expected. Use ''--compute_type int8'' and ''medium'' model for speed; reserve ''large-v3'' for when accuracy matters most. | | ''ModuleNotFoundError'' after opening new terminal | You forgot to activate the venv: ''source ~/dev/whisperx/venv/bin/activate'' (or use the wrapper script which does it for you). | | Models re-downloading every run | Ensure ''HF_TOKEN'' is exported. Models cache in ''~/.cache/huggingface/'' — check write permissions if issues persist. | | ''HF_TOKEN'' not available after ''source ~/.bashrc'' | You used ''HF_TOKEN=...'' instead of ''export HF_TOKEN=...'' in ''~/.bashrc''. Without ''export'', it's a shell variable, not an environment variable. |