whisperx_transcription
Differences
This shows you the differences between two versions of the page.
| Next revision | Previous revision | ||
| whisperx_transcription [2026/10/10 17:18] – created millerjs | whisperx_transcription [2026/10/10 17:20] (current) – [Local Speech-to-Text with WhisperX] millerjs | ||
|---|---|---|---|
| Line 1: | Line 1: | ||
| - | hold | + | ====== Local Speech-to-Text with WhisperX ====== |
| + | This is running on a SFF Optiplex 7070 with an i5-9500 (6 cores / 6 threads, AVX2, 32 GB RAM, Debian 13). No GPU, no cloud, just the CPU. A 50-minute recording takes about 48 minutes end-to-end with the large-v3 model in int8 mode. Not fast, but the machine is usually idle so I just kick it off and walk away. At some point I'd like to put a low profile GPU in here to help with accelerating these kinds of tasks, but this works fine for now. | ||
| + | |||
| + | ===== System Prerequisites ===== | ||
| + | Debian 13 (Trixie) ships everything you need in the main repos. FFmpeg handles decoding the audio files (m4a, mp3, wav, whatever), and the rest is standard Python tooling. | ||
| + | |||
| + | < | ||
| + | sudo apt update | ||
| + | sudo apt install -y ffmpeg | ||
| + | sudo apt install -y python3 python3-venv python3-dev python3-pip build-essential pkg-config | ||
| + | </ | ||
| + | |||
| + | Verify: | ||
| + | |||
| + | < | ||
| + | $ ffmpeg -version | ||
| + | ffmpeg version 7.1.5-0+deb13u1 | ||
| + | |||
| + | $ python3 --version | ||
| + | Python 3.13.5 | ||
| + | </ | ||
| + | |||
| + | ===== Python Virtual Environment ===== | ||
| + | WhisperX pulls in PyTorch, pyannote, faster-whisper, | ||
| + | |||
| + | I staged everything in '' | ||
| + | |||
| + | < | ||
| + | mkdir -p ~/ | ||
| + | cd ~/ | ||
| + | python3 -m venv venv | ||
| + | source venv/ | ||
| + | pip install --upgrade pip | ||
| + | </ | ||
| + | |||
| + | ===== Install WhisperX ===== | ||
| + | Still in the activated venv: | ||
| + | |||
| + | < | ||
| + | pip install whisperx | ||
| + | </ | ||
| + | |||
| + | This takes a few minutes. When it's done, verify the CLI works: | ||
| + | |||
| + | < | ||
| + | $ whisperx --help | ||
| + | usage: whisperx [-h] [--model MODEL] [--model_cache_only MODEL_CACHE_ONLY] | ||
| + | [--model_dir MODEL_DIR] [--device DEVICE] | ||
| + | ... | ||
| + | </ | ||
| + | |||
| + | If '' | ||
| + | |||
| + | ===== Hugging Face Setup (required for diarization) ===== | ||
| + | This is the step people most often miss, and skipping it produces a confusing '' | ||
| + | |||
| + | ==== Create an account ==== | ||
| + | Go to [[https:// | ||
| + | |||
| + | ==== Accept the model licenses ==== | ||
| + | Visit each of these two pages and click "Agree and access repository": | ||
| + | |||
| + | - [[https:// | ||
| + | - [[https:// | ||
| + | |||
| + | You have to accept **both**. If you only accept the first one, you'll get a 403 later. | ||
| + | |||
| + | ==== Generate an access token ==== | ||
| + | - Go to [[https:// | ||
| + | - Click "New token" | ||
| + | - Name it something like '' | ||
| + | - Type: **Read** (sufficient — you only need to download models) | ||
| + | - Copy the token — it starts with '' | ||
| + | |||
| + | ==== Make the token available ==== | ||
| + | Add it to your '' | ||
| + | |||
| + | < | ||
| + | echo ' | ||
| + | source ~/.bashrc | ||
| + | </ | ||
| + | |||
| + | Verify it's actually exported: | ||
| + | |||
| + | < | ||
| + | $ echo " | ||
| + | set (hf_ZQj...) | ||
| + | </ | ||
| + | |||
| + | ==== Verify token access ==== | ||
| + | Before you kick off a long transcription, | ||
| + | |||
| + | <code python test_hf_token.py> | ||
| + | import sys, os | ||
| + | from huggingface_hub import HfApi | ||
| + | |||
| + | token = os.environ[" | ||
| + | api = HfApi(token=token) | ||
| + | |||
| + | models = [ | ||
| + | " | ||
| + | " | ||
| + | ] | ||
| + | |||
| + | all_ok = True | ||
| + | for m in models: | ||
| + | try: | ||
| + | info = api.model_info(m) | ||
| + | print(f" | ||
| + | except Exception as e: | ||
| + | all_ok = False | ||
| + | print(f" | ||
| + | |||
| + | sys.exit(0 if all_ok else 1) | ||
| + | </ | ||
| + | |||
| + | < | ||
| + | $ source ~/ | ||
| + | $ python3 test_hf_token.py | ||
| + | PASS: pyannote/ | ||
| + | PASS: pyannote/ | ||
| + | </ | ||
| + | |||
| + | Do this before running a transcription. It catches token and license issues in seconds instead of after a 20-minute transcription+alignment run that dies at the very end. | ||
| + | |||
| + | ===== The Diarization Problem (and how to fix it) ===== | ||
| + | This is where things got interesting. WhisperX 3.8.6 (with pyannote.audio 4.0.7) has **two separate code paths** that both try to access a model called '' | ||
| + | |||
| + | ==== Layer 1: the --diarize_model flag ==== | ||
| + | WhisperX 3.8.6 defaults to '' | ||
| + | |||
| + | < | ||
| + | --diarize_model pyannote/ | ||
| + | </ | ||
| + | |||
| + | ==== Layer 2: the PLDA patch ==== | ||
| + | Even with the correct '' | ||
| + | |||
| + | The fix is to patch the pipeline to only load PLDA when '' | ||
| + | |||
| + | '' | ||
| + | |||
| + | Find this in '' | ||
| + | |||
| + | <code python> | ||
| + | self.plda = plda | ||
| + | self._plda = get_plda(plda, | ||
| + | |||
| + | self.klustering = clustering | ||
| + | </ | ||
| + | |||
| + | Replace it with: | ||
| + | |||
| + | <code python> | ||
| + | self.plda = plda | ||
| + | self.klustering = clustering | ||
| + | |||
| + | # Only load PLDA when VBxClustering is used; AgglomerativeClustering doesn' | ||
| + | # Loading it unconditionally triggers a download from the gated community-1 repo. | ||
| + | if self.klustering == " | ||
| + | self._plda = get_plda(plda, | ||
| + | else: | ||
| + | self._plda = None | ||
| + | </ | ||
| + | |||
| + | Then verify the pipeline loads cleanly: | ||
| + | |||
| + | < | ||
| + | $ source venv/ | ||
| + | $ export HF_TOKEN=$(grep ' | ||
| + | $ python3 -c "from pyannote.audio import Pipeline; p = Pipeline.from_pretrained(' | ||
| + | OK: AgglomerativeClustering None | ||
| + | </ | ||
| + | |||
| + | If you see '' | ||
| + | |||
| + | **Note:** This patch lives inside the venv. If you ever recreate the venv or force-reinstall pyannote-audio, | ||
| + | |||
| + | ===== Running a Transcription ===== | ||
| + | Once everything is set up and the patch is applied, you're ready to transcribe. Drop your audio file somewhere and run: | ||
| + | |||
| + | < | ||
| + | source ~/ | ||
| + | export HF_TOKEN=$(grep ' | ||
| + | |||
| + | whisperx recording.m4a \ | ||
| + | --model large-v3 \ | ||
| + | --language en \ | ||
| + | --diarize \ | ||
| + | --diarize_model pyannote/ | ||
| + | --hf_token " | ||
| + | --min_speakers 2 \ | ||
| + | --max_speakers 3 \ | ||
| + | --compute_type int8 \ | ||
| + | --output_dir ./ | ||
| + | </ | ||
| + | |||
| + | What each flag does: | ||
| + | |||
| + | ^ Flag ^ Purpose ^ | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | |||
| + | The first time you run with a given model size, WhisperX downloads the model (~3 GB for large-v3). The pyannote models (~100 MB) and the wav2vec2 alignment model (~360 MB) also download on first run. Subsequent runs use cached models from '' | ||
| + | |||
| + | ===== Output Files ===== | ||
| + | WhisperX creates several files in the output directory: | ||
| + | |||
| + | ^ File ^ Description ^ | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | |||
| + | The SRT file looks like this: | ||
| + | |||
| + | < | ||
| + | 1 | ||
| + | 00: | ||
| + | [SPEAKER_00]: | ||
| + | |||
| + | 2 | ||
| + | 00: | ||
| + | [SPEAKER_01]: | ||
| + | |||
| + | 3 | ||
| + | 00: | ||
| + | Early part of the day. | ||
| + | </ | ||
| + | |||
| + | Speaker labels appear at the start of each speaker' | ||
| + | |||
| + | ===== Performance on the i5-9500 ===== | ||
| + | I ran a 50.8-minute meeting recording through the full pipeline. Here's how the time broke down: | ||
| + | |||
| + | ^ Stage ^ Time ^ | ||
| + | | Voice activity detection | ~2 sec | | ||
| + | | Transcription (large-v3, int8) | ~18 min | | ||
| + | | Forced alignment (wav2vec2) | ~3 min | | ||
| + | | Diarization (pyannote 3.1) | ~27 min | | ||
| + | | **Total** | **~48 min** | | ||
| + | |||
| + | So roughly realtime for a 50-minute file. The diarization step is the surprise — it takes longer than the transcription itself on CPU. The Whisper model is well-optimized via CTranslate2' | ||
| + | |||
| + | Memory usage peaked around 7 GB (well within the 32 GB available). CPU was pinned at 390% (about 4 of the 6 cores fully utilized). | ||
| + | |||
| + | ===== Convenience Wrapper ===== | ||
| + | I got tired of typing all those flags, so I made a wrapper script. It activates the venv, checks that '' | ||
| + | |||
| + | <code bash transcribe.sh> | ||
| + | # | ||
| + | # Usage: transcribe.sh < | ||
| + | |||
| + | set -euo pipefail | ||
| + | |||
| + | source " | ||
| + | |||
| + | AUDIO_FILE=" | ||
| + | MODEL=" | ||
| + | MIN_SPK=" | ||
| + | MAX_SPK=" | ||
| + | |||
| + | if [[ -z " | ||
| + | echo " | ||
| + | echo " | ||
| + | exit 1 | ||
| + | fi | ||
| + | |||
| + | whisperx " | ||
| + | --model " | ||
| + | --language en \ | ||
| + | --diarize \ | ||
| + | --diarize_model pyannote/ | ||
| + | --hf_token " | ||
| + | --min_speakers " | ||
| + | --max_speakers " | ||
| + | --compute_type int8 \ | ||
| + | --output_dir ./ | ||
| + | </ | ||
| + | |||
| + | Put it somewhere in your '' | ||
| + | |||
| + | < | ||
| + | $ transcribe.sh meeting.m4a | ||
| + | $ transcribe.sh meeting.m4a medium 2 2 # quick test: medium model, exactly 2 speakers | ||
| + | </ | ||
| + | |||
| + | ===== Batch Transcription ===== | ||
| + | For running through a pile of recordings, kick it off in the background and walk away: | ||
| + | |||
| + | < | ||
| + | nohup bash -c ' | ||
| + | for f in ~/ | ||
| + | echo "=== $(date) — Transcribing: | ||
| + | transcribe.sh " | ||
| + | done | ||
| + | echo "=== $(date) — All done ===" | ||
| + | ' > ~/ | ||
| + | |||
| + | # Check progress: | ||
| + | tail -f ~/ | ||
| + | </ | ||
| + | |||
| + | Or use '' | ||
| + | |||
| + | ===== Troubleshooting ===== | ||
| + | ^ Problem ^ Fix ^ | ||
| + | | '' | ||
| + | | '' | ||
| + | | '' | ||
| + | | Out of memory / killed | Use '' | ||
| + | | Diarization puts everything on one speaker | Try narrowing '' | ||
| + | | Wrong language detected | Always specify '' | ||
| + | | Slow on CPU | Expected. Use '' | ||
| + | | '' | ||
| + | | Models re-downloading every run | Ensure '' | ||
| + | | '' | ||
whisperx_transcription.1791652690.txt.gz · Last modified: by millerjs
