Table of Contents

Local Speech-to-Text with WhisperX

This is running on a SFF Optiplex 7070 with an i5-9500 (6 cores / 6 threads, AVX2, 32 GB RAM, Debian 13). No GPU, no cloud, just the CPU. A 50-minute recording takes about 48 minutes end-to-end with the large-v3 model in int8 mode. Not fast, but the machine is usually idle so I just kick it off and walk away. At some point I'd like to put a low profile GPU in here to help with accelerating these kinds of tasks, but this works fine for now.

System Prerequisites

Debian 13 (Trixie) ships everything you need in the main repos. FFmpeg handles decoding the audio files (m4a, mp3, wav, whatever), and the rest is standard Python tooling.

sudo apt update
sudo apt install -y ffmpeg
sudo apt install -y python3 python3-venv python3-dev python3-pip build-essential pkg-config

Verify:

$ ffmpeg -version
ffmpeg version 7.1.5-0+deb13u1

$ python3 --version
Python 3.13.5

Python Virtual Environment

WhisperX pulls in PyTorch, pyannote, faster-whisper, and a pile of other dependencies — about 2 GB worth. Keep it all isolated in a venv so it doesn't trash your system Python.

I staged everything in ~/dev/whisperx but you can put it wherever you want.

mkdir -p ~/dev/whisperx
cd ~/dev/whisperx
python3 -m venv venv
source venv/bin/activate
pip install --upgrade pip

Install WhisperX

Still in the activated venv:

pip install whisperx

This takes a few minutes. When it's done, verify the CLI works:

$ whisperx --help
usage: whisperx [-h] [--model MODEL] [--model_cache_only MODEL_CACHE_ONLY]
                [--model_dir MODEL_DIR] [--device DEVICE]
...

If whisperx isn't found on your PATH, use python -m whisperx –help instead.

Hugging Face Setup (required for diarization)

This is the step people most often miss, and skipping it produces a confusing 401 Unauthorized or 403 Forbidden error later. The pyannote diarization models are gated on Hugging Face — you have to create an account and explicitly accept the licenses before you can download them.

Create an account

Go to huggingface.co/join and sign up. It's free.

Accept the model licenses

Visit each of these two pages and click “Agree and access repository”:

  1. pyannote/speaker-diarization-3.1 — the main diarization pipeline
  2. pyannote/segmentation-3.0 — a sub-model used by the above

You have to accept both. If you only accept the first one, you'll get a 403 later.

Generate an access token

  1. Click “New token”
  2. Name it something like whisperx-local
  3. Type: Read (sufficient — you only need to download models)
  4. Copy the token — it starts with hf_…

Make the token available

Add it to your ~/.bashrc so it persists across sessions. Make sure to use export — without it, you're just setting a shell variable that never makes it into the environment, and you'll spend way too long figuring out why your token “isn't working.”

echo 'export HF_TOKEN=hf_your_token_here' >> ~/.bashrc
source ~/.bashrc

Verify it's actually exported:

$ echo "${HF_TOKEN:+set (${HF_TOKEN:0:6}...)}"
set (hf_ZQj...)

Verify token access

Before you kick off a long transcription, take 5 seconds to verify your token can actually access both models. WhisperX installs huggingface_hub as a dependency, so you don't need to install anything extra. I wrote a small script for this:

test_hf_token.py
import sys, os
from huggingface_hub import HfApi
 
token = os.environ["HF_TOKEN"]
api = HfApi(token=token)
 
models = [
    "pyannote/speaker-diarization-3.1",
    "pyannote/segmentation-3.0",
]
 
all_ok = True
for m in models:
    try:
        info = api.model_info(m)
        print(f"PASS: {m} (sha={info.sha[:8]})")
    except Exception as e:
        all_ok = False
        print(f"FAIL: {m} -> {e}")
 
sys.exit(0 if all_ok else 1)
$ source ~/dev/whisperx/venv/bin/activate
$ python3 test_hf_token.py
PASS: pyannote/speaker-diarization-3.1 (sha=84fd2591)
PASS: pyannote/segmentation-3.0 (sha=e66f3d3b)

Do this before running a transcription. It catches token and license issues in seconds instead of after a 20-minute transcription+alignment run that dies at the very end.

The Diarization Problem (and how to fix it)

This is where things got interesting. WhisperX 3.8.6 (with pyannote.audio 4.0.7) has two separate code paths that both try to access a model called pyannote/speaker-diarization-community-1 — a newer gated model that I hadn't accepted the license for, and that the setup guides I was following didn't even mention. Both will fail with GatedRepoError: 403 and waste your time.

Layer 1: the --diarize_model flag

WhisperX 3.8.6 defaults to pyannote/speaker-diarization-community-1. The model you actually accepted the license for is pyannote/speaker-diarization-3.1. You have to explicitly tell WhisperX to use the right one:

--diarize_model pyannote/speaker-diarization-3.1

Layer 2: the PLDA patch

Even with the correct –diarize_model flag, I hit a second 403 error. The speaker-diarization-3.1 pipeline uses AgglomerativeClustering, which doesn't need PLDA (a speaker verification sub-model). But pyannote.audio 4.0.7's SpeakerDiarization.init unconditionally loads PLDA from pyannote/speaker-diarization-community-1 anyway, regardless of which clustering method is actually being used. So it tries to download plda/xvec_transform.npz from a gated repo you don't have access to, and dies.

The fix is to patch the pipeline to only load PLDA when VBxClustering is actually used. The file you need to edit is:

venv/lib/python3.13/site-packages/pyannote/audio/pipelines/speaker_diarization.py

Find this in SpeakerDiarization.init (around line 230):

        self.plda = plda
        self._plda = get_plda(plda, token=token, cache_dir=cache_dir)
 
        self.klustering = clustering

Replace it with:

        self.plda = plda
        self.klustering = clustering
 
        # Only load PLDA when VBxClustering is used; AgglomerativeClustering doesn't need it.
        # Loading it unconditionally triggers a download from the gated community-1 repo.
        if self.klustering == "VBxClustering":
            self._plda = get_plda(plda, token=token, cache_dir=cache_dir)
        else:
            self._plda = None

Then verify the pipeline loads cleanly:

$ source venv/bin/activate
$ export HF_TOKEN=$(grep '^export HF_TOKEN=' ~/.bashrc | head -1 | sed 's/^export HF_TOKEN=//')
$ python3 -c "from pyannote.audio import Pipeline; p = Pipeline.from_pretrained('pyannote/speaker-diarization-3.1', token='$HF_TOKEN'); print('OK:', p.klustering, p._plda)"
OK: AgglomerativeClustering None

If you see AgglomerativeClustering None, you're good. If you see a 403 error, the patch didn't apply correctly.

Note: This patch lives inside the venv. If you ever recreate the venv or force-reinstall pyannote-audio, you'll need to re-apply it.

Running a Transcription

Once everything is set up and the patch is applied, you're ready to transcribe. Drop your audio file somewhere and run:

source ~/dev/whisperx/venv/bin/activate
export HF_TOKEN=$(grep '^export HF_TOKEN=' ~/.bashrc | head -1 | sed 's/^export HF_TOKEN=//')

whisperx recording.m4a \
  --model large-v3 \
  --language en \
  --diarize \
  --diarize_model pyannote/speaker-diarization-3.1 \
  --hf_token "$HF_TOKEN" \
  --min_speakers 2 \
  --max_speakers 3 \
  --compute_type int8 \
  --output_dir ./transcripts

What each flag does:

Flag Purpose
–model large-v3 Most accurate Whisper model. Options: tiny, base, small, medium, large-v3. Bigger = more accurate but slower.
–language en Skip language auto-detection (saves time). Remove if your recordings aren't in English.
–diarize Enable speaker diarization (the whole point).
–diarize_model Must be pyannote/speaker-diarization-3.1 — see the section above.
–hf_token Your Hugging Face token.
–min_speakers / –max_speakers Hint the expected speaker count. Helps diarization accuracy for small meetings.
–compute_type int8 Quantize to 8-bit. Significantly reduces memory and speeds up CPU inference. Highly recommended for CPU-only.
–output_dir Where to write output files.

The first time you run with a given model size, WhisperX downloads the model (~3 GB for large-v3). The pyannote models (~100 MB) and the wav2vec2 alignment model (~360 MB) also download on first run. Subsequent runs use cached models from ~/.cache/huggingface/.

Output Files

WhisperX creates several files in the output directory:

File Description
recording.srt Subtitle format with speaker labels (most readable)
recording.vtt WebVTT format
recording.tsv Tab-separated, one segment per line
recording.json Full structured JSON with word-level timestamps
recording.txt Plain text, no timestamps

The SRT file looks like this:

1
00:03:06,340 --> 00:03:08,121
[SPEAKER_00]: Hey, good morning.

2
00:03:09,241 --> 00:03:09,882
[SPEAKER_01]: Good morning.

3
00:03:12,403 --> 00:03:13,443
Early part of the day.

Speaker labels appear at the start of each speaker's turn. Segments without a label continue from the previous speaker.

Performance on the i5-9500

I ran a 50.8-minute meeting recording through the full pipeline. Here's how the time broke down:

Stage Time
Voice activity detection ~2 sec
Transcription (large-v3, int8) ~18 min
Forced alignment (wav2vec2) ~3 min
Diarization (pyannote 3.1) ~27 min
Total ~48 min

So roughly realtime for a 50-minute file. The diarization step is the surprise — it takes longer than the transcription itself on CPU. The Whisper model is well-optimized via CTranslate2's int8 path, but pyannote's speaker embedding extraction is pure PyTorch with no quantization.

Memory usage peaked around 7 GB (well within the 32 GB available). CPU was pinned at 390% (about 4 of the 6 cores fully utilized).

Convenience Wrapper

I got tired of typing all those flags, so I made a wrapper script. It activates the venv, checks that HF_TOKEN is set, and runs WhisperX with my usual defaults.

transcribe.sh
#!/bin/bash
# Usage: transcribe.sh <audio_file> [model_size] [min_speakers] [max_speakers]
 
set -euo pipefail
 
source "$HOME/dev/whisperx/venv/bin/activate"
 
AUDIO_FILE="${1:?Usage: transcribe.sh <audio_file> [model] [min_spk] [max_spk]}"
MODEL="${2:-large-v3}"
MIN_SPK="${3:-2}"
MAX_SPK="${4:-3}"
 
if [[ -z "${HF_TOKEN:-}" ]]; then
  echo "ERROR: HF_TOKEN is not set. Export it first:" >&2
  echo "  export HF_TOKEN=hf_your_token_here" >&2
  exit 1
fi
 
whisperx "$AUDIO_FILE" \
  --model "$MODEL" \
  --language en \
  --diarize \
  --diarize_model pyannote/speaker-diarization-3.1 \
  --hf_token "$HF_TOKEN" \
  --min_speakers "$MIN_SPK" \
  --max_speakers "$MAX_SPK" \
  --compute_type int8 \
  --output_dir ./transcripts

Put it somewhere in your $PATH and make it executable. Then transcribing is just:

$ transcribe.sh meeting.m4a
$ transcribe.sh meeting.m4a medium 2 2   # quick test: medium model, exactly 2 speakers

Batch Transcription

For running through a pile of recordings, kick it off in the background and walk away:

nohup bash -c '
  for f in ~/Meetings/*.m4a; do
    echo "=== $(date) — Transcribing: $f ==="
    transcribe.sh "$f"
  done
  echo "=== $(date) — All done ==="
' > ~/transcribe.log 2>&1 &

# Check progress:
tail -f ~/transcribe.log

Or use tmux / screen if you prefer to keep a live session.

Troubleshooting

Problem Fix
401 Unauthorized or 403 Forbidden You didn't accept both HF model licenses, your token is wrong, or you're hitting the community-1 model — see the diarization section above.
GatedRepoError for speaker-diarization-community-1 Two fixes needed: (1) pass –diarize_model pyannote/speaker-diarization-3.1, and (2) apply the PLDA patch. Both are described above.
ffmpeg not found sudo apt install ffmpeg
Out of memory / killed Use –compute_type int8 and/or a smaller model (medium instead of large-v3). With 32 GB RAM this shouldn't happen.
Diarization puts everything on one speaker Try narrowing –min_speakers / –max_speakers, or check if the recording has a lot of crosstalk.
Wrong language detected Always specify –language en (or the appropriate code).
Slow on CPU Expected. Use –compute_type int8 and medium model for speed; reserve large-v3 for when accuracy matters most.
ModuleNotFoundError after opening new terminal You forgot to activate the venv: source ~/dev/whisperx/venv/bin/activate (or use the wrapper script which does it for you).
Models re-downloading every run Ensure HF_TOKEN is exported. Models cache in ~/.cache/huggingface/ — check write permissions if issues persist.
HF_TOKEN not available after source ~/.bashrc You used HF_TOKEN=… instead of export HF_TOKEN=… in ~/.bashrc. Without export, it's a shell variable, not an environment variable.