SSD Nodes Learn Hosting plans →
Guides Matt ConnorBy Matt Connor

Speaker Diarization on a VPS: Who Said What

Add speaker labels to a Whisper transcript on your VPS. Pick between pyannote and NVIDIA Nemotron, get past gated model 401 errors, and time it on your CPU.

What speaker diarization adds to a transcript

Speaker diarization answers one question about a recording: who spoke when. On a VPS you get it by running a diarization model next to faster-whisper. Then you give each transcribed word the speaker whose turn covers that word's timestamp. Whisper writes the words but does not know who said them. A diarization model marks stretches of audio as SPEAKER_00, SPEAKER_01 and so on, but it writes no words. You need both, plus a short script that joins them.

This guide picks up where the guide to self-hosted speech-to-text and TTS (text to speech) stops. You already have faster-whisper turning audio into text. Here you add speaker labels. You also measure whether your VPS can keep up before you commit to a model.

How the two halves fit together

faster-whisper returns a start and end time for every word when you pass word_timestamps=True. The diarization model returns a list of turns. Each turn is a start time, an end time and a speaker label. The join is plain arithmetic: for each word, find the turn that overlaps it the most, and give the word that turn's speaker.

Work at the word level, not the segment level. A Whisper segment is a chunk of text, and it can run across a change of speaker. Labelling whole segments therefore puts the start of one person's reply into the other person's line. Labelling single words keeps the cut where the speaker actually changed.

Overlapping speech is the second problem. When two people talk at once, a normal diarization output holds two turns for the same moment, and one word cannot belong to both. The pyannote community-1 model card adds an output called exclusive_speaker_diarization for this case. It allows only one speaker at any moment, and the card says it is meant to make reconciliation with transcription timestamps simpler. The script below uses it.

Expect a few words at turn boundaries to land on the wrong side. Whisper's word times are estimates, not measured edges. A short reply such as "yes" right at a change of speaker is the usual place to look for errors.

Which speaker diarization model should you run?

As of October 2026, a reader searching for this finds three realistic options. Two are models. The third is glue code that runs one of those models for you. The licence terms and requirements below come from each project's own model card or README.

pyannote community-1 (pyannote/speaker-diarization-community-1). The model weights are under CC-BY-4.0 (Creative Commons Attribution 4.0). That licence allows commercial use as long as you give credit. The pyannote.audio library that runs the model is MIT licensed. The model is gated on Hugging Face: you must log in and accept its conditions, which include sharing your contact information and agreeing to occasional emails about pyannote news and services. You then need a Hugging Face access token. It runs on the CPU (central processing unit) by default, and one line moves it to a GPU (graphics processing unit). The card says multi-channel audio is downmixed to mono, and other sample rates are resampled to 16 kHz when the file is loaded. The pyannote README says ffmpeg must be installed, because its audio decoder, torchcodec, depends on it.

NVIDIA Nemotron 3 Diarization (nvidia/Nemotron-3-Diarization). The card puts it under the OpenMDW License Agreement 1.1 and marks it as not gated, so no token is needed. It installs through NVIDIA's NeMo toolkit. It tracks at most eight speakers. Its input is 16 kHz mono audio. The hardware listed on the card is NVIDIA GPUs, and the card does not mention CPU inference, so plan on a GPU. One trap: a separate repository, nvidia/Nemotron-3-Diarization-preview, sits under an NVIDIA evaluation licence that limits use to internal testing. Check the repository name before you build anything on it.

WhisperX (m-bain/whisperX, BSD-2-Clause). WhisperX is not a diarization model. It runs faster-whisper, adds its own word-alignment step, calls pyannote community-1 for the speakers, and then assigns a speaker to each word. It downloads the same pyannote model, so it needs the same gated access and the same token. Its README documents a CPU mode with --device cpu --compute_type int8.

The decision is short. On a CPU-only VPS, use pyannote community-1, either through your own script or through WhisperX. On a box with an NVIDIA GPU, Nemotron becomes a real option, and it skips the gated-access step entirely. This guide does not quote accuracy or speed figures for any of them. The number that matters is the one you measure on your own audio and your own VPS, and the timing section shows how.

Prepare the VPS

These commands are for Ubuntu 24.04. They install ffmpeg, a Python virtual environment (an isolated folder for Python packages) and the GNU time tool used later for measurement.

sudo apt update
sudo apt install -y ffmpeg python3-venv time curl
python3 -m venv ~/diarize
source ~/diarize/bin/activate
pip install --upgrade pip
pip install faster-whisper pyannote.audio

Check that both libraries import:

python -c "import faster_whisper, pyannote.audio; print('ok')"

It should print ok. Check free disk space with df -h ~ before the install. pyannote.audio pulls in PyTorch. The PyTorch packages on PyPI for Linux bring CUDA libraries with them even on a machine with no GPU, so the virtual environment takes several gigabytes.

Get Hugging Face access to the gated model

The pyannote model will not download until two things are true: your account has accepted the model's conditions, and your request carries a token from that account.

  1. Log in to Hugging Face and open the community-1 model page. Accept the conditions shown there.
  2. Create a token with read access at your token settings.
  3. Load the token into your shell without writing it to your shell history.
read -rs HF_TOKEN && export HF_TOKEN

Paste the token and press Enter. Nothing is echoed. Now test access before any Python is involved:

curl -s -o /dev/null -w '%{http_code}\n' \
  -H "Authorization: Bearer $HF_TOKEN" \
  https://huggingface.co/pyannote/speaker-diarization-community-1/resolve/main/config.yaml

A healthy result is 200 or a 3xx redirect code. The same URL requested with no token returns 401 Unauthorized. So a 401 here means the request did not carry a valid token. Either the variable is empty because you opened a new shell, or the token was mistyped or revoked. If the token is valid and you still get an error code, the account that owns the token has not accepted the conditions. Acceptance belongs to one account and one model, so a token from a different account does not inherit it.

Normalise the audio once

Convert the recording to 16 kHz mono WAV before either model sees it:

ffmpeg -i meeting.m4a -ac 1 -ar 16000 meeting.wav
ffprobe -hide_banner meeting.wav

The ffprobe line for the audio stream should include 16000 Hz, mono. pyannote would resample a different file on its own. Nemotron's card asks for 16 kHz mono input, though, and feeding every stage one file in one known format removes a variable when you debug.

The script: give every word a speaker

Save this as label.py. It runs diarization first, then transcription, then the join. It prints the time each stage took to standard error, so the timings do not mix with the transcript.

import os
import sys
import time

from faster_whisper import WhisperModel
from pyannote.audio import Pipeline

audio = sys.argv[1]

t0 = time.perf_counter()
pipeline = Pipeline.from_pretrained(
    "pyannote/speaker-diarization-community-1",
    token=os.environ["HF_TOKEN"])
output = pipeline(audio)
turns = [(turn.start, turn.end, speaker)
         for turn, speaker in output.exclusive_speaker_diarization]
t1 = time.perf_counter()
last_end = turns[-1][1] if turns else 0.0
print(f"diarization: {t1 - t0:.1f}s, {len(turns)} turns, "
      f"last turn ends at {last_end:.1f}s", file=sys.stderr)

model = WhisperModel("small", device="cpu", compute_type="int8")
segments, info = model.transcribe(audio, word_timestamps=True)
words = [word for segment in segments for word in segment.words]
t2 = time.perf_counter()
print(f"transcription: {t2 - t1:.1f}s, {len(words)} words", file=sys.stderr)


def speaker_for(start, end):
    best, best_overlap = None, 0.0
    for t_start, t_end, speaker in turns:
        overlap = min(end, t_end) - max(start, t_start)
        if overlap > best_overlap:
            best, best_overlap = speaker, overlap
    if best is None and turns:
        middle = (start + end) / 2
        nearest = min(turns, key=lambda t: min(abs(middle - t[0]), abs(middle - t[1])))
        best = nearest[2]
    return best or "UNKNOWN"


current, line, line_start = None, [], 0.0
for word in words:
    speaker = speaker_for(word.start, word.end)
    if speaker != current:
        if line:
            print(f"[{line_start:8.1f}s] {current}:{''.join(line)}")
        current, line, line_start = speaker, [], word.start
    line.append(word.word)
if line:
    print(f"[{line_start:8.1f}s] {current}:{''.join(line)}")

Sometimes a word falls in a short silence between two turns, so no turn overlaps it. The script then gives that word the speaker of the nearest turn. The line words = [...] does more than collect words. transcribe() returns a generator, and the faster-whisper README states that transcription only starts when you iterate over the segments. Without that line, the timer would report near zero for transcription, and the real work would be hidden inside the join.

Run it and read the output

python label.py meeting.wav > meeting.txt
head meeting.txt

Each line starts with the time the speaker began, then the speaker label, then what they said. Rename SPEAKER_00 and SPEAKER_01 to real names afterwards with a text editor or sed. The model has no way to know names.

If you know how many people are on the recording, say so. Change the pipeline call to pipeline(audio, num_speakers=2), or give a range with min_speakers=2, max_speakers=5. Both options are on the model card. Without them, the model has to estimate the count while it groups voices. Two similar voices can then merge into one label, or one voice can split into two.

How long does it take on your VPS? Time a ten-minute sample

Do not trust published speed figures for this job. They come from other hardware, other audio and other model sizes. Cut a ten-minute sample from a recording like the ones you will really process, and time it:

ffmpeg -i meeting.m4a -t 600 -ac 1 -ar 16000 sample.wav
/usr/bin/time -v python label.py sample.wav > sample.txt

Run it twice and use the second run. The first run downloads the models into ~/.cache/huggingface, and download time is not processing time.

Read the output in two parts. The diarization: and transcription: lines from the script tell you which stage costs more. From /usr/bin/time, take Elapsed (wall clock) time for the total and Maximum resident set size (kbytes) for the peak memory. Compare that peak with the total shown by free -m.

Divide the elapsed seconds by 600 to get the real-time factor. A factor of 0.5 means one hour of audio takes 30 minutes. A factor of 2 means one hour of audio takes two hours.

When does a GPU stop being optional?

Your measured real-time factor answers this question. Apply four rules.

  • If you want Nemotron, you need an NVIDIA GPU from the start, because NVIDIA GPUs are the only hardware its model card lists.
  • If you want live captions, the factor must sit clearly below 1 on the CPU. Above 1, the transcript falls further behind the speaker every minute.
  • For batch work, multiply the hours of audio you receive per day by the factor. If the result is more compute hours than the box has free in a day, the queue only grows, and a GPU is the fix.
  • Look at the stage split before you buy. If transcription dominates, try a smaller Whisper model first, which costs nothing. If diarization dominates, a smaller Whisper model will not help, and a GPU is what moves that number.

When you reach that point, renting a VPS with a GPU covers what to check before you pay. For the wider picture of which speech, text and image models a box of a given size can host, see which AI models you can self-host. If the same server will also run a language model to summarise each meeting, work out how big an LLM fits in your RAM after you subtract the peak memory you just measured.

WhisperX: the same result in one command

WhisperX does transcription, alignment and speaker assignment in one command. Give it its own virtual environment, because it brings its own version requirements for the libraries the script above uses.

python3 -m venv ~/whisperx
source ~/whisperx/bin/activate
pip install whisperx
whisperx meeting.wav --model small --diarize --hf_token "$HF_TOKEN" \
  --compute_type int8 --device cpu --min_speakers 2 --max_speakers 2

It writes the transcript files, with speaker labels, into the current directory. The trade is control. WhisperX decides how words are aligned and how speakers are assigned. Your own script lets you change the rule for boundary words or swap the diarization model. There is also a security point. A token passed as a command-line argument is visible to other users on the same machine in the output of ps while the job runs. On a shared box, use the Python route, which reads the token from the environment.

Nemotron on a GPU box

On a VPS with a supported NVIDIA GPU, install the NeMo toolkit with the command from the model card. The card uses uv, so install that into the virtual environment first.

python3 -m venv ~/nemo
source ~/nemo/bin/activate
pip install uv
uv pip install 'nemo-toolkit[asr]'

The inference code from the card:

from nemo.collections.asr.models import SortformerEncLabelModel

diar_model = SortformerEncLabelModel.from_pretrained("nvidia/Nemotron-3-Diarization")
diar_model.eval()
predicted_segments = diar_model.diarize(audio=["meeting.wav"], batch_size=1)
print(predicted_segments)

The card describes the raw model output as per-frame speaker activity for eight speaker slots. Print what diarize() returns and confirm it holds a start, an end and a speaker for each segment before you write the join. Once you have those, the speaker_for() function from label.py works unchanged. Remember the ceiling: a panel of ten people is beyond what this model tracks.

Failure modes, with what you will see

401 from Hugging Face. The download request carried no valid token. Run the curl check from the access section. In a new terminal HF_TOKEN is empty again, so run read -rs HF_TOKEN && export HF_TOKEN once more.

KeyError: 'HF_TOKEN'. Python could not find the variable at all. This happens in a new shell. It also happens under sudo, which by default does not pass your environment to the command it runs. Run the script as your normal user.

Speaker labels drift later in the file. The first minutes look right and the errors grow with time. This is the sign of a sample-rate mismatch. A timestamp is a sample count divided by the sample rate, so a wrong rate scales every time by the same factor, and the error grows with distance from the start. It happens when you load audio yourself and pass pyannote a {"waveform": ..., "sample_rate": ...} dictionary with the wrong sample_rate. Compare the last turn ends at value printed by the script with the duration ffprobe reports for the file. They should be close. If they differ by a ratio such as 48000 to 16000, pass the file path instead, or pass the true rate.

Everything is one speaker. The model estimated the speaker count and merged the voices. Pass num_speakers, or min_speakers and max_speakers, as shown above.

The process prints Killed. The kernel ran out of memory and stopped Python. Confirm with sudo dmesg | grep -i "killed process". Use a smaller Whisper model, add swap, or move to a bigger plan. Check the peak memory with the timing run first.

Import fails with an error that names torchcodec. Check ffmpeg -version. The pyannote README states that torchcodec needs ffmpeg installed on the machine, and a minimal server image does not include it.

FAQ

Do I need a Hugging Face account to run pyannote speaker diarization?

Yes, to download pyannote/speaker-diarization-community-1. The model is gated. You log in, accept its conditions on the model page, and create a read token for the download. The pyannote card also documents offline use: after one download, you can point Pipeline.from_pretrained() at a local copy of the model directory. NVIDIA's Nemotron-3-Diarization is not gated and needs no token.

Can speaker diarization run on a VPS without a GPU?

pyannote community-1 runs on the CPU by default, and WhisperX documents a CPU mode with --device cpu --compute_type int8. Whether that is fast enough depends on your VPS and your audio. Time a ten-minute sample with /usr/bin/time -v, divide the elapsed seconds by 600, and compare that factor with how much audio you need to process per day. Nemotron 3 Diarization lists only NVIDIA GPUs on its model card.

Why does Whisper not label speakers on its own?

Whisper is trained to turn audio into text. It has no output for speaker identity. Speaker labels come from a separate diarization model that groups stretches of audio by voice. A script then gives each transcribed word the speaker whose turn overlaps it most. WhisperX packages those steps into one command.

Can I use pyannote community-1 or Nemotron in a commercial product?

As of October 2026, the pyannote community-1 model card lists CC-BY-4.0, which allows commercial use with attribution, and the pyannote.audio library is MIT licensed. The nvidia/Nemotron-3-Diarization card lists the OpenMDW License Agreement 1.1. The separate Nemotron-3-Diarization-preview repository is under an evaluation licence that limits use to internal testing. Read the licence on the exact repository you download before you ship.