Self-Host VoiceStudio for VPS: ElevenLabs Wey You Own
Run VoiceStudio for your own VPS with Docker, no more ElevenLabs dollar. Voice cloning, safe access, CPU speed test, and wetin e really support for Yoruba, Hausa, Igbo and Pidgin.
Wetin VoiceStudio be, and why you go self-host am for VPS
VoiceStudio na open-source app wey dey do text-to-speech and voice cloning, and you fit self-host am for your own VPS instead of paying ElevenLabs dollar every month. Text-to-speech (TTS) mean say computer dey read your text out loud. You go run am as one Docker container, lock am to one fixed version, and open the web page through Tailscale or through nginx with TLS (transport layer security, the "s" for https).
VoiceStudio na the app and the web page. The voice itself dey come from one model wey dem call OmniVoice, wey the k2-fsa team build. OmniVoice na "zero-shot" TTS model. Zero-shot mean say e fit copy one voice from short audio clip, without any extra training. VoiceStudio fit run other engines too, but OmniVoice na the default.
Plenty of us for Naija dey pay for voice tools with dollar card. The rate dey change and the card dey decline, but the subscription still dey count every month. When the engine dey your own server, the only bill na the VPS. The text wey you type no comot your server go another company. If you wan see the other open-source voice options first, check self-hosted speech-to-text and text-to-speech engines for VPS.
Wetin you need before you start
- One VPS wey get x86-64 CPU. The VoiceStudio Docker image na
linux/amd64only. E no get ARM version. - Ubuntu 24.04 with Docker Engine already installed. If Docker never dey, follow how to install Docker and Docker Compose for VPS first.
- Enough free disk. The model weights alone na about 2.4 GB, and the image get PyTorch inside, so e no small.
- GPU (graphics card) no compulsory. VoiceStudio dey work for CPU only. E just slow pass.
Check the machine first:
uname -m
docker --version
df -h /uname -m must print x86_64. If e print aarch64, your VPS na ARM, and this image no go run there. Docker go either refuse the pull with no matching manifest, or the container go die with exec format error. The two errors get one cause: the image na amd64 code, and ARM CPU no fit run am directly. Change to x86-64 VPS before you continue.
Why we no dey use the curl | sh installer here
VoiceStudio get one-command installer for desktop. For server, the Docker route better. Everything dey inside one container and one volume, so to remove am na two commands. You fit also lock the exact version wey you dey run. The installer no give you that control.
Lock VoiceStudio to one version, no use :stable
As of October 2026, VoiceStudio don ship five releases inside two weeks: v0.5.2 come out on 10 September 2026, and v0.5.6 come out on 23 September 2026. The project dey move fast. That one good, but e mean say tag like :stable dey change under you.
The image dey for ghcr.io/debpalash/voicestudio. The tags wey dey move na :latest (preview build from the main branch), :stable (the newest stable release) and :0.5 (newest patch inside 0.5). If you use any of them, two pulls for two different days fit give you two different apps. One of them fit break your setup, and you no go know wetin change.
The registry dey publish exact version tags. For October 2026, :0.5.6 dey there, and e point to the same image as :stable. Use :0.5.6:
docker pull ghcr.io/debpalash/voicestudio:0.5.6
docker images --digests ghcr.io/debpalash/voicestudioThe second command print one DIGEST column wey start with sha256:. Digest na the fingerprint of the image content. Version tag fit still move if the maintainer push again, but digest no fit move. Write the digest down. If you want the strictest lock, use ghcr.io/debpalash/voicestudio@sha256:<your digest> as the image name for docker run.
Make the API key and keep am safe
The VoiceStudio docs say make you set OMNIVOICE_API_KEY before you run the container. Na this key dey protect the API and the web page. Make one long random key and save am inside file wey only you fit read:
(umask 077; python3 -c 'import secrets; print(secrets.token_urlsafe(32))' > ~/.omnivoice-key)
ls -l ~/.omnivoice-key
export OMNIVOICE_API_KEY="$(cat ~/.omnivoice-key)"ls -l must show -rw-------. That mean na only your user fit read the file. We save am for file because export dey only last for the current SSH session. If you open new session and run docker run without export again, the variable go empty.
When you log in for browser, the browser go send the key once, and the server go give am short session wey last maximum 8 hours. The browser no dey keep the master key.
Run VoiceStudio with Docker
docker run -d --name omnivoice \
--restart unless-stopped \
-p 127.0.0.1:3900:3900 \
-e OMNIVOICE_API_KEY="$OMNIVOICE_API_KEY" \
-e OMNIVOICE_ANALYTICS_DISABLED=1 \
-v omnivoice-data:/app/omnivoice_data \
ghcr.io/debpalash/voicestudio:0.5.6Wetin each part dey do:
-p 127.0.0.1:3900:3900publish port 3900 only for the server itself. Nobody for internet fit reach am directly. Na the project default, and na the most important line here.-v omnivoice-data:/app/omnivoice_datakeep your projects, your saved voices, settings and the downloaded model inside one Docker volume. Remove the container, the volume still dey.OMNIVOICE_ANALYTICS_DISABLED=1off the usage analytics. The project say analytics na opt-in and no carry your content, but on server we prefer am off.--restart unless-stoppedbring the container back after reboot.
The first start dey slow, because the container dey download about 2.4 GB of weights from Hugging Face. Watch am:
docker logs -f omnivoicePress Ctrl+C to stop watching. The container no go stop. When the download finish, check health:
curl -s http://127.0.0.1:3900/health
sudo ss -tlnp | grep 3900The health answer na small JSON with "status": "ok", and the device field go tell you whether e dey use CPU or GPU. ss must show 127.0.0.1:3900. If you see 0.0.0.0:3900, the port dey open for the whole internet. Read the UFW section below before you do anything else.
How fast VoiceStudio dey for CPU VPS? Time am yourself
The VoiceStudio README talk say CPU-only dey fully usable, just slower. The project performance notes measure dia numbers for Apple laptop, no be VPS. I no go give you RAM number or seconds-per-sentence wey nobody measure for your type of server. Different CPU, different result. Better make you measure your own.
The useful number na RTF (real-time factor). RTF na the time wey e take to generate the audio, divided by how long the audio be. RTF of 0.5 mean say 10 seconds of audio take 5 seconds to make. RTF above 1 mean say you go wait longer than the audio itself.
VoiceStudio get API wey follow the same shape as OpenAI speech API, for /v1/audio/speech. Use am to time one paragraph. The text below na English because na speed we dey measure here, no be accent:
export OMNIVOICE_API_KEY="$(cat ~/.omnivoice-key)"
time curl -s http://127.0.0.1:3900/v1/audio/speech \
-H "Authorization: Bearer $OMNIVOICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"tts-1","voice":"alloy","input":"Good morning. This is a short test paragraph. I want to see how long my server takes to read it out loud.","response_format":"wav"}' \
--output test.wav
file test.wav
python3 -c "import wave; w = wave.open('test.wav'); print(round(w.getnframes() / w.getframerate(), 1), 'seconds of audio')"time print real at the end. Na the generation time be that. file must say RIFF (little-endian) data, WAVE audio. If e say JSON or text instead, the server send error, no be audio. Open am with cat test.wav to read the message. The last line print how many seconds of audio you get. Divide real by that number, and you get your RTF. If python complain unknown format, the WAV na float format wey the wave module no sabi read. Download the file and check the length for any audio player.
Run the same command two times. The first run fit include model loading, so the second run na the honest number. While e dey generate, open second SSH window and run docker stats omnivoice --no-stream. The MEM USAGE column show how much RAM the container dey use. Use that real number to pick your plan, with the method for how to size a VPS for the workload wey you wan run.
How to open the VoiceStudio web page safely
Port 3900 dey only for localhost, so your browser no fit reach am yet. You get two correct ways. Tailscale na the safer one. Nginx with TLS dey for when you need normal public URL.
Option 1: Open VoiceStudio through Tailscale serve
Tailscale na private network between your devices. If Tailscale don dey your VPS and you don log in, one command go put VoiceStudio for your private network with HTTPS:
sudo tailscale serve --bg 3900
tailscale serve statustailscale serve status print one address like https://your-vps.your-tailnet.ts.net. Open am for phone or laptop wey dey the same tailnet, then paste your API key for the login page. If Tailscale tell you say HTTPS never on for your tailnet, e go print one link. Open the link, enable am, then run the command again.
No use tailscale funnel here. Funnel publish the page for the whole internet, and that one na exactly wetin we dey avoid. The difference between the two na the full topic of Tailscale serve versus Tailscale funnel.
Option 2: Put VoiceStudio behind nginx with TLS
Use this one if you need normal domain like voice.example.com. First point the domain DNS A record to your VPS IP. Then install nginx and certbot. Certbot na the tool wey dey collect free TLS certificate from Let's Encrypt.
sudo apt update
sudo apt install -y nginx certbot python3-certbot-nginxCreate /etc/nginx/sites-available/voicestudio:
server {
listen 80;
server_name voice.example.com;
client_max_body_size 100m;
location / {
proxy_pass http://127.0.0.1:3900;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_read_timeout 600s;
}
}Enable am, test am, then collect the certificate:
sudo ln -s /etc/nginx/sites-available/voicestudio /etc/nginx/sites-enabled/
sudo nginx -t
sudo systemctl reload nginx
sudo ufw allow 'Nginx Full'
sudo certbot --nginx -d voice.example.comnginx -t must print syntax is ok and test is successful. Certbot go add the HTTPS block and the redirect from http for you.
Some lines for that config get reason. client_max_body_size 100m dey there because voice clips and audiobook files big, and nginx default na 1 MB, so upload go fail with 413 Request Entity Too Large. proxy_read_timeout 600s dey there because long text for CPU fit take more than one minute, and nginx default wait na 60 seconds, so you go see 504 Gateway Time-out even though the server still dey work. The Upgrade and Connection lines allow WebSocket, wey the dictation feature use for /ws/transcribe. For line-by-line explanation of these directives, read nginx reverse proxy config explained.
Know wetin you dey accept here. With nginx, the login page dey for public internet, and na only the API key dey protect am. The VoiceStudio docs self advise make you keep the backend for private network like Tailscale. Also, no on the "network sharing" PIN for public setup. Na six-digit PIN, and the docs talk am plainly: person fit guess am by trying many times.
No change the port to 0.0.0.0, because UFW no go block am
The VoiceStudio docs show how to change the port to 0.0.0.0:3900:3900 for home LAN. For VPS, 0.0.0.0 mean the whole internet. And UFW no go save you.
The reason: Docker write its own iptables rules for every published port. The packet wey come for port 3900 dey redirected to the container inside the nat table, before e reach the INPUT chain wey UFW dey check. So sudo ufw deny 3900 no dey touch am at all. You fit see am yourself:
sudo iptables -t nat -L DOCKER -nWith 0.0.0.0, you go see DNAT line for tcp dpt:3900 wey no get any IP limit. From your laptop, curl http://YOUR_VPS_IP:3900/health go answer, even when sudo ufw status no show any rule for 3900. On top of that, the traffic na plain HTTP, so your API key go travel without encryption. Keep 127.0.0.1. The full story dey for why Docker published ports bypass UFW.
Yoruba, Hausa, Igbo and Nigerian Pidgin: wetin OmniVoice really support
I check the OmniVoice language list (the docs/languages.md file for the k2-fsa/OmniVoice project) for October 2026. E list 646 languages. The four big Nigerian ones dey inside: Yoruba (yo), Hausa (ha), Igbo (ig) and Nigerian Pidgin (pcm). The list also show how many hours of audio dem use train each language:
The data behind this chart
[
{
"label": "Hausa",
"training_hours": 17.75
},
{
"label": "Yoruba",
"training_hours": 15.66
},
{
"label": "Igbo",
"training_hours": 13.69
},
{
"label": "Nigerian Pidgin",
"training_hours": 11.04
}
]Hausa get the most, 17.75 hours. Nigerian Pidgin get the least of the four, 11.04 hours. Now look the same list for bigger languages:
The data behind this chart
[
{
"label": "English",
"training_hours": "206,061"
},
{
"label": "French",
"training_hours": "23,675"
},
{
"label": "Swahili",
"training_hours": 418.41
},
{
"label": "Nigerian Pidgin",
"training_hours": 11.04
}
]English get 206,061 hours. Swahili get 418.41 hours. Our four languages get small fraction of that.
Hours no be quality score. I no go promise you say the Pidgin voice go sound like person wey grow for Warri or Lagos, because I no fit measure that one for you. Na your ear go judge. Test with sentence wey you sabi well, and compare am with the English output from the same voice. For Yoruba and Igbo, try the same sentence with tone marks (like ọ, ẹ, ṣ) and without them, then listen which one read better. Where the generate screen get language option, set am to the language wey you dey write.
Licence: wetin you fit do and wetin you no fit do
Licence matter here, because the app and the model no carry the same rules. As of October 2026:
- VoiceStudio na AGPL-3.0. To use am as e be, for yourself, no problem. If you change the code and give other people access to your changed version over network, you must give them your source code too.
- OmniVoice code na Apache-2.0. That one free for commercial use.
- OmniVoice model weights (the trained model wey dey actually make the voice) na different matter. The Hugging Face model card list am as CC-BY-NC, because of the data wey dem use train am. NC mean non-commercial. So to sell voice-over or use am for paid client work with the default model no clear at all. Read the model card yourself before you collect money.
- Every other engine wey VoiceStudio fit run get its own licence. Check each one before commercial use.
Consent: clone only voice wey you get permission for
Voice-clone fraud dey real. Somebody collect your uncle voice from WhatsApp voice note, clone am, then call your mama say "I dey police station, send money now". The tool wey you dey install fit do exactly that, so the rule simple: clone your own voice, or voice of person wey give you permission. Get the permission for writing, and keep am.
No clone pastor, politician, celebrity or your boss for joke. Even the OmniVoice model card forbid voice cloning without permission, impersonation and fraud. One small habit wey help: agree secret family word wey everybody go ask for any call wey dey beg for money.
How to update and back up VoiceStudio
To update, pick the new exact version tag from the release page, pull am, remove the old container, then run again with the same key and the same volume. Replace 0.5.6 below with the new version:
export OMNIVOICE_API_KEY="$(cat ~/.omnivoice-key)"
docker pull ghcr.io/debpalash/voicestudio:0.5.6
docker stop omnivoice
docker rm omnivoiceThen run the same docker run command from before, with the new tag. Your voices and projects dey inside the omnivoice-data volume, so dem go still dey. Read the release notes before every update. The project dey change fast.
To back up the volume without the big model cache (you fit download the model again anytime):
docker run --rm -v omnivoice-data:/data -v "$PWD":/backup ubuntu:24.04 \
tar czf /backup/omnivoice-data.tgz --exclude=./huggingface -C /data .
ls -lh omnivoice-data.tgzCopy that file comot the VPS. Your saved voice samples dey inside, so treat am like private data.
Because the API follow the OpenAI speech shape, other tools fit call am too. Release v0.5.4 add n8n workflow export, so if you already get n8n running for your VPS with Docker and HTTPS, you fit make automation wey turn text to audio without any paid API.
When VoiceStudio no work: the error and the reason
Bind for 127.0.0.1:3900 failed: port is already allocated. Another container or program don hold port 3900. Run sudo ss -tlnp | grep 3900 to see who. Stop that one, or change the left side of the port mapping, like 127.0.0.1:3901:3900.
Conflict. The container name "/omnivoice" is already in use. You don run the container before. Run docker rm -f omnivoice, then run again. The volume no go lose anything.
curl or the login page return 401. 401 mean the key wey you send no match the key wey the container start with. This one happen most when OMNIVOICE_API_KEY was empty during docker run, because you open new SSH session and forget the export. Run docker exec omnivoice printenv OMNIVOICE_API_KEY and compare am with cat ~/.omnivoice-key. If dem no match, remove the container and run again after export.
The health check no answer after many minutes. The first start still dey download weights. docker logs -f omnivoice show the progress. If the log stop with disk error, check df -h /, because 2.4 GB of weights plus the image fit fill small disk.
Nginx show 413 Request Entity Too Large or 504 Gateway Time-out. The first one na client_max_body_size, the second one na proxy_read_timeout. Both dey explained for the nginx section.
FAQ
VoiceStudio fit run for VPS wey no get GPU?
Yes. The VoiceStudio README talk say CPU-only dey fully usable, just slower. How slow depend on your CPU, so time one paragraph with the /v1/audio/speech API and time curl, then divide the generation time by the audio length. That number na your real-time factor for that exact server. Note say the Docker image na x86-64 only, so ARM VPS no go work.
OmniVoice support Yoruba, Hausa, Igbo and Nigerian Pidgin?
Yes. As of October 2026, the OmniVoice language list get all four: Yoruba (yo), Hausa (ha), Igbo (ig) and Nigerian Pidgin (pcm). Each one get between 11.04 and 17.75 hours of training audio, while English get 206,061 hours. Support no mean say e go sound natural. Test with sentence wey you sabi well and judge with your own ear.
I fit use VoiceStudio voices for commercial work?
Check the model licence first. VoiceStudio app na AGPL-3.0 and OmniVoice code na Apache-2.0, but the OmniVoice model weights for Hugging Face carry CC-BY-NC licence, wey mean non-commercial. Every other engine get its own licence. Read the model card before you sell any audio, and clone only voices wey you get written permission for.
Why my VoiceStudio port open for internet even though UFW dey block am?
Docker add its own iptables rules for published ports, and those rules redirect the traffic before UFW check am. So if you publish the port as 0.0.0.0:3900 or just 3900:3900, the whole internet fit reach am, whatever UFW talk. Publish am as 127.0.0.1:3900:3900, then reach the page through Tailscale serve or nginx with TLS.