Building a voice AI assistant on a computer is easier than it sounds. You need three blocks: speech-to-text, a language model, and text-to-speech.
Every voice assistant works the same way. First, the computer listens and converts speech to text. Then a language model generates an answer. Finally, text-to-speech reads the answer aloud.
For speech-to-text, the Whisper family from OpenAI is the common choice. For the language model, local runners like Ollama and llama.cpp are popular. For speech output, Piper and Coqui TTS work well.
You can test each block separately before connecting them. This makes debugging much easier.
Whisper comes in several sizes: tiny, base, small, medium, and large. Smaller models are faster but less accurate. Larger models need more RAM and a better CPU or GPU.
faster-whisper is an optimized reimplementation. In late 2024, it became the go-to choice for real-time transcription on CPU. Vosk is another option — it is lightweight and works well on weak hardware.
| Model | Key difference | Price/access |
|---|---|---|
| Whisper tiny | Fastest, lowest accuracy | Free, open-source |
| Whisper base | Balanced speed and accuracy | Free, open-source |
| faster-whisper | Up to 4x faster than Whisper | Free, open-source |
| Vosk | Lightweight, works on weak CPUs | Free, open-source |
Piper is a fast neural TTS that runs on a CPU. As of 2025, it supports more than 20 languages. Coqui TTS offers more natural voices, including voice cloning, but needs more resources.
For quick testing, edge-tts uses Microsoft's cloud voices. It needs internet, but the quality is very high. Latency matters for a voice assistant. Piper can synthesize a short sentence in under a second on a modern CPU.
Install the dependencies first: `pip install faster-whisper sounddevice ollama`. Here is a minimal script that connects the microphone, Whisper, and Ollama:
import sounddevice as sd
from faster_whisper import WhisperModel
import ollama
stt = WhisperModel("base", device="cpu")
def listen():
audio = sd.rec(int(16000 * 3), samplerate=16000, channels=1)
sd.wait()
text, _ = stt.transcribe(audio.flatten())
return text.strip()
while True:
prompt = listen()
print("You:", prompt)
answer = ollama.chat(model="llama3.2", messages=[{"role": "user", "content": prompt}])
print("AI:", answer["message"]["content"])
This script records three seconds of audio, transcribes it, and sends the text to a local LLM. Add Piper at the end to speak the answer. The full pipeline is about 50 lines.
Ollama is the easiest way to run a language model. Install it, then run `ollama run llama3.2`. The model downloads automatically and runs on your computer.
In September 2024, Meta released Llama 3.2, and Ollama added support within days. If you want more control, llama.cpp offers quantized models that use less RAM. For a full beginner setup, see our guide: How to run AI models locally on your PC: a beginner's guide.
A fully local assistant never sends your voice or text to the cloud. That is the main reason to build one. Start with a small model, then scale up.
Once it works on a PC, the same ideas apply to other devices. You can run smaller models on a phone or a smartwatch: How to Run an AI Model Locally on Android: Apps and Setup and How to Use AI Chat on Wear OS Smartwatches: Installation, Setup, and Best Apps.
Checklist for your first build:
What is the easiest way to make a voice assistant on a PC?
Combine faster-whisper for speech-to-text, Ollama for the LLM, and Piper for speech output in a Python script. Test each part separately first.
Do I need a powerful GPU?
No. Small Whisper models, Piper, and quantized LLMs run on a CPU. A GPU helps with larger models.
What is the difference between Whisper and Vosk?
Whisper is more accurate but heavier. Vosk is lighter and faster on weak hardware.
Can the assistant work offline?
Yes, if all three components run locally. Only cloud TTS like edge-tts requires internet.