Local Voice Agent
ActiveA local voice assistant in Python. I connect speech recognition, an LLM and speech synthesis, working on streaming, natural interruptions and initial tool calls.
- Updated
- 2026-09-18
The idea
I want to talk to a voice assistant running on my own hardware. What interests me is how individual local models become a system that listens, responds and handles interruptions during a conversation.
The voice agent is a personal project in active development. Its repository is private and there is currently no public demo.
A modular speech pipeline
The application connects microphone capture, speech recognition, conversation management, a language model and audio output. Each component has its own interface so I can swap backends and test them independently.
- Speech recognition with Faster-Whisper and WebRTC VAD for detecting speech activity.
- Response generation with a local Qwen3 model through Ollama.
- Speech synthesis with Qwen3-TTS and a separate TTS backend.
- Audio playback with streaming and bounded queues.
- Conversation management with history and an initial planner and tool layer.
My focus is on integrating these components and understanding their runtime behaviour. I do not train the underlying speech or language models myself.
Handling interruptions
A voice assistant should not keep talking when I ask a new question. That is why I am working with barge-in: interrupting an ongoing response by speaking.
Model generation, speech synthesis and audio playback need to be cancelled together. Buffered audio must not continue playing afterwards. Acoustic echo cancellation through PipeWire also helps prevent the assistant from treating its own speaker output as a new input.
Conversation history needs to reflect this. An interrupted response is not stored as though I had heard it in full.
Measuring instead of guessing
I record timings across the pipeline, including the interval from the end of an utterance to its transcript, the first LLM token and audio output. Smoke tests help isolate streaming, cancellation and behaviour during backend failures.
The current implementation includes the local speech pipeline, streaming, conversation history, barge-in and initial tool calls. Further agent capabilities and broader use remain development goals.