voice-assistant — Jervis Voice Assistant
A fully-local, private voice assistant. Say "Hey Jervis" and make a request; it answers, controls your house, manages your shopping list, and speaks back through your speakers — driven by a single LLM "brain" that calls tools.
System map
flowchart LR
ESP32[ESP32-S3<br/>wake word] -->|audio| WHISPER[Whisper STT]
WHISPER -->|text| AGENT[Jervis Agent<br/>:8501]
AGENT --> HA[Home Assistant<br/>control · todo · timers]
AGENT --> OMNI[OmniRoute<br/>LLM gateway]
AGENT --> MOPIDY[Mopidy / Spotify<br/>music]
AGENT --> APPRISE[Apprise<br/>notifications]
AGENT --> TTS[Piper TTS]
TTS --> SNAP[Snapcast<br/>multi-room]
SNAP --> SPK[House Speakers]
style AGENT fill:#bbf,stroke:#333,stroke-width:2px
🗺️ Interactive version (pan/zoom/search, dark/light, export): voice-assistant-map.html — hosted on
maps.lab.audasmedia.com.au. Mermaid source:docs/voice-assistant.mmd.
Docs
docs/overview.md— the whole architecture in one page (diagram, hosts/IPs/ports, component map).docs/agent-api.md— the command contract: send text, get a spoken reply. Use from any AI model, script, N8N, ESP32, or web app.docs/clients.md— type-it, web page, and Android gateways.
Quick use
# speak a command (reply is spoken through speakers too)
curl -X POST http://192.168.20.13:8501/voice \
-H "Content-Type: application/json" \
-d '{"text":"add chicken to the shopping list"}'
Related repos
sam/voice_assistant— ESP32 firmware (wake word, mic)sam/voice_bridge— STT bridge + whispersam/speech_piper— TTS → Snapcastsam/voice-assistant— this overview + API + clients
Secrets live in
.env(gitignored) — never commit them.
Description