voice-assistant — Jervis Voice Assistant

A fully-local, private voice assistant. Say "Hey Jervis" and make a request; it answers, controls your house, manages your shopping list, and speaks back through your speakers — driven by a single LLM "brain" that calls tools.

System map

Interactive map

flowchart LR
    ESP32[ESP32-S3<br/>wake word] -->|audio| WHISPER[Whisper STT]
    WHISPER -->|text| AGENT[Jervis Agent<br/>:8501]
    AGENT --> HA[Home Assistant<br/>control · todo · timers]
    AGENT --> OMNI[OmniRoute<br/>LLM gateway]
    AGENT --> MOPIDY[Mopidy / Spotify<br/>music]
    AGENT --> APPRISE[Apprise<br/>notifications]
    AGENT --> TTS[Piper TTS]
    TTS --> SNAP[Snapcast<br/>multi-room]
    SNAP --> SPK[House Speakers]
    style AGENT fill:#bbf,stroke:#333,stroke-width:2px

🗺️ Interactive version (pan/zoom/search, dark/light, export): voice-assistant-map.html — hosted on maps.lab.audasmedia.com.au. Mermaid source: docs/voice-assistant.mmd.

Docs

  • docs/overview.md — the whole architecture in one page (diagram, hosts/IPs/ports, component map).
  • docs/agent-api.md — the command contract: send text, get a spoken reply. Use from any AI model, script, N8N, ESP32, or web app.
  • docs/clients.md — type-it, web page, and Android gateways.

Quick use

# speak a command (reply is spoken through speakers too)
curl -X POST http://192.168.20.13:8501/voice \
  -H "Content-Type: application/json" \
  -d '{"text":"add chicken to the shopping list"}'
  • sam/voice_assistant — ESP32 firmware (wake word, mic)
  • sam/voice_bridge — STT bridge + whisper
  • sam/speech_piper — TTS → Snapcast
  • sam/voice-assistant — this overview + API + clients

Secrets live in .env (gitignored) — never commit them.

Description
Jervis voice assistant — overview + agent API contract + clients
Readme 192 KiB