Files
obsidian-vault/300 areas/360 Dev-Ops Network Computers/Home Voice Assistant System.md

105 lines
8.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
created: 2026-08-30
modified: 2026-10-08
type: area
status: active
tags:
- voice
- system
- ai
- project
aliases:
- Jervis
- Voice Assistant
---
# Home Voice Assistant System (Jervis)
> **Human-readable overview of the local voice-assistant project.**
> - **AI/project tracking:** Vikunja → *Jervis Voice Assistant (Whisper ESP HA Flow)* (project id 5)
> - **AI documentation (for bots):** Outline → *Voice Assistant (Whisper ESP HA)* collection → [Project Overview](/doc/project-overview-genev9IHFT)
> - **API/docs (Gitea):** `sam/voice-assistant` → overview / agent-api / clients
> - **Tasks/progress:** see the Vikunja project; docs live in Outline; this note is the human map.
## What it is (TL;DR)
A private, local voice assistant. Say **"Hey Jervis"** (ESP32 wake word) then a natural-language request — a local **LLM agent** decides intent and calls tools: it controls Home Assistant, manages shopping/tasks/calendar, plays music, answers questions, and speaks back through the house speakers (MiniMax T2A v2 via `tts_router`, with piper as the fallback → Snapcast).
**Command it from anywhere:** `POST http://192.168.20.13:8501/voice {"text": "add milk to the shopping list"}` — the reply is spoken.
## MiniMax voice migration (2026-10-08)
Speech-to-text and text-to-speech now run on **MiniMax cloud APIs**, with a
**local fallback** so the house still works offline. Wake-word detection stays
local; only the spoken command is sent to the cloud.
- **STT**: `voice_stt_shim` (`:5001`) — cloud primary MiniMax `asr-1.0` on the
funded `MINIMAX_SUBSCRIPTION_KEY`, local `base.en` fallback (`:5002`). The
primary moved from OpenRouter `whisper-1` to MiniMax on 2026-10-08: MiniMax
tied OpenRouter on accuracy (WER 0.0151, 21/22 exact) and tripped the live 3 s
timeout 0/22 vs OpenRouter 4/22. Rollback is one compose line.
- **TTS**: `tts_router` behind the unchanged `speak_direct.sh` — MiniMax T2A v2
with piper fallback.
- **Personality**: the agent emits a speech plan (inline sound tags + one
emotion/delivery config). Kill switch: `VOICE_PERSONALITY=0`.
- **Result**: about 239 MiB of resident RAM freed, better voices, and voice
input survives an internet outage.
- Full project note: [[MiniMax Voice — Cloud STT and TTS with Local Fallback]].
- Interactive map: [Archify Map](https://maps.lab.audasmedia.com.au/minimax_voice_tts_stt/docs/minimax-voice-map.html).
## Architecture
```
ESP32 (wake word, mic) → MQTT voice/audio_stream
→ voice_bridge (OpenWakeWord, local)
→ voice_stt_shim (:5001) — cloud STT + local base.en fallback (:5002)
→ Jervis agent (FastAPI :8501, LLM tool-calling)
├─ Home Assistant (lights, timers, calendar, HA todo lists;
│ announcements via Jervis: script.jervis_say → MQTT voice/announce → announce mode, no tools)
├─ OmniRoute LLM (voice-fast: OpenRouter → opencode-go → deepseek flash)
├─ Mopidy → Spotify → Snapcast (music)
├─ Apprise (family WhatsApp/Telegram/ntfy/email)
├─ Vikunja (project tasks)
└─ reply → speak_direct.sh → tts_router (MiniMax T2A / piper) → Snapcast
```
## Components (host : port)
| Component | Host | Port | Role |
| ------------------------------ | ------- | ----------------- | ------------------------------------- |
| ESP32-S3 (voice_assistant.ino) | devices | — | wake word "Hey Jervis", mic, LED |
| voice_bridge | .13 | docker | wake word + audio buffer |
| voice_stt_shim | .13 | :5001 | cloud STT + local fallback chain |
| voice_whisper_fallback | .13 | :5002 | local base.en STT fallback |
| tts_router | .13 | — | MiniMax T2A / piper routing |
| **Jervis agent** | .13 | **:8501** | LLM tool-calling brain (systemd) |
| OmniRoute | .13 | :20129 | LLM gateway (voice-fast route) |
| Home Assistant | .30 | :8123 | control, timers, calendar, todo lists |
| Mopidy + Spotify | .13 | :6600 / librespot | music |
| piper_tts | .13 | :10200 | TTS fallback (piper) → Snapcast |
| Snapcast | .13 | :1780 | multi-room audio |
| Apprise | .35 | :8210 | notifications |
| Vikunja | .35 | vikunja.lab | project tasks/stages (tracker) |
| Outline | .13 | :3000 | project documentation (AI) |
## Status (2026-10-08)
- **Voice working end-to-end**: ESP32 wake word → STT (`voice_stt_shim`, MiniMax `asr-1.0` primary + local `base.en` fallback) → agent → TTS (`tts_router`, MiniMax T2A v2 + piper fallback) → Snapcast (fixed docker piper, HA→.13 SSH auth, resampler, MQTT client-id).
- **HA integration**: lights/scenes via Assist; shopping lists via HA todo (`add_shopping_item`); timers + spoken reminders; calendar alerts spoken 30 min before; add-calendar tool.
- **HA announcements route through Jervis, not piper**: `script.jervis_say` publishes the instruction to MQTT `homeassistant/voice/announce`; the agent runs it in **announce mode** (`POST /announce`, `_run(text, announce=True)`) with **no tools at all**, so an announcement cannot trigger a tool call.
- **Voice reminders**: helper `input_text.voice_reminder_message` (0–255 chars) feeds `script.jervis_say`; `automation.voice_reminder_speak` now triggers on the `timer.finished` event, so cancelling a timer no longer announces.
- **Calendar + piper announcements**: `automation.speak_calendar_event_reminder_30_min_before` and `script.piper_tts_announcement` were repointed off the missing `rest_command.jervis_say` to `script.jervis_say`; the loaded action was proven with a WebSocket `trace/get`.
- **Doorbell TTS repaired**: `doorbell_tts_to_wav.sh` now uses the Docker piper plus `ffmpeg`; the old native piper and `sox` are gone.
- **Weather script**: `script.n8n_voice_query_weather` argument mismatch is fixed (`data.text` / `data.source`); the n8n weather webhook target is dead, so a weather query cannot complete through that script as written.
- **Music**: play/pause/skip/volume via Mopidy → Spotify → Snapcast (credentials.json device-auth).
- **Q&A + web search**: Tavily-backed `web_search`; fast LLM route `voice-fast` (~2–3s).
- **Family messaging**: Apprise (sam/finn/harry/joanna; WhatsApp/email/Telegram/ntfy).
- **AI tooling**: `project-ops` skill gives every pi agent Vikunja + Outline access (keys in `10-secrets.conf`).
## Tools & systems (employment/skills)
FastAPI · Python · Docker · MQTT · Home Assistant (REST/Assist/automations) · MiniMax T2A & Speech-to-Text · OpenRouter Whisper · faster-whisper · piper TTS · Snapcast · Mopidy / librespot / Spotify · OmniRoute LLM routing · Tavily web search · Apprise · Vikunja & Outline APIs · NixOS (systemd services, flakes) · git-filter-repo (secret hygiene) · ESP32 (Arduino/ESP-IDF) · Archify.
## Key configs & links
- **Vikunja project**: `https://vikunja.lab.audasmedia.com.au` → *Jervis Voice Assistant (Whisper ESP HA Flow)*
- **Outline collection**: `https://outline.lab.audasmedia.com.au` → *Voice Assistant (Whisper ESP HA)*
- **MiniMax Voice project**: `/home/sam/paseo/projects/minimax_voice_tts_stt/` on `.13`; map at `https://maps.lab.audasmedia.com.au/minimax_voice_tts_stt/docs/`
- **Gitea**: `sam/voice-assistant` (overview / agent-api / clients docs)
- **Repo layout / plan**: `~/chats/homeassistant` (plan.md, AGENT.md); agent at `/home/sam/voice-agent/`
- **Secrets**: `~/.config/environment.d/10-secrets.conf` on .27/.13/.51 (VIKUNJA_TOKEN, OL_API_KEY, etc.) — see also the `secrets_fixes` note.
See also: [[Docker Containers]], [[Home Network Map Overview]], [[MiniMax Voice — Cloud STT and TTS with Local Fallback]].