105 lines
8.1 KiB
Markdown
105 lines
8.1 KiB
Markdown
---
|
||
created: 2026-08-30
|
||
modified: 2026-10-08
|
||
type: area
|
||
status: active
|
||
tags:
|
||
- voice
|
||
- system
|
||
- ai
|
||
- project
|
||
aliases:
|
||
- Jervis
|
||
- Voice Assistant
|
||
---
|
||
|
||
# Home Voice Assistant System (Jervis)
|
||
|
||
> **Human-readable overview of the local voice-assistant project.**
|
||
> - **AI/project tracking:** Vikunja → *Jervis Voice Assistant (Whisper ESP HA Flow)* (project id 5)
|
||
> - **AI documentation (for bots):** Outline → *Voice Assistant (Whisper ESP HA)* collection → [Project Overview](/doc/project-overview-genev9IHFT)
|
||
> - **API/docs (Gitea):** `sam/voice-assistant` → overview / agent-api / clients
|
||
> - **Tasks/progress:** see the Vikunja project; docs live in Outline; this note is the human map.
|
||
|
||
## What it is (TL;DR)
|
||
A private, local voice assistant. Say **"Hey Jervis"** (ESP32 wake word) then a natural-language request — a local **LLM agent** decides intent and calls tools: it controls Home Assistant, manages shopping/tasks/calendar, plays music, answers questions, and speaks back through the house speakers (MiniMax T2A v2 via `tts_router`, with piper as the fallback → Snapcast).
|
||
|
||
**Command it from anywhere:** `POST http://192.168.20.13:8501/voice {"text": "add milk to the shopping list"}` — the reply is spoken.
|
||
|
||
## MiniMax voice migration (2026-10-08)
|
||
Speech-to-text and text-to-speech now run on **MiniMax cloud APIs**, with a
|
||
**local fallback** so the house still works offline. Wake-word detection stays
|
||
local; only the spoken command is sent to the cloud.
|
||
|
||
- **STT**: `voice_stt_shim` (`:5001`) — cloud primary MiniMax `asr-1.0` on the
|
||
funded `MINIMAX_SUBSCRIPTION_KEY`, local `base.en` fallback (`:5002`). The
|
||
primary moved from OpenRouter `whisper-1` to MiniMax on 2026-10-08: MiniMax
|
||
tied OpenRouter on accuracy (WER 0.0151, 21/22 exact) and tripped the live 3 s
|
||
timeout 0/22 vs OpenRouter 4/22. Rollback is one compose line.
|
||
- **TTS**: `tts_router` behind the unchanged `speak_direct.sh` — MiniMax T2A v2
|
||
with piper fallback.
|
||
- **Personality**: the agent emits a speech plan (inline sound tags + one
|
||
emotion/delivery config). Kill switch: `VOICE_PERSONALITY=0`.
|
||
- **Result**: about 239 MiB of resident RAM freed, better voices, and voice
|
||
input survives an internet outage.
|
||
- Full project note: [[MiniMax Voice — Cloud STT and TTS with Local Fallback]].
|
||
- Interactive map: [Archify Map](https://maps.lab.audasmedia.com.au/minimax_voice_tts_stt/docs/minimax-voice-map.html).
|
||
|
||
## Architecture
|
||
```
|
||
ESP32 (wake word, mic) → MQTT voice/audio_stream
|
||
→ voice_bridge (OpenWakeWord, local)
|
||
→ voice_stt_shim (:5001) — cloud STT + local base.en fallback (:5002)
|
||
→ Jervis agent (FastAPI :8501, LLM tool-calling)
|
||
├─ Home Assistant (lights, timers, calendar, HA todo lists;
|
||
│ announcements via Jervis: script.jervis_say → MQTT voice/announce → announce mode, no tools)
|
||
├─ OmniRoute LLM (voice-fast: OpenRouter → opencode-go → deepseek flash)
|
||
├─ Mopidy → Spotify → Snapcast (music)
|
||
├─ Apprise (family WhatsApp/Telegram/ntfy/email)
|
||
├─ Vikunja (project tasks)
|
||
└─ reply → speak_direct.sh → tts_router (MiniMax T2A / piper) → Snapcast
|
||
```
|
||
|
||
## Components (host : port)
|
||
| Component | Host | Port | Role |
|
||
| ------------------------------ | ------- | ----------------- | ------------------------------------- |
|
||
| ESP32-S3 (voice_assistant.ino) | devices | — | wake word "Hey Jervis", mic, LED |
|
||
| voice_bridge | .13 | docker | wake word + audio buffer |
|
||
| voice_stt_shim | .13 | :5001 | cloud STT + local fallback chain |
|
||
| voice_whisper_fallback | .13 | :5002 | local base.en STT fallback |
|
||
| tts_router | .13 | — | MiniMax T2A / piper routing |
|
||
| **Jervis agent** | .13 | **:8501** | LLM tool-calling brain (systemd) |
|
||
| OmniRoute | .13 | :20129 | LLM gateway (voice-fast route) |
|
||
| Home Assistant | .30 | :8123 | control, timers, calendar, todo lists |
|
||
| Mopidy + Spotify | .13 | :6600 / librespot | music |
|
||
| piper_tts | .13 | :10200 | TTS fallback (piper) → Snapcast |
|
||
| Snapcast | .13 | :1780 | multi-room audio |
|
||
| Apprise | .35 | :8210 | notifications |
|
||
| Vikunja | .35 | vikunja.lab | project tasks/stages (tracker) |
|
||
| Outline | .13 | :3000 | project documentation (AI) |
|
||
|
||
## Status (2026-10-08)
|
||
- **Voice working end-to-end**: ESP32 wake word → STT (`voice_stt_shim`, MiniMax `asr-1.0` primary + local `base.en` fallback) → agent → TTS (`tts_router`, MiniMax T2A v2 + piper fallback) → Snapcast (fixed docker piper, HA→.13 SSH auth, resampler, MQTT client-id).
|
||
- **HA integration**: lights/scenes via Assist; shopping lists via HA todo (`add_shopping_item`); timers + spoken reminders; calendar alerts spoken 30 min before; add-calendar tool.
|
||
- **HA announcements route through Jervis, not piper**: `script.jervis_say` publishes the instruction to MQTT `homeassistant/voice/announce`; the agent runs it in **announce mode** (`POST /announce`, `_run(text, announce=True)`) with **no tools at all**, so an announcement cannot trigger a tool call.
|
||
- **Voice reminders**: helper `input_text.voice_reminder_message` (0–255 chars) feeds `script.jervis_say`; `automation.voice_reminder_speak` now triggers on the `timer.finished` event, so cancelling a timer no longer announces.
|
||
- **Calendar + piper announcements**: `automation.speak_calendar_event_reminder_30_min_before` and `script.piper_tts_announcement` were repointed off the missing `rest_command.jervis_say` to `script.jervis_say`; the loaded action was proven with a WebSocket `trace/get`.
|
||
- **Doorbell TTS repaired**: `doorbell_tts_to_wav.sh` now uses the Docker piper plus `ffmpeg`; the old native piper and `sox` are gone.
|
||
- **Weather script**: `script.n8n_voice_query_weather` argument mismatch is fixed (`data.text` / `data.source`); the n8n weather webhook target is dead, so a weather query cannot complete through that script as written.
|
||
- **Music**: play/pause/skip/volume via Mopidy → Spotify → Snapcast (credentials.json device-auth).
|
||
- **Q&A + web search**: Tavily-backed `web_search`; fast LLM route `voice-fast` (~2–3s).
|
||
- **Family messaging**: Apprise (sam/finn/harry/joanna; WhatsApp/email/Telegram/ntfy).
|
||
- **AI tooling**: `project-ops` skill gives every pi agent Vikunja + Outline access (keys in `10-secrets.conf`).
|
||
|
||
## Tools & systems (employment/skills)
|
||
FastAPI · Python · Docker · MQTT · Home Assistant (REST/Assist/automations) · MiniMax T2A & Speech-to-Text · OpenRouter Whisper · faster-whisper · piper TTS · Snapcast · Mopidy / librespot / Spotify · OmniRoute LLM routing · Tavily web search · Apprise · Vikunja & Outline APIs · NixOS (systemd services, flakes) · git-filter-repo (secret hygiene) · ESP32 (Arduino/ESP-IDF) · Archify.
|
||
|
||
## Key configs & links
|
||
- **Vikunja project**: `https://vikunja.lab.audasmedia.com.au` → *Jervis Voice Assistant (Whisper ESP HA Flow)*
|
||
- **Outline collection**: `https://outline.lab.audasmedia.com.au` → *Voice Assistant (Whisper ESP HA)*
|
||
- **MiniMax Voice project**: `/home/sam/paseo/projects/minimax_voice_tts_stt/` on `.13`; map at `https://maps.lab.audasmedia.com.au/minimax_voice_tts_stt/docs/`
|
||
- **Gitea**: `sam/voice-assistant` (overview / agent-api / clients docs)
|
||
- **Repo layout / plan**: `~/chats/homeassistant` (plan.md, AGENT.md); agent at `/home/sam/voice-agent/`
|
||
- **Secrets**: `~/.config/environment.d/10-secrets.conf` on .27/.13/.51 (VIKUNJA_TOKEN, OL_API_KEY, etc.) — see also the `secrets_fixes` note.
|
||
|
||
See also: [[Docker Containers]], [[Home Network Map Overview]], [[MiniMax Voice — Cloud STT and TTS with Local Fallback]]. |