sam-4screen-desktop 2026-10-8:22:6:28

This commit is contained in:
2026-10-08 22:06:28 +11:00
parent 057487f876
commit 303355eb16
5 changed files with 51 additions and 13 deletions

View File

@@ -22,7 +22,7 @@ aliases:
> - **Tasks/progress:** see the Vikunja project; docs live in Outline; this note is the human map.
## What it is (TL;DR)
A private, local voice assistant. Say **"Hey Jervis"** (ESP32 wake word) then a natural-language request — a local **LLM agent** decides intent and calls tools: it controls Home Assistant, manages shopping/tasks/calendar, plays music, answers questions, and speaks back through the house speakers (piper → Snapcast).
A private, local voice assistant. Say **"Hey Jervis"** (ESP32 wake word) then a natural-language request — a local **LLM agent** decides intent and calls tools: it controls Home Assistant, manages shopping/tasks/calendar, plays music, answers questions, and speaks back through the house speakers (MiniMax T2A v2 via `tts_router`, with piper as the fallback → Snapcast).
**Command it from anywhere:** `POST http://192.168.20.13:8501/voice {"text": "add milk to the shopping list"}` — the reply is spoken.
@@ -31,8 +31,11 @@ Speech-to-text and text-to-speech now run on **MiniMax cloud APIs**, with a
**local fallback** so the house still works offline. Wake-word detection stays
local; only the spoken command is sent to the cloud.
- **STT**: `voice_stt_shim` (`:5001`) — cloud primary (OpenRouter `whisper-1`;
MiniMax `asr-1.0` ready), local `base.en` fallback (`:5002`).
- **STT**: `voice_stt_shim` (`:5001`) — cloud primary MiniMax `asr-1.0` on the
funded `MINIMAX_SUBSCRIPTION_KEY`, local `base.en` fallback (`:5002`). The
primary moved from OpenRouter `whisper-1` to MiniMax on 2026-10-08: MiniMax
tied OpenRouter on accuracy (WER 0.0151, 21/22 exact) and tripped the live 3 s
timeout 0/22 vs OpenRouter 4/22. Rollback is one compose line.
- **TTS**: `tts_router` behind the unchanged `speak_direct.sh` — MiniMax T2A v2
with piper fallback.
- **Personality**: the agent emits a speech plan (inline sound tags + one
@@ -48,7 +51,8 @@ ESP32 (wake word, mic) → MQTT voice/audio_stream
→ voice_bridge (OpenWakeWord, local)
→ voice_stt_shim (:5001) — cloud STT + local base.en fallback (:5002)
→ Jervis agent (FastAPI :8501, LLM tool-calling)
├─ Home Assistant (lights, timers, calendar, HA todo lists)
├─ Home Assistant (lights, timers, calendar, HA todo lists;
│ announcements via Jervis: script.jervis_say → MQTT voice/announce → announce mode, no tools)
├─ OmniRoute LLM (voice-fast: OpenRouter → opencode-go → deepseek flash)
├─ Mopidy → Spotify → Snapcast (music)
├─ Apprise (family WhatsApp/Telegram/ntfy/email)
@@ -68,15 +72,20 @@ ESP32 (wake word, mic) → MQTT voice/audio_stream
| OmniRoute | .13 | :20129 | LLM gateway (voice-fast route) |
| Home Assistant | .30 | :8123 | control, timers, calendar, todo lists |
| Mopidy + Spotify | .13 | :6600 / librespot | music |
| piper_tts | .13 | :10200 | TTS → Snapcast |
| piper_tts | .13 | :10200 | TTS fallback (piper) → Snapcast |
| Snapcast | .13 | :1780 | multi-room audio |
| Apprise | .35 | :8210 | notifications |
| Vikunja | .35 | vikunja.lab | project tasks/stages (tracker) |
| Outline | .13 | :3000 | project documentation (AI) |
## Status (2026-10-08)
- **Voice working end-to-end**: ESP32 wake word → whisper → agent → piper → Snapcast (fixed docker piper, HA→.13 SSH auth, resampler, MQTT client-id).
- **HA integration**: lights/scenes via Assist; shopping lists via HA todo (`add_shopping_item`); timers + spoken reminders (HA `timer` → automation → piper); calendar alerts spoken 30 min before (HA automation); add-calendar tool.
- **Voice working end-to-end**: ESP32 wake word → STT (`voice_stt_shim`, MiniMax `asr-1.0` primary + local `base.en` fallback) → agent → TTS (`tts_router`, MiniMax T2A v2 + piper fallback) → Snapcast (fixed docker piper, HA→.13 SSH auth, resampler, MQTT client-id).
- **HA integration**: lights/scenes via Assist; shopping lists via HA todo (`add_shopping_item`); timers + spoken reminders; calendar alerts spoken 30 min before; add-calendar tool.
- **HA announcements route through Jervis, not piper**: `script.jervis_say` publishes the instruction to MQTT `homeassistant/voice/announce`; the agent runs it in **announce mode** (`POST /announce`, `_run(text, announce=True)`) with **no tools at all**, so an announcement cannot trigger a tool call.
- **Voice reminders**: helper `input_text.voice_reminder_message` (0–255 chars) feeds `script.jervis_say`; `automation.voice_reminder_speak` now triggers on the `timer.finished` event, so cancelling a timer no longer announces.
- **Calendar + piper announcements**: `automation.speak_calendar_event_reminder_30_min_before` and `script.piper_tts_announcement` were repointed off the missing `rest_command.jervis_say` to `script.jervis_say`; the loaded action was proven with a WebSocket `trace/get`.
- **Doorbell TTS repaired**: `doorbell_tts_to_wav.sh` now uses the Docker piper plus `ffmpeg`; the old native piper and `sox` are gone.
- **Weather script**: `script.n8n_voice_query_weather` argument mismatch is fixed (`data.text` / `data.source`); the n8n weather webhook target is dead, so a weather query cannot complete through that script as written.
- **Music**: play/pause/skip/volume via Mopidy → Spotify → Snapcast (credentials.json device-auth).
- **Q&A + web search**: Tavily-backed `web_search`; fast LLM route `voice-fast` (~2–3s).
- **Family messaging**: Apprise (sam/finn/harry/joanna; WhatsApp/email/Telegram/ntfy).