sam-4screen-desktop 2026-10-8:16:21:28

This commit is contained in:
2026-10-08 16:21:28 +11:00
parent ebba19b8c3
commit 057487f876
4 changed files with 134 additions and 9 deletions

View File

@@ -210,8 +210,9 @@
"templater-obsidian:Templater": false "templater-obsidian:Templater": false
} }
}, },
"active": "26813e16e829b798", "active": "8018aa52f0591df9",
"lastOpenFiles": [ "lastOpenFiles": [
"010 inbox/MiniMax Voice — Cloud STT and TTS with Local Fallback.md",
"300 areas/355 AI Research and Techniques/Ai memory management.md", "300 areas/355 AI Research and Techniques/Ai memory management.md",
"200 projects/Tools Software WebUI/Family Console.md", "200 projects/Tools Software WebUI/Family Console.md",
"200 projects/Tools Software WebUI/Family Home Lab.md", "200 projects/Tools Software WebUI/Family Home Lab.md",
@@ -239,7 +240,6 @@
"400 Personal Family/410 Business Ideas/Business Ideas Display Screens.md", "400 Personal Family/410 Business Ideas/Business Ideas Display Screens.md",
"010 inbox/Business Idea - Franchises display screens.md", "010 inbox/Business Idea - Franchises display screens.md",
"400 Personal Family/410 Business Ideas/Business Idea - Franchises display screens.md", "400 Personal Family/410 Business Ideas/Business Idea - Franchises display screens.md",
"010 inbox/Business Ideas Display Screens.md",
"010 inbox/_WATCHER", "010 inbox/_WATCHER",
"300 areas/355 AI Research and Techniques", "300 areas/355 AI Research and Techniques",
"300 areas/355 Ai Research and Techniques", "300 areas/355 Ai Research and Techniques",

View File

@@ -0,0 +1,101 @@
---
title: MiniMax Voice
type: note
status: active
proposed_folder: 200 projects/Tools Software WebUI
proposed_toc: "100 Table of Contents/Projects.md → Tools & Software"
proposed_tags:
- migration
- home-assistant
- project
- esp32
- architecture
- minimax
- voice
- ai
- showcase
confidence: 0.66
approved: false
proposed_for: 324fa7828911
created: "2026-10-08"
tags:
- project
- voice
- ai
- showcase
aliases:
- MiniMax Voice
skills:
- API integration (MiniMax T2A, Speech-to-Text)
- Python standard-library services
- Docker and Docker Compose
- Audio DSP (PCM decode, linear resampling)
- MQTT and Snapcast
- Benchmarking and A/B measurement
tools:
- MiniMax cloud APIs
- OpenRouter Whisper
- faster-whisper
- piper TTS
- Snapcast
- Home Assistant
- Archify
- NixOS
---
# MiniMax Voice — Overview
## 💡 What is this project?
A migration of the **Jervis** home voice assistant to **MiniMax cloud** speech
services, while keeping a working **local fallback** so the house keeps working
without internet.
Wake-word detection stays on the local machine. Only the spoken command is sent
to the cloud. This keeps cost low and keeps the continuous microphone stream
private.
The assistant hears the wake word, transcribes the spoken command, decides an
intent with an LLM, calls a tool, and speaks the reply through the house
speakers.
## 🌟 Key Highlights (Portfolio & Employment)
- **Concept**: Replace a heavy local speech model with a cloud API, without ever
losing voice input when the internet is down.
- **Tools & Skills Showcase**: MiniMax and OpenRouter APIs, Python standard
library, Docker, MQTT, Snapcast, audio resampling, drop-in service
compatibility, graceful fallback design, reproducible benchmarks.
- **Practical Value**: Frees a few hundred MB of RAM on the home server, improves
voice quality, and keeps the house working during outages.
## 📐 Architecture & Flow
```mermaid
flowchart LR
ESP[ESP32-S3 wake word + mic] --> BRIDGE[voice_bridge]
BRIDGE --> SHIM[voice_stt_shim :5001]
SHIM -->|primary| CLOUD[Cloud STT]
SHIM -->|fallback| LOCAL[voice_whisper_fallback :5002]
SHIM --> AGENT[voice-agent Jervis :8501]
AGENT --> HA[Home Assistant]
AGENT --> TTS[tts_router]
TTS -->|fallback| PIPER[piper_tts]
TTS --> SNAP[Snapcast :4953]
PIPER --> SNAP
SNAP --> SPEAKERS[House speakers]
```
## 🔗 Related Notes & Links
- Interactive Map: [Archify Map](https://maps.lab.audasmedia.com.au/minimax_voice_tts_stt/docs/minimax-voice-map.html)
- Project workspace: `/home/sam/paseo/projects/minimax_voice_tts_stt/` on `.13`
- Vikunja: *Jervis Voice Assistant (Whisper ESP HA Flow)* project
- Outline: *Voice Assistant (Whisper ESP HA)* collection
- Related: [[Home Voice Assistant System]]
<!--
NOTES FOR THE AGENT WRITING THIS NOTE
* `type: note` is intentional — the obsidian-sorter assigns the real type.
* Do NOT choose a folder. The sorter proposes it and Sam approves.
-->

View File

@@ -20023,3 +20023,6 @@ What the watcher did, newest last.
### 2026-10-03 09:34:31 ### 2026-10-03 09:34:31
- **re-proposing** `Family Holiday Camping Wilsons Prom.md` — you edited it - **re-proposing** `Family Holiday Camping Wilsons Prom.md` — you edited it
- **proposed** `Family Holiday Camping Wilsons Prom.md` → `400 Personal Family/470 Holidays Travel` (folder already set — kept) (confidence 0.70) - **proposed** `Family Holiday Camping Wilsons Prom.md` → `400 Personal Family/470 Holidays Travel` (folder already set — kept) (confidence 0.70)
### 2026-10-08 16:16:18
- **proposed** `MiniMax Voice — Cloud STT and TTS with Local Fallback.md` → `200 projects/Tools Software WebUI` (confidence 0.66)

View File

@@ -1,6 +1,6 @@
--- ---
created: 2026-08-30 created: 2026-08-30
modified: 2026-09-05 modified: 2026-10-08
type: area type: area
status: active status: active
tags: tags:
@@ -26,24 +26,44 @@ A private, local voice assistant. Say **"Hey Jervis"** (ESP32 wake word) then a
**Command it from anywhere:** `POST http://192.168.20.13:8501/voice {"text": "add milk to the shopping list"}` — the reply is spoken. **Command it from anywhere:** `POST http://192.168.20.13:8501/voice {"text": "add milk to the shopping list"}` — the reply is spoken.
## MiniMax voice migration (2026-10-08)
Speech-to-text and text-to-speech now run on **MiniMax cloud APIs**, with a
**local fallback** so the house still works offline. Wake-word detection stays
local; only the spoken command is sent to the cloud.
- **STT**: `voice_stt_shim` (`:5001`) — cloud primary (OpenRouter `whisper-1`;
MiniMax `asr-1.0` ready), local `base.en` fallback (`:5002`).
- **TTS**: `tts_router` behind the unchanged `speak_direct.sh` — MiniMax T2A v2
with piper fallback.
- **Personality**: the agent emits a speech plan (inline sound tags + one
emotion/delivery config). Kill switch: `VOICE_PERSONALITY=0`.
- **Result**: about 239 MiB of resident RAM freed, better voices, and voice
input survives an internet outage.
- Full project note: [[MiniMax Voice — Cloud STT and TTS with Local Fallback]].
- Interactive map: [Archify Map](https://maps.lab.audasmedia.com.au/minimax_voice_tts_stt/docs/minimax-voice-map.html).
## Architecture ## Architecture
``` ```
ESP32 (wake word, mic) → MQTT voice/audio_stream ESP32 (wake word, mic) → MQTT voice/audio_stream
→ voice_bridge + whisper (STT → text) → voice_bridge (OpenWakeWord, local)
→ voice_stt_shim (:5001) — cloud STT + local base.en fallback (:5002)
→ Jervis agent (FastAPI :8501, LLM tool-calling) → Jervis agent (FastAPI :8501, LLM tool-calling)
├─ Home Assistant (lights, timers, calendar, HA todo lists) ├─ Home Assistant (lights, timers, calendar, HA todo lists)
├─ OmniRoute LLM (voice-fast: OpenRouter → opencode-go → deepseek flash) ├─ OmniRoute LLM (voice-fast: OpenRouter → opencode-go → deepseek flash)
├─ Mopidy → Spotify → Snapcast (music) ├─ Mopidy → Spotify → Snapcast (music)
├─ Apprise (family WhatsApp/Telegram/ntfy/email) ├─ Apprise (family WhatsApp/Telegram/ntfy/email)
├─ Vikunja (project tasks) ├─ Vikunja (project tasks)
└─ reply → piper → Snapcast speakers └─ reply → speak_direct.sh → tts_router (MiniMax T2A / piper) → Snapcast
``` ```
## Components (host : port) ## Components (host : port)
| Component | Host | Port | Role | | Component | Host | Port | Role |
| ------------------------------ | ------- | ----------------- | ------------------------------------- | | ------------------------------ | ------- | ----------------- | ------------------------------------- |
| ESP32-S3 (voice_assistant.ino) | devices | — | wake word "Hey Jervis", mic, LED | | ESP32-S3 (voice_assistant.ino) | devices | — | wake word "Hey Jervis", mic, LED |
| voice_bridge / voice_whisper | .13 | docker / :5000 | speech → text (faster-whisper) | | voice_bridge | .13 | docker | wake word + audio buffer |
| voice_stt_shim | .13 | :5001 | cloud STT + local fallback chain |
| voice_whisper_fallback | .13 | :5002 | local base.en STT fallback |
| tts_router | .13 | — | MiniMax T2A / piper routing |
| **Jervis agent** | .13 | **:8501** | LLM tool-calling brain (systemd) | | **Jervis agent** | .13 | **:8501** | LLM tool-calling brain (systemd) |
| OmniRoute | .13 | :20129 | LLM gateway (voice-fast route) | | OmniRoute | .13 | :20129 | LLM gateway (voice-fast route) |
| Home Assistant | .30 | :8123 | control, timers, calendar, todo lists | | Home Assistant | .30 | :8123 | control, timers, calendar, todo lists |
@@ -54,7 +74,7 @@ ESP32 (wake word, mic) → MQTT voice/audio_stream
| Vikunja | .35 | vikunja.lab | project tasks/stages (tracker) | | Vikunja | .35 | vikunja.lab | project tasks/stages (tracker) |
| Outline | .13 | :3000 | project documentation (AI) | | Outline | .13 | :3000 | project documentation (AI) |
## Status (2026-09-05) ## Status (2026-10-08)
- **Voice working end-to-end**: ESP32 wake word → whisper → agent → piper → Snapcast (fixed docker piper, HA→.13 SSH auth, resampler, MQTT client-id). - **Voice working end-to-end**: ESP32 wake word → whisper → agent → piper → Snapcast (fixed docker piper, HA→.13 SSH auth, resampler, MQTT client-id).
- **HA integration**: lights/scenes via Assist; shopping lists via HA todo (`add_shopping_item`); timers + spoken reminders (HA `timer` → automation → piper); calendar alerts spoken 30 min before (HA automation); add-calendar tool. - **HA integration**: lights/scenes via Assist; shopping lists via HA todo (`add_shopping_item`); timers + spoken reminders (HA `timer` → automation → piper); calendar alerts spoken 30 min before (HA automation); add-calendar tool.
- **Music**: play/pause/skip/volume via Mopidy → Spotify → Snapcast (credentials.json device-auth). - **Music**: play/pause/skip/volume via Mopidy → Spotify → Snapcast (credentials.json device-auth).
@@ -63,13 +83,14 @@ ESP32 (wake word, mic) → MQTT voice/audio_stream
- **AI tooling**: `project-ops` skill gives every pi agent Vikunja + Outline access (keys in `10-secrets.conf`). - **AI tooling**: `project-ops` skill gives every pi agent Vikunja + Outline access (keys in `10-secrets.conf`).
## Tools & systems (employment/skills) ## Tools & systems (employment/skills)
FastAPI · Python · Docker · MQTT · Home Assistant (REST/Assist/automations) · Whisper · piper TTS · Snapcast · Mopidy / librespot / Spotify · OmniRoute LLM routing · Tavily web search · Apprise · Vikunja & Outline APIs · NixOS (systemd services, flakes) · git-filter-repo (secret hygiene) · ESP32 (Arduino/ESP-IDF). FastAPI · Python · Docker · MQTT · Home Assistant (REST/Assist/automations) · MiniMax T2A & Speech-to-Text · OpenRouter Whisper · faster-whisper · piper TTS · Snapcast · Mopidy / librespot / Spotify · OmniRoute LLM routing · Tavily web search · Apprise · Vikunja & Outline APIs · NixOS (systemd services, flakes) · git-filter-repo (secret hygiene) · ESP32 (Arduino/ESP-IDF) · Archify.
## Key configs & links ## Key configs & links
- **Vikunja project**: `https://vikunja.lab.audasmedia.com.au` → *Jervis Voice Assistant (Whisper ESP HA Flow)* - **Vikunja project**: `https://vikunja.lab.audasmedia.com.au` → *Jervis Voice Assistant (Whisper ESP HA Flow)*
- **Outline collection**: `https://outline.lab.audasmedia.com.au` → *Voice Assistant (Whisper ESP HA)* - **Outline collection**: `https://outline.lab.audasmedia.com.au` → *Voice Assistant (Whisper ESP HA)*
- **MiniMax Voice project**: `/home/sam/paseo/projects/minimax_voice_tts_stt/` on `.13`; map at `https://maps.lab.audasmedia.com.au/minimax_voice_tts_stt/docs/`
- **Gitea**: `sam/voice-assistant` (overview / agent-api / clients docs) - **Gitea**: `sam/voice-assistant` (overview / agent-api / clients docs)
- **Repo layout / plan**: `~/chats/homeassistant` (plan.md, AGENT.md); agent at `/home/sam/voice-agent/` - **Repo layout / plan**: `~/chats/homeassistant` (plan.md, AGENT.md); agent at `/home/sam/voice-agent/`
- **Secrets**: `~/.config/environment.d/10-secrets.conf` on .27/.13/.51 (VIKUNJA_TOKEN, OL_API_KEY, etc.) — see also the `secrets_fixes` note. - **Secrets**: `~/.config/environment.d/10-secrets.conf` on .27/.13/.51 (VIKUNJA_TOKEN, OL_API_KEY, etc.) — see also the `secrets_fixes` note.
See also: [[Docker Containers]], [[Home Network Map Overview]]. See also: [[Docker Containers]], [[Home Network Map Overview]], [[MiniMax Voice — Cloud STT and TTS with Local Fallback]].