--- title: MiniMax Voice type: note status: active proposed_folder: 200 projects/Tools Software WebUI proposed_toc: "100 Table of Contents/Projects.md → Tools & Software" proposed_tags: - migration - home-assistant - project - architecture - minimax - voice - ai - showcase confidence: 0.71 approved: false proposed_for: 636ce731505f created: "2026-10-08" tags: - project - voice - ai - showcase aliases: - MiniMax Voice skills: - API integration (MiniMax T2A, Speech-to-Text) - Python standard-library services - Docker and Docker Compose - Audio DSP (PCM decode, linear resampling) - MQTT and Snapcast - Benchmarking and A/B measurement tools: - MiniMax cloud APIs - OpenRouter Whisper - faster-whisper - piper TTS - Snapcast - Home Assistant - Archify - NixOS --- # MiniMax Voice — Overview ## 💡 What is this project? A migration of the **Jervis** home voice assistant to **MiniMax cloud** speech services, while keeping a working **local fallback** so the house keeps working without internet. Wake-word detection stays on the local machine. Only the spoken command is sent to the cloud. This keeps cost low and keeps the continuous microphone stream private. The assistant hears the wake word, transcribes the spoken command, decides an intent with an LLM, calls a tool, and speaks the reply through the house speakers. ## 🌟 Key Highlights (Portfolio & Employment) - **Concept**: Replace a heavy local speech model with a cloud API, without ever losing voice input when the internet is down. - **Tools & Skills Showcase**: MiniMax and OpenRouter APIs, Python standard library, Docker, MQTT, Snapcast, audio resampling, drop-in service compatibility, graceful fallback design, reproducible benchmarks. - **Practical Value**: Frees a few hundred MB of RAM on the home server, improves voice quality, and keeps the house working during outages. ## 📐 Architecture & Flow ```mermaid flowchart LR ESP[ESP32-S3 wake word + mic] --> BRIDGE[voice_bridge] BRIDGE --> SHIM[voice_stt_shim :5001] SHIM -->|primary| CLOUD[Cloud STT] SHIM -->|fallback| LOCAL[voice_whisper_fallback :5002] SHIM --> AGENT[voice-agent Jervis :8501] AGENT --> HA[Home Assistant] AGENT --> TTS[tts_router] TTS -->|fallback| PIPER[piper_tts] TTS --> SNAP[Snapcast :4953] PIPER --> SNAP SNAP --> SPEAKERS[House speakers] ``` ## 🏠 Home Assistant integration Home Assistant announcements no longer call piper directly. `script.jervis_say` publishes an instruction to MQTT `homeassistant/voice/announce`, and the agent runs it in **announce mode** with **no tools**, so an announcement cannot trigger a tool call. Voice reminders use the helper `input_text.voice_reminder_message` (0–255 characters) and trigger on the `timer.finished` event, so cancelling a timer is silent. The calendar reminder and `script.piper_tts_announcement` were repointed from the missing `rest_command.jervis_say` onto `script.jervis_say`. The doorbell script `doorbell_tts_to_wav.sh` now uses the Docker piper plus `ffmpeg`. The live STT primary is MiniMax `asr-1.0` on the funded `MINIMAX_SUBSCRIPTION_KEY`, with the local `base.en` fallback unchanged. ## 🔗 Related Notes & Links - Interactive Map: [Archify Map](https://maps.lab.audasmedia.com.au/minimax_voice_tts_stt/docs/minimax-voice-map.html) - Project workspace: `/home/sam/paseo/projects/minimax_voice_tts_stt/` on `.13` - Vikunja: *Jervis Voice Assistant (Whisper ESP HA Flow)* project - Outline: *Voice Assistant (Whisper ESP HA)* collection - Related: [[Home Voice Assistant System]]