Files
obsidian-vault/010 inbox/MiniMax Voice — Cloud STT and TTS with Local Fallback.md

102 lines
3.0 KiB
Markdown

---
title: MiniMax Voice
type: note
status: active
proposed_folder: 200 projects/Tools Software WebUI
proposed_toc: "100 Table of Contents/Projects.md → Tools & Software"
proposed_tags:
- migration
- home-assistant
- project
- esp32
- architecture
- minimax
- voice
- ai
- showcase
confidence: 0.66
approved: false
proposed_for: 324fa7828911
created: "2026-10-08"
tags:
- project
- voice
- ai
- showcase
aliases:
- MiniMax Voice
skills:
- API integration (MiniMax T2A, Speech-to-Text)
- Python standard-library services
- Docker and Docker Compose
- Audio DSP (PCM decode, linear resampling)
- MQTT and Snapcast
- Benchmarking and A/B measurement
tools:
- MiniMax cloud APIs
- OpenRouter Whisper
- faster-whisper
- piper TTS
- Snapcast
- Home Assistant
- Archify
- NixOS
---
# MiniMax Voice — Overview
## 💡 What is this project?
A migration of the **Jervis** home voice assistant to **MiniMax cloud** speech
services, while keeping a working **local fallback** so the house keeps working
without internet.
Wake-word detection stays on the local machine. Only the spoken command is sent
to the cloud. This keeps cost low and keeps the continuous microphone stream
private.
The assistant hears the wake word, transcribes the spoken command, decides an
intent with an LLM, calls a tool, and speaks the reply through the house
speakers.
## 🌟 Key Highlights (Portfolio & Employment)
- **Concept**: Replace a heavy local speech model with a cloud API, without ever
losing voice input when the internet is down.
- **Tools & Skills Showcase**: MiniMax and OpenRouter APIs, Python standard
library, Docker, MQTT, Snapcast, audio resampling, drop-in service
compatibility, graceful fallback design, reproducible benchmarks.
- **Practical Value**: Frees a few hundred MB of RAM on the home server, improves
voice quality, and keeps the house working during outages.
## 📐 Architecture & Flow
```mermaid
flowchart LR
ESP[ESP32-S3 wake word + mic] --> BRIDGE[voice_bridge]
BRIDGE --> SHIM[voice_stt_shim :5001]
SHIM -->|primary| CLOUD[Cloud STT]
SHIM -->|fallback| LOCAL[voice_whisper_fallback :5002]
SHIM --> AGENT[voice-agent Jervis :8501]
AGENT --> HA[Home Assistant]
AGENT --> TTS[tts_router]
TTS -->|fallback| PIPER[piper_tts]
TTS --> SNAP[Snapcast :4953]
PIPER --> SNAP
SNAP --> SPEAKERS[House speakers]
```
## 🔗 Related Notes & Links
- Interactive Map: [Archify Map](https://maps.lab.audasmedia.com.au/minimax_voice_tts_stt/docs/minimax-voice-map.html)
- Project workspace: `/home/sam/paseo/projects/minimax_voice_tts_stt/` on `.13`
- Vikunja: *Jervis Voice Assistant (Whisper ESP HA Flow)* project
- Outline: *Voice Assistant (Whisper ESP HA)* collection
- Related: [[Home Voice Assistant System]]
<!--
NOTES FOR THE AGENT WRITING THIS NOTE
* `type: note` is intentional — the obsidian-sorter assigns the real type.
* Do NOT choose a folder. The sorter proposes it and Sam approves.
-->