102 lines
3.0 KiB
Markdown
102 lines
3.0 KiB
Markdown
---
|
|
title: MiniMax Voice
|
|
type: note
|
|
status: active
|
|
proposed_folder: 200 projects/Tools Software WebUI
|
|
proposed_toc: "100 Table of Contents/Projects.md → Tools & Software"
|
|
proposed_tags:
|
|
- migration
|
|
- home-assistant
|
|
- project
|
|
- esp32
|
|
- architecture
|
|
- minimax
|
|
- voice
|
|
- ai
|
|
- showcase
|
|
confidence: 0.66
|
|
approved: false
|
|
proposed_for: 324fa7828911
|
|
created: "2026-10-08"
|
|
tags:
|
|
- project
|
|
- voice
|
|
- ai
|
|
- showcase
|
|
aliases:
|
|
- MiniMax Voice
|
|
skills:
|
|
- API integration (MiniMax T2A, Speech-to-Text)
|
|
- Python standard-library services
|
|
- Docker and Docker Compose
|
|
- Audio DSP (PCM decode, linear resampling)
|
|
- MQTT and Snapcast
|
|
- Benchmarking and A/B measurement
|
|
tools:
|
|
- MiniMax cloud APIs
|
|
- OpenRouter Whisper
|
|
- faster-whisper
|
|
- piper TTS
|
|
- Snapcast
|
|
- Home Assistant
|
|
- Archify
|
|
- NixOS
|
|
---
|
|
|
|
# MiniMax Voice — Overview
|
|
|
|
## 💡 What is this project?
|
|
|
|
A migration of the **Jervis** home voice assistant to **MiniMax cloud** speech
|
|
services, while keeping a working **local fallback** so the house keeps working
|
|
without internet.
|
|
|
|
Wake-word detection stays on the local machine. Only the spoken command is sent
|
|
to the cloud. This keeps cost low and keeps the continuous microphone stream
|
|
private.
|
|
|
|
The assistant hears the wake word, transcribes the spoken command, decides an
|
|
intent with an LLM, calls a tool, and speaks the reply through the house
|
|
speakers.
|
|
|
|
## 🌟 Key Highlights (Portfolio & Employment)
|
|
|
|
- **Concept**: Replace a heavy local speech model with a cloud API, without ever
|
|
losing voice input when the internet is down.
|
|
- **Tools & Skills Showcase**: MiniMax and OpenRouter APIs, Python standard
|
|
library, Docker, MQTT, Snapcast, audio resampling, drop-in service
|
|
compatibility, graceful fallback design, reproducible benchmarks.
|
|
- **Practical Value**: Frees a few hundred MB of RAM on the home server, improves
|
|
voice quality, and keeps the house working during outages.
|
|
|
|
## 📐 Architecture & Flow
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
ESP[ESP32-S3 wake word + mic] --> BRIDGE[voice_bridge]
|
|
BRIDGE --> SHIM[voice_stt_shim :5001]
|
|
SHIM -->|primary| CLOUD[Cloud STT]
|
|
SHIM -->|fallback| LOCAL[voice_whisper_fallback :5002]
|
|
SHIM --> AGENT[voice-agent Jervis :8501]
|
|
AGENT --> HA[Home Assistant]
|
|
AGENT --> TTS[tts_router]
|
|
TTS -->|fallback| PIPER[piper_tts]
|
|
TTS --> SNAP[Snapcast :4953]
|
|
PIPER --> SNAP
|
|
SNAP --> SPEAKERS[House speakers]
|
|
```
|
|
|
|
## 🔗 Related Notes & Links
|
|
|
|
- Interactive Map: [Archify Map](https://maps.lab.audasmedia.com.au/minimax_voice_tts_stt/docs/minimax-voice-map.html)
|
|
- Project workspace: `/home/sam/paseo/projects/minimax_voice_tts_stt/` on `.13`
|
|
- Vikunja: *Jervis Voice Assistant (Whisper ESP HA Flow)* project
|
|
- Outline: *Voice Assistant (Whisper ESP HA)* collection
|
|
- Related: [[Home Voice Assistant System]]
|
|
|
|
<!--
|
|
NOTES FOR THE AGENT WRITING THIS NOTE
|
|
* `type: note` is intentional — the obsidian-sorter assigns the real type.
|
|
* Do NOT choose a folder. The sorter proposes it and Sam approves.
|
|
-->
|