Files
obsidian-vault/010 inbox/MiniMax Voice — Cloud STT and TTS with Local Fallback.md

3.0 KiB

title, type, status, proposed_folder, proposed_toc, proposed_tags, confidence, approved, proposed_for, created, tags, aliases, skills, tools
title type status proposed_folder proposed_toc proposed_tags confidence approved proposed_for created tags aliases skills tools
MiniMax Voice note active 200 projects/Tools Software WebUI 100 Table of Contents/Projects.md → Tools & Software
migration
home-assistant
project
esp32
architecture
minimax
voice
ai
showcase
0.66 false 324fa7828911 2026-10-08
project
voice
ai
showcase
MiniMax Voice
API integration (MiniMax T2A, Speech-to-Text)
Python standard-library services
Docker and Docker Compose
Audio DSP (PCM decode, linear resampling)
MQTT and Snapcast
Benchmarking and A/B measurement
MiniMax cloud APIs
OpenRouter Whisper
faster-whisper
piper TTS
Snapcast
Home Assistant
Archify
NixOS

MiniMax Voice — Overview

💡 What is this project?

A migration of the Jervis home voice assistant to MiniMax cloud speech services, while keeping a working local fallback so the house keeps working without internet.

Wake-word detection stays on the local machine. Only the spoken command is sent to the cloud. This keeps cost low and keeps the continuous microphone stream private.

The assistant hears the wake word, transcribes the spoken command, decides an intent with an LLM, calls a tool, and speaks the reply through the house speakers.

🌟 Key Highlights (Portfolio & Employment)

  • Concept: Replace a heavy local speech model with a cloud API, without ever losing voice input when the internet is down.
  • Tools & Skills Showcase: MiniMax and OpenRouter APIs, Python standard library, Docker, MQTT, Snapcast, audio resampling, drop-in service compatibility, graceful fallback design, reproducible benchmarks.
  • Practical Value: Frees a few hundred MB of RAM on the home server, improves voice quality, and keeps the house working during outages.

📐 Architecture & Flow

flowchart LR
    ESP[ESP32-S3 wake word + mic] --> BRIDGE[voice_bridge]
    BRIDGE --> SHIM[voice_stt_shim :5001]
    SHIM -->|primary| CLOUD[Cloud STT]
    SHIM -->|fallback| LOCAL[voice_whisper_fallback :5002]
    SHIM --> AGENT[voice-agent Jervis :8501]
    AGENT --> HA[Home Assistant]
    AGENT --> TTS[tts_router]
    TTS -->|fallback| PIPER[piper_tts]
    TTS --> SNAP[Snapcast :4953]
    PIPER --> SNAP
    SNAP --> SPEAKERS[House speakers]
  • Interactive Map: Archify Map
  • Project workspace: /home/sam/paseo/projects/minimax_voice_tts_stt/ on .13
  • Vikunja: Jervis Voice Assistant (Whisper ESP HA Flow) project
  • Outline: Voice Assistant (Whisper ESP HA) collection
  • Related: Home Voice Assistant System