Files
obsidian-vault/010 inbox/MiniMax Voice — Cloud STT and TTS with Local Fallback.md

3.7 KiB
Raw Blame History

title, type, status, proposed_folder, proposed_toc, proposed_tags, confidence, approved, proposed_for, created, tags, aliases, skills, tools
title type status proposed_folder proposed_toc proposed_tags confidence approved proposed_for created tags aliases skills tools
MiniMax Voice note active 200 projects/Tools Software WebUI 100 Table of Contents/Projects.md → Tools & Software
migration
home-assistant
project
architecture
minimax
voice
ai
showcase
0.71 false 636ce731505f 2026-10-08
project
voice
ai
showcase
MiniMax Voice
API integration (MiniMax T2A, Speech-to-Text)
Python standard-library services
Docker and Docker Compose
Audio DSP (PCM decode, linear resampling)
MQTT and Snapcast
Benchmarking and A/B measurement
MiniMax cloud APIs
OpenRouter Whisper
faster-whisper
piper TTS
Snapcast
Home Assistant
Archify
NixOS

MiniMax Voice — Overview

💡 What is this project?

A migration of the Jervis home voice assistant to MiniMax cloud speech services, while keeping a working local fallback so the house keeps working without internet.

Wake-word detection stays on the local machine. Only the spoken command is sent to the cloud. This keeps cost low and keeps the continuous microphone stream private.

The assistant hears the wake word, transcribes the spoken command, decides an intent with an LLM, calls a tool, and speaks the reply through the house speakers.

🌟 Key Highlights (Portfolio & Employment)

  • Concept: Replace a heavy local speech model with a cloud API, without ever losing voice input when the internet is down.
  • Tools & Skills Showcase: MiniMax and OpenRouter APIs, Python standard library, Docker, MQTT, Snapcast, audio resampling, drop-in service compatibility, graceful fallback design, reproducible benchmarks.
  • Practical Value: Frees a few hundred MB of RAM on the home server, improves voice quality, and keeps the house working during outages.

📐 Architecture & Flow

flowchart LR
    ESP[ESP32-S3 wake word + mic] --> BRIDGE[voice_bridge]
    BRIDGE --> SHIM[voice_stt_shim :5001]
    SHIM -->|primary| CLOUD[Cloud STT]
    SHIM -->|fallback| LOCAL[voice_whisper_fallback :5002]
    SHIM --> AGENT[voice-agent Jervis :8501]
    AGENT --> HA[Home Assistant]
    AGENT --> TTS[tts_router]
    TTS -->|fallback| PIPER[piper_tts]
    TTS --> SNAP[Snapcast :4953]
    PIPER --> SNAP
    SNAP --> SPEAKERS[House speakers]

🏠 Home Assistant integration

Home Assistant announcements no longer call piper directly. script.jervis_say publishes an instruction to MQTT homeassistant/voice/announce, and the agent runs it in announce mode with no tools, so an announcement cannot trigger a tool call. Voice reminders use the helper input_text.voice_reminder_message (0–255 characters) and trigger on the timer.finished event, so cancelling a timer is silent. The calendar reminder and script.piper_tts_announcement were repointed from the missing rest_command.jervis_say onto script.jervis_say. The doorbell script doorbell_tts_to_wav.sh now uses the Docker piper plus ffmpeg. The live STT primary is MiniMax asr-1.0 on the funded MINIMAX_SUBSCRIPTION_KEY, with the local base.en fallback unchanged.

  • Interactive Map: Archify Map
  • Project workspace: /home/sam/paseo/projects/minimax_voice_tts_stt/ on .13
  • Vikunja: Jervis Voice Assistant (Whisper ESP HA Flow) project
  • Outline: Voice Assistant (Whisper ESP HA) collection
  • Related: Home Voice Assistant System