Files
obsidian-vault/300 areas/360 Dev-Ops Network Computers/Error Tracking.md

9.0 KiB

created, modified, title, summary, type, client, project, status, priority
created modified title summary type client project status priority
2026-09-06 2026-10-06 Error Tracking Log of network/system errors encountered and their fixes area sam devops active 5

Error Tracking

A running log of errors encountered across the network, with root causes and fixes. Append new entries as they occur.


2026-09-06 — Router: "Another management entity is configuring the Broadband Router"

Symptom

While in the router admin (192.168.20.1), saw:

Broadband Router Error — Another management entity is currently configuring the Broadband Router. Please close the Broadband Router Configuration window, and try again in a few minutes.

Error appeared repeatedly; also coincided with .13 losing all network connectivity (wired + Wi-Fi + Tailscale unreachable from .27).

Root cause

Not a second admin session and not a rogue device — .13's own flapping Wi-Fi adapter (wlp4s0) was seen by the router as a "management entity":

  • .13 has two MACs / two IPs on the router table (root of the confusion):
    • Wired enp42s0 = MAC 34:5a:60:a4:cc:8c → 192.168.20.13 (LAN)
    • Wi-Fi wlp4s0 = MAC a0:f3:c1:d0:e3:2b → 192.168.20.46 (2.4G)
  • Wi-Fi kept going DORMANT / re-registering and the default route flipped onto it → router interpreted the churn as a second configuring device.
  • sudo reboot on .13 cleared it. Error has not returned since.

Verification / non-causes

  • Router login log showed no login except .27 → no foreign admin session.
  • Client table showed only known devices (ESP32s, Pis, .27/.35/.51/.13, etc.).
  • Not caused by Langfuse/Clickhouse removal (unrelated; those never touch networking).

Potential fix — pin the Wi-Fi MAC address

If the error (or Wi-Fi flakiness on .13) recurs, the likely fix is to stop the Wi-Fi adapter's MAC churn:

  • On .13, disable MAC address randomization in NetworkManager for wlp4s0, forcing a stable MAC so the router sees one consistent identity per interface:
    # /etc/NetworkManager/conf.d/ (or via home-manager)
    [connection]
    wifi.cloned-mac-address=permanent   # or "stable"
    # per-interface: wifi.mac-randomization-disabled
    
  • On NixOS this is done in hardware-configuration.nix / networking options (e.g. networking.wireless or NetworkManager connection settings in home-manager), or simply disable the Wi-Fi interface entirely if wired is preferred: sudo nmcli radio wifi off.
  • .13 network layout: enp42s0 = wired .13, wlp4s0 = Wi-Fi .46 (2.4G). Default route prefers Wi-Fi (lower metric) — worth reviewing if flakiness returns.
  • 192.168.80.0/20 bridge on .13 is normal Docker (wherewoof-db_default subnet) — not an issue.

2026-09-06 — Voice agent: "feed the fish" did nothing, LED red

Symptom

The voice command "feed the fish" returned no action. The ESP32 LED went red.

Root cause

Two failures blocked the request path.

  • OmniRoute combo voice-fast used old model IDs. The new IDs are openai/*. The combo returned 503 "all accounts inactive".
  • The agent MQTT intake connected once. A network flap broke that connection. The agent did not retry, so it lost the ESP32 text feed.

Fix

  • Recreated the voice-fast combo with three healthy OpenRouter models.
  • Changed app.py MQTT intake to retry and reconnect. Backup: app.py.bak-mqtt-20260906.
  • Run sudo systemctl restart voice-agent to load the change.
  • DeepSeek key is valid but the account balance is zero. Add credit at platform.deepseek.com.

2026-09-09 — Garage S3 web returns 404/403 for site images (anonymous hosting)

Symptom

A Go + htmx site (being built for kontra.day) tried to serve uploaded images through Garage S3 web (:3902) and got 404 on the bucket root and 403 Forbidden: Garage does not support anonymous access yet in docker logs family-home-lab-garage-1.

Root cause

Two things misled the diagnosis:

  1. Wrong claim (from an earlier AI): it said [s3_web] was absent and the web server wasn't running. False — [s3_web] IS in garage.toml (bind_addr = "[::]:3902", root_domain = ".web.lab.audasmedia.com.au", index = "index.html") and the log shows Web server listening on http://[::]:3902. Verified: ss -tln shows 3902 bound (Garage web server), 3903 admin, 3909 garage-webui.
  2. Real issue: Garage S3 Web is not anonymous static hosting. It requires per-bucket web access + a read-scoped access key. Without it, every request gets 403 Forbidden: Garage does not support anonymous access yet, and the bucket root 404s with no index/perm.

Fix / correct pattern

  • Do NOT point site/app code at :3902 for anonymous browsing.
  • Use the S3 API (:3900, region garage) with a read-scoped key for uploads/reads.
  • Serve images through the app or Caddy (reverse-proxy to Garage with the read key, or a Caddy file_server static dir) — not via :3902 directly.
  • If web-serving a bucket is truly needed, enable per-bucket web access + a guest/read key in Garage (not a global switch).

Verification / non-causes

  • [s3_web] present; web server confirmed running (log line + port 3902 bound) → the earlier "no s3_web / not running" diagnosis was wrong.
  • Port map: 3900 S3 API · 3902 S3 web · 3903 admin API · 3909 garage-webui. No port conflict.
  • Documented in the Garage section of Backup System — Borg, Kopia & Restic and the kontra.day site handoff.

2026-09-14 — Voice: ESP32 green, never blue; then no response

Symptom

The ESP32 LED turned green but never blue. The bridge, whisper, and broker were up.

Root cause

Three failures in one chain.

  • The voice_bridge OpenWakeWord process was wedged. It used 0.01% CPU and ignored audio.
  • The ESP32 MQTT client was in a stale state. The broker dropped it for keepalive timeouts. It published no audio.
  • The voice-agent process ran for 7 days without the MQTT reconnect patch. Journald restarted and lost the agent stdout socket, so the agent logs vanished. The agent never re-subscribed to the .30 topic.

Fix

  • Restart the bridge container: docker restart voice_bridge
  • Power-cycle the ESP32
  • Restart the agent service: sudo systemctl restart voice-agent
  • Verify: journal shows [mqtt] listening on 192.168.20.30:1883 homeassistant/voice/text

Note

The agent MQTT intake now retries and reconnects. A future broker flap heals itself. Report the LED state to exercise the chain end to end. Keepalive timeouts on ESP32 suggest a weak WiFi link; consider a stable keepalive or fixed install if it recurs.


2026-10-05 — Music stack silent: Mopidy has no Spotify, "NixOS-Speaker" missing from Spotify

Symptom

  • No audio from any Snapcast speaker.
  • The Spotify client did not list the NixOS-Speaker Connect device.
  • The Mopidy / Iris Spotify library was empty.

Root cause

One network fault caused two failures.

  • The .13 reboot at 20:16 broke DNS. NetworkManager assumed enp42s0 and wrote a profile without the gateway. The IPv4 default route dropped.
  • librespot could not reach clienttoken.spotify.com. systemd restarted it 6,361 times between 20:16 and 05:32.
  • mopidy refreshes its Spotify OAuth token once, at backend start. The refresh failed at 20:16:14 and never retried. mopidy.service stayed active with a dead Spotify backend for 10.5 hours.

idle on all Snapcast streams is normal when nothing plays. It is not a fault indicator.

Fix

Restart the FIFO writers first, then Snapserver.

# Mopidy Spotify. This alone restored the outage.
sudo systemctl restart mopidy

# Use this instead if Snapcast is also silent.
sudo systemctl restart librespot mopidy && sleep 3 && sudo systemctl restart snapserver

If Spotify stays empty after the restart, the stored refresh token was revoked. Re-run the mopidy-spotify OAuth flow (mopidy spotify).

Verify

systemctl is-active snapserver mopidy librespot
journalctl -u mopidy -b | grep -iE 'Logged into Spotify|OAuth'
journalctl -u librespot -b | grep -c 'Main process exited'   # expect 0

Journal signatures on .13:

Log line Meaning
OAuth token refresh failed: Unknown error. Spotify backend dead. Restart Mopidy.
Failed to load Spotify user profile Same fault, next line.
Logged into Spotify Web API as guzzbinkles Healthy.
Refreshed 152 Spotify playlists Healthy, library loaded.

Notes

  • Root cause fixed in commit 47f956e: networkmanager.unmanaged = [ "enp42s0" ] plus name_servers in resolvconf.extraConfig.
  • Ports on .13: 1704 Snapcast, 1705 control, 1780 web/JSON-RPC, 4953 announcements, 57621 librespot Connect handshake.
  • Open, low impact: .13 publishes no _spotify-connect._tcp mDNS record. librespot's libmdns and avahi-daemon both bind UDP 5353. Spotify Connect still worked through Spotify's cloud, so no change was made.
  • Open, cosmetic: boot ordering still starts snapserver before its FIFO writers. It logs End of file, length: 0 at boot and self-heals.