Files
family_home_lab/ON_OFF.md

22 KiB
Raw Blame History

ON_OFF.md — Turning services on and off from the Family Console

Target: implement on/off control for RAM-heavy services from console.lab.audasmedia.com.au Repo: sam/family_home_lab (this repo) · deployed to .13:/home/sam/Docker/Containers/family-home-lab Status: plan for the on/off controller; console UX additions (back-to-top, collapsible panels) and the LMMS Xvfb fix are already implemented. Last updated: 2026-10-07

Confirmed decisions

  1. Togglable services default to OFF. The RAM is the point — freeing it immediately matters more than a 30–60 s wait on first use. The UI must make Start the obvious primary action and show the freed/used RAM on every card. Consequence: "off" must actually hold across deployments — see §4.6. Do not skip that section.

0. Why

.13 has ~46 containers plus several native services. Several are heavy and idle almost all day. A 2026-10 audit found roughly 2.4 GB of addressable memory, most of it sitting idle:

Service Idle RAM
Media editors (GIMP, Kdenlive, Audacity, LMMS) ~1.5 GB incl. a 570 MB LMMS Xvfb
Prefect (server + worker + photo-dashboard) ~650 MB
Paperclip ~495 MB
n8n ~286 MB
WorldMonitor (4 containers) ~120 MB

The goal: let a family member start a tool when they need it and stop it when they don't, from the console they already use — free RAM without free-of-charge surprise.


1. Current state (verified 2026-10-06)

1.1 What exists

Thing Path / detail
Console portal (FastAPI + Jinja + HTMX) portal/ — main.py, tools.py, auth.py, database.py, config.py
Tool catalogue portal/tools.py — a Tool dataclass + _cat() returning the list (~335 lines)
Templates portal/templates/ — dashboard.html, admin.html, voice.html, files.html, partials/
Design system DESIGN.md — CSS tokens, sticker palette, status colours already defined
Routes /, /tool/{id}, /files, /voice, /admin, /admin/users, /admin/pi, /account/password
Auth session login; roles via admin: bool on each Tool; users seeded from .env
Portal container family-home-lab-portal-1 on .13:8500
DB Postgres (pgvector/pgvector:pg16) inside the same compose project
Deployment edit here on .27 → copy to .13:/home/sam/Docker/Containers/family-home-lab → docker compose up -d

1.2 The thing that lies today

Tool already has a status field:

@dataclass(frozen=True)
class Tool:
    ...
    status: str = "online"   # <-- a hardcoded string, never updated

It is a literal default, not live state. The dashboard currently reports "online" for every tool regardless of reality. You cannot show on/off state until this becomes real, so fixing it is part 1 of this work, not a nice-to-have.

1.3 The blast radius

The portal's own compose project contains postgres, redis, rabbitmq, garage, worker, transcriber, transcriber-mus. If the console can stop those, it can kill itself. Any design must make the console's own dependencies untouchable.


2. Issues to solve

  1. Two different control planes. Targets are a mix of Docker containers and systemd user services. A container can reach the Docker API; it cannot reach systemctl --user on the host without extra plumbing (D-Bus mount, or a helper). One mechanism will not cover both.

    Containers: gimp, video-editor, audio-editor, lmms, prefect-server, n8n, worldmonitor*, voice_whisper, headroom, … systemd user services: paperclipai, prefect-worker, photo-dashboard, where-woof, chrome-pi, engram, voice-agent, lan-mouse. systemd system services: snapserver, librespot, mopidy, caddy.

  2. /var/run/docker.sock is root-equivalent. The portal is internet-reachable via Caddy and protected only by a family login. Mounting the socket into it means one XSS or one weak password equals full host root. Do not do this.

  3. Self-destruct risk. The console must not be able to stop postgres, redis, rabbitmq, garage, portal, worker, or any service the console itself depends on. This needs to be enforced by an allowlist, not by UI politeness.

  4. The dashboard has no live state (§1.2). Needs a real status source before toggles make sense.

  5. State drift. Every service uses restart: unless-stopped, which is helpful — a docker stop sticks. But any future docker compose up -d silently restarts everything, undoing the user's choices. Toggle state must be persisted and reconciled, and drift must be visible.

  6. Startup is slow and the UI will hang. LinuxServer Webtop containers take 30–60 s to become usable. The endpoint must return immediately and the card must show starting, then poll.

  7. Ordering on start. worker/transcriber need redis+rabbitmq+postgres. Media containers need the shared-media volume mounted. Starting things in the wrong order produces confusing errors.

  8. Users cannot make informed choices. Nothing shows what a service costs. The card should show approximate RAM so the trade-off is visible.

  9. Shared services affect everyone. Stopping GIMP stops it for Harry and Finn too. Needs either admin-only gating or a visible "who's using this" indicator.

  10. Failure feedback. If a container enters a crash loop the user must see "failed" with a reason, not a spinner.

  11. Distinguish "stopped on purpose" from "crashed". These need different UI treatment. The design system already has the colours: green = online, orange = busy/starting, faint grey = offline.

  12. Idle auto-off is tempting but risky. Auto-stopping a media editor mid-edit would be hostile. If implemented, it must detect actual use (selkies has live websocket connections) and only apply to genuinely abandoned sessions.

  13. ✅ DONE (2026-10-07): LMMS Xvfb framebuffer fixed. MAX_RES: 2560x1440 added to music/docker-compose.yml; the webtop clamped the virtual screen from 15360x8640 (~506MB/plane) to 2560x1440 (~14MB/plane). Deployed + verified on .13.


3. Targets

Safety ratings: ✅ safe · ⚠️ needs care · ❌ never

Service Kind Idle RAM Safe? Who Notes
gimp container ~233 MB ✅ all users Webtop; slow start (~40 s)
video-editor (kdenlive) container ~237 MB ✅ all users Webtop
audio-editor (audacity) container ~201 MB ✅ all users Webtop
lmms container ~798 MB + 570 MB Xvfb ✅ all users Xvfb fixed 2026-10-07 (MAX_RES)
prefect-server container ~411 MB ✅ admin Only needed while ingesting
prefect-worker systemd user ~129 MB ✅ admin Start with the server
photo-dashboard systemd user ~112 MB ✅ admin Pairs with prefect
paperclipai systemd user ~495 MB ✅ admin Idle since 2026-09-26
n8n container ~286 MB ⚠️ admin Check for scheduled workflows first
worldmonitor +3 containers ~120 MB ✅ admin Group toggle (4 containers)
engram systemd user ~1 MB ✅ admin Dead/empty — candidate for removal not toggling
headroom container ~18 MB n/a — Leave running (JEV proxy, in the pi path)
voice_whisper container 372–857 MB ❌ — House voice depends on it
postgres, redis, rabbitmq, garage, portal, worker containers — ❌ — Console's own dependencies
snapserver, librespot, mopidy systemd system — ❌ — House audio; FIFO ordering trap (see below)

⚠️ Snapcast ordering trap: snapserver opens its FIFOs at start. If you restart snapserver before librespot/mopidy, the streams come up dead (End of file, length: 0) and audio is silent. See /home/sam/chats/snapcast/snapcast-findings.md. Keep these out of the togglable set.


4. Design

4.1 Architecture — a host-side service controller

Do not give the portal the Docker socket. Introduce one small controller that owns all privileged actions and exposes a narrow, allowlisted API.

 Browser (HTMX)
      │
      ▼
 portal container (.13:8500)                    ← no docker.sock, no host access
      │  HTTP  →  http://host.docker.internal:8091
      ▼
 service-controller  (systemd USER unit on .13, runs as `sam`, 127.0.0.1:8091)
      ├─ Docker API      (sam is in the `docker` group)  → containers
      └─ systemctl --user                                → user services
      ▲
      └─ ALLOWLIST + metadata lives here (single source of truth)

Why a host-side systemd user unit:

  • systemctl --user needs the user manager, which a container does not have. Running on the host covers user services natively.
  • sam is already in the docker group, so the same process can also drive containers — no privileged container, no socket in the portal.
  • host.docker.internal is reachable from the portal via Compose: extra_hosts: ["host.docker.internal:host-gateway"].

The allowlist is the security boundary. Anything not listed cannot be started or stopped, and there is no generic passthrough. Reject unknown ids with 404 — never forward caller-supplied names.

Alternative considered and rejected: tecnativa/docker-socket-proxy. It is a fine tool but only covers Docker, and we need systemd user services too. It would still leave the portal holding a token that can act on containers.

4.2 Controller API (localhost only)

Method Path Purpose
GET /services List allowlisted services with live state + metadata
GET /services/{id} One service: state, since, error, ram_mb
POST /services/{id}/start Start. Returns immediately (202)
POST /services/{id}/stop Stop
GET /health Liveness for the portal

Response shape per service:

{
  "id": "gimp",
  "kind": "container",
  "target": "family-home-lab-gimp-1",
  "group": "media",
  "label": "GIMP",
  "ram_mb": 233,
  "togglable": true,
  "state": "running",          // running | stopped | starting | stopping | failed | unknown
  "since": "2026-10-06T05:12:00+11:00",
  "error": null
}

Bind to 127.0.0.1 only. Add a shared secret header (X-Controller-Token) read from ~/.config/environment.d/10-secrets.conf so a stray process on the LAN cannot drive it.

4.3 State and drift (§2.5)

  • Persist desired state per service in the portal's Postgres (service_state table: id, desired, updated_by, updated_at).
  • The controller reports actual state. The UI shows both when they disagree and offers "reconcile".
  • On portal startup, compare desired vs actual and surface drift rather than silently "fixing" it.
  • Document that docker compose up -d re-starts everything. Consider a deploy/ note or a post-deploy reconcile step so deployments don't silently turn the whole lab back on.

4.4 Portal changes

Extend the Tool dataclass (portal/tools.py) — keep backwards compatibility:

@dataclass(frozen=True)
class Tool:
    ...existing fields...
    # NEW — optional on/off control
    service_id: str | None = None      # controller id, e.g. "gimp"
    ram_mb: int | None = None          # shown on the card
    togglable: bool = False            # only True if allowlisted

Then tag the relevant entries in _cat(), e.g.:

Tool("gimp", "GIMP", "image", "Photo editing & retouching",
     "https://gimp.lab.audasmedia.com.au",
     service_id="gimp", ram_mb=233, togglable=True),

New routes (portal/main.py):

Route Behaviour
GET /api/services Proxy the controller's list; HTMX polls this
POST /service/{id}/start Admin (or owner) only → controller → return the card partial
POST /service/{id}/stop Same
GET /partials/service/{id} The card fragment, for HTMX swaps

All mutating routes must enforce auth and the admin flag unless explicitly opened to users.

Template — portal/templates/partials/ gets a service_card.html with a status pill and a button. Use the tokens already in DESIGN.md:

  • accent-green #1aae39 → running
  • accent-orange #dd5b00 → starting / stopping
  • ink-faint #a39e98 → stopped
  • add a red-ish failed state (reuse accent-orange-deep #793400 pending a design decision)

Use HTMX polling while a service is in a transitional state (hx-trigger="every 2s"), and stop polling once it settles — otherwise the dashboard hammers the controller.

4.6 Making "off by default" actually hold (revised after verified audit 2026-10-07)

First, a correction to earlier drafts of this section. The togglable services are NOT all in one giant project. Verified on .13 via compose labels:

Service Compose project Can a console up -d touch it?
lmms music ❌ No — separate project
prefect-server prefect ❌ No
n8n n8n_data ❌ No
gimp, video-editor, audio-editor family-home-lab ✅ Yes — same project as the portal

So only the three media editors share the portal's project. The drift risk is limited to them.

The mechanism, plainly: starting the console container does NOTHING to other containers — Docker never cascades starts. The ONLY thing that re-starts a stopped container is an explicit start: docker start <c>, a restart policy, or docker compose up -d run inside that project's folder (which re-creates/restarts every service defined in that file). So:

  • docker stop gimp sticks — no timer revives it. ✔
  • BUT cd …/family-home-lab && docker compose up -d (the normal portal redeploy) silently wakes gimp, audio-editor, video-editor too, because they're in the same file. ✘
  • lmms/prefect/n8n are safe — different folders, different projects, unreachable. ✔

Resolution (user decision, 2026-10-07): no project migration, no profiles — the exposure is just three containers. Instead:

  1. The controller records desired state (§4.3 service_state) and on the portal's next deploy re-asserts it: after any docker compose up -d in family-home-lab, anything the family stopped is stopped again. This is small and targeted — not a full "reconcile everything" system.
  2. The deploy notes (deploy/DEPLOYMENT.md) will say: a portal redeploy also restarts the three media editors; run the one-line re-assert if anyone had stopped them.
  3. The UI shows actual state, so if a deploy ever overrides a stop, the card says "running" and the user or admin can stop it again — nothing is hidden.

4.5 Idle auto-off (optional, later)

Only if wanted: stop a media container after N minutes with no active selkies websocket (docker exec <c> ss -tn | grep -c ESTAB), never on a fixed timer alone. Default off; opt-in per service. Show a countdown on the card so nobody is surprised mid-edit.


5. Implementation plan

Phase 0 — Truth first (no toggles yet)

  • Add a GET /api/status source. Start simple: the portal asks the controller for container states and renders real "online/offline" pills.
  • Fix Tool.status misuse — make status live, or remove the field and derive it.
  • Add deploy/DEPLOYMENT.md note on the .27 → .13 copy step.
  • Decide and record the mechanism from §4.6 for making "off by default" stick.

Done when: the dashboard tells the truth about what is running.

Phase 1 — The controller

  • Write service-controller (small Python/FastAPI or Go binary) in this repo under deploy/service-controller/.
  • Define services.toml — the allowlist with id, kind, target, group, label, ram_mb, and default_state = "stopped" (§4.6).
  • Implement the §4.6 plan: record desired state in the controller; re-assert after portal redeploys that docker compose up -d on the main project leaves those services stopped.
  • Implement GET /services, GET /health, and start/stop for containers only.
  • Run as a systemd user unit family-service-controller.service, bound to 127.0.0.1:8091, with a token from 10-secrets.conf.
  • Add extra_hosts: ["host.docker.internal:host-gateway"] to the portal service in docker-compose.yml.
  • Test with curl from the host and from inside the portal container.

Done when: curl -XPOST localhost:8091/services/gimp/stop stops GIMP, and an unlisted id 404s.

Phase 2 — systemd user services

  • Extend the controller to systemctl --user start|stop <unit> using kind = "systemd-user".
  • Verify from the container over host.docker.internal.
  • Confirm the controller cannot be used to stop voice-agent or anything not allowlisted.

Done when: Paperclip and the Prefect worker toggle from the same API.

Phase 3 — UI

  • Extend Tool (§4.4) and tag the targets from §3.
  • Add partials/service_card.html with the status pill + button.
  • Wire HTMX: hx-post to start/stop, swap the card, poll while transitional.
  • Admin-gate the routes; show RAM on the card.
  • Handle failed with the error text and a "view logs" link (dozzle is already on .35).

Done when: a family member can stop GIMP from the dashboard and see the RAM free up on .13.

Phase 4 — Drift and safety nets

  • service_state table + desired/actual comparison + "reconcile" action.
  • Warn in the admin area when a deploy has restarted everything.
  • Group toggles: the WorldMonitor 4-container stack, and Prefect server+worker+dashboard.
  • Document the ordering rules and make grouped start respect them.

Phase 5 — Optional

  • Idle auto-off (§4.5), default off.
  • Per-user visibility: show "in use by Harry" for shared media containers.
  • RAM freed counter — a running total of what the family has saved.

6. Safety rules (non-negotiable)

  1. Never give the portal docker.sock or host filesystem access.
  2. Allowlist only. No generic docker/systemctl passthrough. Unknown id ⇒ 404.
  3. Never make the console's own dependencies togglable: portal, postgres, redis, rabbitmq, garage, worker, transcriber*.
  4. Never make house-critical services togglable: snapserver, librespot, mopidy, voice_whisper, voice_bridge, piper_tts, mosquitto, pihole, omniroute.
  5. Controller binds to 127.0.0.1 and requires a token.
  6. Mutating routes are admin-gated by default.
  7. Every toggle is logged with who, what, when.

7. Test plan

Level Test Pass
Unit Controller rejects an unlisted id 404, no side effect
Unit Controller rejects a missing/bad token 401
Unit Start/stop a container, assert state State matches docker inspect
Unit Start/stop a systemd user service State matches systemctl --user is-active
Security Attempt to stop postgres via the API 404/403, container keeps running
Security Portal container has no docker.sock ls /var/run/docker.sock fails inside it
UI Card shows correct state after a manual docker stop Pills update
UI Start a Webtop container Shows starting, then running, no hang
UI Force a crash loop Shows failed with a reason
Perf Poll interval does not hammer the controller ≤1 req/2 s per transitional card
Regression Dashboard, files, voice, admin pages still work All load
Regression Snapcast still plays after toggling media tools Audio unaffected

8. Open questions

  1. Who may toggle what? ✅ DECIDED (2026-10-07): per-user allowed sets — Sam sees everything, others get a curated subset. Not blanket admin-only; users toggle what's available to them.
  2. Auto-off ✅ DECIDED: some services, not all (opt-in per service, only when abandoned, never mid-edit).
  3. n8n ✅ DECIDED: MUST remain on. Mark togglable: False — never in the toggle set, never auto-off.
  4. paperclipai ✅ DECIDED: still in use — keep it togglable (not retired).
  5. Prefect group ✅ DECIDED: prefect-server is shared infra (like n8n) — keep ON; the prefect-worker + photo-dashboard are the pipeline's on/off switch, kept ON while the Google Photos ingestion is being worked on, admin-guarded so they can't be stopped mid-ingestion.
  6. Failed-state colour ✅ DECIDED: yes — add a proper error-red token to DESIGN.md.
  7. History ✅ DECIDED: single bounded audit log (append-only, capped length — e.g. last 500 entries), no DB table.
  8. LMMS Xvfb ✅ DONE (2026-10-07): MAX_RES: 2560x1440 in music/docker-compose.yml — framebuffer ~506MB/plane → ~14MB/plane. Deployed + verified on .13.
  9. Controller language ✅ DECIDED: Python/FastAPI (matches portal; Go's RAM edge is moot for an on-demand tool).
  10. Deploy reconcile ✅ DECIDED: no project migration — only gimp/audio-editor/video-editor share the portal project; controller records desired state and re-asserts it after a portal redeploy; UI always shows actual state (see revised §4.6).

9. Reference

# --- current container states / memory ---
docker ps --format '{{.Names}}\t{{.Status}}'
docker stats --no-stream --format '{{.Name}} {{.MemUsage}}'

# --- systemd user services (the ones a container cannot reach) ---
systemctl --user list-units --type=service --state=running
systemctl --user is-active paperclipai prefect-worker photo-dashboard where-woof chrome-pi engram

# --- the portal ---
cd /home/sam/Docker/Containers/family-home-lab && docker compose ps
docker logs -f family-home-lab-portal-1

# --- deploy ---
# edit in /home/sam/home_network/custom_tools/family_home_lab (on .27)
# copy to .13:/home/sam/Docker/Containers/family-home-lab, then: docker compose up -d

Files to touch

File Change
portal/tools.py extend Tool; tag togglable entries
portal/main.py /api/services, `/service/{id}/start
portal/templates/partials/service_card.html new
portal/templates/dashboard.html render the card partial
portal/database.py service_state table (Phase 4)
docker-compose.yml extra_hosts for the portal
deploy/service-controller/ new — the controller + services.toml
DESIGN.md one error/failed colour token (pending Q6)
deploy/DEPLOYMENT.md controller install + drift note