Files
family_home_lab/ON_OFF.md

22 KiB
Raw Blame History

ON_OFF.md — Turning services on and off from the Family Console

Target: implement on/off control for RAM-heavy services from console.lab.audasmedia.com.au Repo: sam/family_home_lab (this repo) · deployed to .13:/home/sam/Docker/Containers/family-home-lab Status: plan only. Nothing implemented. Last updated: 2026-10-06

Confirmed decisions

  1. Togglable services default to OFF. The RAM is the point — freeing it immediately matters more than a 30–60 s wait on first use. The UI must make Start the obvious primary action and show the freed/used RAM on every card. Consequence: "off" must actually hold across deployments — see §4.6. Do not skip that section.

0. Why

.13 has ~46 containers plus several native services. Several are heavy and idle almost all day. A 2026-10 audit found roughly 2.4 GB of addressable memory, most of it sitting idle:

Service Idle RAM
Media editors (GIMP, Kdenlive, Audacity, LMMS) ~1.5 GB incl. a 570 MB LMMS Xvfb
Prefect (server + worker + photo-dashboard) ~650 MB
Paperclip ~495 MB
n8n ~286 MB
WorldMonitor (4 containers) ~120 MB

The goal: let a family member start a tool when they need it and stop it when they don't, from the console they already use — free RAM without free-of-charge surprise.


1. Current state (verified 2026-10-06)

1.1 What exists

Thing Path / detail
Console portal (FastAPI + Jinja + HTMX) portal/ — main.py, tools.py, auth.py, database.py, config.py
Tool catalogue portal/tools.py — a Tool dataclass + _cat() returning the list (~335 lines)
Templates portal/templates/ — dashboard.html, admin.html, voice.html, files.html, partials/
Design system DESIGN.md — CSS tokens, sticker palette, status colours already defined
Routes /, /tool/{id}, /files, /voice, /admin, /admin/users, /admin/pi, /account/password
Auth session login; roles via admin: bool on each Tool; users seeded from .env
Portal container family-home-lab-portal-1 on .13:8500
DB Postgres (pgvector/pgvector:pg16) inside the same compose project
Deployment edit here on .27 → copy to .13:/home/sam/Docker/Containers/family-home-lab → docker compose up -d

1.2 The thing that lies today

Tool already has a status field:

@dataclass(frozen=True)
class Tool:
    ...
    status: str = "online"   # <-- a hardcoded string, never updated

It is a literal default, not live state. The dashboard currently reports "online" for every tool regardless of reality. You cannot show on/off state until this becomes real, so fixing it is part 1 of this work, not a nice-to-have.

1.3 The blast radius

The portal's own compose project contains postgres, redis, rabbitmq, garage, worker, transcriber, transcriber-mus. If the console can stop those, it can kill itself. Any design must make the console's own dependencies untouchable.


2. Issues to solve

  1. Two different control planes. Targets are a mix of Docker containers and systemd user services. A container can reach the Docker API; it cannot reach systemctl --user on the host without extra plumbing (D-Bus mount, or a helper). One mechanism will not cover both.

    Containers: gimp, video-editor, audio-editor, lmms, prefect-server, n8n, worldmonitor*, voice_whisper, headroom, … systemd user services: paperclipai, prefect-worker, photo-dashboard, where-woof, chrome-pi, engram, voice-agent, lan-mouse. systemd system services: snapserver, librespot, mopidy, caddy.

  2. /var/run/docker.sock is root-equivalent. The portal is internet-reachable via Caddy and protected only by a family login. Mounting the socket into it means one XSS or one weak password equals full host root. Do not do this.

  3. Self-destruct risk. The console must not be able to stop postgres, redis, rabbitmq, garage, portal, worker, or any service the console itself depends on. This needs to be enforced by an allowlist, not by UI politeness.

  4. The dashboard has no live state (§1.2). Needs a real status source before toggles make sense.

  5. State drift. Every service uses restart: unless-stopped, which is helpful — a docker stop sticks. But any future docker compose up -d silently restarts everything, undoing the user's choices. Toggle state must be persisted and reconciled, and drift must be visible.

  6. Startup is slow and the UI will hang. LinuxServer Webtop containers take 30–60 s to become usable. The endpoint must return immediately and the card must show starting, then poll.

  7. Ordering on start. worker/transcriber need redis+rabbitmq+postgres. Media containers need the shared-media volume mounted. Starting things in the wrong order produces confusing errors.

  8. Users cannot make informed choices. Nothing shows what a service costs. The card should show approximate RAM so the trade-off is visible.

  9. Shared services affect everyone. Stopping GIMP stops it for Harry and Finn too. Needs either admin-only gating or a visible "who's using this" indicator.

  10. Failure feedback. If a container enters a crash loop the user must see "failed" with a reason, not a spinner.

  11. Distinguish "stopped on purpose" from "crashed". These need different UI treatment. The design system already has the colours: green = online, orange = busy/starting, faint grey = offline.

  12. Idle auto-off is tempting but risky. Auto-stopping a media editor mid-edit would be hostile. If implemented, it must detect actual use (selkies has live websocket connections) and only apply to genuinely abandoned sessions.

  13. Known separate bug — do not confuse with this work. The lmms container runs Xvfb at 15360x8640, costing a 570 MB framebuffer. That is a config fix (reduce the virtual resolution), not a toggle. Fix it separately; it is the cheapest single win on the box.


3. Targets

Safety ratings: ✅ safe · ⚠️ needs care · ❌ never

Service Kind Idle RAM Safe? Who Notes
gimp container ~233 MB ✅ all users Webtop; slow start (~40 s)
video-editor (kdenlive) container ~237 MB ✅ all users Webtop
audio-editor (audacity) container ~201 MB ✅ all users Webtop
lmms container ~798 MB + 570 MB Xvfb ✅ all users Fix Xvfb resolution first (§2.13)
prefect-server container ~411 MB ✅ admin Only needed while ingesting
prefect-worker systemd user ~129 MB ✅ admin Start with the server
photo-dashboard systemd user ~112 MB ✅ admin Pairs with prefect
paperclipai systemd user ~495 MB ✅ admin Idle since 2026-09-26
n8n container ~286 MB ⚠️ admin Check for scheduled workflows first
worldmonitor +3 containers ~120 MB ✅ admin Group toggle (4 containers)
engram systemd user ~1 MB ✅ admin Dead/empty — candidate for removal not toggling
headroom container ~18 MB n/a — Leave running (JEV proxy, in the pi path)
voice_whisper container 372–857 MB ❌ — House voice depends on it
postgres, redis, rabbitmq, garage, portal, worker containers — ❌ — Console's own dependencies
snapserver, librespot, mopidy systemd system — ❌ — House audio; FIFO ordering trap (see below)

⚠️ Snapcast ordering trap: snapserver opens its FIFOs at start. If you restart snapserver before librespot/mopidy, the streams come up dead (End of file, length: 0) and audio is silent. See /home/sam/chats/snapcast/snapcast-findings.md. Keep these out of the togglable set.


4. Design

4.1 Architecture — a host-side service controller

Do not give the portal the Docker socket. Introduce one small controller that owns all privileged actions and exposes a narrow, allowlisted API.

 Browser (HTMX)
      │
      ▼
 portal container (.13:8500)                    ← no docker.sock, no host access
      │  HTTP  →  http://host.docker.internal:8091
      ▼
 service-controller  (systemd USER unit on .13, runs as `sam`, 127.0.0.1:8091)
      ├─ Docker API      (sam is in the `docker` group)  → containers
      └─ systemctl --user                                → user services
      ▲
      └─ ALLOWLIST + metadata lives here (single source of truth)

Why a host-side systemd user unit:

  • systemctl --user needs the user manager, which a container does not have. Running on the host covers user services natively.
  • sam is already in the docker group, so the same process can also drive containers — no privileged container, no socket in the portal.
  • host.docker.internal is reachable from the portal via Compose: extra_hosts: ["host.docker.internal:host-gateway"].

The allowlist is the security boundary. Anything not listed cannot be started or stopped, and there is no generic passthrough. Reject unknown ids with 404 — never forward caller-supplied names.

Alternative considered and rejected: tecnativa/docker-socket-proxy. It is a fine tool but only covers Docker, and we need systemd user services too. It would still leave the portal holding a token that can act on containers.

4.2 Controller API (localhost only)

Method Path Purpose
GET /services List allowlisted services with live state + metadata
GET /services/{id} One service: state, since, error, ram_mb
POST /services/{id}/start Start. Returns immediately (202)
POST /services/{id}/stop Stop
GET /health Liveness for the portal

Response shape per service:

{
  "id": "gimp",
  "kind": "container",
  "target": "family-home-lab-gimp-1",
  "group": "media",
  "label": "GIMP",
  "ram_mb": 233,
  "togglable": true,
  "state": "running",          // running | stopped | starting | stopping | failed | unknown
  "since": "2026-10-06T05:12:00+11:00",
  "error": null
}

Bind to 127.0.0.1 only. Add a shared secret header (X-Controller-Token) read from ~/.config/environment.d/10-secrets.conf so a stray process on the LAN cannot drive it.

4.3 State and drift (§2.5)

  • Persist desired state per service in the portal's Postgres (service_state table: id, desired, updated_by, updated_at).
  • The controller reports actual state. The UI shows both when they disagree and offers "reconcile".
  • On portal startup, compare desired vs actual and surface drift rather than silently "fixing" it.
  • Document that docker compose up -d re-starts everything. Consider a deploy/ note or a post-deploy reconcile step so deployments don't silently turn the whole lab back on.

4.4 Portal changes

Extend the Tool dataclass (portal/tools.py) — keep backwards compatibility:

@dataclass(frozen=True)
class Tool:
    ...existing fields...
    # NEW — optional on/off control
    service_id: str | None = None      # controller id, e.g. "gimp"
    ram_mb: int | None = None          # shown on the card
    togglable: bool = False            # only True if allowlisted

Then tag the relevant entries in _cat(), e.g.:

Tool("gimp", "GIMP", "image", "Photo editing & retouching",
     "https://gimp.lab.audasmedia.com.au",
     service_id="gimp", ram_mb=233, togglable=True),

New routes (portal/main.py):

Route Behaviour
GET /api/services Proxy the controller's list; HTMX polls this
POST /service/{id}/start Admin (or owner) only → controller → return the card partial
POST /service/{id}/stop Same
GET /partials/service/{id} The card fragment, for HTMX swaps

All mutating routes must enforce auth and the admin flag unless explicitly opened to users.

Template — portal/templates/partials/ gets a service_card.html with a status pill and a button. Use the tokens already in DESIGN.md:

  • accent-green #1aae39 → running
  • accent-orange #dd5b00 → starting / stopping
  • ink-faint #a39e98 → stopped
  • add a red-ish failed state (reuse accent-orange-deep #793400 pending a design decision)

Use HTMX polling while a service is in a transitional state (hx-trigger="every 2s"), and stop polling once it settles — otherwise the dashboard hammers the controller.

4.6 Making "off by default" actually hold

This is the direct consequence of the confirmed decision, and it is easy to get wrong.

Every service currently uses restart: unless-stopped, and the togglable ones live in the same compose project as the portal. That means:

  • A plain docker stop does stick — good, unless-stopped means it stays down until something explicitly starts it.
  • But any future docker compose up -d on the project starts them all again, silently undoing the family's choices. Deployments happen (see deploy/DEPLOYMENT.md), so this will occur.

Three ways to handle it, in order of preference:

Option How Trade-off
A. Separate compose project Move togglable services into docker-compose.media.yml (its own project name, e.g. family-media). The main project's up -d can never touch them. Cleanest and most robust. Needs a one-time migration of the containers.
B. Compose profiles Tag each with profiles: ["media"]. Plain docker compose up -d skips them; --profile media starts them. Simple, but behaviour when a container already exists and the profile is off should be verified before relying on it.
C. Reconcile on deploy Keep as-is, and have the portal (or a deploy hook) restore the previous desired state after any up -d. Works, but reactive — there is a window where everything is on and RAM is consumed.

Recommendation: do A in Phase 1. It makes "off by default" a structural property rather than something the UI has to keep re-asserting. It also means the controller's docker start/docker stop never fights compose.

Whichever is chosen, keep the §4.3 desired-state record so drift is detectable and reported.

4.5 Idle auto-off (optional, later)

Only if wanted: stop a media container after N minutes with no active selkies websocket (docker exec <c> ss -tn | grep -c ESTAB), never on a fixed timer alone. Default off; opt-in per service. Show a countdown on the card so nobody is surprised mid-edit.


5. Implementation plan

Phase 0 — Truth first (no toggles yet)

  • Add a GET /api/status source. Start simple: the portal asks the controller for container states and renders real "online/offline" pills.
  • Fix Tool.status misuse — make status live, or remove the field and derive it.
  • Add deploy/DEPLOYMENT.md note on the .27 → .13 copy step.
  • Decide and record the mechanism from §4.6 for making "off by default" stick.

Done when: the dashboard tells the truth about what is running.

Phase 1 — The controller

  • Write service-controller (small Python/FastAPI or Go binary) in this repo under deploy/service-controller/.
  • Define services.toml — the allowlist with id, kind, target, group, label, ram_mb, and default_state = "stopped" (§4.6).
  • Implement the §4.6 option (recommended: separate family-media compose project) and confirm that docker compose up -d on the main project leaves those services stopped.
  • Implement GET /services, GET /health, and start/stop for containers only.
  • Run as a systemd user unit family-service-controller.service, bound to 127.0.0.1:8091, with a token from 10-secrets.conf.
  • Add extra_hosts: ["host.docker.internal:host-gateway"] to the portal service in docker-compose.yml.
  • Test with curl from the host and from inside the portal container.

Done when: curl -XPOST localhost:8091/services/gimp/stop stops GIMP, and an unlisted id 404s.

Phase 2 — systemd user services

  • Extend the controller to systemctl --user start|stop <unit> using kind = "systemd-user".
  • Verify from the container over host.docker.internal.
  • Confirm the controller cannot be used to stop voice-agent or anything not allowlisted.

Done when: Paperclip and the Prefect worker toggle from the same API.

Phase 3 — UI

  • Extend Tool (§4.4) and tag the targets from §3.
  • Add partials/service_card.html with the status pill + button.
  • Wire HTMX: hx-post to start/stop, swap the card, poll while transitional.
  • Admin-gate the routes; show RAM on the card.
  • Handle failed with the error text and a "view logs" link (dozzle is already on .35).

Done when: a family member can stop GIMP from the dashboard and see the RAM free up on .13.

Phase 4 — Drift and safety nets

  • service_state table + desired/actual comparison + "reconcile" action.
  • Warn in the admin area when a deploy has restarted everything.
  • Group toggles: the WorldMonitor 4-container stack, and Prefect server+worker+dashboard.
  • Document the ordering rules and make grouped start respect them.

Phase 5 — Optional

  • Idle auto-off (§4.5), default off.
  • Per-user visibility: show "in use by Harry" for shared media containers.
  • RAM freed counter — a running total of what the family has saved.

6. Safety rules (non-negotiable)

  1. Never give the portal docker.sock or host filesystem access.
  2. Allowlist only. No generic docker/systemctl passthrough. Unknown id ⇒ 404.
  3. Never make the console's own dependencies togglable: portal, postgres, redis, rabbitmq, garage, worker, transcriber*.
  4. Never make house-critical services togglable: snapserver, librespot, mopidy, voice_whisper, voice_bridge, piper_tts, mosquitto, pihole, omniroute.
  5. Controller binds to 127.0.0.1 and requires a token.
  6. Mutating routes are admin-gated by default.
  7. Every toggle is logged with who, what, when.

7. Test plan

Level Test Pass
Unit Controller rejects an unlisted id 404, no side effect
Unit Controller rejects a missing/bad token 401
Unit Start/stop a container, assert state State matches docker inspect
Unit Start/stop a systemd user service State matches systemctl --user is-active
Security Attempt to stop postgres via the API 404/403, container keeps running
Security Portal container has no docker.sock ls /var/run/docker.sock fails inside it
UI Card shows correct state after a manual docker stop Pills update
UI Start a Webtop container Shows starting, then running, no hang
UI Force a crash loop Shows failed with a reason
Perf Poll interval does not hammer the controller ≤1 req/2 s per transitional card
Regression Dashboard, files, voice, admin pages still work All load
Regression Snapcast still plays after toggling media tools Audio unaffected

8. Open questions

  1. Who may toggle what? Admin-only for everything, or should any family member be able to stop the media tools (which are shared)?
  2. Auto-off — wanted at all? If yes, what idle window (30 min?) and should it warn first?
  3. n8n — are there scheduled workflows that must keep running? If so it stays out of scope.
  4. paperclipai — is it still in use, or should it be retired rather than toggled?
  5. Prefect group — should server + worker + photo-dashboard always be one toggle?
  6. Failed-state colour — you have accent-orange-deep #793400. Add a proper error red token to DESIGN.md, or reuse an existing token?
  7. History — do you want a "who turned this off and when" audit view, or is a log line enough?
  8. Should the LMMS Xvfb resolution be fixed as part of this (§2.13), or separately? It is ~500 MB for one config line.
  9. Controller language — Python/FastAPI to match the portal, or a single Go binary for a smaller footprint on .13?
  10. Deploy reconcile — after docker compose up -d, should the portal auto-restore the previous on/off state, or just warn?

9. Reference

# --- current container states / memory ---
docker ps --format '{{.Names}}\t{{.Status}}'
docker stats --no-stream --format '{{.Name}} {{.MemUsage}}'

# --- systemd user services (the ones a container cannot reach) ---
systemctl --user list-units --type=service --state=running
systemctl --user is-active paperclipai prefect-worker photo-dashboard where-woof chrome-pi engram

# --- the portal ---
cd /home/sam/Docker/Containers/family-home-lab && docker compose ps
docker logs -f family-home-lab-portal-1

# --- deploy ---
# edit in /home/sam/home_network/custom_tools/family_home_lab (on .27)
# copy to .13:/home/sam/Docker/Containers/family-home-lab, then: docker compose up -d

Files to touch

File Change
portal/tools.py extend Tool; tag togglable entries
portal/main.py /api/services, `/service/{id}/start
portal/templates/partials/service_card.html new
portal/templates/dashboard.html render the card partial
portal/database.py service_state table (Phase 4)
docker-compose.yml extra_hosts for the portal
deploy/service-controller/ new — the controller + services.toml
DESIGN.md one error/failed colour token (pending Q6)
deploy/DEPLOYMENT.md controller install + drift note