# ON_OFF.md — Turning services on and off from the Family Console **Target:** implement on/off control for RAM-heavy services from `console.lab.audasmedia.com.au` **Repo:** `sam/family_home_lab` (this repo) · deployed to `.13:/home/sam/Docker/Containers/family-home-lab` **Status:** plan for the on/off controller; console UX additions (back-to-top, collapsible panels) and the LMMS Xvfb fix are already implemented. **Last updated:** 2026-10-07 ### Confirmed decisions 1. **Togglable services default to OFF.** The RAM is the point — freeing it immediately matters more than a 30–60 s wait on first use. The UI must make *Start* the obvious primary action and show the freed/used RAM on every card. **Consequence:** "off" must actually hold across deployments — see §4.6. Do not skip that section. --- ## 0. Why `.13` has ~46 containers plus several native services. Several are heavy and idle almost all day. A 2026-10 audit found roughly **2.4 GB of addressable memory**, most of it sitting idle: | Service | Idle RAM | |---|---| | Media editors (GIMP, Kdenlive, Audacity, LMMS) | ~1.5 GB incl. a 570 MB LMMS Xvfb | | Prefect (server + worker + photo-dashboard) | ~650 MB | | Paperclip | ~495 MB | | n8n | ~286 MB | | WorldMonitor (4 containers) | ~120 MB | The goal: let a family member start a tool when they need it and stop it when they don't, from the console they already use — free RAM without free-of-charge surprise. --- ## 1. Current state (verified 2026-10-06) ### 1.1 What exists | Thing | Path / detail | |---|---| | Console portal (FastAPI + Jinja + HTMX) | `portal/` — `main.py`, `tools.py`, `auth.py`, `database.py`, `config.py` | | Tool catalogue | `portal/tools.py` — a `Tool` dataclass + `_cat()` returning the list (~335 lines) | | Templates | `portal/templates/` — `dashboard.html`, `admin.html`, `voice.html`, `files.html`, `partials/` | | Design system | `DESIGN.md` — CSS tokens, sticker palette, **status colours already defined** | | Routes | `/`, `/tool/{id}`, `/files`, `/voice`, `/admin`, `/admin/users`, `/admin/pi`, `/account/password` | | Auth | session login; roles via `admin: bool` on each `Tool`; users seeded from `.env` | | Portal container | `family-home-lab-portal-1` on `.13:8500` | | DB | Postgres (`pgvector/pgvector:pg16`) inside the same compose project | | Deployment | edit here on `.27` → copy to `.13:/home/sam/Docker/Containers/family-home-lab` → `docker compose up -d` | ### 1.2 The thing that lies today `Tool` already has a status field: ```python @dataclass(frozen=True) class Tool: ... status: str = "online" # <-- a hardcoded string, never updated ``` **It is a literal default, not live state.** The dashboard currently reports "online" for every tool regardless of reality. You cannot show on/off state until this becomes real, so fixing it is part 1 of this work, not a nice-to-have. ### 1.3 The blast radius The portal's own compose project contains `postgres`, `redis`, `rabbitmq`, `garage`, `worker`, `transcriber`, `transcriber-mus`. **If the console can stop those, it can kill itself.** Any design must make the console's own dependencies untouchable. --- ## 2. Issues to solve 1. **Two different control planes.** Targets are a mix of Docker containers and **systemd user services**. A container can reach the Docker API; it cannot reach `systemctl --user` on the host without extra plumbing (D-Bus mount, or a helper). One mechanism will not cover both. Containers: `gimp`, `video-editor`, `audio-editor`, `lmms`, `prefect-server`, `n8n`, `worldmonitor*`, `voice_whisper`, `headroom`, … systemd user services: `paperclipai`, `prefect-worker`, `photo-dashboard`, `where-woof`, `chrome-pi`, `engram`, `voice-agent`, `lan-mouse`. systemd system services: `snapserver`, `librespot`, `mopidy`, `caddy`. 2. **`/var/run/docker.sock` is root-equivalent.** The portal is internet-reachable via Caddy and protected only by a family login. Mounting the socket into it means one XSS or one weak password equals full host root. **Do not do this.** 3. **Self-destruct risk.** The console must not be able to stop `postgres`, `redis`, `rabbitmq`, `garage`, `portal`, `worker`, or any service the console itself depends on. This needs to be enforced by an allowlist, not by UI politeness. 4. **The dashboard has no live state** (§1.2). Needs a real status source before toggles make sense. 5. **State drift.** Every service uses `restart: unless-stopped`, which is helpful — a `docker stop` sticks. But **any future `docker compose up -d` silently restarts everything**, undoing the user's choices. Toggle state must be persisted and reconciled, and drift must be visible. 6. **Startup is slow and the UI will hang.** LinuxServer Webtop containers take 30–60 s to become usable. The endpoint must return immediately and the card must show *starting*, then poll. 7. **Ordering on start.** `worker`/`transcriber` need `redis`+`rabbitmq`+`postgres`. Media containers need the `shared-media` volume mounted. Starting things in the wrong order produces confusing errors. 8. **Users cannot make informed choices.** Nothing shows what a service costs. The card should show approximate RAM so the trade-off is visible. 9. **Shared services affect everyone.** Stopping GIMP stops it for Harry and Finn too. Needs either admin-only gating or a visible "who's using this" indicator. 10. **Failure feedback.** If a container enters a crash loop the user must see "failed" with a reason, not a spinner. 11. **Distinguish "stopped on purpose" from "crashed".** These need different UI treatment. The design system already has the colours: green = online, orange = busy/starting, faint grey = offline. 12. **Idle auto-off is tempting but risky.** Auto-stopping a media editor mid-edit would be hostile. If implemented, it must detect actual use (selkies has live websocket connections) and only apply to genuinely abandoned sessions. 13. ✅ **DONE (2026-10-07): LMMS Xvfb framebuffer fixed.** `MAX_RES: 2560x1440` added to `music/docker-compose.yml`; the webtop clamped the virtual screen from `15360x8640` (~506MB/plane) to `2560x1440` (~14MB/plane). Deployed + verified on .13. --- ## 3. Targets Safety ratings: ✅ safe · ⚠️ needs care · ❌ never | Service | Kind | Idle RAM | Safe? | Who | Notes | |---|---|---|---|---|---| | `gimp` | container | ~233 MB | ✅ | all users | Webtop; slow start (~40 s) | | `video-editor` (kdenlive) | container | ~237 MB | ✅ | all users | Webtop | | `audio-editor` (audacity) | container | ~201 MB | ✅ | all users | Webtop | | `lmms` | container | ~798 MB + 570 MB Xvfb | ✅ | all users | Xvfb fixed 2026-10-07 (MAX_RES) | | `prefect-server` | container | ~411 MB | ✅ | admin | Only needed while ingesting | | `prefect-worker` | systemd user | ~129 MB | ✅ | admin | Start with the server | | `photo-dashboard` | systemd user | ~112 MB | ✅ | admin | Pairs with prefect | | `paperclipai` | systemd user | ~495 MB | ✅ | admin | Idle since 2026-09-26 | | `n8n` | container | ~286 MB | ⚠️ | admin | Check for scheduled workflows first | | `worldmonitor` +3 | containers | ~120 MB | ✅ | admin | Group toggle (4 containers) | | `engram` | systemd user | ~1 MB | ✅ | admin | Dead/empty — candidate for removal not toggling | | `headroom` | container | ~18 MB | n/a | — | Leave running (JEV proxy, in the pi path) | | `voice_whisper` | container | 372–857 MB | ❌ | — | **House voice depends on it** | | `postgres`, `redis`, `rabbitmq`, `garage`, `portal`, `worker` | containers | — | ❌ | — | **Console's own dependencies** | | `snapserver`, `librespot`, `mopidy` | systemd system | — | ❌ | — | House audio; FIFO ordering trap (see below) | > ⚠️ **Snapcast ordering trap:** snapserver opens its FIFOs at start. If you restart `snapserver` > before `librespot`/`mopidy`, the streams come up dead (`End of file, length: 0`) and audio is > silent. See `/home/sam/chats/snapcast/snapcast-findings.md`. Keep these out of the togglable set. --- ## 4. Design ### 4.1 Architecture — a host-side service controller Do **not** give the portal the Docker socket. Introduce one small controller that owns all privileged actions and exposes a narrow, allowlisted API. ``` Browser (HTMX) │ ▼ portal container (.13:8500) ← no docker.sock, no host access │ HTTP → http://host.docker.internal:8091 ▼ service-controller (systemd USER unit on .13, runs as `sam`, 127.0.0.1:8091) ├─ Docker API (sam is in the `docker` group) → containers └─ systemctl --user → user services ▲ └─ ALLOWLIST + metadata lives here (single source of truth) ``` **Why a host-side systemd *user* unit:** - `systemctl --user` needs the user manager, which a container does not have. Running on the host covers user services natively. - `sam` is already in the `docker` group, so the same process can also drive containers — no privileged container, no socket in the portal. - `host.docker.internal` is reachable from the portal via Compose: `extra_hosts: ["host.docker.internal:host-gateway"]`. **The allowlist is the security boundary.** Anything not listed cannot be started or stopped, and there is no generic passthrough. Reject unknown ids with 404 — never forward caller-supplied names. **Alternative considered and rejected:** `tecnativa/docker-socket-proxy`. It is a fine tool but only covers Docker, and we need systemd user services too. It would still leave the portal holding a token that can act on containers. ### 4.2 Controller API (localhost only) | Method | Path | Purpose | |---|---|---| | `GET` | `/services` | List allowlisted services with live state + metadata | | `GET` | `/services/{id}` | One service: `state`, `since`, `error`, `ram_mb` | | `POST` | `/services/{id}/start` | Start. Returns immediately (202) | | `POST` | `/services/{id}/stop` | Stop | | `GET` | `/health` | Liveness for the portal | Response shape per service: ```json { "id": "gimp", "kind": "container", "target": "family-home-lab-gimp-1", "group": "media", "label": "GIMP", "ram_mb": 233, "togglable": true, "state": "running", // running | stopped | starting | stopping | failed | unknown "since": "2026-10-06T05:12:00+11:00", "error": null } ``` Bind to **`127.0.0.1` only**. Add a shared secret header (`X-Controller-Token`) read from `~/.config/environment.d/10-secrets.conf` so a stray process on the LAN cannot drive it. ### 4.3 State and drift (§2.5) - Persist **desired state** per service in the portal's Postgres (`service_state` table: `id`, `desired`, `updated_by`, `updated_at`). - The controller reports **actual state**. The UI shows both when they disagree and offers "reconcile". - On portal startup, compare desired vs actual and surface drift rather than silently "fixing" it. - Document that `docker compose up -d` re-starts everything. Consider a `deploy/` note or a post-deploy reconcile step so deployments don't silently turn the whole lab back on. ### 4.4 Portal changes **Extend the `Tool` dataclass** (`portal/tools.py`) — keep backwards compatibility: ```python @dataclass(frozen=True) class Tool: ...existing fields... # NEW — optional on/off control service_id: str | None = None # controller id, e.g. "gimp" ram_mb: int | None = None # shown on the card togglable: bool = False # only True if allowlisted ``` Then tag the relevant entries in `_cat()`, e.g.: ```python Tool("gimp", "GIMP", "image", "Photo editing & retouching", "https://gimp.lab.audasmedia.com.au", service_id="gimp", ram_mb=233, togglable=True), ``` **New routes** (`portal/main.py`): | Route | Behaviour | |---|---| | `GET /api/services` | Proxy the controller's list; HTMX polls this | | `POST /service/{id}/start` | Admin (or owner) only → controller → return the **card partial** | | `POST /service/{id}/stop` | Same | | `GET /partials/service/{id}` | The card fragment, for HTMX swaps | All mutating routes must enforce auth **and** the `admin` flag unless explicitly opened to users. **Template** — `portal/templates/partials/` gets a `service_card.html` with a status pill and a button. Use the tokens already in `DESIGN.md`: - `accent-green` `#1aae39` → running - `accent-orange` `#dd5b00` → starting / stopping - `ink-faint` `#a39e98` → stopped - add a red-ish **failed** state (reuse `accent-orange-deep` `#793400` pending a design decision) Use HTMX polling while a service is in a transitional state (`hx-trigger="every 2s"`), and stop polling once it settles — otherwise the dashboard hammers the controller. ### 4.6 Making "off by default" actually hold (revised after verified audit 2026-10-07) **First, a correction to earlier drafts of this section.** The togglable services are NOT all in one giant project. Verified on .13 via compose labels: | Service | Compose project | Can a console `up -d` touch it? | |---|---|---| | `lmms` | `music` | ❌ No — separate project | | `prefect-server` | `prefect` | ❌ No | | `n8n` | `n8n_data` | ❌ No | | `gimp`, `video-editor`, `audio-editor` | `family-home-lab` | ✅ **Yes** — same project as the portal | So only the **three media editors** share the portal's project. The drift risk is limited to them. **The mechanism, plainly:** starting the console container does NOTHING to other containers — Docker never cascades starts. The ONLY thing that re-starts a stopped container is an explicit start: `docker start `, a restart policy, or `docker compose up -d` run **inside that project's folder** (which re-creates/restarts every service defined in *that* file). So: - `docker stop gimp` sticks — no timer revives it. ✔ - BUT `cd …/family-home-lab && docker compose up -d` (the normal portal redeploy) silently wakes `gimp`, `audio-editor`, `video-editor` too, because they're in the same file. ✘ - `lmms`/`prefect`/`n8n` are safe — different folders, different projects, unreachable. ✔ **Resolution (user decision, 2026-10-07):** no project migration, no profiles — the exposure is just three containers. Instead: 1. The **controller records desired state** (§4.3 `service_state`) and on the portal's next deploy re-asserts it: after any `docker compose up -d` in `family-home-lab`, anything the family stopped is stopped again. This is small and targeted — not a full "reconcile everything" system. 2. The **deploy notes** (`deploy/DEPLOYMENT.md`) will say: a portal redeploy also restarts the three media editors; run the one-line re-assert if anyone had stopped them. 3. The UI shows **actual state**, so if a deploy ever overrides a stop, the card says "running" and the user or admin can stop it again — nothing is hidden. ### 4.5 Idle auto-off (optional, later) Only if wanted: stop a media container after N minutes with **no active selkies websocket** (`docker exec ss -tn | grep -c ESTAB`), never on a fixed timer alone. Default **off**; opt-in per service. Show a countdown on the card so nobody is surprised mid-edit. --- ## 5. Implementation plan ### Phase 0 — Truth first (no toggles yet) - [ ] Add a `GET /api/status` source. Start simple: the portal asks the controller for container states and renders real "online/offline" pills. - [ ] Fix `Tool.status` misuse — make status live, or remove the field and derive it. - [ ] Add `deploy/DEPLOYMENT.md` note on the `.27 → .13` copy step. - [ ] Decide and record the mechanism from §4.6 for making "off by default" stick. **Done when:** the dashboard tells the truth about what is running. ### Phase 1 — The controller - [ ] Write `service-controller` (small Python/FastAPI or Go binary) in this repo under `deploy/service-controller/`. - [ ] Define `services.toml` — the allowlist with `id`, `kind`, `target`, `group`, `label`, `ram_mb`, and `default_state = "stopped"` (§4.6). - [ ] Implement the §4.6 plan: record desired state in the controller; re-assert after portal redeploys that `docker compose up -d` on the main project leaves those services stopped. - [ ] Implement `GET /services`, `GET /health`, and start/stop for **containers only**. - [ ] Run as a systemd user unit `family-service-controller.service`, bound to `127.0.0.1:8091`, with a token from `10-secrets.conf`. - [ ] Add `extra_hosts: ["host.docker.internal:host-gateway"]` to the portal service in `docker-compose.yml`. - [ ] Test with `curl` from the host **and** from inside the portal container. **Done when:** `curl -XPOST localhost:8091/services/gimp/stop` stops GIMP, and an unlisted id 404s. ### Phase 2 — systemd user services - [ ] Extend the controller to `systemctl --user start|stop ` using `kind = "systemd-user"`. - [ ] Verify from the container over `host.docker.internal`. - [ ] Confirm the controller cannot be used to stop `voice-agent` or anything not allowlisted. **Done when:** Paperclip and the Prefect worker toggle from the same API. ### Phase 3 — UI - [ ] Extend `Tool` (§4.4) and tag the targets from §3. - [ ] Add `partials/service_card.html` with the status pill + button. - [ ] Wire HTMX: `hx-post` to start/stop, swap the card, poll while transitional. - [ ] Admin-gate the routes; show RAM on the card. - [ ] Handle `failed` with the error text and a "view logs" link (dozzle is already on `.35`). **Done when:** a family member can stop GIMP from the dashboard and see the RAM free up on `.13`. ### Phase 4 — Drift and safety nets - [ ] `service_state` table + desired/actual comparison + "reconcile" action. - [ ] Warn in the admin area when a deploy has restarted everything. - [ ] Group toggles: the WorldMonitor 4-container stack, and Prefect server+worker+dashboard. - [ ] Document the ordering rules and make grouped start respect them. ### Phase 5 — Optional - [ ] Idle auto-off (§4.5), default off. - [ ] Per-user visibility: show "in use by Harry" for shared media containers. - [ ] RAM freed counter — a running total of what the family has saved. --- ## 6. Safety rules (non-negotiable) 1. **Never** give the portal `docker.sock` or host filesystem access. 2. **Allowlist only.** No generic `docker`/`systemctl` passthrough. Unknown id ⇒ 404. 3. **Never** make the console's own dependencies togglable: `portal`, `postgres`, `redis`, `rabbitmq`, `garage`, `worker`, `transcriber*`. 4. **Never** make house-critical services togglable: `snapserver`, `librespot`, `mopidy`, `voice_whisper`, `voice_bridge`, `piper_tts`, `mosquitto`, `pihole`, `omniroute`. 5. Controller binds to `127.0.0.1` and requires a token. 6. Mutating routes are admin-gated by default. 7. Every toggle is **logged** with who, what, when. --- ## 7. Test plan | Level | Test | Pass | |---|---|---| | Unit | Controller rejects an unlisted id | 404, no side effect | | Unit | Controller rejects a missing/bad token | 401 | | Unit | Start/stop a container, assert state | State matches `docker inspect` | | Unit | Start/stop a systemd user service | State matches `systemctl --user is-active` | | Security | Attempt to stop `postgres` via the API | 404/403, container keeps running | | Security | Portal container has no `docker.sock` | `ls /var/run/docker.sock` fails inside it | | UI | Card shows correct state after a manual `docker stop` | Pills update | | UI | Start a Webtop container | Shows *starting*, then *running*, no hang | | UI | Force a crash loop | Shows *failed* with a reason | | Perf | Poll interval does not hammer the controller | ≤1 req/2 s per transitional card | | Regression | Dashboard, files, voice, admin pages still work | All load | | Regression | Snapcast still plays after toggling media tools | Audio unaffected | --- ## 8. Open questions 1. **Who may toggle what?** ✅ DECIDED (2026-10-07): per-user allowed sets — Sam sees everything, others get a curated subset. Not blanket admin-only; users toggle what's available to *them*. 2. **Auto-off** ✅ DECIDED: some services, not all (opt-in per service, only when abandoned, never mid-edit). 3. **`n8n`** ✅ DECIDED: **MUST remain on.** Mark `togglable: False` — never in the toggle set, never auto-off. 4. **`paperclipai`** ✅ DECIDED: **still in use** — keep it togglable (not retired). 5. **Prefect group** ✅ DECIDED: `prefect-server` is shared infra (like n8n) — keep ON; the `prefect-worker` + `photo-dashboard` are the pipeline's on/off switch, kept ON while the Google Photos ingestion is being worked on, admin-guarded so they can't be stopped mid-ingestion. 6. **Failed-state colour** ✅ DECIDED: yes — add a proper error-red token to `DESIGN.md`. 7. **History** ✅ DECIDED: single bounded audit log (append-only, capped length — e.g. last 500 entries), no DB table. 8. ~~LMMS Xvfb~~ ✅ **DONE (2026-10-07):** `MAX_RES: 2560x1440` in `music/docker-compose.yml` — framebuffer ~506MB/plane → ~14MB/plane. Deployed + verified on .13. 9. **Controller language** ✅ DECIDED: **Python/FastAPI** (matches portal; Go's RAM edge is moot for an on-demand tool). 10. **Deploy reconcile** ✅ DECIDED: no project migration — only `gimp`/`audio-editor`/`video-editor` share the portal project; controller records desired state and re-asserts it after a portal redeploy; UI always shows actual state (see revised §4.6). --- ## 9. Reference ```bash # --- current container states / memory --- docker ps --format '{{.Names}}\t{{.Status}}' docker stats --no-stream --format '{{.Name}} {{.MemUsage}}' # --- systemd user services (the ones a container cannot reach) --- systemctl --user list-units --type=service --state=running systemctl --user is-active paperclipai prefect-worker photo-dashboard where-woof chrome-pi engram # --- the portal --- cd /home/sam/Docker/Containers/family-home-lab && docker compose ps docker logs -f family-home-lab-portal-1 # --- deploy --- # edit in /home/sam/home_network/custom_tools/family_home_lab (on .27) # copy to .13:/home/sam/Docker/Containers/family-home-lab, then: docker compose up -d ``` **Files to touch** | File | Change | |---|---| | `portal/tools.py` | extend `Tool`; tag togglable entries | | `portal/main.py` | `/api/services`, `/service/{id}/start|stop`, card partial | | `portal/templates/partials/service_card.html` | new | | `portal/templates/dashboard.html` | render the card partial | | `portal/database.py` | `service_state` table (Phase 4) | | `docker-compose.yml` | `extra_hosts` for the portal | | `deploy/service-controller/` | new — the controller + `services.toml` | | `DESIGN.md` | one error/failed colour token (pending Q6) | | `deploy/DEPLOYMENT.md` | controller install + drift note |