ON_OFF: record all decisions (per-user toggles, n8n/paperclip/prefect, bounded audit, python controller, MAX_RES done); revise 4.6 with verified compose-project audit

This commit is contained in:
2026-10-07 16:22:13 +11:00
parent b37c5a5134
commit bb2c6e5d0d

View File

@@ -2,8 +2,9 @@
**Target:** implement on/off control for RAM-heavy services from `console.lab.audasmedia.com.au`
**Repo:** `sam/family_home_lab` (this repo) · deployed to `.13:/home/sam/Docker/Containers/family-home-lab`
**Status:** plan only. Nothing implemented.
**Last updated:** 2026-10-06
**Status:** plan for the on/off controller; console UX additions (back-to-top, collapsible panels)
and the LMMS Xvfb fix are already implemented.
**Last updated:** 2026-10-07
### Confirmed decisions
@@ -119,9 +120,9 @@ must make the console's own dependencies untouchable.
If implemented, it must detect actual use (selkies has live websocket connections) and only apply
to genuinely abandoned sessions.
13. **Known separate bug — do not confuse with this work.** The `lmms` container runs Xvfb at
`15360x8640`, costing a **570 MB framebuffer**. That is a config fix (reduce the virtual resolution),
not a toggle. Fix it separately; it is the cheapest single win on the box.
13. ✅ **DONE (2026-10-07): LMMS Xvfb framebuffer fixed.** `MAX_RES: 2560x1440` added to
`music/docker-compose.yml`; the webtop clamped the virtual screen from `15360x8640`
(~506MB/plane) to `2560x1440` (~14MB/plane). Deployed + verified on .13.
---
@@ -134,7 +135,7 @@ Safety ratings: ✅ safe · ⚠️ needs care · ❌ never
| `gimp` | container | ~233 MB | ✅ | all users | Webtop; slow start (~40 s) |
| `video-editor` (kdenlive) | container | ~237 MB | ✅ | all users | Webtop |
| `audio-editor` (audacity) | container | ~201 MB | ✅ | all users | Webtop |
| `lmms` | container | ~798 MB + 570 MB Xvfb | ✅ | all users | Fix Xvfb resolution first (§2.13) |
| `lmms` | container | ~798 MB + 570 MB Xvfb | ✅ | all users | Xvfb fixed 2026-10-07 (MAX_RES) |
| `prefect-server` | container | ~411 MB | ✅ | admin | Only needed while ingesting |
| `prefect-worker` | systemd user | ~129 MB | ✅ | admin | Start with the server |
| `photo-dashboard` | systemd user | ~112 MB | ✅ | admin | Pairs with prefect |
@@ -274,31 +275,40 @@ button. Use the tokens already in `DESIGN.md`:
Use HTMX polling while a service is in a transitional state (`hx-trigger="every 2s"`), and stop
polling once it settles — otherwise the dashboard hammers the controller.
### 4.6 Making "off by default" actually hold
### 4.6 Making "off by default" actually hold (revised after verified audit 2026-10-07)
This is the direct consequence of the confirmed decision, and it is easy to get wrong.
**First, a correction to earlier drafts of this section.** The togglable services are NOT all in
one giant project. Verified on .13 via compose labels:
Every service currently uses `restart: unless-stopped`, and the togglable ones live in the **same
compose project as the portal**. That means:
- A plain `docker stop` **does** stick — good, `unless-stopped` means it stays down until something
explicitly starts it.
- But **any future `docker compose up -d` on the project starts them all again**, silently undoing
the family's choices. Deployments happen (see `deploy/DEPLOYMENT.md`), so this *will* occur.
Three ways to handle it, in order of preference:
| Option | How | Trade-off |
| Service | Compose project | Can a console `up -d` touch it? |
|---|---|---|
| **A. Separate compose project** | Move togglable services into `docker-compose.media.yml` (its own project name, e.g. `family-media`). The main project's `up -d` can never touch them. | Cleanest and most robust. Needs a one-time migration of the containers. |
| **B. Compose profiles** | Tag each with `profiles: ["media"]`. Plain `docker compose up -d` skips them; `--profile media` starts them. | Simple, but behaviour when a container already exists and the profile is off should be **verified** before relying on it. |
| **C. Reconcile on deploy** | Keep as-is, and have the portal (or a deploy hook) restore the previous desired state after any `up -d`. | Works, but reactive — there is a window where everything is on and RAM is consumed. |
| `lmms` | `music` | ❌ No — separate project |
| `prefect-server` | `prefect` | ❌ No |
| `n8n` | `n8n_data` | ❌ No |
| `gimp`, `video-editor`, `audio-editor` | `family-home-lab` | ✅ **Yes** — same project as the portal |
**Recommendation:** do **A** in Phase 1. It makes "off by default" a structural property rather than
something the UI has to keep re-asserting. It also means the controller's `docker start`/`docker stop`
never fights compose.
So only the **three media editors** share the portal's project. The drift risk is limited to them.
Whichever is chosen, keep the §4.3 desired-state record so drift is detectable and reported.
**The mechanism, plainly:** starting the console container does NOTHING to other containers — Docker
never cascades starts. The ONLY thing that re-starts a stopped container is an explicit start:
`docker start <c>`, a restart policy, or `docker compose up -d` run **inside that project's folder**
(which re-creates/restarts every service defined in *that* file). So:
- `docker stop gimp` sticks — no timer revives it. ✔
- BUT `cd …/family-home-lab && docker compose up -d` (the normal portal redeploy) silently wakes
`gimp`, `audio-editor`, `video-editor` too, because they're in the same file. ✘
- `lmms`/`prefect`/`n8n` are safe — different folders, different projects, unreachable. ✔
**Resolution (user decision, 2026-10-07):** no project migration, no profiles — the exposure is just
three containers. Instead:
1. The **controller records desired state** (§4.3 `service_state`) and on the portal's next deploy
re-asserts it: after any `docker compose up -d` in `family-home-lab`, anything the family stopped
is stopped again. This is small and targeted — not a full "reconcile everything" system.
2. The **deploy notes** (`deploy/DEPLOYMENT.md`) will say: a portal redeploy also restarts the three
media editors; run the one-line re-assert if anyone had stopped them.
3. The UI shows **actual state**, so if a deploy ever overrides a stop, the card says "running" and
the user or admin can stop it again — nothing is hidden.
### 4.5 Idle auto-off (optional, later)
@@ -326,7 +336,7 @@ per service. Show a countdown on the card so nobody is surprised mid-edit.
`deploy/service-controller/`.
- [ ] Define `services.toml` — the allowlist with `id`, `kind`, `target`, `group`, `label`,
`ram_mb`, and `default_state = "stopped"` (§4.6).
- [ ] Implement the §4.6 option (recommended: separate `family-media` compose project) and confirm
- [ ] Implement the §4.6 plan: record desired state in the controller; re-assert after portal redeploys
that `docker compose up -d` on the main project leaves those services stopped.
- [ ] Implement `GET /services`, `GET /health`, and start/stop for **containers only**.
- [ ] Run as a systemd user unit `family-service-controller.service`, bound to `127.0.0.1:8091`,
@@ -405,21 +415,24 @@ per service. Show a countdown on the card so nobody is surprised mid-edit.
## 8. Open questions
1. **Who may toggle what?** Admin-only for everything, or should any family member be able to stop
the *media* tools (which are shared)?
2. **Auto-off** — wanted at all? If yes, what idle window (30 min?) and should it warn first?
3. **`n8n`** — are there scheduled workflows that must keep running? If so it stays out of scope.
4. **`paperclipai`** — is it still in use, or should it be retired rather than toggled?
5. **Prefect group** — should server + worker + photo-dashboard always be one toggle?
6. **Failed-state colour** — you have `accent-orange-deep #793400`. Add a proper error red token to
`DESIGN.md`, or reuse an existing token?
7. **History** — do you want a "who turned this off and when" audit view, or is a log line enough?
8. **Should the LMMS Xvfb resolution be fixed as part of this** (§2.13), or separately? It is ~500 MB
for one config line.
9. **Controller language** — Python/FastAPI to match the portal, or a single Go binary for a smaller
footprint on `.13`?
10. **Deploy reconcile** — after `docker compose up -d`, should the portal auto-restore the previous
on/off state, or just warn?
1. **Who may toggle what?** ✅ DECIDED (2026-10-07): per-user allowed sets — Sam sees everything, others
get a curated subset. Not blanket admin-only; users toggle what's available to *them*.
2. **Auto-off** ✅ DECIDED: some services, not all (opt-in per service, only when abandoned, never mid-edit).
3. **`n8n`** ✅ DECIDED: **MUST remain on.** Mark `togglable: False` — never in the toggle set, never auto-off.
4. **`paperclipai`** ✅ DECIDED: **still in use** — keep it togglable (not retired).
5. **Prefect group** ✅ DECIDED: `prefect-server` is shared infra (like n8n) — keep ON; the
`prefect-worker` + `photo-dashboard` are the pipeline's on/off switch, kept ON while the Google
Photos ingestion is being worked on, admin-guarded so they can't be stopped mid-ingestion.
6. **Failed-state colour** ✅ DECIDED: yes — add a proper error-red token to `DESIGN.md`.
7. **History** ✅ DECIDED: single bounded audit log (append-only, capped length — e.g. last 500
entries), no DB table.
8. ~~LMMS Xvfb~~ ✅ **DONE (2026-10-07):** `MAX_RES: 2560x1440` in `music/docker-compose.yml` —
framebuffer ~506MB/plane → ~14MB/plane. Deployed + verified on .13.
9. **Controller language** ✅ DECIDED: **Python/FastAPI** (matches portal; Go's RAM edge is moot for an
on-demand tool).
10. **Deploy reconcile** ✅ DECIDED: no project migration — only `gimp`/`audio-editor`/`video-editor`
share the portal project; controller records desired state and re-asserts it after a portal
redeploy; UI always shows actual state (see revised §4.6).
---