ON_OFF: record all decisions (per-user toggles, n8n/paperclip/prefect, bounded audit, python controller, MAX_RES done); revise 4.6 with verified compose-project audit

This commit is contained in:
2026-10-07 16:22:13 +11:00
parent b37c5a5134
commit bb2c6e5d0d

View File

@@ -2,8 +2,9 @@
**Target:** implement on/off control for RAM-heavy services from `console.lab.audasmedia.com.au` **Target:** implement on/off control for RAM-heavy services from `console.lab.audasmedia.com.au`
**Repo:** `sam/family_home_lab` (this repo) · deployed to `.13:/home/sam/Docker/Containers/family-home-lab` **Repo:** `sam/family_home_lab` (this repo) · deployed to `.13:/home/sam/Docker/Containers/family-home-lab`
**Status:** plan only. Nothing implemented. **Status:** plan for the on/off controller; console UX additions (back-to-top, collapsible panels)
**Last updated:** 2026-10-06 and the LMMS Xvfb fix are already implemented.
**Last updated:** 2026-10-07
### Confirmed decisions ### Confirmed decisions
@@ -119,9 +120,9 @@ must make the console's own dependencies untouchable.
If implemented, it must detect actual use (selkies has live websocket connections) and only apply If implemented, it must detect actual use (selkies has live websocket connections) and only apply
to genuinely abandoned sessions. to genuinely abandoned sessions.
13. **Known separate bug — do not confuse with this work.** The `lmms` container runs Xvfb at 13. ✅ **DONE (2026-10-07): LMMS Xvfb framebuffer fixed.** `MAX_RES: 2560x1440` added to
`15360x8640`, costing a **570 MB framebuffer**. That is a config fix (reduce the virtual resolution), `music/docker-compose.yml`; the webtop clamped the virtual screen from `15360x8640`
not a toggle. Fix it separately; it is the cheapest single win on the box. (~506MB/plane) to `2560x1440` (~14MB/plane). Deployed + verified on .13.
--- ---
@@ -134,7 +135,7 @@ Safety ratings: ✅ safe · ⚠️ needs care · ❌ never
| `gimp` | container | ~233 MB | ✅ | all users | Webtop; slow start (~40 s) | | `gimp` | container | ~233 MB | ✅ | all users | Webtop; slow start (~40 s) |
| `video-editor` (kdenlive) | container | ~237 MB | ✅ | all users | Webtop | | `video-editor` (kdenlive) | container | ~237 MB | ✅ | all users | Webtop |
| `audio-editor` (audacity) | container | ~201 MB | ✅ | all users | Webtop | | `audio-editor` (audacity) | container | ~201 MB | ✅ | all users | Webtop |
| `lmms` | container | ~798 MB + 570 MB Xvfb | ✅ | all users | Fix Xvfb resolution first (§2.13) | | `lmms` | container | ~798 MB + 570 MB Xvfb | ✅ | all users | Xvfb fixed 2026-10-07 (MAX_RES) |
| `prefect-server` | container | ~411 MB | ✅ | admin | Only needed while ingesting | | `prefect-server` | container | ~411 MB | ✅ | admin | Only needed while ingesting |
| `prefect-worker` | systemd user | ~129 MB | ✅ | admin | Start with the server | | `prefect-worker` | systemd user | ~129 MB | ✅ | admin | Start with the server |
| `photo-dashboard` | systemd user | ~112 MB | ✅ | admin | Pairs with prefect | | `photo-dashboard` | systemd user | ~112 MB | ✅ | admin | Pairs with prefect |
@@ -274,31 +275,40 @@ button. Use the tokens already in `DESIGN.md`:
Use HTMX polling while a service is in a transitional state (`hx-trigger="every 2s"`), and stop Use HTMX polling while a service is in a transitional state (`hx-trigger="every 2s"`), and stop
polling once it settles — otherwise the dashboard hammers the controller. polling once it settles — otherwise the dashboard hammers the controller.
### 4.6 Making "off by default" actually hold ### 4.6 Making "off by default" actually hold (revised after verified audit 2026-10-07)
This is the direct consequence of the confirmed decision, and it is easy to get wrong. **First, a correction to earlier drafts of this section.** The togglable services are NOT all in
one giant project. Verified on .13 via compose labels:
Every service currently uses `restart: unless-stopped`, and the togglable ones live in the **same | Service | Compose project | Can a console `up -d` touch it? |
compose project as the portal**. That means:
- A plain `docker stop` **does** stick — good, `unless-stopped` means it stays down until something
explicitly starts it.
- But **any future `docker compose up -d` on the project starts them all again**, silently undoing
the family's choices. Deployments happen (see `deploy/DEPLOYMENT.md`), so this *will* occur.
Three ways to handle it, in order of preference:
| Option | How | Trade-off |
|---|---|---| |---|---|---|
| **A. Separate compose project** | Move togglable services into `docker-compose.media.yml` (its own project name, e.g. `family-media`). The main project's `up -d` can never touch them. | Cleanest and most robust. Needs a one-time migration of the containers. | | `lmms` | `music` | ❌ No — separate project |
| **B. Compose profiles** | Tag each with `profiles: ["media"]`. Plain `docker compose up -d` skips them; `--profile media` starts them. | Simple, but behaviour when a container already exists and the profile is off should be **verified** before relying on it. | | `prefect-server` | `prefect` | ❌ No |
| **C. Reconcile on deploy** | Keep as-is, and have the portal (or a deploy hook) restore the previous desired state after any `up -d`. | Works, but reactive — there is a window where everything is on and RAM is consumed. | | `n8n` | `n8n_data` | ❌ No |
| `gimp`, `video-editor`, `audio-editor` | `family-home-lab` | ✅ **Yes** — same project as the portal |
**Recommendation:** do **A** in Phase 1. It makes "off by default" a structural property rather than So only the **three media editors** share the portal's project. The drift risk is limited to them.
something the UI has to keep re-asserting. It also means the controller's `docker start`/`docker stop`
never fights compose.
Whichever is chosen, keep the §4.3 desired-state record so drift is detectable and reported. **The mechanism, plainly:** starting the console container does NOTHING to other containers — Docker
never cascades starts. The ONLY thing that re-starts a stopped container is an explicit start:
`docker start <c>`, a restart policy, or `docker compose up -d` run **inside that project's folder**
(which re-creates/restarts every service defined in *that* file). So:
- `docker stop gimp` sticks — no timer revives it. ✔
- BUT `cd …/family-home-lab && docker compose up -d` (the normal portal redeploy) silently wakes
`gimp`, `audio-editor`, `video-editor` too, because they're in the same file. ✘
- `lmms`/`prefect`/`n8n` are safe — different folders, different projects, unreachable. ✔
**Resolution (user decision, 2026-10-07):** no project migration, no profiles — the exposure is just
three containers. Instead:
1. The **controller records desired state** (§4.3 `service_state`) and on the portal's next deploy
re-asserts it: after any `docker compose up -d` in `family-home-lab`, anything the family stopped
is stopped again. This is small and targeted — not a full "reconcile everything" system.
2. The **deploy notes** (`deploy/DEPLOYMENT.md`) will say: a portal redeploy also restarts the three
media editors; run the one-line re-assert if anyone had stopped them.
3. The UI shows **actual state**, so if a deploy ever overrides a stop, the card says "running" and
the user or admin can stop it again — nothing is hidden.
### 4.5 Idle auto-off (optional, later) ### 4.5 Idle auto-off (optional, later)
@@ -326,7 +336,7 @@ per service. Show a countdown on the card so nobody is surprised mid-edit.
`deploy/service-controller/`. `deploy/service-controller/`.
- [ ] Define `services.toml` — the allowlist with `id`, `kind`, `target`, `group`, `label`, - [ ] Define `services.toml` — the allowlist with `id`, `kind`, `target`, `group`, `label`,
`ram_mb`, and `default_state = "stopped"` (§4.6). `ram_mb`, and `default_state = "stopped"` (§4.6).
- [ ] Implement the §4.6 option (recommended: separate `family-media` compose project) and confirm - [ ] Implement the §4.6 plan: record desired state in the controller; re-assert after portal redeploys
that `docker compose up -d` on the main project leaves those services stopped. that `docker compose up -d` on the main project leaves those services stopped.
- [ ] Implement `GET /services`, `GET /health`, and start/stop for **containers only**. - [ ] Implement `GET /services`, `GET /health`, and start/stop for **containers only**.
- [ ] Run as a systemd user unit `family-service-controller.service`, bound to `127.0.0.1:8091`, - [ ] Run as a systemd user unit `family-service-controller.service`, bound to `127.0.0.1:8091`,
@@ -405,21 +415,24 @@ per service. Show a countdown on the card so nobody is surprised mid-edit.
## 8. Open questions ## 8. Open questions
1. **Who may toggle what?** Admin-only for everything, or should any family member be able to stop 1. **Who may toggle what?** ✅ DECIDED (2026-10-07): per-user allowed sets — Sam sees everything, others
the *media* tools (which are shared)? get a curated subset. Not blanket admin-only; users toggle what's available to *them*.
2. **Auto-off** — wanted at all? If yes, what idle window (30 min?) and should it warn first? 2. **Auto-off** ✅ DECIDED: some services, not all (opt-in per service, only when abandoned, never mid-edit).
3. **`n8n`** — are there scheduled workflows that must keep running? If so it stays out of scope. 3. **`n8n`** ✅ DECIDED: **MUST remain on.** Mark `togglable: False` — never in the toggle set, never auto-off.
4. **`paperclipai`** — is it still in use, or should it be retired rather than toggled? 4. **`paperclipai`** ✅ DECIDED: **still in use** — keep it togglable (not retired).
5. **Prefect group** — should server + worker + photo-dashboard always be one toggle? 5. **Prefect group** ✅ DECIDED: `prefect-server` is shared infra (like n8n) — keep ON; the
6. **Failed-state colour** — you have `accent-orange-deep #793400`. Add a proper error red token to `prefect-worker` + `photo-dashboard` are the pipeline's on/off switch, kept ON while the Google
`DESIGN.md`, or reuse an existing token? Photos ingestion is being worked on, admin-guarded so they can't be stopped mid-ingestion.
7. **History** — do you want a "who turned this off and when" audit view, or is a log line enough? 6. **Failed-state colour** ✅ DECIDED: yes — add a proper error-red token to `DESIGN.md`.
8. **Should the LMMS Xvfb resolution be fixed as part of this** (§2.13), or separately? It is ~500 MB 7. **History** ✅ DECIDED: single bounded audit log (append-only, capped length — e.g. last 500
for one config line. entries), no DB table.
9. **Controller language** — Python/FastAPI to match the portal, or a single Go binary for a smaller 8. ~~LMMS Xvfb~~ ✅ **DONE (2026-10-07):** `MAX_RES: 2560x1440` in `music/docker-compose.yml` —
footprint on `.13`? framebuffer ~506MB/plane → ~14MB/plane. Deployed + verified on .13.
10. **Deploy reconcile** — after `docker compose up -d`, should the portal auto-restore the previous 9. **Controller language** ✅ DECIDED: **Python/FastAPI** (matches portal; Go's RAM edge is moot for an
on/off state, or just warn? on-demand tool).
10. **Deploy reconcile** ✅ DECIDED: no project migration — only `gimp`/`audio-editor`/`video-editor`
share the portal project; controller records desired state and re-asserts it after a portal
redeploy; UI always shows actual state (see revised §4.6).
--- ---