Files
family_home_lab/ON_OFF.md

459 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ON_OFF.md — Turning services on and off from the Family Console
**Target:** implement on/off control for RAM-heavy services from `console.lab.audasmedia.com.au`
**Repo:** `sam/family_home_lab` (this repo) · deployed to `.13:/home/sam/Docker/Containers/family-home-lab`
**Status:** plan only. Nothing implemented.
**Last updated:** 2026-10-06
### Confirmed decisions
1. **Togglable services default to OFF.** The RAM is the point — freeing it immediately matters more
than a 30–60 s wait on first use. The UI must make *Start* the obvious primary action and show the
freed/used RAM on every card.
**Consequence:** "off" must actually hold across deployments — see §4.6. Do not skip that section.
---
## 0. Why
`.13` has ~46 containers plus several native services. Several are heavy and idle almost all day.
A 2026-10 audit found roughly **2.4 GB of addressable memory**, most of it sitting idle:
| Service | Idle RAM |
|---|---|
| Media editors (GIMP, Kdenlive, Audacity, LMMS) | ~1.5 GB incl. a 570 MB LMMS Xvfb |
| Prefect (server + worker + photo-dashboard) | ~650 MB |
| Paperclip | ~495 MB |
| n8n | ~286 MB |
| WorldMonitor (4 containers) | ~120 MB |
The goal: let a family member start a tool when they need it and stop it when they don't, from the
console they already use — free RAM without free-of-charge surprise.
---
## 1. Current state (verified 2026-10-06)
### 1.1 What exists
| Thing | Path / detail |
|---|---|
| Console portal (FastAPI + Jinja + HTMX) | `portal/` — `main.py`, `tools.py`, `auth.py`, `database.py`, `config.py` |
| Tool catalogue | `portal/tools.py` — a `Tool` dataclass + `_cat()` returning the list (~335 lines) |
| Templates | `portal/templates/` — `dashboard.html`, `admin.html`, `voice.html`, `files.html`, `partials/` |
| Design system | `DESIGN.md` — CSS tokens, sticker palette, **status colours already defined** |
| Routes | `/`, `/tool/{id}`, `/files`, `/voice`, `/admin`, `/admin/users`, `/admin/pi`, `/account/password` |
| Auth | session login; roles via `admin: bool` on each `Tool`; users seeded from `.env` |
| Portal container | `family-home-lab-portal-1` on `.13:8500` |
| DB | Postgres (`pgvector/pgvector:pg16`) inside the same compose project |
| Deployment | edit here on `.27` → copy to `.13:/home/sam/Docker/Containers/family-home-lab` → `docker compose up -d` |
### 1.2 The thing that lies today
`Tool` already has a status field:
```python
@dataclass(frozen=True)
class Tool:
...
status: str = "online" # <-- a hardcoded string, never updated
```
**It is a literal default, not live state.** The dashboard currently reports "online" for every tool
regardless of reality. You cannot show on/off state until this becomes real, so fixing it is part 1
of this work, not a nice-to-have.
### 1.3 The blast radius
The portal's own compose project contains `postgres`, `redis`, `rabbitmq`, `garage`, `worker`,
`transcriber`, `transcriber-mus`. **If the console can stop those, it can kill itself.** Any design
must make the console's own dependencies untouchable.
---
## 2. Issues to solve
1. **Two different control planes.** Targets are a mix of Docker containers and **systemd user
services**. A container can reach the Docker API; it cannot reach `systemctl --user` on the host
without extra plumbing (D-Bus mount, or a helper). One mechanism will not cover both.
Containers: `gimp`, `video-editor`, `audio-editor`, `lmms`, `prefect-server`, `n8n`,
`worldmonitor*`, `voice_whisper`, `headroom`, …
systemd user services: `paperclipai`, `prefect-worker`, `photo-dashboard`, `where-woof`,
`chrome-pi`, `engram`, `voice-agent`, `lan-mouse`.
systemd system services: `snapserver`, `librespot`, `mopidy`, `caddy`.
2. **`/var/run/docker.sock` is root-equivalent.** The portal is internet-reachable via Caddy and
protected only by a family login. Mounting the socket into it means one XSS or one weak password
equals full host root. **Do not do this.**
3. **Self-destruct risk.** The console must not be able to stop `postgres`, `redis`, `rabbitmq`,
`garage`, `portal`, `worker`, or any service the console itself depends on. This needs to be
enforced by an allowlist, not by UI politeness.
4. **The dashboard has no live state** (§1.2). Needs a real status source before toggles make sense.
5. **State drift.** Every service uses `restart: unless-stopped`, which is helpful — a `docker stop`
sticks. But **any future `docker compose up -d` silently restarts everything**, undoing the user's
choices. Toggle state must be persisted and reconciled, and drift must be visible.
6. **Startup is slow and the UI will hang.** LinuxServer Webtop containers take 30–60 s to become
usable. The endpoint must return immediately and the card must show *starting*, then poll.
7. **Ordering on start.** `worker`/`transcriber` need `redis`+`rabbitmq`+`postgres`. Media containers
need the `shared-media` volume mounted. Starting things in the wrong order produces confusing errors.
8. **Users cannot make informed choices.** Nothing shows what a service costs. The card should show
approximate RAM so the trade-off is visible.
9. **Shared services affect everyone.** Stopping GIMP stops it for Harry and Finn too. Needs either
admin-only gating or a visible "who's using this" indicator.
10. **Failure feedback.** If a container enters a crash loop the user must see "failed" with a reason,
not a spinner.
11. **Distinguish "stopped on purpose" from "crashed".** These need different UI treatment. The design
system already has the colours: green = online, orange = busy/starting, faint grey = offline.
12. **Idle auto-off is tempting but risky.** Auto-stopping a media editor mid-edit would be hostile.
If implemented, it must detect actual use (selkies has live websocket connections) and only apply
to genuinely abandoned sessions.
13. **Known separate bug — do not confuse with this work.** The `lmms` container runs Xvfb at
`15360x8640`, costing a **570 MB framebuffer**. That is a config fix (reduce the virtual resolution),
not a toggle. Fix it separately; it is the cheapest single win on the box.
---
## 3. Targets
Safety ratings: ✅ safe · ⚠️ needs care · ❌ never
| Service | Kind | Idle RAM | Safe? | Who | Notes |
|---|---|---|---|---|---|
| `gimp` | container | ~233 MB | ✅ | all users | Webtop; slow start (~40 s) |
| `video-editor` (kdenlive) | container | ~237 MB | ✅ | all users | Webtop |
| `audio-editor` (audacity) | container | ~201 MB | ✅ | all users | Webtop |
| `lmms` | container | ~798 MB + 570 MB Xvfb | ✅ | all users | Fix Xvfb resolution first (§2.13) |
| `prefect-server` | container | ~411 MB | ✅ | admin | Only needed while ingesting |
| `prefect-worker` | systemd user | ~129 MB | ✅ | admin | Start with the server |
| `photo-dashboard` | systemd user | ~112 MB | ✅ | admin | Pairs with prefect |
| `paperclipai` | systemd user | ~495 MB | ✅ | admin | Idle since 2026-09-26 |
| `n8n` | container | ~286 MB | ⚠️ | admin | Check for scheduled workflows first |
| `worldmonitor` +3 | containers | ~120 MB | ✅ | admin | Group toggle (4 containers) |
| `engram` | systemd user | ~1 MB | ✅ | admin | Dead/empty — candidate for removal not toggling |
| `headroom` | container | ~18 MB | n/a | — | Leave running (JEV proxy, in the pi path) |
| `voice_whisper` | container | 372–857 MB | ❌ | — | **House voice depends on it** |
| `postgres`, `redis`, `rabbitmq`, `garage`, `portal`, `worker` | containers | — | ❌ | — | **Console's own dependencies** |
| `snapserver`, `librespot`, `mopidy` | systemd system | — | ❌ | — | House audio; FIFO ordering trap (see below) |
> ⚠️ **Snapcast ordering trap:** snapserver opens its FIFOs at start. If you restart `snapserver`
> before `librespot`/`mopidy`, the streams come up dead (`End of file, length: 0`) and audio is
> silent. See `/home/sam/chats/snapcast/snapcast-findings.md`. Keep these out of the togglable set.
---
## 4. Design
### 4.1 Architecture — a host-side service controller
Do **not** give the portal the Docker socket. Introduce one small controller that owns all privileged
actions and exposes a narrow, allowlisted API.
```
Browser (HTMX)
│
▼
portal container (.13:8500) ← no docker.sock, no host access
│ HTTP → http://host.docker.internal:8091
▼
service-controller (systemd USER unit on .13, runs as `sam`, 127.0.0.1:8091)
├─ Docker API (sam is in the `docker` group) → containers
└─ systemctl --user → user services
▲
└─ ALLOWLIST + metadata lives here (single source of truth)
```
**Why a host-side systemd *user* unit:**
- `systemctl --user` needs the user manager, which a container does not have. Running on the host
covers user services natively.
- `sam` is already in the `docker` group, so the same process can also drive containers — no
privileged container, no socket in the portal.
- `host.docker.internal` is reachable from the portal via Compose:
`extra_hosts: ["host.docker.internal:host-gateway"]`.
**The allowlist is the security boundary.** Anything not listed cannot be started or stopped, and
there is no generic passthrough. Reject unknown ids with 404 — never forward caller-supplied names.
**Alternative considered and rejected:** `tecnativa/docker-socket-proxy`. It is a fine tool but only
covers Docker, and we need systemd user services too. It would still leave the portal holding a
token that can act on containers.
### 4.2 Controller API (localhost only)
| Method | Path | Purpose |
|---|---|---|
| `GET` | `/services` | List allowlisted services with live state + metadata |
| `GET` | `/services/{id}` | One service: `state`, `since`, `error`, `ram_mb` |
| `POST` | `/services/{id}/start` | Start. Returns immediately (202) |
| `POST` | `/services/{id}/stop` | Stop |
| `GET` | `/health` | Liveness for the portal |
Response shape per service:
```json
{
"id": "gimp",
"kind": "container",
"target": "family-home-lab-gimp-1",
"group": "media",
"label": "GIMP",
"ram_mb": 233,
"togglable": true,
"state": "running", // running | stopped | starting | stopping | failed | unknown
"since": "2026-10-06T05:12:00+11:00",
"error": null
}
```
Bind to **`127.0.0.1` only**. Add a shared secret header (`X-Controller-Token`) read from
`~/.config/environment.d/10-secrets.conf` so a stray process on the LAN cannot drive it.
### 4.3 State and drift (§2.5)
- Persist **desired state** per service in the portal's Postgres (`service_state` table: `id`,
`desired`, `updated_by`, `updated_at`).
- The controller reports **actual state**. The UI shows both when they disagree and offers
"reconcile".
- On portal startup, compare desired vs actual and surface drift rather than silently "fixing" it.
- Document that `docker compose up -d` re-starts everything. Consider a `deploy/` note or a
post-deploy reconcile step so deployments don't silently turn the whole lab back on.
### 4.4 Portal changes
**Extend the `Tool` dataclass** (`portal/tools.py`) — keep backwards compatibility:
```python
@dataclass(frozen=True)
class Tool:
...existing fields...
# NEW — optional on/off control
service_id: str | None = None # controller id, e.g. "gimp"
ram_mb: int | None = None # shown on the card
togglable: bool = False # only True if allowlisted
```
Then tag the relevant entries in `_cat()`, e.g.:
```python
Tool("gimp", "GIMP", "image", "Photo editing & retouching",
"https://gimp.lab.audasmedia.com.au",
service_id="gimp", ram_mb=233, togglable=True),
```
**New routes** (`portal/main.py`):
| Route | Behaviour |
|---|---|
| `GET /api/services` | Proxy the controller's list; HTMX polls this |
| `POST /service/{id}/start` | Admin (or owner) only → controller → return the **card partial** |
| `POST /service/{id}/stop` | Same |
| `GET /partials/service/{id}` | The card fragment, for HTMX swaps |
All mutating routes must enforce auth **and** the `admin` flag unless explicitly opened to users.
**Template** — `portal/templates/partials/` gets a `service_card.html` with a status pill and a
button. Use the tokens already in `DESIGN.md`:
- `accent-green` `#1aae39` → running
- `accent-orange` `#dd5b00` → starting / stopping
- `ink-faint` `#a39e98` → stopped
- add a red-ish **failed** state (reuse `accent-orange-deep` `#793400` pending a design decision)
Use HTMX polling while a service is in a transitional state (`hx-trigger="every 2s"`), and stop
polling once it settles — otherwise the dashboard hammers the controller.
### 4.6 Making "off by default" actually hold
This is the direct consequence of the confirmed decision, and it is easy to get wrong.
Every service currently uses `restart: unless-stopped`, and the togglable ones live in the **same
compose project as the portal**. That means:
- A plain `docker stop` **does** stick — good, `unless-stopped` means it stays down until something
explicitly starts it.
- But **any future `docker compose up -d` on the project starts them all again**, silently undoing
the family's choices. Deployments happen (see `deploy/DEPLOYMENT.md`), so this *will* occur.
Three ways to handle it, in order of preference:
| Option | How | Trade-off |
|---|---|---|
| **A. Separate compose project** | Move togglable services into `docker-compose.media.yml` (its own project name, e.g. `family-media`). The main project's `up -d` can never touch them. | Cleanest and most robust. Needs a one-time migration of the containers. |
| **B. Compose profiles** | Tag each with `profiles: ["media"]`. Plain `docker compose up -d` skips them; `--profile media` starts them. | Simple, but behaviour when a container already exists and the profile is off should be **verified** before relying on it. |
| **C. Reconcile on deploy** | Keep as-is, and have the portal (or a deploy hook) restore the previous desired state after any `up -d`. | Works, but reactive — there is a window where everything is on and RAM is consumed. |
**Recommendation:** do **A** in Phase 1. It makes "off by default" a structural property rather than
something the UI has to keep re-asserting. It also means the controller's `docker start`/`docker stop`
never fights compose.
Whichever is chosen, keep the §4.3 desired-state record so drift is detectable and reported.
### 4.5 Idle auto-off (optional, later)
Only if wanted: stop a media container after N minutes with **no active selkies websocket**
(`docker exec <c> ss -tn | grep -c ESTAB`), never on a fixed timer alone. Default **off**; opt-in
per service. Show a countdown on the card so nobody is surprised mid-edit.
---
## 5. Implementation plan
### Phase 0 — Truth first (no toggles yet)
- [ ] Add a `GET /api/status` source. Start simple: the portal asks the controller for container
states and renders real "online/offline" pills.
- [ ] Fix `Tool.status` misuse — make status live, or remove the field and derive it.
- [ ] Add `deploy/DEPLOYMENT.md` note on the `.27 → .13` copy step.
- [ ] Decide and record the mechanism from §4.6 for making "off by default" stick.
**Done when:** the dashboard tells the truth about what is running.
### Phase 1 — The controller
- [ ] Write `service-controller` (small Python/FastAPI or Go binary) in this repo under
`deploy/service-controller/`.
- [ ] Define `services.toml` — the allowlist with `id`, `kind`, `target`, `group`, `label`,
`ram_mb`, and `default_state = "stopped"` (§4.6).
- [ ] Implement the §4.6 option (recommended: separate `family-media` compose project) and confirm
that `docker compose up -d` on the main project leaves those services stopped.
- [ ] Implement `GET /services`, `GET /health`, and start/stop for **containers only**.
- [ ] Run as a systemd user unit `family-service-controller.service`, bound to `127.0.0.1:8091`,
with a token from `10-secrets.conf`.
- [ ] Add `extra_hosts: ["host.docker.internal:host-gateway"]` to the portal service in
`docker-compose.yml`.
- [ ] Test with `curl` from the host **and** from inside the portal container.
**Done when:** `curl -XPOST localhost:8091/services/gimp/stop` stops GIMP, and an unlisted id 404s.
### Phase 2 — systemd user services
- [ ] Extend the controller to `systemctl --user start|stop <unit>` using `kind = "systemd-user"`.
- [ ] Verify from the container over `host.docker.internal`.
- [ ] Confirm the controller cannot be used to stop `voice-agent` or anything not allowlisted.
**Done when:** Paperclip and the Prefect worker toggle from the same API.
### Phase 3 — UI
- [ ] Extend `Tool` (§4.4) and tag the targets from §3.
- [ ] Add `partials/service_card.html` with the status pill + button.
- [ ] Wire HTMX: `hx-post` to start/stop, swap the card, poll while transitional.
- [ ] Admin-gate the routes; show RAM on the card.
- [ ] Handle `failed` with the error text and a "view logs" link (dozzle is already on `.35`).
**Done when:** a family member can stop GIMP from the dashboard and see the RAM free up on `.13`.
### Phase 4 — Drift and safety nets
- [ ] `service_state` table + desired/actual comparison + "reconcile" action.
- [ ] Warn in the admin area when a deploy has restarted everything.
- [ ] Group toggles: the WorldMonitor 4-container stack, and Prefect server+worker+dashboard.
- [ ] Document the ordering rules and make grouped start respect them.
### Phase 5 — Optional
- [ ] Idle auto-off (§4.5), default off.
- [ ] Per-user visibility: show "in use by Harry" for shared media containers.
- [ ] RAM freed counter — a running total of what the family has saved.
---
## 6. Safety rules (non-negotiable)
1. **Never** give the portal `docker.sock` or host filesystem access.
2. **Allowlist only.** No generic `docker`/`systemctl` passthrough. Unknown id ⇒ 404.
3. **Never** make the console's own dependencies togglable: `portal`, `postgres`, `redis`,
`rabbitmq`, `garage`, `worker`, `transcriber*`.
4. **Never** make house-critical services togglable: `snapserver`, `librespot`, `mopidy`,
`voice_whisper`, `voice_bridge`, `piper_tts`, `mosquitto`, `pihole`, `omniroute`.
5. Controller binds to `127.0.0.1` and requires a token.
6. Mutating routes are admin-gated by default.
7. Every toggle is **logged** with who, what, when.
---
## 7. Test plan
| Level | Test | Pass |
|---|---|---|
| Unit | Controller rejects an unlisted id | 404, no side effect |
| Unit | Controller rejects a missing/bad token | 401 |
| Unit | Start/stop a container, assert state | State matches `docker inspect` |
| Unit | Start/stop a systemd user service | State matches `systemctl --user is-active` |
| Security | Attempt to stop `postgres` via the API | 404/403, container keeps running |
| Security | Portal container has no `docker.sock` | `ls /var/run/docker.sock` fails inside it |
| UI | Card shows correct state after a manual `docker stop` | Pills update |
| UI | Start a Webtop container | Shows *starting*, then *running*, no hang |
| UI | Force a crash loop | Shows *failed* with a reason |
| Perf | Poll interval does not hammer the controller | ≤1 req/2 s per transitional card |
| Regression | Dashboard, files, voice, admin pages still work | All load |
| Regression | Snapcast still plays after toggling media tools | Audio unaffected |
---
## 8. Open questions
1. **Who may toggle what?** Admin-only for everything, or should any family member be able to stop
the *media* tools (which are shared)?
2. **Auto-off** — wanted at all? If yes, what idle window (30 min?) and should it warn first?
3. **`n8n`** — are there scheduled workflows that must keep running? If so it stays out of scope.
4. **`paperclipai`** — is it still in use, or should it be retired rather than toggled?
5. **Prefect group** — should server + worker + photo-dashboard always be one toggle?
6. **Failed-state colour** — you have `accent-orange-deep #793400`. Add a proper error red token to
`DESIGN.md`, or reuse an existing token?
7. **History** — do you want a "who turned this off and when" audit view, or is a log line enough?
8. **Should the LMMS Xvfb resolution be fixed as part of this** (§2.13), or separately? It is ~500 MB
for one config line.
9. **Controller language** — Python/FastAPI to match the portal, or a single Go binary for a smaller
footprint on `.13`?
10. **Deploy reconcile** — after `docker compose up -d`, should the portal auto-restore the previous
on/off state, or just warn?
---
## 9. Reference
```bash
# --- current container states / memory ---
docker ps --format '{{.Names}}\t{{.Status}}'
docker stats --no-stream --format '{{.Name}} {{.MemUsage}}'
# --- systemd user services (the ones a container cannot reach) ---
systemctl --user list-units --type=service --state=running
systemctl --user is-active paperclipai prefect-worker photo-dashboard where-woof chrome-pi engram
# --- the portal ---
cd /home/sam/Docker/Containers/family-home-lab && docker compose ps
docker logs -f family-home-lab-portal-1
# --- deploy ---
# edit in /home/sam/home_network/custom_tools/family_home_lab (on .27)
# copy to .13:/home/sam/Docker/Containers/family-home-lab, then: docker compose up -d
```
**Files to touch**
| File | Change |
|---|---|
| `portal/tools.py` | extend `Tool`; tag togglable entries |
| `portal/main.py` | `/api/services`, `/service/{id}/start|stop`, card partial |
| `portal/templates/partials/service_card.html` | new |
| `portal/templates/dashboard.html` | render the card partial |
| `portal/database.py` | `service_state` table (Phase 4) |
| `docker-compose.yml` | `extra_hosts` for the portal |
| `deploy/service-controller/` | new — the controller + `services.toml` |
| `DESIGN.md` | one error/failed colour token (pending Q6) |
| `deploy/DEPLOYMENT.md` | controller install + drift note |