22 KiB
ON_OFF.md — Turning services on and off from the Family Console
Target: implement on/off control for RAM-heavy services from console.lab.audasmedia.com.au
Repo: sam/family_home_lab (this repo) · deployed to .13:/home/sam/Docker/Containers/family-home-lab
Status: plan only. Nothing implemented.
Last updated: 2026-10-06
Confirmed decisions
- Togglable services default to OFF. The RAM is the point — freeing it immediately matters more than a 30–60 s wait on first use. The UI must make Start the obvious primary action and show the freed/used RAM on every card. Consequence: "off" must actually hold across deployments — see §4.6. Do not skip that section.
0. Why
.13 has ~46 containers plus several native services. Several are heavy and idle almost all day.
A 2026-10 audit found roughly 2.4 GB of addressable memory, most of it sitting idle:
| Service | Idle RAM |
|---|---|
| Media editors (GIMP, Kdenlive, Audacity, LMMS) | ~1.5 GB incl. a 570 MB LMMS Xvfb |
| Prefect (server + worker + photo-dashboard) | ~650 MB |
| Paperclip | ~495 MB |
| n8n | ~286 MB |
| WorldMonitor (4 containers) | ~120 MB |
The goal: let a family member start a tool when they need it and stop it when they don't, from the console they already use — free RAM without free-of-charge surprise.
1. Current state (verified 2026-10-06)
1.1 What exists
| Thing | Path / detail |
|---|---|
| Console portal (FastAPI + Jinja + HTMX) | portal/ — main.py, tools.py, auth.py, database.py, config.py |
| Tool catalogue | portal/tools.py — a Tool dataclass + _cat() returning the list (~335 lines) |
| Templates | portal/templates/ — dashboard.html, admin.html, voice.html, files.html, partials/ |
| Design system | DESIGN.md — CSS tokens, sticker palette, status colours already defined |
| Routes | /, /tool/{id}, /files, /voice, /admin, /admin/users, /admin/pi, /account/password |
| Auth | session login; roles via admin: bool on each Tool; users seeded from .env |
| Portal container | family-home-lab-portal-1 on .13:8500 |
| DB | Postgres (pgvector/pgvector:pg16) inside the same compose project |
| Deployment | edit here on .27 → copy to .13:/home/sam/Docker/Containers/family-home-lab → docker compose up -d |
1.2 The thing that lies today
Tool already has a status field:
@dataclass(frozen=True)
class Tool:
...
status: str = "online" # <-- a hardcoded string, never updated
It is a literal default, not live state. The dashboard currently reports "online" for every tool regardless of reality. You cannot show on/off state until this becomes real, so fixing it is part 1 of this work, not a nice-to-have.
1.3 The blast radius
The portal's own compose project contains postgres, redis, rabbitmq, garage, worker,
transcriber, transcriber-mus. If the console can stop those, it can kill itself. Any design
must make the console's own dependencies untouchable.
2. Issues to solve
-
Two different control planes. Targets are a mix of Docker containers and systemd user services. A container can reach the Docker API; it cannot reach
systemctl --useron the host without extra plumbing (D-Bus mount, or a helper). One mechanism will not cover both.Containers:
gimp,video-editor,audio-editor,lmms,prefect-server,n8n,worldmonitor*,voice_whisper,headroom, … systemd user services:paperclipai,prefect-worker,photo-dashboard,where-woof,chrome-pi,engram,voice-agent,lan-mouse. systemd system services:snapserver,librespot,mopidy,caddy. -
/var/run/docker.sockis root-equivalent. The portal is internet-reachable via Caddy and protected only by a family login. Mounting the socket into it means one XSS or one weak password equals full host root. Do not do this. -
Self-destruct risk. The console must not be able to stop
postgres,redis,rabbitmq,garage,portal,worker, or any service the console itself depends on. This needs to be enforced by an allowlist, not by UI politeness. -
The dashboard has no live state (§1.2). Needs a real status source before toggles make sense.
-
State drift. Every service uses
restart: unless-stopped, which is helpful — adocker stopsticks. But any futuredocker compose up -dsilently restarts everything, undoing the user's choices. Toggle state must be persisted and reconciled, and drift must be visible. -
Startup is slow and the UI will hang. LinuxServer Webtop containers take 30–60 s to become usable. The endpoint must return immediately and the card must show starting, then poll.
-
Ordering on start.
worker/transcriberneedredis+rabbitmq+postgres. Media containers need theshared-mediavolume mounted. Starting things in the wrong order produces confusing errors. -
Users cannot make informed choices. Nothing shows what a service costs. The card should show approximate RAM so the trade-off is visible.
-
Shared services affect everyone. Stopping GIMP stops it for Harry and Finn too. Needs either admin-only gating or a visible "who's using this" indicator.
-
Failure feedback. If a container enters a crash loop the user must see "failed" with a reason, not a spinner.
-
Distinguish "stopped on purpose" from "crashed". These need different UI treatment. The design system already has the colours: green = online, orange = busy/starting, faint grey = offline.
-
Idle auto-off is tempting but risky. Auto-stopping a media editor mid-edit would be hostile. If implemented, it must detect actual use (selkies has live websocket connections) and only apply to genuinely abandoned sessions.
-
Known separate bug — do not confuse with this work. The
lmmscontainer runs Xvfb at15360x8640, costing a 570 MB framebuffer. That is a config fix (reduce the virtual resolution), not a toggle. Fix it separately; it is the cheapest single win on the box.
3. Targets
Safety ratings: ✅ safe · ⚠️ needs care · ❌ never
| Service | Kind | Idle RAM | Safe? | Who | Notes |
|---|---|---|---|---|---|
gimp |
container | ~233 MB | ✅ | all users | Webtop; slow start (~40 s) |
video-editor (kdenlive) |
container | ~237 MB | ✅ | all users | Webtop |
audio-editor (audacity) |
container | ~201 MB | ✅ | all users | Webtop |
lmms |
container | ~798 MB + 570 MB Xvfb | ✅ | all users | Fix Xvfb resolution first (§2.13) |
prefect-server |
container | ~411 MB | ✅ | admin | Only needed while ingesting |
prefect-worker |
systemd user | ~129 MB | ✅ | admin | Start with the server |
photo-dashboard |
systemd user | ~112 MB | ✅ | admin | Pairs with prefect |
paperclipai |
systemd user | ~495 MB | ✅ | admin | Idle since 2026-09-26 |
n8n |
container | ~286 MB | ⚠️ | admin | Check for scheduled workflows first |
worldmonitor +3 |
containers | ~120 MB | ✅ | admin | Group toggle (4 containers) |
engram |
systemd user | ~1 MB | ✅ | admin | Dead/empty — candidate for removal not toggling |
headroom |
container | ~18 MB | n/a | — | Leave running (JEV proxy, in the pi path) |
voice_whisper |
container | 372–857 MB | ❌ | — | House voice depends on it |
postgres, redis, rabbitmq, garage, portal, worker |
containers | — | ❌ | — | Console's own dependencies |
snapserver, librespot, mopidy |
systemd system | — | ❌ | — | House audio; FIFO ordering trap (see below) |
⚠️ Snapcast ordering trap: snapserver opens its FIFOs at start. If you restart
snapserverbeforelibrespot/mopidy, the streams come up dead (End of file, length: 0) and audio is silent. See/home/sam/chats/snapcast/snapcast-findings.md. Keep these out of the togglable set.
4. Design
4.1 Architecture — a host-side service controller
Do not give the portal the Docker socket. Introduce one small controller that owns all privileged actions and exposes a narrow, allowlisted API.
Browser (HTMX)
│
▼
portal container (.13:8500) ← no docker.sock, no host access
│ HTTP → http://host.docker.internal:8091
▼
service-controller (systemd USER unit on .13, runs as `sam`, 127.0.0.1:8091)
├─ Docker API (sam is in the `docker` group) → containers
└─ systemctl --user → user services
▲
└─ ALLOWLIST + metadata lives here (single source of truth)
Why a host-side systemd user unit:
systemctl --userneeds the user manager, which a container does not have. Running on the host covers user services natively.samis already in thedockergroup, so the same process can also drive containers — no privileged container, no socket in the portal.host.docker.internalis reachable from the portal via Compose:extra_hosts: ["host.docker.internal:host-gateway"].
The allowlist is the security boundary. Anything not listed cannot be started or stopped, and there is no generic passthrough. Reject unknown ids with 404 — never forward caller-supplied names.
Alternative considered and rejected: tecnativa/docker-socket-proxy. It is a fine tool but only
covers Docker, and we need systemd user services too. It would still leave the portal holding a
token that can act on containers.
4.2 Controller API (localhost only)
| Method | Path | Purpose |
|---|---|---|
GET |
/services |
List allowlisted services with live state + metadata |
GET |
/services/{id} |
One service: state, since, error, ram_mb |
POST |
/services/{id}/start |
Start. Returns immediately (202) |
POST |
/services/{id}/stop |
Stop |
GET |
/health |
Liveness for the portal |
Response shape per service:
{
"id": "gimp",
"kind": "container",
"target": "family-home-lab-gimp-1",
"group": "media",
"label": "GIMP",
"ram_mb": 233,
"togglable": true,
"state": "running", // running | stopped | starting | stopping | failed | unknown
"since": "2026-10-06T05:12:00+11:00",
"error": null
}
Bind to 127.0.0.1 only. Add a shared secret header (X-Controller-Token) read from
~/.config/environment.d/10-secrets.conf so a stray process on the LAN cannot drive it.
4.3 State and drift (§2.5)
- Persist desired state per service in the portal's Postgres (
service_statetable:id,desired,updated_by,updated_at). - The controller reports actual state. The UI shows both when they disagree and offers "reconcile".
- On portal startup, compare desired vs actual and surface drift rather than silently "fixing" it.
- Document that
docker compose up -dre-starts everything. Consider adeploy/note or a post-deploy reconcile step so deployments don't silently turn the whole lab back on.
4.4 Portal changes
Extend the Tool dataclass (portal/tools.py) — keep backwards compatibility:
@dataclass(frozen=True)
class Tool:
...existing fields...
# NEW — optional on/off control
service_id: str | None = None # controller id, e.g. "gimp"
ram_mb: int | None = None # shown on the card
togglable: bool = False # only True if allowlisted
Then tag the relevant entries in _cat(), e.g.:
Tool("gimp", "GIMP", "image", "Photo editing & retouching",
"https://gimp.lab.audasmedia.com.au",
service_id="gimp", ram_mb=233, togglable=True),
New routes (portal/main.py):
| Route | Behaviour |
|---|---|
GET /api/services |
Proxy the controller's list; HTMX polls this |
POST /service/{id}/start |
Admin (or owner) only → controller → return the card partial |
POST /service/{id}/stop |
Same |
GET /partials/service/{id} |
The card fragment, for HTMX swaps |
All mutating routes must enforce auth and the admin flag unless explicitly opened to users.
Template — portal/templates/partials/ gets a service_card.html with a status pill and a
button. Use the tokens already in DESIGN.md:
accent-green#1aae39→ runningaccent-orange#dd5b00→ starting / stoppingink-faint#a39e98→ stopped- add a red-ish failed state (reuse
accent-orange-deep#793400pending a design decision)
Use HTMX polling while a service is in a transitional state (hx-trigger="every 2s"), and stop
polling once it settles — otherwise the dashboard hammers the controller.
4.6 Making "off by default" actually hold
This is the direct consequence of the confirmed decision, and it is easy to get wrong.
Every service currently uses restart: unless-stopped, and the togglable ones live in the same
compose project as the portal. That means:
- A plain
docker stopdoes stick — good,unless-stoppedmeans it stays down until something explicitly starts it. - But any future
docker compose up -don the project starts them all again, silently undoing the family's choices. Deployments happen (seedeploy/DEPLOYMENT.md), so this will occur.
Three ways to handle it, in order of preference:
| Option | How | Trade-off |
|---|---|---|
| A. Separate compose project | Move togglable services into docker-compose.media.yml (its own project name, e.g. family-media). The main project's up -d can never touch them. |
Cleanest and most robust. Needs a one-time migration of the containers. |
| B. Compose profiles | Tag each with profiles: ["media"]. Plain docker compose up -d skips them; --profile media starts them. |
Simple, but behaviour when a container already exists and the profile is off should be verified before relying on it. |
| C. Reconcile on deploy | Keep as-is, and have the portal (or a deploy hook) restore the previous desired state after any up -d. |
Works, but reactive — there is a window where everything is on and RAM is consumed. |
Recommendation: do A in Phase 1. It makes "off by default" a structural property rather than
something the UI has to keep re-asserting. It also means the controller's docker start/docker stop
never fights compose.
Whichever is chosen, keep the §4.3 desired-state record so drift is detectable and reported.
4.5 Idle auto-off (optional, later)
Only if wanted: stop a media container after N minutes with no active selkies websocket
(docker exec <c> ss -tn | grep -c ESTAB), never on a fixed timer alone. Default off; opt-in
per service. Show a countdown on the card so nobody is surprised mid-edit.
5. Implementation plan
Phase 0 — Truth first (no toggles yet)
- Add a
GET /api/statussource. Start simple: the portal asks the controller for container states and renders real "online/offline" pills. - Fix
Tool.statusmisuse — make status live, or remove the field and derive it. - Add
deploy/DEPLOYMENT.mdnote on the.27 → .13copy step. - Decide and record the mechanism from §4.6 for making "off by default" stick.
Done when: the dashboard tells the truth about what is running.
Phase 1 — The controller
- Write
service-controller(small Python/FastAPI or Go binary) in this repo underdeploy/service-controller/. - Define
services.toml— the allowlist withid,kind,target,group,label,ram_mb, anddefault_state = "stopped"(§4.6). - Implement the §4.6 option (recommended: separate
family-mediacompose project) and confirm thatdocker compose up -don the main project leaves those services stopped. - Implement
GET /services,GET /health, and start/stop for containers only. - Run as a systemd user unit
family-service-controller.service, bound to127.0.0.1:8091, with a token from10-secrets.conf. - Add
extra_hosts: ["host.docker.internal:host-gateway"]to the portal service indocker-compose.yml. - Test with
curlfrom the host and from inside the portal container.
Done when: curl -XPOST localhost:8091/services/gimp/stop stops GIMP, and an unlisted id 404s.
Phase 2 — systemd user services
- Extend the controller to
systemctl --user start|stop <unit>usingkind = "systemd-user". - Verify from the container over
host.docker.internal. - Confirm the controller cannot be used to stop
voice-agentor anything not allowlisted.
Done when: Paperclip and the Prefect worker toggle from the same API.
Phase 3 — UI
- Extend
Tool(§4.4) and tag the targets from §3. - Add
partials/service_card.htmlwith the status pill + button. - Wire HTMX:
hx-postto start/stop, swap the card, poll while transitional. - Admin-gate the routes; show RAM on the card.
- Handle
failedwith the error text and a "view logs" link (dozzle is already on.35).
Done when: a family member can stop GIMP from the dashboard and see the RAM free up on .13.
Phase 4 — Drift and safety nets
service_statetable + desired/actual comparison + "reconcile" action.- Warn in the admin area when a deploy has restarted everything.
- Group toggles: the WorldMonitor 4-container stack, and Prefect server+worker+dashboard.
- Document the ordering rules and make grouped start respect them.
Phase 5 — Optional
- Idle auto-off (§4.5), default off.
- Per-user visibility: show "in use by Harry" for shared media containers.
- RAM freed counter — a running total of what the family has saved.
6. Safety rules (non-negotiable)
- Never give the portal
docker.sockor host filesystem access. - Allowlist only. No generic
docker/systemctlpassthrough. Unknown id ⇒ 404. - Never make the console's own dependencies togglable:
portal,postgres,redis,rabbitmq,garage,worker,transcriber*. - Never make house-critical services togglable:
snapserver,librespot,mopidy,voice_whisper,voice_bridge,piper_tts,mosquitto,pihole,omniroute. - Controller binds to
127.0.0.1and requires a token. - Mutating routes are admin-gated by default.
- Every toggle is logged with who, what, when.
7. Test plan
| Level | Test | Pass |
|---|---|---|
| Unit | Controller rejects an unlisted id | 404, no side effect |
| Unit | Controller rejects a missing/bad token | 401 |
| Unit | Start/stop a container, assert state | State matches docker inspect |
| Unit | Start/stop a systemd user service | State matches systemctl --user is-active |
| Security | Attempt to stop postgres via the API |
404/403, container keeps running |
| Security | Portal container has no docker.sock |
ls /var/run/docker.sock fails inside it |
| UI | Card shows correct state after a manual docker stop |
Pills update |
| UI | Start a Webtop container | Shows starting, then running, no hang |
| UI | Force a crash loop | Shows failed with a reason |
| Perf | Poll interval does not hammer the controller | ≤1 req/2 s per transitional card |
| Regression | Dashboard, files, voice, admin pages still work | All load |
| Regression | Snapcast still plays after toggling media tools | Audio unaffected |
8. Open questions
- Who may toggle what? Admin-only for everything, or should any family member be able to stop the media tools (which are shared)?
- Auto-off — wanted at all? If yes, what idle window (30 min?) and should it warn first?
n8n— are there scheduled workflows that must keep running? If so it stays out of scope.paperclipai— is it still in use, or should it be retired rather than toggled?- Prefect group — should server + worker + photo-dashboard always be one toggle?
- Failed-state colour — you have
accent-orange-deep #793400. Add a proper error red token toDESIGN.md, or reuse an existing token? - History — do you want a "who turned this off and when" audit view, or is a log line enough?
- Should the LMMS Xvfb resolution be fixed as part of this (§2.13), or separately? It is ~500 MB for one config line.
- Controller language — Python/FastAPI to match the portal, or a single Go binary for a smaller
footprint on
.13? - Deploy reconcile — after
docker compose up -d, should the portal auto-restore the previous on/off state, or just warn?
9. Reference
# --- current container states / memory ---
docker ps --format '{{.Names}}\t{{.Status}}'
docker stats --no-stream --format '{{.Name}} {{.MemUsage}}'
# --- systemd user services (the ones a container cannot reach) ---
systemctl --user list-units --type=service --state=running
systemctl --user is-active paperclipai prefect-worker photo-dashboard where-woof chrome-pi engram
# --- the portal ---
cd /home/sam/Docker/Containers/family-home-lab && docker compose ps
docker logs -f family-home-lab-portal-1
# --- deploy ---
# edit in /home/sam/home_network/custom_tools/family_home_lab (on .27)
# copy to .13:/home/sam/Docker/Containers/family-home-lab, then: docker compose up -d
Files to touch
| File | Change |
|---|---|
portal/tools.py |
extend Tool; tag togglable entries |
portal/main.py |
/api/services, `/service/{id}/start |
portal/templates/partials/service_card.html |
new |
portal/templates/dashboard.html |
render the card partial |
portal/database.py |
service_state table (Phase 4) |
docker-compose.yml |
extra_hosts for the portal |
deploy/service-controller/ |
new — the controller + services.toml |
DESIGN.md |
one error/failed colour token (pending Q6) |
deploy/DEPLOYMENT.md |
controller install + drift note |