Correct .13/.27 notes from live 2026-10-05 inspection

- Home Network Map: .13 is memory-bound (swap 81%); KDE Plasma not Niri;
  GPU dead with no driver bound; Open WebUI->Outline on :3000; full .13
  container table; Pi-hole primary .35 (confirmed from nebulasync config);
  .27 airflow row removed
- Filesystem Drive Map: .13 device letters corrected (root sda2, storage
  sdb2, data sdc2, ubuntu_storage sdd1); add sda3 swap; Maxtor not mounted
- Docker Containers: garage v2.1.0; paseo is native not Docker; engram
  container dead + native service empty; airflow dead with 24GB logs;
  add missing containers and native/user service lists
- Websites on .13: add :8000 Caddy detail and media-tool ports
- New note: Resource Use - .13 CPU & RAM
This commit is contained in:
2026-10-05 14:48:08 +11:00
parent 7c5151f4bd
commit 7ed646ac0a
5 changed files with 273 additions and 34 deletions

View File

@@ -6,7 +6,7 @@ client: sam
project: ai
status: active
priority: 5
last_verified: 2026-06-19
last_verified: 2026-10-05
tags:
- network
- docker
@@ -36,12 +36,12 @@ id: 1778553013-ARYX
- **nebula-sync** — Pi-hole Gravity sync
- **pocketbase** — PocketBase chat app backend
- **doorbell_media** — Doorbell camera/media (nginx)
- **engram** — Journaling app (ghcr.io/gentleman-programming/engram)
- **headroom** — AI headroom dashboard (ghcr.io/chopratejas/headroom)
- **engram** — ⚠️ **container dead.** Created 2026-06-13, never started (its `engram-proxy` socat sidecar was never created either). The live engram is now a **native user systemd service** (`/home/sam/.local/bin/engram serve`, `:7437`, data `~/.engram/`). Live API reports **0 sessions / 0 observations / 0 prompts — unused**. See [[Resource Use — .13 CPU & RAM]]
- **headroom** — JEV / context-compression proxy (`:8787`). Runs `headroom proxy --proxy-extension fast_jev` with `OPENAI_TARGET_API_URL=http://192.168.20.13:20129/v1` (i.e. it fronts OmniRoute). 🟡 Idle — ~18 MB resident, ~516 MB swapped
- **n8n_data** — Workflow automation (n8n)
- **paseo** — Multi-agent orchestration GUI (ghcr.io/getpaseo/paseo:latest + pi 0.82.1 + omni ext baked in). Port **6767**. Drives pi agents spawned via `pi --mode rpc`, LLM via OmniRoute. Web UI `http://192.168.20.13:6767` (password). Config: `/home/sam/Docker/Containers/paseo/` (compose, `.env` 600, `paseo-home` volume, `workspace`). Docs: [[Paseo Pi GUI Tool]] · Gitea `sam/paseo` · maps.lab `/paseo/docs/`
- **t3_stack_react** — T3 stack React app
- **sams-home-network** — Home network management app
- **paseo** — ⚠️ **Not a container any more.** Paseo runs as a **native user systemd service** (`Paseo Daemon`, `:6767`, ~151 MB). Drives pi agents spawned via `pi --mode rpc`, LLM via OmniRoute. Web UI `http://192.168.20.13:6767` (password). Docs: [[Paseo Pi GUI Tool]] · Gitea `sam/paseo` · maps.lab `/paseo/docs/`
- **t3_stack_react** — T3 stack React app — *not running*
- **sams-home-network** — Home network management app — *not running*
- **worldmonitor** — Geopolitical news & 3D globe dashboard. Next.js app, Redis cache, ais-relay. Port 3002. LLM features route through OmniRoute.
### `/home/sam/deployment/ai-resume/`
@@ -50,19 +50,20 @@ id: 1778553013-ARYX
- **ai-resume-backend** — AI Resume backend API
### `/home/sam/deployment/lite_llm/`
- **litellm** — LiteLLM AI proxy (ghcr.io/berriai/litellm)
- **litellm** — LiteLLM AI proxy (ghcr.io/berriai/litellm, `:4000`). 🟡 **Idle and fully swapped out** (~4 MB resident, ~526 MB swapped). Duplicate of OmniRoute for LLM traffic — candidate to stop. ⚠️ Its `docker-compose.yml` contains an OpenRouter API key in plaintext
### `/home/sam/deployment/langfuse/` — ⛔ REMOVED 2026-09-06
- ~~**langfuse** — AI observability & tracing (langfuse/langfuse:3)~~ — **removed** (containers + all data volumes deleted). Was the resource-heavy Langfuse stack (web/worker/clickhouse/postgres/redis/minio) that pi-langfuse traced into. Compose file left in place but unused.
### `/home/sam/deployment/airflow/`
- **airflow** — Data pipeline orchestration (apache/airflow:2.10.5)
### `/home/sam/deployment/airflow/` — ⛔ DEAD
- ~~**airflow** — Data pipeline orchestration (apache/airflow:2.10.5)~~ — **all 7 containers exited 4 weeks ago** (webserver, scheduler, worker, triggerer, init, postgres, redis). `restart: no`. **24 GB of logs** in `logs/` (6,162 files). Nothing running depends on it.
- ⚠️ A **second, separate** Airflow install exists at `/home/sam/deployment/ai-resume/airflow/` (own compose + `dags/gitea_ingestion_dag.py`). Its containers are not running either; the ai-resume `knowledge-service` and `langgraph-service` still run independently of it.
### `/home/sam/deployment/trigger_dev/`
- **trigger_dev** — Trigger.dev automation platform
### `/home/sam/deployment/garage/`
- **garage** — S3-compatible object storage (dxflrs/garage:v1.0.0). On `.13:3900/3902`. Buckets per user + `shared-media`. S3 API root domain `.s3.lab.audasmedia.com.au`.
- **garage** — S3-compatible object storage (dxflrs/garage:**v2.1.0** — live; this note previously said v1.0.0). On `.13:3900-3903`. Buckets per user + `shared-media`. S3 API root domain `.s3.lab.audasmedia.com.au`.
- **S3 viewing** — `offsite.lab.audasmedia.com.au` (→ `.13:8095`, basic-auth) shows storage status at `/status.html`. Garage​web UI intended at `garage-ui.lab.audasmedia.com.au` (resolves; confirm live).
### `/home/sam/Docker/Containers/family-home-lab/` — Family Home Lab core
@@ -73,7 +74,27 @@ id: 1778553013-ARYX
### `/home/sam/Docker/Containers/dsh/` — DeepSeek Harness
- **dsh-sam / dsh-jo / dsh-harry / dsh-finn** — per-user chat (ports 3081–3084). Memory, web/doc summarise, image ingest + generate. Hardened (no shell, read-only rootfs).
Also on .13: **where-woof**, **prefect**, **omniroute**, ~~langfuse (removed 2026-09-06)~~, airflow, µstacks above.
Also on .13: **where-woof** (`:3020`), **prefect** (`:4200`), **omniroute** (`:20128/20129`), ~~langfuse (removed 2026-09-06)~~, ~~airflow (dead — 24 GB logs)~~, µstacks above.
### Containers missing from this note (live-verified 2026-10-05)
| Container(s) | Project dir | Ports |
|---|---|---|
| `outline`, `outline_postgres`, `outline_redis` | `/home/sam/Docker/Containers/outline` | 3000 |
| `wherewoof-db`, `wherewoof-admin`, `wherewoof-minio` | `/home/sam/Docker/Containers/wherewoof-*` | 5434, 3031, 9010/9011 |
| `kontra` | `/home/sam/Docker/Containers/kontra` | 8600 |
| `metube` | `/home/sam/Docker/Containers/metube` | 8086 |
| `prefect-server` | `/home/sam/Docker/Containers/prefect` | 4200 |
| `worldmonitor-redis-rest`, `worldmonitor-ais-relay` | `/home/sam/Docker/Containers/worldmonitor` | 8079 (localhost) |
| `omniroute-redis` | `/home/sam/Docker/Containers/omniroute` | internal |
| `family-home-lab-garage-1`, `-garage-webui-1`, `-transcriber-mus-1` | `/home/sam/Docker/Containers/family-home-lab` | 3900-3903, 3909 |
| `lmms` (`music` project) | `/home/sam/Docker/Containers/music` | 8085 |
| `ai-resume-db-1` | `/home/sam/deployment/ai-resume` | 5432 |
### Native (non-Docker) services on .13
**System:** `caddy :8000` (serves `/var/www` resume/portfolio + sprinklers, sequence, chat) · `snapserver :1704/1705/1780` · `librespot` · `mopidy :6680` · `postgresql :5433` (database `paperclip`) · `voice-agent :8501` (Jervis) · `offsite-web :8095` · `tailscaled` · `docker`
**User:** `paperclipai :3100` (~495 MB) · `Paseo Daemon :6767` · `where-woof :3020` · `photo-dashboard :8092` · `prefect-worker` (pool `photo-pool`) · `engram :7437` · `chrome-pi :9222` (headless Chrome/CDP) · `lan-mouse` · KDE Plasma session (~2.4 GB)
### `/home/sam/Docker/Containers/outline/` — Outline wiki
- **outline** — Project documentation/wiki (on `.13:3000`, Docker stack: outline + outline_postgres + outline_redis, Garage S3 storage). URL: `outline.lab.audasmedia.com.au`. Has REST API + MCP; project-ops skill covers docs access.

View File

@@ -1,6 +1,6 @@
---
created: 2026-05-28
modified: 2026-08-29
modified: 2026-10-05
type: area
status: active
tags:
@@ -124,14 +124,17 @@ aliases: [drive-map, filesystem-map]
### Local Drives
> **⚠️ Corrected 2026-10-05 (live-verified on `.13`).** The device letters below were wrong. NixOS uses `fstab` UUIDs, so letters shift when the USB drive is replugged — always check `lsblk`/`df` before using a device path.
| Logical Name | Device | Size | Used | Avail | Mount | Label | UUID | Note |
|-------------|--------|------|------|-------|-------|-------|------|------|
| **System Root** | `sda2` | 907G | 424G | 437G (50%) | `/` | root | `0d57bb68-...` | NixOS system |
| **Boot** | `sda1` | 1G | — | — | `/boot` | — | `4D80-F99E` | EFI boot |
| **Storage** | `sdc2` | 931G | 2.1M | **870G (1%)** | `/mnt/storage` | archive_storage | — | **EMPTY** — reformatted ext4 2026-05-30 (was NTFS Windows drive) |
| **Data** | `sdb2` | 1.8T | 2.1M | **1.7T (1%)** | `/mnt/data` | Data | — | **EMPTY** — reformatted ext4 2026-05-30. **Planned Takeout landing zone** |
| **Ubuntu Storage** | `sde1` | 2.7T | 874G | **1.9T (33%)** | `/mnt/ubuntu_storage_3TB` | ubuntu_storage_3 | `037a542c-...` | **USB drive** — holds `archive/` (photo master), `timeshift`, `backup/`. Backup target for .27 + .13 Borg repos. ⚠️ USB — unplugged Aug 21-26 2026 caused backup failures; replugged, letter shifted `sdd1→sde1` |
| **MaxtorBackup** | `sdd1` | 1.4T | ~0 | — | `/mnt/maxtor_backup` | MaxtorBackup | `b0fa7768-...` | Old backup drive — only NixOS migration tarballs, no photos |
| **System Root** | `sda2` | 907G | **459G** | 402G (54%) | `/` | root | `0d57bb68-...` | NixOS system |
| **Boot** | `sda1` | 1G | 139M | 884M | `/boot` | — | `4D80-F99E` | EFI boot |
| **Swap** | `sda3` | 8.8G | **7.1G** | 1.7G | `[SWAP]` | — | — | **81% full — the main resource problem on `.13`** |
| **Storage** | `sdb2` | 916G | 2.1M | **870G (1%)** | `/mnt/storage` | archive_storage | — | **EMPTY** — reformatted ext4 2026-05-30 (was NTFS Windows drive) |
| **Data** | `sdc2` | 1.8T | **357G** | 1.4T (21%) | `/mnt/data` | Data | — | Takeout landing zone |
| **Ubuntu Storage** | `sdd1` | 2.7T | **963G** | **1.8T (36%)** | `/mnt/ubuntu_storage_3TB` | ubuntu_storage_3 | `037a542c-...` | **USB drive** — holds `archive/` (photo master), `timeshift`, `backup/`. Backup target for .27 + .13 Borg repos. ⚠️ USB — unplugged Aug 21-26 2026 caused backup failures. Letter has been `sdd1`, `sde1` and back again |
| **MaxtorBackup** | `sde1` | 1.4T | — | — | **not mounted** | MaxtorBackup | `b0fa7768-...` | Old backup drive — only NixOS migration tarballs, no photos. **Not currently mounted** (was reported mounted on 2026-08-06) |
### Key Photo Location (moved 2026-05-30; reconciled 2026-08-27)
@@ -233,10 +236,12 @@ No SSH access available. Contents fully visible via .35's CIFS mount.
1. ✅ **Immich photos recovered** — found on My Passport, now mounted at `/mnt/hd`
2. ✅ **`/mnt/hd/` backup added** — `/host_fs/mnt/hd` added to Backrest (Restic) plan
3. **.23 has no backup** — single point of failure for .35's backup repos
4. **MaxtorBackup on .13 now mounted** — `/mnt/maxtor_backup` empty (4KB), verified 2026-08-06
4. ~~**MaxtorBackup on .13 now mounted** — `/mnt/maxtor_backup` empty (4KB), verified 2026-08-06~~ **Not mounted as of 2026-10-05.**
### Next Steps
- [ ] Confirm Backrest backup of `/host_fs/mnt/hd` runs successfully on Sunday
- [ ] Consider if `.23` file-server needs its own backup
- [ ] Check MaxtorBackup (`/dev/sde1` on .13) if needed
- [ ] **Reclaim 24 GB of dead Airflow logs** on `.13` — `/home/sam/deployment/airflow/logs` (6,162 files, stack stopped 4 weeks; nothing depends on it)
- [ ] **`/` on `.13` is 54% full and its swap is 81% full** — see [[Resource Use — .13 CPU & RAM]]

View File

@@ -1,6 +1,6 @@
---
created: 2026-07-30
modified: 2026-09-06
modified: 2026-10-05
type: area
status: active
tags:
@@ -17,7 +17,7 @@ aliases: []
|---------|---------------|
| Subnet | `192.168.20.0/24` |
| Gateway | `192.168.20.1` |
| DNS | Pi-hole on .13 (primary) + .35 (secondary) |
| DNS | Pi-hole on **.35 (primary)** + **.13 (replica)** — confirmed live from `nebulasync` config: `PRIMARY=http://192.168.20.35`, `REPLICAS=http://192.168.20.13` |
---
@@ -50,7 +50,9 @@ Mermaid + Archify (diagrams), Vikunja (tasks) and Outline (docs) are documented
| **knowledge-service** | Custom Python API (`:8080`) |
| **langgraph-service** | LangGraph agent framework (`:8090`) |
| **opencode-brain** | OpenCode AI service (`:5000`) |
| **airflow** | Workflow orchestration |
| **wherewoof-admin** | Where Woof admin UI — **exited** (6 weeks) |
> Corrected 2026-10-05: `airflow` does **not** run on `.27`. The only Airflow installs are on `.13` (see below).
### .13 — nixos-desktop (Server)
@@ -61,13 +63,21 @@ Mermaid + Archify (diagrams), Vikunja (tasks) and Outline (docs) are documented
| **OS** | NixOS |
| **SSH** | ✅ `sam@192.168.20.13` |
| **Tailscale** | ✅ `100.114.62.46` (nixos-desktop) |
| **Role** | Docker host, OmniRoute LLM proxy, always-on server. 15.5GB RAM. GPU: GTX 760 (dead — see below).
| **Key services** | Open WebUI v0.11.0 `:3000` (NixOS native, not Docker) |
| **Role** | Docker host, OmniRoute LLM proxy, always-on server. 15 GiB RAM. GPU: GTX 760 (dead — see below).
| **Key services** | Outline wiki `:3000` (Docker). Open WebUI was removed 2026-08-26 |
| **NixOS** | 26.11 (Zokor) |
| **CPU** | AMD Ryzen 5 5600 — 6 cores / 12 threads |
| **Disks** | `sda2` 907G `/` (54% used) · `sdb2` `/mnt/storage` · `sdc2` `/mnt/data` · `sdd1` `/mnt/ubuntu_storage_3TB` (USB) · `sde1` Maxtor (unmounted) |
| **Live state**<br>(2026-10-05) | load avg **0.27** (12 threads) · RAM 6.8 GiB used of 15 GiB · **swap 7.1 GiB of 8.8 GiB used (81%)** |
> **`.13` is memory-bound, not CPU-bound.** The CPU is almost idle; the swap partition is 81% full. See [[Resource Use — .13 CPU & RAM]] and the project inventory at `~/chats/sys_config/resource_use_cpu_ram_13/SYSTEMS-INVENTORY.md`.
#### GPU upgrade — dead GTX 760 → AMD RX 6600 (recommended)
**Status (2026-08-25):** GTX 760 (Kepler, PCI `10de:11c2`) no longer works. `nvidiaPackages.stable` dropped Kepler support after driver branch 470; dmesg shows `NVRM: does not include the required GPU ... probe failed (-1)`. The legacy 470 driver is EOL and very unlikely to build on kernel 6.18 (~10–25% odds even with kernel pinning).
**Verified 2026-10-05:** no driver is bound at all. The only DRM device is `/dev/dri/card0` with `DRIVER=simple-framebuffer`; `/proc/driver/nvidia` does not exist. The desktop renders through software (llvmpipe/swrast) — this is why `.13`'s KDE processes use ~2.4 GB. Note `hardware.nvidia.open = true` in `configuration.nix` can never work on Kepler (open modules need Turing or newer).
**Recommended replacement: AMD RX 6600** (~$180–210 USD)
- **Zero NixOS driver pain:** in-kernel `amdgpu` — config becomes just `services.xserver.videoDrivers = [ "amdgpu" ];`. No legacy branches, no kernel-version roulette, ever.
- **No PSU gamble:** many models need no PCIe power connector (132 W). Safe with the new PSU regardless of wattage headroom.
@@ -82,14 +92,25 @@ Mermaid + Archify (diagrams), Vikunja (tasks) and Outline (docs) are documented
| **~~Langfuse~~** | ~~LLM observability — traces, evals, cost tracking~~ **⛔ REMOVED 2026-09-06** (was port `:3001`, bumped from 3000 by Open WebUI). Stack + data deleted — resource hog. |
| **worldmonitor** | Geopolitical news dashboard (port 3002) |
| **mosquitto** | MQTT broker |
| **pihole** | DNS ad-blocking (primary) |
| **nebula-sync** | Pi-hole Gravity sync |
| **headroom** | Context compression proxy (port 8787) |
| **engram** | Journaling service (port 7437) |
| **pihole** | DNS ad-blocking — **replica** of `.35` (web UI `:8080`) |
| **nebula-sync** | Pulls Pi-hole Gravity from `.35:8095` → `.13:8080` |
| **headroom** | JEV/context-compression proxy (port 8787) — `--proxy-extension fast_jev` → OmniRoute |
| **~~engram~~** | ⚠️ The *container* was created 2026-06-13 and **never started** (dead). The live engram is a **native user systemd service** `/home/sam/.local/bin/engram serve` on `:7437`, with data at `~/.engram/`. Its API reports **0 sessions, 0 observations, 0 prompts — it is empty and unused** |
| **n8n** | Workflow automation |
| **voice_bridge + voice_whisper** | MQTT audio bridge |
| **outline + outline_postgres + outline_redis** | Wiki / project documentation (`:3000`) |
| **voice_bridge + voice_whisper** | MQTT audio bridge + Whisper STT (`:5000`, model `small.en`, ~877 MB — largest container) |
| **piper_tts** | Text-to-speech |
| **Plus:** pocketbase, doorbell_media, litellm, airflow, trigger_dev, garage, **garage-webui**, t3_stack_react, sams-home-network | |
| **prefect-server** | Photo-pipeline workflow server (`:4200`) — pair with the `prefect-worker` user service |
| **family-home-lab** (12 containers) | portal `:8500`, transcriber, gimp `:8087`, video `:8083`, audio `:8084`, lmms `:8085`, postgres, rabbitmq, redis, garage `:3900-3903`, garage-webui `:3909` |
| **dsh** (4 containers) | DeepSeek Harness per family member (`:3081-3084`) |
| **wherewoof** (3 containers) | `wherewoof-db :5434`, `wherewoof-admin :3031`, `wherewoof-minio :9010/9011` |
| **ai-resume** (4 containers) | backend `:8001`, db, knowledge-service `:8082`, langgraph-service `:8091` — the last two also duplicate `.27` |
| **worldmonitor** (4 containers) | dashboard `:3002`, redis, redis-rest, ais-relay |
| **Plus:** pocketbase `:8090`, doorbell_media `:8088`, metube `:8086`, kontra `:8600`, lite_llm `:4000`, lite_llm| |
| **⛔ Dead stacks** | `airflow` — 7 containers exited 4 weeks ago, **24 GB of logs** in `/home/sam/deployment/airflow/logs` · `engram` container (never started) · `trigger_dev` (not running) |
**Native services on `.13` (not Docker):** `caddy :8000` (serves `/var/www` resume + portfolio sites), `snapserver`, `librespot`, `mopidy`, `postgresql :5433` (`paperclip` DB), `voice-agent :8501` (Jervis), `offsite-web :8095`, `tailscaled`.
**User services:** `paperclipai :3100` (~495 MB), `Paseo Daemon :6767`, `where-woof :3020` (Go), `photo-dashboard :8092`, `prefect-worker`, `engram :7437`, `chrome-pi :9222`, `lan-mouse`, KDE Plasma session (~2.4 GB).
See [[Docker Containers]] for full container list.
@@ -251,8 +272,16 @@ Host phone
| `8006` | Proxmox VE web UI | .28 |
| `8007` | Proxmox Backup Server web UI | .48 |
| `8123` | Home Assistant web UI | .30 |
| `3000` | Open WebUI | .13 |
| `3000` | Outline wiki (was Open WebUI — removed 2026-08-26) | .13 |
| ~~`3001`~~ | ~~Langfuse — removed 2026-09-06~~ | .13 |
| `8000` | Caddy static sites (`/var/www`) | .13 |
| `5433` | Native PostgreSQL (`paperclip` DB) | .13 |
| `3100` | Paperclip agent harness | .13 |
| `6767` | Paseo Daemon (native, not Docker) | .13 |
| `3020` | Where Woof frontend (Go) | .13 |
| `8092` | photo-pipeline dashboard | .13 |
| `8095` | offsite status page | .13 |
| `9222` | Headless Chrome (pi browser harness, localhost) | .13 |
| `2222` | Gitea SSH (git remotes) | .35 |
| `3001` | Gitea web UI | .35 |
| `3090` | Archon | .27 |
@@ -290,6 +319,11 @@ Host phone
|-----|-----|-------|
| Snapcast | `http://192.168.20.13:1780/` | Multi-room audio control — needs Caddy DNS |
| ai-resume-backend | `http://192.168.20.13:8001` | Resume AI backend (Docker) |
| Outline wiki | `http://192.168.20.13:3000` | Project docs/wiki (Docker) |
| Caddy static sites | `http://192.168.20.13:8000` | `/var/www` — resume/portfolio + sprinklers, sequence, chat |
| Paperclip | `http://192.168.20.13:3100` | Agent harness manager (~495 MB) |
| Paseo | `http://192.168.20.13:6767` | Multi-agent GUI (native user service, not Docker) |
| Where Woof | `http://192.168.20.13:3020` | Go/HTMX frontend |
### .35 — caddy-server (Reverse Proxy)
@@ -356,8 +390,8 @@ Account: `samuelrolfe@gmail.com`
| Server | IP | Role |
|--------|----|------|
| Pi-hole (primary) | `192.168.20.35` | DNS ad-blocking, local DNS for `.home.lab` domains |
| Pi-hole (secondary) | `192.168.20.13` | DNS ad-blocking, failover (nebula-sync from .35) |
| Pi-hole (primary) | `192.168.20.35` | DNS ad-blocking, local DNS for `.home.lab` domains — **confirmed 2026-10-05** |
| Pi-hole (replica) | `192.168.20.13` | DNS ad-blocking, failover (nebula-sync pulls from .35 → .13) |
### Local domains needing DNS records

View File

@@ -0,0 +1,166 @@
---
created: 2026-10-05
modified: 2026-10-05
type: area
status: active
tags:
- dev-ops
- resource-usage
- machines
aliases: [resource-use-13]
---
# Resource Use — .13 CPU & RAM
> Live-verified **2026-10-05**. Read-only investigation — no configuration was changed.
> Full inventory: `~/chats/sys_config/resource_use_cpu_ram_13/SYSTEMS-INVENTORY.md`
## Headline
`.13` is **not CPU-bound. It is memory-bound.**
| Metric | Value |
|---|---|
| CPU load average (1/5/15 min) | **0.27 / 0.42 / 0.37** on 12 threads — the CPU is nearly idle |
| RAM | 6.8 GiB used of 15 GiB · 9.0 GiB buff/cache · 8.7 GiB available |
| **Swap** | **7.1 GiB used of 8.8 GiB — 81 % full** |
| Uptime | 28 days |
| Root disk | 459 GB of 907 GB (54 %) |
The swap partition is the constraint. A full swap explains the stalls. It also wears the SSD.
## Why the swap is full
Linux swapped out the *coldest* anonymous pages of long-running processes. Three groups dominate:
| Group | Combined RAM + swap | Detail |
|---|---|---|
| **KDE Plasma desktop session** | ~2.4 GB | ~30 processes: plasmashell (301 MB RSS + 870 MB swap), kscreenlocker (117+128), kwin_wayland (59+141), kdeconnectd (54+178), xdg-desktop-portal-kde (11+118) |
| **Docker containers** | ~3.5 GB RSS | `voice_whisper` 877 MB, `omniroute` 827 MB, `prefect-server` 411 MB, `n8n` 286 MB, `pihole` 205 MB, `outline` 187 MB |
| **Idle LLM gateways** | ~1.05 GB swap | `litellm` (~4 MB resident, **526 MB swapped**) and `headroom` (~18 MB resident, **516 MB swapped**) |
### Why the Plasma session is so large
Two reasons, and they reinforce each other:
1. **There is no working GPU driver.** The only DRM device is `/dev/dri/card0` with `DRIVER=simple-framebuffer`. `/proc/driver/nvidia` does not exist. Plasma and KWin fall back to **software rasterisation** (llvmpipe/swrast — 33 mesa mappings in `plasmashell`). Software GL keeps much larger CPU-side buffers than a GPU path does.
> `hardware.nvidia.open = true` in `configuration.nix` can never work on the installed Kepler card. Open kernel modules need Turing or newer. `hardware.nvidia.open` must be `false` for the legacy driver branch.
2. **28 days of uptime with no restart.** Plasma and KDE Connect accumulate memory over weeks. `kdeconnectd` alone holds 178 MB of swap with only 54 MB resident.
Also worth noting: **`kscreenlocker_greet` has been running for 19 days.** A monitor is attached (`card0-Unknown-1: connected`) and the session on `tty2` has been logged in since 2026-09-06, but the screen has been locked and untouched for ~19 days. That process costs 117 MB resident + 128 MB swap for nothing.
### Why `voice_whisper` is 877 MB
It is a Flask API that loads a Faster-Whisper model **at import time and holds it forever**:
```python
model_size = "small.en"
model = WhisperModel(model_size, device="cpu", compute_type="int8")
```
- Model: `Systran/faster-whisper-small.en`, cached in the image (image size 1.01 GB)
- Process: `VmRSS 877 MB`, `VmSwap 88 MB`, `VmSize 2.4 GB`
- Device: **CPU** — no GPU acceleration is possible while the GPU is dead
- Port 5000, no memory limit set (unlike `voice_bridge`, which is capped at 500 MB)
The model is resident 24/7 for a voice assistant that is used occasionally.
### What "swapped" means for `litellm` and `headroom`
`VmSwap` is the amount of a process's memory that the kernel has written to the swap partition.
**Swapped memory is not using RAM right now** — so these two are not consuming RAM.
They are consuming *swap space*, and that is still the scarce resource on `.13`.
| Service | VmRSS (in RAM) | VmSwap (on disk) |
|---|---|---|
| `litellm` | 4 MB | **526 MB** |
| `headroom` | 18 MB | **516 MB** |
Both have been idle for weeks, so the kernel paged almost all of them out. Together they occupy
~1.05 GB of an 8.8 GB swap partition — about 12 % of it — while doing no work.
`litellm` is a second LLM gateway alongside OmniRoute and is functionally redundant.
`headroom` is the **JEV** proxy (`--proxy-extension fast_jev`, forwarding to OmniRoute on `:20129`),
so it is only needed when JEV-based determination is in use. Stop them and the kernel reclaims the swap.
## Dead stacks
| Stack | State | Size |
|---|---|---|
| `airflow` (`/home/sam/deployment/airflow/`) | 7 containers **exited 4 weeks ago**, `restart: no` | **24 GB of logs** (6,162 files) |
| `ai-resume/airflow` | separate compose, containers not running | small |
| `engram` container | created 2026-06-13, **never started** | image 36 MB |
| `trigger_dev`, `t3_stack_react`, `sams-home-network` | not running | — |
| `nixos_backup.tar.gz` in `/home/sam` | stale migration tarball | **6.6 GB** |
Nothing running depends on any of these. The `ai-resume` `knowledge-service` and `langgraph-service`
still run and do not need Airflow.
## Engram — why it is unused
It is running correctly and is simply **empty**:
```
$ curl -s http://localhost:7437/stats
{"total_sessions":0,"total_observations":0,"total_prompts":0,"projects":null}
```
- Live service: user systemd unit `engram.service` → `/home/sam/.local/bin/engram serve`, `:7437`
- Data: `/home/sam/.engram/engram.db` — 4 KB, tables `sessions` and `observations`, both with 0 rows
- CPU used in 4 weeks of uptime: **1.1 seconds**
- The Docker container of the same name was never started and can be removed
- **No pi extension, skill or MCP server references engram.** The client side that would write
sessions and observations was never wired up (or was removed)
- The `.27` → `.13:7437` SSH tunnel documented in [[Home Network Map Overview]] is not established
So engram is a working service with no clients. Either wire up a client, or stop the unit and
remove the container.
## Reduction candidates
Ordered by expected relief. **Nothing has been changed.**
| Candidate | Est. relief | Risk | Note |
|---|---|---|---|
| Stop `litellm` | ~526 MB swap | Low | Redundant with OmniRoute |
| Stop `headroom` when JEV is not in use | ~516 MB swap | Low | Needed only for `fast_jev` |
| Unlock / log out / disable the locked KDE session on `.13` | ~2.4 GB | Medium | Use `systemctl isolate multi-user` or switch `videoDrivers` once the GPU is replaced |
| Stop `voice_whisper` | ~877 MB | High | Breaks the voice pipeline; move to `.27` (has GPU + 41 GB free) or switch to `base.en` |
| Start Prefect + `photo-dashboard` on demand | ~640 MB | Low | Weekly batch job |
| Stop the media-tool containers (gimp, kdenlive, audacity, lmms) | ~170 MB + 4 selkies X servers | Low | Idle all day |
| Stop `n8n` | 286 MB | Medium | Depends on whether scheduled workflows matter |
| Remove dead Airflow stack + 24 GB logs | 24 GB disk | Low | Nothing depends on it |
| Remove the `engram` container | 36 MB + swap | Low | Never started |
| Delete `nixos_backup.tar.gz` | 6.6 GB disk | Low | Stale migration tarball |
| Repair or replace the GPU | — | — | Fixes the root cause of the KDE memory cost |
## What must stay on `.13`
| Service | Why |
|---|---|
| Pi-hole `:53` (replica of `.35`) | DNS for the LAN |
| OmniRoute `:20128/20129` | **Every pi agent on `.27`, `.13`, `.51` routes LLM traffic through it** |
| Mosquitto `:1883` | Voice pipeline and Home Assistant on `.30` |
| Borg backup target `/mnt/ubuntu_storage_3TB` | `.27` pushes here daily at 04:00 |
| Caddy `:8000` + `/var/www` | Resume and portfolio sites |
| Docker daemon | Everything above |
## Open questions
1. Is the KDE session on `.13` used interactively, or is `.13` headless in practice?
2. Are Paperclip and Paseo needed 24/7, or only on demand?
3. Is `litellm` still needed alongside OmniRoute?
4. Should the weekly Prefect photo pipeline move to `.27`?
5. Should `knowledge-service` / `langgraph-service` be de-duplicated (they run on `.13` **and** `.27`)?
6. Wire up engram, or retire it?
## Related
- [[Home Network Map Overview]]
- [[Docker Containers]]
- [[Filesystem Drive Map]]
- [[Paperclip Agent Harness Manager]]
- [[Paseo Pi GUI Tool]]
- [[Family Home Lab]]

View File

@@ -1,6 +1,6 @@
---
created: 2026-08-05 13:11
modified: 2026-08-26 12:20
modified: 2026-10-05 12:20
type: area
status: active
tags:
@@ -48,4 +48,17 @@ LAN-only. Source of truth: this repo (`sam/family_home_lab` on Gitea) →
- Code is in this repo and pushed to **Gitea**: `sam/family_home_lab` (SSH
`ssh://git@gitea.lab.audasmedia.com.au:2222/sam/family_home_lab.git`).
## Verified 2026-10-05
- `.13` also runs its **own Caddy on port 8000**, serving the static sites from `/var/www`:
`sam-developer`, `sam-devops`, `sam-iot-electronics`, `sam-pursuits`, `sprinklers`,
`sequence` and `chat` (see `/etc/caddy/caddy_config`). All return HTTP 200.
It does **not** serve ports 80/443 — the public entry point is the `.35` Caddy.
- These portfolio/resume sites are **static files**. They cost almost nothing to keep running
(Caddy itself is ~21 MB). The expensive part of `.13` is elsewhere — see
[[Resource Use — .13 CPU & RAM]].
- Live media-tool ports on `.13`: gimp `:8087`, video (kdenlive) `:8083`, audio (audacity) `:8084`,
lmms `:8085`. These containers are idle all day and are good candidates for start-on-demand.
- `wikijs.lab` and `lynx.lab` are hosted on `.35` (confirmed as running containers there), not `.13`. `photo-filter.lab` was not verified.
*Full inventory: [[Home Network Map Overview]] and repo `docs/websites.md`.*