Files
obsidian-vault/300 areas/360 Dev-Ops Network Computers/Resource Use — .13 CPU & RAM.md
Sam Rolfe 7ed646ac0a Correct .13/.27 notes from live 2026-10-05 inspection
- Home Network Map: .13 is memory-bound (swap 81%); KDE Plasma not Niri;
  GPU dead with no driver bound; Open WebUI->Outline on :3000; full .13
  container table; Pi-hole primary .35 (confirmed from nebulasync config);
  .27 airflow row removed
- Filesystem Drive Map: .13 device letters corrected (root sda2, storage
  sdb2, data sdc2, ubuntu_storage sdd1); add sda3 swap; Maxtor not mounted
- Docker Containers: garage v2.1.0; paseo is native not Docker; engram
  container dead + native service empty; airflow dead with 24GB logs;
  add missing containers and native/user service lists
- Websites on .13: add :8000 Caddy detail and media-tool ports
- New note: Resource Use - .13 CPU & RAM
2026-10-05 14:48:08 +11:00

167 lines
7.8 KiB
Markdown

---
created: 2026-10-05
modified: 2026-10-05
type: area
status: active
tags:
- dev-ops
- resource-usage
- machines
aliases: [resource-use-13]
---
# Resource Use — .13 CPU & RAM
> Live-verified **2026-10-05**. Read-only investigation — no configuration was changed.
> Full inventory: `~/chats/sys_config/resource_use_cpu_ram_13/SYSTEMS-INVENTORY.md`
## Headline
`.13` is **not CPU-bound. It is memory-bound.**
| Metric | Value |
|---|---|
| CPU load average (1/5/15 min) | **0.27 / 0.42 / 0.37** on 12 threads — the CPU is nearly idle |
| RAM | 6.8 GiB used of 15 GiB · 9.0 GiB buff/cache · 8.7 GiB available |
| **Swap** | **7.1 GiB used of 8.8 GiB — 81 % full** |
| Uptime | 28 days |
| Root disk | 459 GB of 907 GB (54 %) |
The swap partition is the constraint. A full swap explains the stalls. It also wears the SSD.
## Why the swap is full
Linux swapped out the *coldest* anonymous pages of long-running processes. Three groups dominate:
| Group | Combined RAM + swap | Detail |
|---|---|---|
| **KDE Plasma desktop session** | ~2.4 GB | ~30 processes: plasmashell (301 MB RSS + 870 MB swap), kscreenlocker (117+128), kwin_wayland (59+141), kdeconnectd (54+178), xdg-desktop-portal-kde (11+118) |
| **Docker containers** | ~3.5 GB RSS | `voice_whisper` 877 MB, `omniroute` 827 MB, `prefect-server` 411 MB, `n8n` 286 MB, `pihole` 205 MB, `outline` 187 MB |
| **Idle LLM gateways** | ~1.05 GB swap | `litellm` (~4 MB resident, **526 MB swapped**) and `headroom` (~18 MB resident, **516 MB swapped**) |
### Why the Plasma session is so large
Two reasons, and they reinforce each other:
1. **There is no working GPU driver.** The only DRM device is `/dev/dri/card0` with `DRIVER=simple-framebuffer`. `/proc/driver/nvidia` does not exist. Plasma and KWin fall back to **software rasterisation** (llvmpipe/swrast — 33 mesa mappings in `plasmashell`). Software GL keeps much larger CPU-side buffers than a GPU path does.
> `hardware.nvidia.open = true` in `configuration.nix` can never work on the installed Kepler card. Open kernel modules need Turing or newer. `hardware.nvidia.open` must be `false` for the legacy driver branch.
2. **28 days of uptime with no restart.** Plasma and KDE Connect accumulate memory over weeks. `kdeconnectd` alone holds 178 MB of swap with only 54 MB resident.
Also worth noting: **`kscreenlocker_greet` has been running for 19 days.** A monitor is attached (`card0-Unknown-1: connected`) and the session on `tty2` has been logged in since 2026-09-06, but the screen has been locked and untouched for ~19 days. That process costs 117 MB resident + 128 MB swap for nothing.
### Why `voice_whisper` is 877 MB
It is a Flask API that loads a Faster-Whisper model **at import time and holds it forever**:
```python
model_size = "small.en"
model = WhisperModel(model_size, device="cpu", compute_type="int8")
```
- Model: `Systran/faster-whisper-small.en`, cached in the image (image size 1.01 GB)
- Process: `VmRSS 877 MB`, `VmSwap 88 MB`, `VmSize 2.4 GB`
- Device: **CPU** — no GPU acceleration is possible while the GPU is dead
- Port 5000, no memory limit set (unlike `voice_bridge`, which is capped at 500 MB)
The model is resident 24/7 for a voice assistant that is used occasionally.
### What "swapped" means for `litellm` and `headroom`
`VmSwap` is the amount of a process's memory that the kernel has written to the swap partition.
**Swapped memory is not using RAM right now** — so these two are not consuming RAM.
They are consuming *swap space*, and that is still the scarce resource on `.13`.
| Service | VmRSS (in RAM) | VmSwap (on disk) |
|---|---|---|
| `litellm` | 4 MB | **526 MB** |
| `headroom` | 18 MB | **516 MB** |
Both have been idle for weeks, so the kernel paged almost all of them out. Together they occupy
~1.05 GB of an 8.8 GB swap partition — about 12 % of it — while doing no work.
`litellm` is a second LLM gateway alongside OmniRoute and is functionally redundant.
`headroom` is the **JEV** proxy (`--proxy-extension fast_jev`, forwarding to OmniRoute on `:20129`),
so it is only needed when JEV-based determination is in use. Stop them and the kernel reclaims the swap.
## Dead stacks
| Stack | State | Size |
|---|---|---|
| `airflow` (`/home/sam/deployment/airflow/`) | 7 containers **exited 4 weeks ago**, `restart: no` | **24 GB of logs** (6,162 files) |
| `ai-resume/airflow` | separate compose, containers not running | small |
| `engram` container | created 2026-06-13, **never started** | image 36 MB |
| `trigger_dev`, `t3_stack_react`, `sams-home-network` | not running | — |
| `nixos_backup.tar.gz` in `/home/sam` | stale migration tarball | **6.6 GB** |
Nothing running depends on any of these. The `ai-resume` `knowledge-service` and `langgraph-service`
still run and do not need Airflow.
## Engram — why it is unused
It is running correctly and is simply **empty**:
```
$ curl -s http://localhost:7437/stats
{"total_sessions":0,"total_observations":0,"total_prompts":0,"projects":null}
```
- Live service: user systemd unit `engram.service` → `/home/sam/.local/bin/engram serve`, `:7437`
- Data: `/home/sam/.engram/engram.db` — 4 KB, tables `sessions` and `observations`, both with 0 rows
- CPU used in 4 weeks of uptime: **1.1 seconds**
- The Docker container of the same name was never started and can be removed
- **No pi extension, skill or MCP server references engram.** The client side that would write
sessions and observations was never wired up (or was removed)
- The `.27` → `.13:7437` SSH tunnel documented in [[Home Network Map Overview]] is not established
So engram is a working service with no clients. Either wire up a client, or stop the unit and
remove the container.
## Reduction candidates
Ordered by expected relief. **Nothing has been changed.**
| Candidate | Est. relief | Risk | Note |
|---|---|---|---|
| Stop `litellm` | ~526 MB swap | Low | Redundant with OmniRoute |
| Stop `headroom` when JEV is not in use | ~516 MB swap | Low | Needed only for `fast_jev` |
| Unlock / log out / disable the locked KDE session on `.13` | ~2.4 GB | Medium | Use `systemctl isolate multi-user` or switch `videoDrivers` once the GPU is replaced |
| Stop `voice_whisper` | ~877 MB | High | Breaks the voice pipeline; move to `.27` (has GPU + 41 GB free) or switch to `base.en` |
| Start Prefect + `photo-dashboard` on demand | ~640 MB | Low | Weekly batch job |
| Stop the media-tool containers (gimp, kdenlive, audacity, lmms) | ~170 MB + 4 selkies X servers | Low | Idle all day |
| Stop `n8n` | 286 MB | Medium | Depends on whether scheduled workflows matter |
| Remove dead Airflow stack + 24 GB logs | 24 GB disk | Low | Nothing depends on it |
| Remove the `engram` container | 36 MB + swap | Low | Never started |
| Delete `nixos_backup.tar.gz` | 6.6 GB disk | Low | Stale migration tarball |
| Repair or replace the GPU | — | — | Fixes the root cause of the KDE memory cost |
## What must stay on `.13`
| Service | Why |
|---|---|
| Pi-hole `:53` (replica of `.35`) | DNS for the LAN |
| OmniRoute `:20128/20129` | **Every pi agent on `.27`, `.13`, `.51` routes LLM traffic through it** |
| Mosquitto `:1883` | Voice pipeline and Home Assistant on `.30` |
| Borg backup target `/mnt/ubuntu_storage_3TB` | `.27` pushes here daily at 04:00 |
| Caddy `:8000` + `/var/www` | Resume and portfolio sites |
| Docker daemon | Everything above |
## Open questions
1. Is the KDE session on `.13` used interactively, or is `.13` headless in practice?
2. Are Paperclip and Paseo needed 24/7, or only on demand?
3. Is `litellm` still needed alongside OmniRoute?
4. Should the weekly Prefect photo pipeline move to `.27`?
5. Should `knowledge-service` / `langgraph-service` be de-duplicated (they run on `.13` **and** `.27`)?
6. Wire up engram, or retire it?
## Related
- [[Home Network Map Overview]]
- [[Docker Containers]]
- [[Filesystem Drive Map]]
- [[Paperclip Agent Harness Manager]]
- [[Paseo Pi GUI Tool]]
- [[Family Home Lab]]