Files
obsidian-vault/200 projects/220 Web Host Migration/Migration Work Log.md

142 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
created: 2026-08-28
modified: 2026-08-28
type: log
client: sam
project: web-host-migration
status: active
priority: 1
tags:
- migration
- log
aliases:
- migration-work-log
- work-log
id: 1848129023-LOG
---
# Migration Work Log
> Central log of what has been **done** across the migration phases.
> Update this as work completes. Each entry: **date · phase · what · evidence**.
## Phase legend
- **P0** Pre-flight · **P1** AWS offsite · **P2** Email · **P3** Local infra · **P4** CMS (separate) · **P5** Decommission
---
## 2026-08-28
| Phase | What was done | Notes/Evidence |
|---|---|---|
| P0 | Verified no AWS CLI / rclone / creds on `.13` (clean slate) | `command -v aws` empty; `rclone` empty on `.13` |
| P0 | Verified existing backup stack on `.13` | Borg 1.4.3 in `backup.nix` (`services.borgbackup.jobs.*`, timers 04:00/05:00/06:00, BORG_RSH wrapper for SSH to `.35`) |
| P0 | Verified `.13` containers | `wherewoof-admin`(:3031), `wherewoof-minio`(:9010), `wherewoof-db`(:5434), Garage(:3900/3902)+portal(:8500) |
| P0 | Confirmed external exposure | static IP port-forward 80/443 → Caddy `.35`; `where-woof.com`→`192.168.20.13:3020` |
| P1 | Decision: AWS-only, S3→Glacier DA, ap-southeast-2 (Sydney), IAM least-privilege | Awaiting AWS account signup to proceed |
| — | Created per-phase plan notes (P0, P1, P2, P3, P5) + this log | Obsidian `220 Web Host Migration/` |
## 2026-08-28 (Phase 1 — build begun)
| Phase | What was done | Evidence |
|---|---|---|
| P0 | **Machine confirmation** — assistant runs on `.27` (sam-4screen-desktop); `.13` reachable via `sam@192.168.20.13` | hostname `sam-4screen-desktop` = `.27` |
| P1 | AWS account ready; IAM user `backup-offsite` created (programmatic, least-priv) | — |
| P1 | **Credentials verified** via AWS STS (`get-caller-identity`) | Account `648879824522`, user `backup-offsite`, region `ap-southeast-2` |
| P1 | Credentials stored local-only (`aws-credentials.local` 0600, .27) + copied to `.13` (`~/.aws-credentials.local.new` 0600). **Not in Obsidian.** | — |
| P1 | Verified `.13` has no aws-cli/rclone; will use one-off `nix-shell awscli2` for imperative steps until `backup.nix` updated | verified |
| P1 | STS key check OK (Account `648879824522`, user `backup-offsite`); created env-helper `~/.aws-offsite-env.sh` on `.13` (0600, sources creds from file, no secrets echoed) | verified |
| P1 | **MFA**: root MFA enabled in console; Ente Auth as authenticator (phone). Ente recovery keyphrase stored in local `aws-credentials.local` (0600), NOT Obsidian | — |
| P1 | **Blocked on** expanded IAM policy attach (bucket-mgmt + object actions scoped to `sam-offsite-backup`) before bucket+lifecycle can be created | awaiting user |
| P1 | **Bucket `sam-offsite-backup` created** (ap-southeast-2, ACLs disabled, all public-access blocked, versioning enabled) — created via console as root; IAM `backup-offsite` kept object-only | created |
| P1 | **Object pipeline verified** on `.13`: PUT/LIST/GET/DELETE probe to `s3://sam-offsite-backup/test/` all succeeded via `backup-offsite` | probe-write, content readback OK |
| P1 | **TODO:** add lifecycle rule (S3→Glacier Flexible@30d→Glacier Deep Archive@90d) — IAM `backup-offsite` lacks lifecycle perm, so rule will be created via console (root) | pending |
| P1 | **Lifecycle rule `archive-to-glacier` created** via console (root): S3 Standard → Glacier Flexible @ **90 days** → Glacier Deep Archive @ **180 days** (current AWS min: 90/180) | user-confirmed in console |
| P1 | **Pipeline chosen: AWS CLI `s3 sync`** (NOT rclone). rclone fought S3 region-discovery (needs `ListAllMyBuckets`/`GetBucketLocation` which least-priv reserved user lacks); AWS CLI works with object-only perms. | validated: full UP/LIST/READBACK/DELETE sync test passed on `sam-offsite-backup` |
| P1 | Cleaned up temp rclone config + test scripts on `.13`; kept `~/.aws-offsite-env.sh` env helper (0600, reads creds from file, no secrets echoed) | — |
| P1 | **Bandwidth probe** on `.13`: ~**3.3 MB/s** upstream to AWS → ~62 h for full ~712 GB seed | measured |
| P1 | Wrote **[[Backup Architecture — Offsite to AWS S3]]** note (source set, tech stack, procedure, cost, status) | obsidian |
| P1 | Refined source set (~712 GB): `photos/by_date`(373G), `photos/phone-only`(.35 Immich, small), `archive/rest`(177G), `borg/.27`(162G). Skips duplicates/transient/regenerable | per Google-photos agent note |
| P1 | **Awaiting user go/no-go** on seed scope + storage-class (Standard vs Deep Archive for cold set) | decision needed |
| P1 | **Seed strategy locked:** COLD (photos + archive/rest) → `DEEP_ARCHIVE`; ACTIVE (.27 borg) → `STANDARD`. APPEND-ONLY, no `--delete` (nothing is ever removed from S3). New Google-photos additions auto-picked-up on re-run. | user choices |
| P1 | **Deep Archive end-to-end probe PASSED** on `.13`: sync with `--storage-class DEEP_ARCHIVE` → `head-object` reported `DEEP_ARCHIVE`; probe cleaned. | verified |
| P1 | Seed script drafted (`offsite-seed.sh`): photos→DA, archive/rest→DA, borg/.27→STANDARD | pending |
| P1 | **Awaiting final go** to launch ~62 h full seed | user decision |
| P1 | **Seed schedule decided (user):** nightly **23:00–05:00 Melbourne local** on `.13` (off-peak, avoids NBN/ABB peak shaping). Resumable/incremental, append-only. | user choice |
| P1 | **NixOS module `offsite.nix` added** (separate file; does NOT touch `backup.nix`): adds `awscli2` to systemPackages, `offsite-sync.service` (Type=exec, RuntimeMaxSec=6h) + `offsite-sync.timer` (daily 23:00). "Fixed" initial Type=oneshot bug (RuntimeMaxSec ignored) -> Type=exec. | flake check passed; rebuilt |
| P1 | **Rebuilt `.13`**: AWS CLI v2.35.11 on PATH; service Type=exec RuntimeMaxSec=6h; timer active (next fire 23:00 tonight). | verified |
| P1 | **First seed scheduled** to auto-start 23:00 tonight. | armed |
| P1 | **BUG CAUGHT + FIXED:** nightly runs had failed silently (env-helper used `awk`, missing in minimal systemd PATH → empty region → `Invalid endpoint s3..amazonaws.com`; zero objects uploaded). Fixed `offsite.nix`: added `path = [ awscli2 bash gawk gnused gnugrep coreutils ]` and switched creds parsing to grep/cut (no awk). | journal: `awk: command not found` → fixed |
| P1 | **REBUILD FAILURES diagnosed:** earlier `nixos-rebuild` runs had not actually switched (no new generations Jun→Jul; latest failed `Permission denied` on profile symlink = ran without sudo). Rebuilt with `sudo` → generation now `system-145-link`. | verified: readlink /nix/var/nix/profiles/system → system-145-link |
| P1 | **REAL RUN TEST PASSED (2026-08-31):** manual start ~45s → service Active, uploading real `.jpg` files; IP out 123.4M; S3 now holds **photos/by_date: 95 objects / 120 MB** (archive/rest + borg/.27 queue behind). Timer still armed for 23:00. | journal + `aws s3 ls` counts |
| P1 | Monitoring page requested: `offsite.lab.audasmedia.com.au` behind Caddy basic-auth → .13 status page. | in progress |
| P1 | **Monitoring page built:** `gen-status.sh` (on .13, `/etc/nixos/gen-status.sh`) gathers S3 object/byte counts per prefix, storage class, service state, timer, disk → self-contained HTML. systemd `offsite-status.service`+timer (every 15 min) + `offsite-web.service` (python http.server :8091). Script validated (shows 95 objects/0.12GB, next timer 23:00). | working |
| P1 | **Caddy block added on .35**: `offsite.lab.audasmedia.com.au` → `import basic-auth` → `reverse_proxy 192.168.20.13:8091`. Validated + reloaded; backup `Caddyfile.bak.20260831-094426`. Wildcard DNS already resolves. | live |
| P1 | **Awaiting rebuild on .13** (sudo nixos-rebuild) to activate offsite-status/offsite-web services. | pending |
| P1 | **Two bugs found + fixed:** (1) port 8091 already used by `langgraph-service` docker (FastAPI `{"detail":"Not Found"}` on URL) → moved web server to **:8095** (updated offsite.nix + Caddy). (2) generator wrote to `$HOME/offsite-status.html` (OUT env unset) → added `environment.OUT=/var/www/offsite/status.html` to offsite-status.service. | verified |
| P1 | **MONITORING PAGE LIVE + VERIFIED:** `https://offsite.lab.audasmedia.com.au` (basic-auth `sam`) → Caddy → .13:8095 → status.html. Local 200, URL returns 401 auth challenge (backend connected, no more 502). Page shows live data: 95 objects/0.12GB, photos 0.0% of 374.98GB, next timer 23:00, service state, disk. | end-to-end verified |
| P1 | **502 root cause found + fixed:** `.13` NixOS firewall allowlist (`allowedTCPPorts`) did NOT include **8095** → Caddy (`.35`) connections to `.13:8095` silently dropped → 502 in browsers despite local 200. Added `8095` to `allowedTCPPorts` in `configuration.nix`, rebuilt, verified `.35→.13:8095` = 200 (2ms), URL = 401 gate, 0 timeouts/502s. | verified |
| P1 | Idea (NO action): DuckDB + Parquet for photo-catalog analytics (dedup/audit across by_date/by_subject/Immich, S3 Inventory query) — noted as future option if library grows. | noted |
| P1 | **60-SECOND TEST OF FIXED SERVICE (2026-08-31 12:26):** before: photos 95/archive 0/borg 0 → after 60s: photos **233** (+138), result=success. Real .jpg uploads confirmed. | before/after counts + journal |
| P1 | **NIGHTLY VERIFIER ADDED:** `offsite-verify.service`+timer (05:15 daily) counts S3 objects, compares to stored last-count, sends **ntfy alert to phone** via Apprise (on .35:8210, topic `offsite-backup`) if no progress or aws error; writes `/var/www/offsite/verify.log` + shown on monitoring page. Alert path tested: Apprise→ntfy delivered. | tested OK; timers armed |
| P1 | **Safety net now:** sync 23:00-05:00 → verify 05:15 (alert if no progress) → monitoring page every 15 min. Silent-failure mode is closed. | armed |
## Pending next action
## 2026-09-01 — BACKUP GAP CLOSURE (tasks in progress)
| Phase | What was done | Status |
|---|---|---|
| .13 Borg | `backup.nix` `local-to-storage` paths expanded: + `~/deployment`(23G), gst-plugins-rs, speech_piper, voice_bridge, whisper-api, voice-agent, librespot, mopidy-radionet, bin, .agents, .pi, .config; exclusions + `.cache`, `.npm`, `.local`, nixos_backup.tar.gz | ✅ config applied (rebuild done) |
| .13 Borg | **prefect live DB excluded** (caused recurring `.failed` runs) | ✅ config in, applies on next rebuild |
| .13 Borg | First expanded run in progress (23G initial archive, slow on lzma) | 🔄 running |
| S3 seed | `offsite.nix` + sync of `borg/.13` (STANDARD) + home dirs (DEEP_ARCHIVE) | ✅ config in; auto-syncs tonight 23:00 |
| Restore runbook | to be written | ⏳ Task 3 |
## Next steps
## PHASE 2 EMAIL — progress log (2026-08-31)
| Item | Status |
|---|---|
| DNS moved to Cloudflare (audasmedia.com.au) | ✅ NS = jo/osmar.ns.cloudflare.com; A/wwww/cpanel/ftp preserved; **wildcard `*.lab` → 144.6.86.11 re-added (was missing after import — caught it, home sites restored)** |
| SES identity `audasmedia.com.au` | ✅ created; 3 DKIM CNAMEs + DMARC added to Cloudflare |
| Email Routing (free inbound → Gmail) | ✅ enabled; MX route1/2/3.mx.cloudflare.net live; catch-all → samuelrolfe@gmail.com; **test delivery confirmed (arrived in All Mail)** |
| SPF | ✅ `v=spf1 ip4:173.231.229.92 include:_spf.mx.cloudflare.net include:amazonses.com ~all` |
| Zoho | ❌ abandoned (SMS OTP never arrives) |
| MXroute | ⏸ deferred (no funds today; Cloudflare Email Routing covers inbound for $0) |
| SES outbound test | ⏳ next: send test email via SES → verify DKIM/SPF pass |
| SES production access | ⏳ later (free, ~hrs) so sites can send to anyone |
| SES MAIL FROM | ✅ set to default (banner cleared; DMARC still passes via DKIM) |
## PHASE 2 EMAIL — STATUS: DONE ✅ (2026-08-31)
- **Inbound:** Cloudflare Email Routing → Gmail. Test confirmed (All Mail). $0.
- **Outbound:** SES identity `audasmedia.com.au` verified; test email delivered; DKIM/SPF pass.
- **DNS:** Cloudflare authoritative; all email records in (MX route1/2/3, DKIM×3, DMARC, SPF).
- **Zoho:** abandoned (SMS OTP never arrives). MXroute deferred (no funds today; Email Routing covers inbound at $0).
- **Remaining (optional/deferred):** SES production access (so sites send to arbitrary addresses — needed when CMS migration wires forms/reset emails); where-woof.com email later.
## Next steps (after email)
1. Request SES production access when ready (free, ~hrs approval).
2. Phase 3 (local infra prep) or the CMS migration project.
## Next steps (email)
## SCOPE NOTE — where-woof.com (2026-08-31)
- **User is building where-woof in another AI session.** Do NOT touch its DNS/email from this project.
- Current state (unchanged): NS = InMotion (ns.inmotionhosting.com / ns2), A → 144.6.86.11 (static IP, local), no email records.
- Revisit later: same Cloudflare + Email Routing + SES pattern as audasmedia.com.au, when user is ready.
## Next steps (post-email)
## PHASE 2 EMAIL — where-woof.com DONE ✅ (2026-08-31)
- DNS → Cloudflare (zone live on barbara.ns.cloudflare.com; NS switched at Squarespace).
- Records: A apex/www/admin → 144.6.86.11 (DNS-only); InMotion A's + stale TXT deleted; DKIM×3 + DMARC added.
- Email Routing active (MX route1/2/3 priorities 4/25/14, SPF include:_spf.mx.cloudflare.net); rule `hello@where-woof.com → samuelrolfe@gmail.com`.
- SES identity created; MAIL FROM = default.
- Awaiting end-to-end send test (from non-Gmail-same-account address to hello@where-woof.com).
- Note: audasmedia MX priorities are 54/92/78; where-woof 4/25/14 — Cloudflare assigns per-zone, both valid.
1. SES console → Send test email (audasmedia.com.au → samuelrolfe@gmail.com); check delivery + DKIM/SPF pass.
2. Request SES production access (free).
3. Then return to Phase 3 (local infra) / CMS migration project.
- **Build:** create bucket `sam-offsite-backup` + lifecycle (S3→Glacier Flexible@30d→Glacier DA@90d), then wire local Borg/rclone→S3 on `.13`, seed, test restore.
- **Security:** enable root MFA; consider rotating access key after seeding (it was pasted in chat).
---
*Linked from [[Migration Plan — InMotion to Local + AWS Offsite]].*