Family Home Lab: portal, dsh (chat+plugins), transcriber, music/media tools, home dash

This commit is contained in:
2026-08-26 11:30:34 +10:00
commit 1110dbc978
62 changed files with 4237 additions and 0 deletions

8
docs/README.md Normal file
View File

@@ -0,0 +1,8 @@
# Docs
Future phase — document/library tools (e.g. paperless-ngx, BookStack, or an internal wiki).
- Category colour assigned: teal `#2a9d99`.
- No service ships from this folder yet; design decisions belong in plan.md/AGENT.md
when a tool is chosen.
- User permission flag (`can_docs`) already exists in the portal user model, default off.

60
docs/audio-plan.md Normal file
View File

@@ -0,0 +1,60 @@
# Audio & Music — Implementation Plan (Draft)
> For Sam's son (music student). Goal: give the family lab a remote music-production
> workspace (DAWs) and an **audio → sheet-music / MIDI transcription** pipeline.
> Installing extra software in containers is approved.
## 1. Feasibility of the proposed tools (verified)
| Tool | What it is | Remote-app container? | Fit |
|---|---|---|---|
| **Zrythm** | Open-source DAW (GPLv3), Pro-grade: mixed, automation, piano roll, module lanes | ❌ no public image **exists**. Installable into a LinuxServer **webtop** container via its official installer (zrythm.org `install.sh` / apt repo). | ✅ son's main DAW |
| **LMMS** | FL-Studio-style open-source DAW (beats / MIDI / virtual instruments) | ✅ `apt install lmms` inside a webtop container (Ubuntu). | ✅ secondary / beats |
| **Spotify Basic Pitch** | Audio→**MIDI** neural net (polyphonic, pitch bends). `pip install basic-pitch`. Best on a single instrument. | Runs as a **Celery task** (no UI). | ✅ fast, lightweight first-pass MIDI |
| **MuScriptor** (Kyutai + Mirelo) | State-of-the-art **multi-instrument** transcription → **MIDI + MusicXML + engraved PDF + guitar tabs** in one run. `muscriptor transcribe audio.wav --format sheets`. | CLI / heavy model (1.4B). | ✅ the "sheet music" engine you described |
| **MuseScore** | Open-source notation/engraver (MuScriptor uses it internally). | `apt install musescore`, also headless CLI. | ✅ fallback: MIDI→MusicXML→PDF locally |
## 2. Key caveats to decide on
- **MuScriptor licensing:** code is **MIT**, but model weights are **CC BY-NC 4.0** (non-commercial). That's fine for personal/family use — but not for anything commercial.
- **MuScriptor resources:** 1.4B-parameter model. Feasible on CPU but **slow**; a GPU makes it practical. We have no GPU confirmed on `.13` → treat MuScriptor as best-effort/beta, queue it in Celery, or throttle.
- **Basic Pitch scale:** fast on CPU, but "best on one instrument at a time". Good default for quick single-part transcription.
- **Webtop note:** the LinuxServer webtop terminal grants root inside the container. Safe on the trusted LAN behind Caddy; don't expose it to the public internet.
## 3. Proposed architecture
```
Family-music user (son/family)
│ console.lab.audasmedia.com.au (portal)
├── "Music (LMMS)" → lmms.lab.audasmedia.com.au (webtop:8086? ) remote DAW
├── "Zrythm" → zrythm.lab.audasmedia.com.au (webtop:8087? ) remote DAW
└── "Transcriber" → console form: pick audio (Garage) / upload
│ enqueue Celery task
▼
worker (existing family-home-lab-worker)
├─ Basic Pitch → MIDI (fast, single-instrument)
└─ MuScriptor → MIDI+MusicXML+PDF+tabs (beta, slow)
▼
writes results back to Garage bucket + shared-media
▼
Transcriber page lists/plays/downloads results
```
- Input & output live in Garage (`sam`/…/`shared-media` buckets), backed by Borg.
- Uses the **existing** RabbitMQ+Celery worker — no new queue.
- DAWs are separate webtop containers on **free host ports** with their own Caddy rows.
## 4. Blocker found on the DAWs (ports)
Originlab snippet used **:3001** for zrythm — but that port is **taken by langfuse**.
We'll use fresh free ports (e.g. **8085/8086/8087**) and correct Caddy/nearby URLs.
## 5. Suggested build order (each is independently useful)
1. **Transcriber (Basic Pitch)** — quick win: upload/pick audio → MIDI; show result in console. (Easy, CPU, no license.)
2. **LMMS remote DAW** — `apt install lmms` on a webtop container, free port + Caddy row.
3. **Zrythm remote DAW** — custom webtop image with Zrythm installed.
4. **MuScriptor sheet-music** — add the multi-instrument + MusicXML/PDF stage (needs HF login/license; beta if no GPU).
## 6. Open decisions (need your call)
- Which transcriber to prioritize: **Basic Pitch** (fast, MIDI) vs **MuScriptor** (full sheet music, heavier) — or pair them.
- LMMS and/or Zrythm both? (I suggest both — they cover different use cases.)
- Confirm we may use the webtop image which gives container-root to the terminal.

45
docs/dsh-plugin-plan.md Normal file
View File

@@ -0,0 +1,45 @@
# dsh plugins — plan (docs / summarize / web / images)
Goal: give each family dsh instance practical "plugin-style" capabilities:
handle documents, summarize web pages & docs, ingest images, output images,
and edit images from text. Decide how to deliver them on our lightweight
FastAPI+SSE dsh.
## Approach: extend our dsh (not adopt the official harness)
Our dsh is intentionally a small FastAPI+SSE chat (OmniRoute `auto/best-chat`).
The official DeepSeek Harness + its ~11k-plugin ecosystem is heavier and
prerelease. We keep our app and add plugin-style capabilities natively, phased.
Capabilities map:
| # | Capability | Mechanism | Phase |
|---|------------|-----------|-------|
| 1 | Summarize a web page | dsh fetches URL (httpx) → text → LLM summary. Add `/api/tool/web` or a "summarize URL" box. | 1 |
| 2 | Summarize a document | read a file from `/workspace` (txt/md; PDF via pypdf) → LLM summary. Add "Summarize file" button + `/api/tool/docs`. | 1 |
| 3 | Handle / attach docs | pass referenced workspace file contents into the prompt (file picker, cap size). | 1–2 |
| 4 | Ingest images (understand) | need a **vision** model on OmniRoute (e.g. qwen-vl / gpt-4o). Send image as data-URI in OpenAI vision content format. `DSH_LLM_VISION_MODEL`. | 2 |
| 5 | Output images (generate) | OmniRoute is LLM-only → wire an image API (OpenRouter image, or internal ComfyUI/SD). `/api/tool/image`. | 3 |
| 6 | Edit image from text | image-editing model (instruction-based) on the image API. | 3 |
## Phase 1 — text capabilities (no new infra)
- [ ] Web summarize: `POST /api/tool/web {url}` → httpx fetch → strip HTML → LLM summary (stream). UI: a "Paste URL to summarize" box.
- [ ] Docs summarize: `POST /api/tool/docs {path}` (relative to /workspace) → read txt/md (pypdf for PDFs) → LLM summary.
- [ ] Chat context: pick a workspace file → prepend its contents to the message (cap ~8k tokens).
## Phase 2 — image ingest (vision)
- [ ] Find/configure a vision model on OmniRoute (qwen2.5-vl, gpt-4o, or similar). Set `DSH_LLM_VISION_MODEL`.
- [ ] Image upload in chat.html (`<input type=file accept=image/*>` + base64) → send OpenAI vision `content` array with `image_url` data-URI.
- [ ] Handle docs as images (scan/photo) → describe/OCR via the vision model.
## Phase 3 — image output & edit
- [ ] Wire an image-generation API (OpenRouter `gpt-image-1`/`flux`, or internal ComfyUI/SD webui).
- [ ] `/api/tool/image` — text→image; stream/provide a URL or return a base64 image to display in chat.
- [ ] Image edit — instruction/ref edit endpoint (gpt-image edit or SD img2img) from an uploaded image + text.
- [ ] Save images into the user's S3 bucket (`image/…`) so they're kept.
## Open questions
- Vision + image models on OmniRoute: confirm availability/ids (check `/v1/models`).
- Kids' instances: should image output be gated (cost/appropriateness)? Probably gate image-gen to sam initially.
- Marker: keep everything streaming + iframe-friendly CSP.
## Sources
- docs/dsh-plugins.md — shortlist from github topic: WeKnora (docs→RAG), modlens (vision), open-design.

60
docs/dsh-plugins.md Normal file
View File

@@ -0,0 +1,60 @@
# DeepSeek Harness (dsh) — plugin shortlist
Pulled 2026-08-25 from https://github.com/topics/dsh-plugin (the topic feed is
noisy, so this is the curated subset relevant to the family home lab).
**Start here:** `awesome-dsh-plugin/awesome-dsh-plugin` — the curated index of
dsh plugins (categorized, maintained). Everything below is notable from the
topic feed itself.
## Knowledge / Docs / Summarize (matches "Docs, Links — summarize" todo)
- **Tencent/WeKnora** — open-source knowledge platform: turns raw documents
into a queryable RAG layer. Good fit for family docs/wiki summarisation.
- **volcengine/OpenViking** — self-evolving context/memory DB for agents
(unify memory + knowledge).
- **distilly** — distill "how they think" into reusable skills for any agent
(good for teaching the kids' assistants skills).
- **nocobase/nocobase** — AI + no-code platform (if we want a plugin that
builds data apps).
## Image / Media (matches "Image ingest/create" todo)
- **liustack/modlens** — vision plugin: lets dsh "see" images (ingest/understand).
- **freestylefly/awesome-gpt-image-2** — industrial prompt engine + 530+
reverse-engineered templates for image generation.
- **Nagi-ovo/voyager** — enhancement suite (Gemini/AI Studio/Claude/ChatGPT,
multimodal).
- **nexu-io/open-design** — design plugin (open-source alternative to Claude
Design; good for Sam's portfolio work).
## Memory
- **EverMind-AI/EverOS** — portable, local-first, Markdown-native memory layer
for agents.
- **MemTensor/MemOS** — ultra-persistent memory OS for LLM agents.
- **agentscope-ai/ReMe** — "Remember Me, Refine Me" memory management kit.
## UI / Desktop
- **anywhere-labs/dsh-desktop** — modern desktop client for the dsh plugin
ecosystem ("everything is a plugin, the desktop is a plugin").
- **zhu1090093659/dsh-web** — dsh Web plugin aggregator pack.
- **crafter-station/petdex** — public gallery of animated pets for agents
(fun for the kids' instances).
## Dev / Agent tooling (power users / Sam)
- **esengine/DeepSeek-Reasonix** — DeepSeek-native coding agent for the
terminal.
- **ruvnet/ruflo** — original agent "meta-harness" (multi-player swarms —
advanced).
- **tt-a1i/archify** — agent skill for architecture/workflow/sequence
diagrams.
- **yjh051108/dsh-routing-suite** — injector + router-standard kit for dsh.
- **xiaobright/dsh-anchored-standard** — two-phase dsh preset bootstrap.
## Notes
- Topic feed total is ~11k repos; most are non-dsh / gamed. Use the
awesome-dsh-plugin index before installing anything.
- Our own dsh is a lightweight FastAPI+SSE chat; plugins from the ecosystem
target the official DeepSeek Harness client — wiring them in means either
adopting the official harness or adapting individual ideas into our app.
- todo.txt already tracks: "DeepSeek Plugins: Docs/Links summarize" and
"Image ingest/create" — pick candidates from Knowledge/Docs and Image/Media
above.

55
docs/websites.md Normal file
View File

@@ -0,0 +1,55 @@
# Family Home Lab — website / service inventory
> For another AI session: the services behind the home lab and their addresses.
> All public *.lab.audasmedia.com.au are reachable off-LAN (own logins except
> the no-login tools that sit behind Caddy basic-auth). *.home.lab are LAN-only.
## Public — *.lab.audasmedia.com.au
console.lab # Family Home Lab portal (login)
photo.lab / video.lab / audio.lab / lmms.lab # media editor tools (basic-auth)
dsh-sam.lab dsh-jo.lab dsh-harry.lab dsh-finn.lab # DeepSeek chat (basic-auth)
chat.lab # family group chat
pb.chat.lab # chat admin (admin-only, basic-auth)
wikijs.lab lynx.lab sequence.lab # portfolio / resume
sam-developer.lab sam-devops.lab sam-iot-electronics.lab sam-pursuits.lab
omniroute.lab # AI router dashboard (basic-auth)
photo-filter.lab # photo utility (basic-auth)
silverbullet.lab hedgedoc.lab flatnotes.lab vikunja.lab bookstack.lab dokuwiki.lab trillium.lab # notes/docs
affine.lab penpot.lab # design / docs
nocodb.lab homebox.lab toora.lab linkace.lab # data / apps
gitea.lab # git
apprise.lab ntfy.lab # notifications
immich.lab jellyfin.lab jellyseerr.lab audiobookshelf.lab # media
nextcloud.lab paperless.lab pingvin.lab filebrowser.lab # storage / files
shopping.lab owl.lab # household
grafana.lab uptimekuma.lab worldmonitor.lab # monitoring
n8n.lab prefect.lab openweb.lab # automation / AI
firefly.lab # personal finance
vaultwarden.lab # password manager
sonarr.lab radarr.lab readarr.lab lidarr.lab sabnzbd.lab qbittorrent.lab spotweb.lab headphones.lab # media download stack
freshrss.lab ghost.lab kanboard.lab # rss / blog / tasks
proxmox.lab proxmox_backup.lab # virtualization (own login)
supabase.lab # dev (LangChain)
## Internal — *.home.lab (LAN only)
proxmox.home proxmox_backup.home portainer.home # admin
homeassistant.home # home automation
grafana.home influxdb.home librenms.home watchyourlan.home phpipam.home
dozzle.home uptimekuma.home kopia.home restic.home # infra / monitoring / backup
mopidy.home snapcast.home # audio
archon.home langfuse.home supabase.home openweb.home # AI / dev
pihole.home pihole2.home # DNS
sams.home # home file server
games.home # retro games (emulatorjs)
react-server.home laravel-server.home t3_stack_react.home sams_home_network.home # dev demos
## Other public domains
where-woof.com admin.where-woof.com # family location
homeassistant.lab.quickweb.com.au # HA mirror
## auth model
- No-login tools behind Caddy basic-auth: photo/video/audio/lmms, dsh-*, omniroute,
photo-filter, dozzle, sams, pb.chat, ntfy, plus apps left on basic-auth (toora,
homebox, phpipam, watchyourlan, librenms, kopia, backrest, filebrowser, ...).
- Apps with own login (basic-auth REMOVED so their mobile apps/APIs work): homeassistant,
nextcloud, immich, audiobookshelf, vikunja, proxmox(+backup), portainer, grafana,
uptimekuma, supabase.