Compare commits

...

4 Commits

Author SHA1 Message Date
d278d18bb3 Correct system-architect specs from live 2026-10-05 inspection
- .13: KDE Plasma 6 (not Niri); RAM 15GB; disk letters corrected
  (root is /dev/sda2, storage sdb2, data sdc2, ubuntu_storage sdd1);
  add sda3 swap (81% full)
- .13 GPU: GTX 760 dead, no driver bound, software rendering
- .27: 4 cores/8 threads, Quadro P620 + HD 630, add swap + NTFS disk
- add Resource Pressure on .13 section and must-stay-on list
- note dead stacks (airflow 24GB logs, engram container) and caveats
2026-10-05 14:47:54 +11:00
8111abea38 image-maker: don't hardcode a per-machine secrets path
The previous version pointed at /home/sam/.secrets, which exists on .13 but
NOT on .27 — so the subagent would work on one machine and fail on another.
It had the same bug before in reverse (pointed at environment.d/10-secrets.conf,
which was missing the key on .13).

Now: require the OPENROUTER_API_KEY environment variable, document
~/.config/environment.d/10-secrets.conf as the canonical source (the only one
systemd user services read, which is what Paperclip workers get), and fall back
to ~/.secrets only if it exists.
2026-09-27 18:43:45 +10:00
3a8a6e98b1 fix(cost): per-model pricing so pi/Paperclip report real spend
The omni provider's models.json entries carried no 'cost' field, so pi's
calculateCost() returned 0 for every run and Paperclip cost events/budgets
all read $0.00 despite real spend.

- scripts/omni-pricing.mjs: pulls authoritative prices (OmniRoute /api/pricing
  for deepseek/*, OpenRouter /api/v1/models for openrouter/*) and injects
  { cost: { input, output, cacheRead, cacheWrite } } USD per 1M tokens.
- Priced: deepseek-v4-flash, deepseek-v4-pro, gemini-3.1-flash-lite, gemini-2.5-flash.
- Must be re-run after every /omni sync, which rewrites models.json wholesale.
2026-09-27 18:37:24 +10:00
3f36f599ce chore(models.json): refresh from /omni sync (476 omni models) 2026-09-27 18:37:14 +10:00
4 changed files with 285 additions and 39 deletions

View File

@@ -68,13 +68,22 @@ The image comes back as base64 in `choices[0].message.images[0].image_url.url`.
## API Key
`OPENROUTER_API_KEY` lives in **`/home/sam/.secrets`** (a plain `KEY=value` file). Load it with:
`OPENROUTER_API_KEY` must be present as an **environment variable**. Do not assume a fixed file path — the canonical location differs per machine.
**Canonical source on all machines:** `~/.config/environment.d/10-secrets.conf` (plain `KEY=value`, no `export`). This is the only location that **systemd user services read**, which is what matters when this subagent runs under Paperclip rather than in an interactive shell.
Check availability first:
```bash
set -a; . /home/sam/.secrets; set +a
echo "${OPENROUTER_API_KEY:+key loaded}"
```
It is NOT in `~/.config/environment.d/10-secrets.conf` — older versions of this file claimed that and the call failed.
If it is empty, source it for the current shell:
```bash
set -a; . ~/.config/environment.d/10-secrets.conf; set +a
```
**Do not hardcode other paths.** `.13` additionally keeps keys in `~/.secrets` (an `export`-style file), but `.27` has no `~/.secrets` at all — anything that sources that path fails there. Prefer the environment variable; fall back to `~/.secrets` only if the variable is unset *and* the file exists.
Write detailed, specific prompts. Save images to the user's current working directory or a specified path. Tell the user where you saved the file.

View File

@@ -63,6 +63,17 @@
"api": "omni-prompt-tools",
"apiKey": "sk-fff5c0d55c0bc0bf-b58ed0-76da5716",
"models": [
{
"id": "default-opencode-go-ds-flash",
"name": "Default Opencode Go Ds Flash",
"api": "omni-prompt-tools",
"input": [
"text"
],
"contextWindow": 1000000,
"maxTokens": 384000,
"reasoning": true
},
{
"id": "aug/claude-haiku-4.5",
"name": "Claude Haiku 4.5",
@@ -642,17 +653,6 @@
"maxTokens": 8192,
"reasoning": true
},
{
"id": "default-opencode-go-ds-flash",
"name": "Default Opencode Go Ds Flash",
"api": "omni-prompt-tools",
"input": [
"text"
],
"contextWindow": 1000000,
"maxTokens": 384000,
"reasoning": true
},
{
"id": "deepseek/deepseek-v4-flash",
"name": "V4 Flash",
@@ -662,7 +662,13 @@
],
"contextWindow": 1000000,
"maxTokens": 384000,
"reasoning": true
"reasoning": true,
"cost": {
"input": 0.07,
"output": 0.28,
"cacheRead": 0.014,
"cacheWrite": 0.07
}
},
{
"id": "deepseek/deepseek-v4-pro",
@@ -673,7 +679,13 @@
],
"contextWindow": 1000000,
"maxTokens": 384000,
"reasoning": true
"reasoning": true,
"cost": {
"input": 0.435,
"output": 0.87,
"cacheRead": 0.0036,
"cacheWrite": 0.435
}
},
{
"id": "ds/deepseek-v4-flash",
@@ -1975,7 +1987,13 @@
],
"contextWindow": 1048576,
"maxTokens": 65535,
"reasoning": true
"reasoning": true,
"cost": {
"input": 0.3,
"output": 2.5,
"cacheRead": 0.03,
"cacheWrite": 0.083333
}
},
{
"id": "openrouter/google/gemini-2.5-flash-image",
@@ -2083,7 +2101,13 @@
],
"contextWindow": 1048576,
"maxTokens": 65536,
"reasoning": true
"reasoning": true,
"cost": {
"input": 0.25,
"output": 1.5,
"cacheRead": 0.025,
"cacheWrite": 0.083333
}
},
{
"id": "openrouter/google/gemini-3.1-flash-lite-image",
@@ -5292,4 +5316,4 @@
]
}
}
}
}

164
scripts/omni-pricing.mjs Normal file
View File

@@ -0,0 +1,164 @@
#!/usr/bin/env node
/**
* Inject per-model pricing into ~/.agents/models.json for the `omni` provider.
*
* WHY THIS EXISTS
* ---------------
* `pi --list-models` / the omni extension's /omni sync regenerates models.json
* from OmniRoute and writes entries WITHOUT a `cost` field. pi's calculateCost()
* then reports $0.00 for every run, so Paperclip's cost events and budgets all
* read zero even though real money is being spent.
*
* The fix is to add { cost: { input, output, cacheRead, cacheWrite } } (USD per
* million tokens) to the model entries we actually use.
*
* Because /omni sync rewrites the whole file, this script must be re-run after
* every sync. That is the whole reason it is a script and not a one-off edit.
*
* PRICE SOURCES (both authoritative, no guessing)
* - deepseek/* -> OmniRoute /api/pricing, group "deepseek"
* - openrouter/* -> OpenRouter /api/v1/models (strip the "openrouter/" prefix)
*
* Model ids that use routing prefixes (aug/, ds/, oc/, tllm/, opencode-go/) do
* NOT map cleanly to pricing groups and are deliberately skipped — add them to
* MANUAL below if you need them.
*
* USAGE
* node omni-pricing.mjs --dry-run # show what would change
* node omni-pricing.mjs # write models.json
*/
import fs from "node:fs";
import path from "node:path";
const MODELS_JSON = path.join(process.env.HOME, ".agents", "models.json");
const OMNIROUTE_PRICING = "http://192.168.20.13:20128/api/pricing";
const OPENROUTER_MODELS = "https://openrouter.ai/api/v1/models";
// Models we actually assign to workers. Keep this list short and deliberate.
const TARGETS = [
"deepseek/deepseek-v4-flash",
"deepseek/deepseek-v4-pro",
"openrouter/google/gemini-3.1-flash-lite",
"openrouter/google/gemini-2.5-flash",
];
// Hand-priced entries: id -> {input, output, cacheRead, cacheWrite}
const MANUAL = {};
const dryRun = process.argv.includes("--dry-run");
function omniKey() {
const cfg = JSON.parse(fs.readFileSync(MODELS_JSON, "utf8"));
return cfg.providers?.omni?.apiKey ?? "";
}
/** Fetch OmniRoute's own price table: { group: { model: {input,output,cached,cache_creation} } } */
async function fetchOmniRoutePricing() {
const res = await fetch(OMNIROUTE_PRICING, {
headers: { Authorization: `Bearer ${omniKey()}` },
});
if (!res.ok) throw new Error(`OmniRoute pricing HTTP ${res.status}`);
return res.json();
}
/** Fetch OpenRouter prices keyed by bare model id (no "openrouter/" prefix). */
async function fetchOpenRouterPricing() {
const res = await fetch(OPENROUTER_MODELS);
if (!res.ok) throw new Error(`OpenRouter models HTTP ${res.status}`);
const { data } = await res.json();
const out = {};
for (const m of data) {
const p = m.pricing ?? {};
out[m.id] = {
input: Number(p.prompt ?? 0) * 1e6,
output: Number(p.completion ?? 0) * 1e6,
cacheRead: Number(p.input_cache_read ?? 0) * 1e6,
cacheWrite: Number(p.input_cache_write ?? 0) * 1e6,
};
}
return out;
}
const r6 = (n) => Math.round(Number(n) * 1e6) / 1e6;
/** Resolve a pi model id to a {input,output,cacheRead,cacheWrite} price, or null. */
function resolvePrice(id, omniPricing, orPricing) {
if (MANUAL[id]) return MANUAL[id];
// deepseek/<name> -> OmniRoute group "deepseek"
if (id.startsWith("deepseek/")) {
const name = id.slice("deepseek/".length);
const row = omniPricing?.deepseek?.[name];
if (row) {
return {
input: r6(row.input),
output: r6(row.output),
cacheRead: r6(row.cached ?? 0),
cacheWrite: r6(row.cache_creation ?? 0),
};
}
}
// openrouter/<bare-id> -> OpenRouter catalogue
if (id.startsWith("openrouter/")) {
const bare = id.slice("openrouter/".length);
if (orPricing[bare]) {
const p = orPricing[bare];
return { input: r6(p.input), output: r6(p.output), cacheRead: r6(p.cacheRead), cacheWrite: r6(p.cacheWrite) };
}
}
return null;
}
const cfg = JSON.parse(fs.readFileSync(MODELS_JSON, "utf8"));
const models = cfg.providers?.omni?.models;
if (!Array.isArray(models)) {
console.error("models.json: providers.omni.models is not an array — aborting");
process.exit(1);
}
const [omniPricing, orPricing] = await Promise.all([
fetchOmniRoutePricing(),
fetchOpenRouterPricing(),
]);
const byId = new Map(models.filter((m) => m && m.id).map((m) => [m.id, m]));
let changed = 0;
const missing = [];
for (const id of TARGETS) {
const entry = byId.get(id);
if (!entry) {
missing.push(id);
continue;
}
const price = resolvePrice(id, omniPricing, orPricing);
if (!price) {
missing.push(id);
continue;
}
const current = JSON.stringify(entry.cost ?? null);
const next = JSON.stringify(price);
if (current !== next) {
entry.cost = price;
changed++;
console.log(
`${dryRun ? "[dry-run] " : ""}${id} -> in $${price.input} out $${price.output} cacheR $${price.cacheRead} cacheW $${price.cacheWrite}`,
);
} else {
console.log(`${id} -> already priced, unchanged`);
}
}
if (missing.length) console.log(`\nnot found / no price (skipped): ${missing.join(", ")}`);
if (dryRun) {
console.log(`\ndry-run: ${changed} entr(y/ies) would change`);
} else if (changed) {
fs.writeFileSync(MODELS_JSON, JSON.stringify(cfg, null, 2) + "\n");
console.log(`\nwrote ${MODELS_JSON} (${changed} changed)`);
} else {
console.log("\nnothing to do");
}

View File

@@ -13,56 +13,83 @@ metadata:
hardware:
cpu:
model: "Intel(R) Core(TM) i7-7700K CPU @ 4.20GHz"
cores: 8
cores: 4
threads: 8
ram:
total_gb: 62.7
total_gb: 62
gpu:
- i915
- nvidia
- Intel HD Graphics 630 (i915) — drives the 4 display outputs
- NVIDIA Quadro P620 2GB (nvidia 580.173.02)
disks:
- device: /dev/nvme0n1p2
size: 869G
mountpoint: /
note: "ext4, 66% used (2026-10-05)"
- device: /dev/nvme0n1p3
size: 69G
mountpoint: "[SWAP]"
- device: efivarfs
size: 256K
mountpoint: /sys/firmware/efi/efivars
- device: /dev/nvme0n1p1
size: 1022M
mountpoint: /boot
- device: /dev/sda2
size: 224G
mountpoint: "(not mounted)"
note: NTFS
verified: 2026-10-05
- name: nixos-desktop
type: Desktop
ip: 192.168.20.13
os: "NixOS, Niri"
role: NixOS with Niri. Second dev. Web facing. Network scripts.
tech: Docker (Mosquitto, N8N, Pi Hole back up with nebulasync from sam-ubuntu1 pihole, pocketbase, speech_piper, voice_bridge), websites at /var/www/, snapcast, gstreamer, librespot, mopidy
os: "NixOS 26.11 (Zokor), KDE Plasma 6 / KWin Wayland (SDDM)"
role: Always-on web-facing server. Docker host + native services. Memory-bound — do not treat as CPU-bound.
tech: Docker (OmniRoute, Outline, n8n, Pi-hole replica + nebulasync, Mosquitto, voice_bridge/voice_whisper, speech_piper, pocketbase, family-home-lab, dsh, worldmonitor, prefect, litellm, headroom/JEV), Caddy :8000 serving /var/www, snapcast, librespot, mopidy, native PostgreSQL :5433 (paperclip), paperclipai, Paseo Daemon, where-woof
hardware:
cpu:
model: "AMD Ryzen 5 5600 6-Core Processor"
cores: 12
cores: 6
threads: 12
ram:
total_gb: 15.5
total_gb: 15
note: "6.8 GiB used; swap 7.1 GiB of 8.8 GiB used (81%) - 2026-10-05"
gpu:
- "NVIDIA GTX 760 (Kepler, PCI 10de:11c2) - DEAD. No driver bound; only /dev/dri/card0 simple-framebuffer. /proc/driver/nvidia absent. Desktop renders via software (llvmpipe). hardware.nvidia.open=true can never work on Kepler (open modules need Turing+). Recommended replacement: AMD RX 6600 (in-kernel amdgpu)."
disks:
- device: /dev/sdb2
- device: /dev/sda2
size: 907G
mountpoint: /
- device: efivarfs
size: 128K
mountpoint: /sys/firmware/efi/efivars
- device: /dev/sdb1
note: "ext4, 459G used (54%) - 2026-10-05"
- device: /dev/sda3
size: 8.8G
mountpoint: "[SWAP]"
note: "81% full - the main resource problem on this machine"
- device: /dev/sda1
size: 1022M
mountpoint: /boot
- device: /dev/sdb2
size: 916G
mountpoint: /mnt/storage
note: empty spare
- device: /dev/sdc2
size: 1.8T
mountpoint: /mnt/data
note: "21% used"
- device: /dev/sdd1
size: 2.7T
mountpoint: /mnt/ubuntu_storage_3TB
- device: /dev/sda2
size: 1.9T
mountpoint: /mnt/data
- device: /dev/sdc2
size: 932G
mountpoint: /mnt/storage
note: "USB drive - photo master + Borg backup target for .27 and .13. Letter shifts on replug (has been sdd1/sde1)."
- device: /dev/sde1
size: 1.4T
mountpoint: "(not mounted)"
note: MaxtorBackup, old migration tarballs
caveats:
- "Device letters are NOT stable - NixOS mounts by UUID. Always check lsblk/df before using a device path."
- "GPU is dead: no hardware acceleration. KDE session costs ~2.4 GB because of software rendering plus 28-day uptime."
- "litellm and headroom are idle and almost entirely swapped out (~1.05 GB of swap combined)."
- "Dead stacks to ignore: airflow (7 containers exited, 24 GB of logs in /home/sam/deployment/airflow/logs), engram container (never started - the live engram is a user systemd service on :7437 and its DB is empty), trigger_dev, t3_stack_react, sams-home-network."
- "There is a SECOND airflow install at /home/sam/deployment/ai-resume/airflow/ - also not running."
verified: 2026-10-05
- name: sam-ubuntu1
type: virtual machine
ip: 192.168.20.35
@@ -184,6 +211,28 @@ Sam's NixOS multi-machine infrastructure reference.
Full hardware specs are in the frontmatter metadata.
## Resource Pressure on .13 (verified 2026-10-05)
`.13` is **memory-bound, not CPU-bound**. Load average is ~0.27 on 12 threads, but the
8.8 GB swap partition is **81% full**. Before proposing anything that adds load to `.13`,
check the swap. Full detail: Obsidian → `300 areas/360 Dev-Ops Network Computers/Resource Use — .13 CPU & RAM.md`.
| Consumer | RAM + swap | Notes |
|---|---|---|
| KDE Plasma session (~30 processes) | ~2.4 GB | Software rendering (GPU dead) + 28-day uptime |
| Docker containers (all) | ~3.5 GB RSS | `voice_whisper` 877 MB is the largest |
| `litellm` + `headroom` | ~1.05 GB *swap* | Idle; both fully paged out; redundant with OmniRoute |
| Dead Airflow logs | 24 GB *disk* | `/home/sam/deployment/airflow/logs` — nothing depends on it |
**Must stay on `.13`:** Pi-hole `:53`, OmniRoute `:20128/20129` (all pi LLM traffic),
Mosquitto `:1883` (Home Assistant on `.30`), the Borg backup target on `/mnt/ubuntu_storage_3TB`
(`.27` pushes daily at 04:00), Caddy `:8000` + `/var/www`.
**Corrections applied 2026-10-05:** `.13` runs **KDE Plasma 6**, not Niri. `.13` disk device
letters were wrong (root is `/dev/sda2`). Home Assistant is on **`.30`, not `.35`** (`.35` is
`sam-ubuntu1`, the Caddy/docker app host). Pi-hole primary is **`.35`**; `.13` is the replica
(confirmed from `nebulasync`: `PRIMARY=http://192.168.20.35`, `REPLICAS=http://192.168.20.13`).
## Pi Agent (uniform across machines)
- **Binary**: nixpkgs `pi-coding-agent` 0.82.1 in `home.packages` on .27/.13/.51 (Nix-managed; npx wrapper + npm-global install removed 2026-08-05). Updates = nixpkgs flake bump + rebuild.