Files
memory/log.md
T
2026-07-24 22:00:15 +00:00

6.6 KiB

Memory Log

[2026-07-24] fix | Velero CSI Volume Snapshots Three-Layer Fix

  • Velero backups PartiallyFailed for 121 days — all PVCs skipped
  • Fix 1: Created VolumeSnapshotClass ceph-rbd-snapclass (rbd.csi.ceph.com)
  • Fix 2: Migrated ArgoCD repoURLs from dead Gitea (10.0.30.105) to new (ssh://git@10.0.30.202:22)
    • Gitea SSH user is git not gitea; git.schoen.codes → Traefik (10.0.30.203), SSH on 10.0.30.202
    • New SSH deploy key generated, added to Gitea repo
  • Fix 3: features: EnableCSI must be under configuration: in Velero Helm values (not top-level)
    • Must be string, not array — array form breaks Helm template
  • Verified: test backup created VolumeSnapshot + VolumeSnapshotContent, uploaded to S3
  • Solution doc: docs/solutions/bug-fixes/2026-07-24-velero-csi-snapshots-three-layer-fix.md
  • Commits: 5ebfa30, 19f35bc, cd18c58, baf85f5, f652bfb

[2026-07-24] fix | VFIO Module Loading Race Condition RCA + Fix

  • RCA: GPU passthrough failed on ms-a2-1 due to missing modprobe config (NOT kernel version)
  • Root cause: Missing softdep amdgpu pre: vfio-pci + incomplete DRM blacklist → race condition (302s vs 2.4s bind time)
  • Fix applied to ms-a2-1: Added blacklist drm, blacklist drm_kms_helper, softdep amdgpu pre: vfio-pci, rebuilt initramfs
  • Verification pending (reboot required, 4 VMs + OSD 13 on ms-a2-1)
  • Solution doc: docs/solutions/bug-fixes/2026-07-24-vfio-pci-module-loading-race-condition.md
  • Also documented: K8s persistent storage architecture (all PVCs on Ceph RBD, 3x replication)
  • Solution doc: docs/solutions/architecture/2026-07-24-k8s-persistent-storage-ceph-rbd.md

[2026-07-24] fix | worker-05 GPU Passthrough — Move to ms-a2-2

  • Problem: worker-05 (VM 102) on ms-a2-1 had kernel BUG with 1002:13c0 GPU — /dev/dri/renderD128 missing
  • ms-a2-1 and ms-a2-2 have identical hardware (AMD Ryzen 9 9955HX, GPU 1002:13c0)
  • Solution: Moved VM 102 from ms-a2-1 to ms-a2-2 (offline migration, Ceph RBD shared storage)
  • Result: renderD128 visible in guest, amdgpu driver loaded, all pods running
  • IaC: main.tf cleaned up — single unified rke2_worker block, worker-05 node=ms-a2-2, worker-06 removed
  • PCI mapping: amd-gpu now includes n5pro + ms-a2-1 + ms-a2-2
  • Commit: 21bd335
  • Also: ms-a2-2 upgraded PVE 9.1.1→9.2.5, kernel 7.0.14-6-pve, VFIO persistent after reboot, HA integrated

[2026-07-24] init | ms-a2-2 Node Initialization

  • New Proxmox node ms-a2-2 (10.0.20.93, nodeid 1) joined cluster
  • Hardware: AMD Ryzen 9 9955HX (32 threads), 91GB RAM, 1.8TB NVMe
  • GPU: AMD Radeon 1002:13c0 (identical to ms-a2-1) — VFIO passthrough configured
  • PCI mapping: amd-gpu-ms-a2-2 created, added to generic amd-gpu mapping
  • Ceph OSD 14: created on nvme0n1 (1.8TB SSD), class ssd, weight 1.82
  • Fluent-bit: installed + configured (hostname adapted to ms-a2-2)
  • APT sources: fixed from enterprise to no-subscription (ceph + pve)
  • LLM Wiki: ip-map.md, ceph.md updated; AGENTS.md updated (9 nodes)
  • Note: PVE 9.1.1 on ms-a2-2 (rest of cluster is 9.2.3) — upgrade pending

[2026-07-24] cleanup | proxmox1 Permanent Removal + Ceph PG Fixes

  • Removed proxmox1 from Ceph CRUSH (empty host bucket, weight 0)
  • Also removed stale empty buckets: proxmox, px-tmp20
  • Updated corosync.conf: removed proxmox1 node entry, config_version 14→15, expected_votes 9→8
  • Removed stale /etc/pve/nodes/{proxmox1,proxmox,pve,px-tmp20}
  • Updated IaC: proxmox_api_url 10.0.20.10→10.0.20.91, template_node proxmox1→n5pro, target_nodes list updated
  • Updated AGENTS.md topology (9→8 nodes), LLM Wiki ip-map.md + ceph.md
  • Commit: 469d29d
  • Additional: Manual pg-upmap for stuck PGs 5.13→[1,6,7], 6.6c→[11,8,7] — all clean+remapped eliminated

[2026-07-24] fix | Ceph Remapped PGs — Weight Imbalance + pg_num Mismatch

  • RCA: 11 active+clean+remapped PGs caused by CRUSH placement failure
  • Root cause 1: Pool 6 (hdd_disk) pg_num=120 ≠ pgp_num=112 (autoscaler mid-merge)
  • Root cause 2: Extreme HDD host weight imbalance (n5pro=10TB vs proxmox6=0.3TB)
  • Fix: Equalized pg_num→112, set nopgchange=true, reweighted osd.7 0.80→1.0
  • Result: 9/11 PGs recovering (backfilling at 36MiB/s), 2 cosmetic clean+remapped remain
  • Created systems/ceph.md with full OSD/pool/CRUSH documentation

[2026-07-24] update | Gitea Actions K8s Runner Bug + DinD Pitfalls

  • Documented Gitea 1.27.0 bug: actions_ready_job queue doesn't assign tasks to K8s-based runners
  • Updated systems/gitea.md with known issue, DinD pitfalls, runner management via DB
  • Solution doc: docs/solutions/bug-fixes/2026-07-24-gitea-actions-k8s-runner-task-assignment-bug.md
  • IaC committed: runner-deployment.yaml + runner-config.yaml (DinD TCP sidecar pattern)

[2026-04-28] init | Memory System Created

  • Created memory directory structure at ~/.hermes/memory
  • Installed qmd 2.1.0 (Query Markup Documents)
  • Created initial index.md with entities, concepts, and sources
  • Created Entity pages: Infrastructure, Email-System, Health-Fitness, Team-Structure, Memory-System
  • Created Concept pages: Network-Architecture, Email-Organization, Monitoring-System, Proxmox-Cluster, qmd-Knowledge-Base
  • Initialized git repo, pushed to origin/main

[2026-04-29] fix | Memory Process Established

  • Corrected stale memory-entry-001.md
  • Established protocol: memory update + git sync on every session

[2026-04-30] ingest | cloud.familie-schoen.com

  • Full port scan, DNS, SSL, HTTP headers, API analysis
  • Identified: Seafile 13.0.19 on nginx, Let's Encrypt SSL

[2026-07-24] restructure | Full Wiki Restructuring

  • Deleted 5 root-duplicate files (dominik-schoen, email-organization, team-structure, llm-wiki-pattern, embedding-model)
  • Deleted obsolete Entities/Memory-System.md, Concepts/qmd-Knowledge-Base.md
  • Deleted raw/articles/ and memory-entry-001/002.md (stale)
  • Renamed Entities/ → entities/, Concepts/ → concepts/ (lowercase)
  • Created new directories: systems/, reference/
  • Moved Proxmox-Cluster, Monitoring-System, seafile to systems/
  • Updated all existing pages with current data (April → July 2026):
    • Infrastructure: 2 Nodes → 9 Nodes, added Ceph/RKE2/Galera/Loki
    • Proxmox-Cluster: PVE 9.2.3, 9 Nodes, Fluent Bit details, PVE Tasks, Ceph Audit
    • Network-Architecture: cleaned up, links to reference pages
    • Monitoring: Prometheus/Grafana/Loki/HolmesGPT (was just "weekly network scans")
  • Created 6 new system pages: ceph-cluster, rke2-kubernetes, galera-maxscale, loki-fluentbit, gitea, hindsight
  • Created 3 reference pages: ip-map, ssh-keys, ports
  • Created 2 concept pages: gitops-workflow, credential-policy
  • Rewrote index.md with new structure + Memory Layer Architecture table
  • Total: 18 pages (was 21 with dupes, now 18 clean unique pages)