Compare commits

..
4 Commits
Author SHA1 Message Date
Dominik Schön 7c36a7600a Auto-sync: 2026-08-31 2026-08-31 22:00:17 +00:00
Dominik Schön 23e3702f0a chore: update skill-impact tracker with retrospective entry 2026-08-30 11:57:19 +00:00
Dominik Schön 5cf6a573f1 retro: compound-learning 30-day retrospective + 2 new patterns
New patterns:
- PAT-008: Ansible default_ipv6 fact missing on fresh VMs
- PAT-009: Stale NBD devices after RBD volume swap

Updated:
- index.md: added PAT-008, PAT-009
- log.md: retrospective entry with 4 solution docs + 2 patterns
- Solution docs dispatched to ~/docs/solutions/
2026-08-30 11:56:47 +00:00
Dominik Schön bbbe4f8985 feat: patterns/ directory + skill-impact tracker (WikiSkill-inspired)
- New patterns/ directory with 7 initial failure-mode patterns (PAT-001..007)
- skill-impact.md audit trail for skill modifications
- _template.md for future pattern creation
- index.md updated with Patterns section
- log.md entry for this change
- Inspired by arXiv:2608.27454 (WikiSkill)
2026-08-30 11:16:19 +00:00
15 changed files with 704 additions and 11 deletions
+11 -11
View File
@@ -1,23 +1,23 @@
# User Drift Report — 2026-08-24
# User Drift Report — 2026-08-31
**Recent window:** Last 14 days (253 msgs)
**Baseline:** Previous 90 days (4727 msgs)
**Recent window:** Last 14 days (382 msgs)
**Baseline:** Previous 90 days (5490 msgs)
## Detected Drifts
- **message_length**: Messages 80% longer (baseline: 1120 chars → recent: 2017)
- **new_focus**: New dominant topics: ai_ml, galera
- **declining_focus**: Topics fading from focus: gitops, ceph
- **message_length**: Messages 41% longer (baseline: 1083 chars → recent: 1531)
- **new_focus**: New dominant topics: ai_ml, galera, security
- **declining_focus**: Topics fading from focus: gitops, ceph, k8s
## Signal Summary
| Metric | Baseline | Recent |
|---|---|---|
| Avg msg length | 1120 | 2017 |
| Frustration rate | 4.0% | 8.3% |
| Correction rate | 7.3% | 12.3% |
| Action-first | 1.8% | 2.0% |
| Top topics | ceph, proxmox, k8s | galera, proxmox, k8s |
| Avg msg length | 1083 | 1531 |
| Frustration rate | 4.0% | 6.0% |
| Correction rate | 6.8% | 8.9% |
| Action-first | 1.7% | 0.0% |
| Top topics | ceph, proxmox, k8s | security, galera, ai_ml |
## Recommendations
+14
View File
@@ -12,6 +12,7 @@
| `systems/` | Software-Systeme und Services |
| `concepts/` | Abstrakte Patterns & Konventionen |
| `reference/` | Quick-Lookup Tabellen (IPs, Keys, Ports) |
| `patterns/` | Rekurrente Failure-Mode Patterns + Skill-Impact Tracker |
## Entities
- [[entities/infrastructure]] — Physikalische Hardware, Cluster-Übersicht
@@ -42,6 +43,19 @@
- [[reference/ssh-keys]] — Key → Zweck → Fingerprint
- [[reference/ports]] — Port → Service → Host
## Patterns
- [[patterns/_README]] — Konzept und Nutzung des Patterns-Verzeichnisses
- [[patterns/galera-ddl-deadlock]] — OPTIMIZE TABLE → TOI Deadlock (PAT-001)
- [[patterns/traefik-reload-unreliable]] — reload unreliable, restart statt reload (PAT-002)
- [[patterns/1password-rate-limit]] — Account-Level Rate-Limit, ESO stoppen (PAT-003)
- [[patterns/vfio-gpu-passthrough]] — softdep + DRM blacklist Race Condition (PAT-004)
- [[patterns/finanzblick-sync-waf]] — POST /sync WAF-blocked, UI Modal Sequence (PAT-005)
- [[patterns/linkedin-react-input]] — nativeInputSetter, Session Expiry (PAT-006)
- [[patterns/ceph-ssd-wear-ec-pool]] — EC Pool unusable + SSD Wear-Level (PAT-007)
- [[patterns/ansible-default-ipv6]] — lablabs.rke2 role fails on IPv6-less VMs (PAT-008)
- [[patterns/k8s-stale-nbd-devices]] — RBD swap leaves stale NBD mappings (PAT-009)
- [[patterns/skill-impact]] — Skill Modification Audit Trail
## Memory Layer Architektur
| Layer | System | Pfad | Rolle |
|-------|--------|------|-------|
+30
View File
@@ -1,5 +1,35 @@
# Memory Log
## [2026-08-30] retro | Compound Learning Retrospective (Last 30 Days)
- Reviewed sessions from Jul 31 Aug 30, 2026
- **4 new solution docs** written by subagents:
- `bug-fixes/2026-08-05-ceph-squid-ec-pool-mark-complete-bug.md` — EC pool mark-complete fails in Squid
- `bug-fixes/2026-08-06-traefik-cross-namespace-routing-pitfall.md` — Two Traefik 404 incidents
- `architecture/2026-08-01-k8s-full-cluster-rebuild-procedure.md` — Full RKE2 rebuild sequence
- `architecture/2026-08-01-immich-data-loss-no-backup.md` — Ceph redundancy ≠ backup
- **2 new patterns** extracted:
- PAT-008: Ansible default_ipv6 fact missing on fresh VMs
- PAT-009: Stale NBD devices after RBD volume swap
- **4 Hindsight entries** indexed with solution summaries
- Key themes: K8s cluster disaster recovery, Ceph EC pool unrecoverability, Traefik routing complexity, backup gap identification
## [2026-08-30] feat | WikiSkill Patterns Directory + Skill-Impact Tracker
- Analysed arXiv:2608.27454 (WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution)
- **New `patterns/` directory** in LLM Wiki — structured failure-mode patterns inspired by WikiSkill's Wiki Layer
- Created 7 initial pattern pages from existing MEMORY.md entries + solution docs:
- PAT-001: Galera DDL TOI Deadlock
- PAT-002: Traefik reload unreliable
- PAT-003: 1Password account-level rate-limit
- PAT-004: VFIO GPU passthrough race condition
- PAT-005: Finanzblick sync WAF-blocked
- PAT-006: LinkedIn React nativeInputSetter
- PAT-007: Ceph EC pool + SSD wear-level
- **Skill-Impact Tracker** (`patterns/skill-impact.md`) — audit trail for skill modifications (accept/reject history)
- **Template** (`patterns/_template.md`) for future pattern creation
- **compound-learning skill patched** — added Phase 4.8 (Pattern Extraction) + Phase 4.9 (Skill-Impact Tracking)
- Updated index.md with Patterns section
- Hindsight: indexed with tags architecture, wikiskill, patterns, skill-evolution
## [2026-08-13] feat | Memory Architecture Enhancement (arxiv 2602.06052v4)
- Implemented 3 new automated memory capabilities from survey paper analysis:
- **Decay & Importance Scoring** (`hindsight_decay_scoring.py`): Monthly cron (1st 03:30). Score = log(1+access)/log(51) × 0.5^(age/90d). Archives score<0.15 with 0 accesses. First run: 200 evaluated, 50 archived.
+51
View File
@@ -0,0 +1,51 @@
---
pattern_id: PAT-003
title: "1Password CLI rate-limit is account-level — stop ESO, wait 1 hour"
category: tooling
severity: high
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: [rke2-kubernetes]
related_solution_docs:
- docs/solutions/architecture/2026-07-24-1password-account-level-rate-limit.md
- docs/solutions/bug-fixes/2026-07-14-1password-token-telegram-corruption.md
related_skills: [1password-cli]
---
# 1Password CLI rate-limit is account-level — stop ESO, wait 1 hour
## Symptom
All 1Password CLI (`op`) calls suddenly fail with rate-limit errors. This affects
both Hermes tools and ExternalSecrets Operator (ESO) in K8s simultaneously.
The outage lasts approximately 1 hour.
## Root Cause
The 1Password CLI enforces rate limits at the **account level**, not per-token or
per-client. When ESO polls aggressively (multiple SecretStores + frequent sync intervals),
it exhausts the quota for the ENTIRE account, blocking all other `op` callers including
Hermes automation.
## Mitigation
1. **Stop ESO immediately** — scale down the ESO deployment to halt API calls:
```bash
kubectl scale deploy -n external-secrets external-secrets-operator --replicas=0
```
2. **Wait 1 hour** — the rate limit resets automatically. No polling needed.
3. **Resume ESO** with reduced polling frequency afterward
## Prevention
- Reduce ESO `refreshInterval` to 10+ minutes (not the default 1 minute)
- Limit the number of SecretStores that reference 1Password
- Consider caching secrets locally to reduce API pressure
- Never run `op` in tight loops — always add delays for batch operations
## Evidence
- Hit twice: Jul 2026 (initial discovery) and during token corruption investigation
- Solution doc: `docs/solutions/architecture/2026-07-24-1password-account-level-rate-limit.md`
- MEMORY.md entry: "1P CLI:3 vaults R/W. Rate-limit=acct-level. Stop ESO, wait 1hr, no polling."
+51
View File
@@ -0,0 +1,51 @@
# Patterns Directory
Inspired by [WikiSkill (arXiv:2608.27454)](https://arxiv.org/abs/2608.27454) — this directory
contains structured **failure-mode patterns** and **successful strategies** extracted from
real operational experience.
## Purpose
Unlike `systems/` (describes what exists) or `concepts/` (abstract conventions),
`patterns/` captures **recurrent problems and their proven solutions** — the "lessons learned"
that compound across incidents.
## Structure
Each pattern is a standalone Markdown file:
```
patterns/
├── _README.md ← this file
├── _template.md ← copy this for new patterns
├── galera-ddl-deadlock.md
├── traefik-reload-unreliable.md
├── ...
└── skill-impact.md ← audit trail of skill modifications (accept/reject history)
```
## How Patterns Are Born
1. **Incident occurs** → problem is diagnosed and fixed
2. **Compound-learning cycle** runs → solution doc created in `~/docs/solutions/`
3. **Pattern extracted** → if the failure mode is recurrent or broadly applicable,
a pattern page is created here
4. **Wiki Maintainer** (currently manual, future: automated) consolidates evidence
from multiple incidents into the pattern page
## Relationship to Other Layers
| Layer | Holds | Retrieval |
|-------|-------|-----------|
| `~/docs/solutions/` | One-time detailed incident docs | `search_files` (keyword) |
| `patterns/` (here) | Recurring failure-mode patterns | Browse `index.md`, cross-linked |
| MEMORY.md (L0) | Compressed pointer if high-priority | System prompt injection |
| Hindsight (L2) | Semantic index of all above | `hindsight_recall` |
## Rules
- **One pattern per file** — don't merge unrelated patterns
- **Evidence-based** — cite real incidents (link to solution docs or session dates)
- **Actionable** — every pattern must have a "Mitigation" or "Prevention" section
- **Never delete without redirect** — if a pattern is obsolete, mark `status: superseded`
and link to the replacement
+30
View File
@@ -0,0 +1,30 @@
---
pattern_id: PAT-XXX
title: "<short descriptive title>"
category: database|infrastructure|storage|networking|tooling|integration|security
severity: low|medium|high
status: active|superseded
first_observed: YYYY-MM
last_updated: YYYY-MM-DD
related_systems: []
related_solution_docs: []
related_skills: []
---
# <Title>
## Symptom
What goes wrong? What are the observable symptoms?
## Root Cause
Why does it happen? Trace the actual cause, not just the symptom.
## Mitigation
What was the fix? Include commands/snippets if relevant.
## Prevention
How to avoid this in the future? (monitoring, lint rule, convention, etc.)
## Evidence
- When was this observed? Which incidents?
- Links to solution docs, session IDs, MEMORY entries
+59
View File
@@ -0,0 +1,59 @@
---
pattern_id: PAT-008
title: "Ansible default_ipv6 fact missing on fresh VMs → lablabs.rke2 role fails"
category: tooling
severity: medium
status: active
first_observed: 2026-07-21
last_updated: 2026-08-30
related_systems: [rke2-kubernetes]
related_solution_docs:
- docs/solutions/bug-fixes/2026-07-21-ansible-rke2-default-ipv6-bug.md
related_skills: []
---
# Ansible default_ipv6 fact missing on fresh VMs → lablabs.rke2 role fails
## Symptom
The `lablabs.rke2` Ansible role fails on freshly provisioned VMs with a Jinja2 error:
```
ansible_facts['default_ipv6']['address']
```
The `default_ipv6` fact doesn't exist on VMs without IPv6 configured, causing the
`meta/argument_specs.yml` validation to crash.
## Root Cause
The `lablabs.rke2` role's `meta/argument_specs.yml` references
`ansible_facts['default_ipv6']['address']` with a hardcoded default. On fresh VMs
without IPv6, `ansible_facts['default_ipv6']` is undefined → Jinja2 raises
`UndefinedError`.
A `pre_task` setting `default_ipv6` via `combine()` doesn't help because the
argument_specs validation runs BEFORE pre_tasks — it re-collects facts and
overwrites the pre_task fix.
## Mitigation
Patch `meta/argument_specs.yml` directly in the role:
```yaml
# Replace:
default: "{{ ansible_facts['default_ipv6']['address'] }}"
# With:
default: "{{ ansible_facts.default_ipv6.address | default(None) }}"
```
Or: enable IPv6 on the target VMs (faster workaround, no role patching needed).
## Prevention
- Pin Ansible roles and patch argument_specs when they assume facts that may not exist
- Test roles on fresh VMs without IPv6 before production use
- Consider forking the role with the fix upstream
## Evidence
- Observed during GPU worker provisioning (worker-04/05, Jul 2026)
- Solution doc: `docs/solutions/bug-fixes/2026-07-21-ansible-rke2-default-ipv6-bug.md`
- Session: @session:default/20260721_115501_78b25032
+60
View File
@@ -0,0 +1,60 @@
---
pattern_id: PAT-007
title: "Ceph EC pool unusable + SSD wear-level failing — monitor pg states"
category: storage
severity: high
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: [ceph-cluster]
related_solution_docs:
- docs/solutions/architecture/2026-07-12-ceph-ec-pool-no-rebalance-headroom.md
- docs/solutions/bug-fixes/2026-07-13-ceph-ratio-ordering-constraint.md
- docs/solutions/bug-fixes/2026-07-13-pveceph-wal-lvm-device-mapper-limitation.md
related_skills: [ceph-cluster-administration, ceph-ec-incomplete-pg-recovery]
---
# Ceph EC pool unusable + SSD wear-level failing — monitor pg states
## Symptom
Multiple Ceph failure modes manifest simultaneously:
1. EC pool (pool 8/media_ec) becomes unusable — PGs stuck in `incomplete` state
2. SSD OSD (osd.3 on px4) reports `Wear Leveling Count` SMART attribute indicating
imminent failure
3. `ceph health` shows `HEALTH_ERR`
## Root Cause
These are compounded issues:
1. **EC pool incomplete PGs**: After OSD loss, EC pools cannot recover without enough
surviving shards. The minimum copies requirement for EC k+m encoding is stricter than
replicated pools.
2. **SSD wear-level failure**: Consumer-grade SSDs in Ceph OSD duty cycle reach wear limits.
Once `Wear Leveling Count` exceeds threshold, the SSD becomes read-only or unreliable.
3. **Weight imbalance**: Unequal OSD weights cause CRUSH placement failures, leaving PGs
perpetually remapped.
## Mitigation
1. **EC pool recovery**: Use the `ceph-ec-incomplete-pg-recovery` skill — specialized
procedure for forcing EC PG recovery after OSD loss
2. **Replace failing SSD**: Mark OSD `out``destroy``zap` → physically replace → recreate
3. **Fix weight imbalance**: Equalize `pg_num``pgp_num`, set `nopgchange=true`,
reweight OSDs proportionally to disk capacity
4. **Remove empty CRUSH buckets**: Clean up stale host entries from decommissioned nodes
## Prevention
- Monitor SMART attributes monthly — alert on wear-level > 80%
- Never use consumer SSDs for Ceph OSDs in production (use enterprise grade with PLP)
- Keep EC pool k+m ratio proportional to available OSDs (don't over-provision redundancy)
- Regular `ceph pg dump` audits for stuck/unactive PGs
- Separate PBS onto its own tier (CephFS), away from RBD pools
## Evidence
- Multiple incidents Jul 2026 (OSD 0+2 destruction, OSD 3 wear-level, EC pool unusable)
- Solution docs: `2026-07-12-ceph-ec-pool-no-rebalance-headroom.md`,
`2026-07-13-ceph-ratio-ordering-constraint.md`
- MEMORY.md entry: "Ceph:HEALTH_ERR. Pool8/media_ec unusable. px4 SSD(osd.3) WearLevel FAILING."
+53
View File
@@ -0,0 +1,53 @@
---
pattern_id: PAT-005
title: "Finanzblick sync requires modal sequence — POST /sync is WAF-blocked"
category: integration
severity: medium
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: []
related_solution_docs: []
related_skills: [finanzblick-cashflow]
---
# Finanzblick sync requires modal sequence — POST /sync is WAF-blocked
## Symptom
Programmatic synchronization with Finanzblick (banking data aggregator) fails when
calling the `POST /sync` API endpoint directly. The request is blocked by the WAF
(Web Application Firewall), returning 403 or connection reset.
Historical fetches (without sync) work fine with the `--no-sync` flag, avoiding 2FA.
## Root Cause
Finanzblick's WAF detects and blocks automated POST requests to the sync endpoint
that don't originate from the legitimate browser session with proper CSRF tokens
and session cookies.
## Mitigation
Sync must be performed via the **UI button + 2FA modal sequence**:
1. Navigate to the Finanzblick web interface in a browser
2. Click the sync button (UI-triggered, not API)
3. Handle the 2FA modal sequence in order:
- PIN modal → click OK
- AUTH modal → click WEITER
- ERR modal → click OK
4. Wait for sync completion
For historical data fetches (no sync needed), use the `--no-sync` flag — this
bypasses 2FA entirely.
## Prevention
- Never attempt direct `POST /sync` calls — always use the UI flow
- The `finanzblick-cashflow` skill encodes this modal sequence
- This skill is USER-OWNED and needs `hermes curator adopt` to manage
## Evidence
- Observed during Finanzblick cashflow analysis sessions (Jul 2026)
- MEMORY.md entry: "FB sync=UI btn+2FA modals(PIN→OK,AUTH→WEITER,ERR→OK). POST /sync=WAF-blocked. --no-sync flag for hist.fetches(no 2FA). fb-cashflow skill=USER-OWNED,needs `hermes curator adopt`."
+52
View File
@@ -0,0 +1,52 @@
---
pattern_id: PAT-001
title: "OPTIMIZE TABLE on Galera triggers TOI DDL → Deadlock"
category: database
severity: high
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: [galera-maxscale]
related_solution_docs:
- docs/solutions/bug-fixes/2026-07-23-cnpg-timeline-corruption-rebuild.md
related_skills: []
---
# OPTIMIZE TABLE on Galera triggers TOI DDL → Deadlock
## Symptom
Running `OPTIMIZE TABLE` on a Galera cluster node causes a cluster-wide stall or deadlock.
The table lock propagates via TOI (Total Order Isolation) to all nodes, blocking all writes
to ALL tables during the operation.
## Root Cause
Galera executes DDL statements (including `OPTIMIZE TABLE`, `ALTER TABLE`, etc.) in TOI mode.
This means the DDL is replicated as a global operation that blocks the entire cluster — not
just the target table. For large tables, the rebuild phase can take minutes, causing apparent
outages.
## Mitigation
1. **Use RSU before DDL**: Switch the node to Rolling Schema Update (RSU) mode before running
`OPTIMIZE TABLE`. This prevents cluster-wide blocking:
```
SET GLOBAL wsrep_OSU_method = 'RSU';
-- run OPTIMIZE TABLE on this node only
SET GLOBAL wsrep_OSU_method = 'TOI'; -- restore
```
2. **Schedule during maintenance window** — even with RSU, the node itself is degraded
3. **Consider pt-online-schema-change** for large tables — avoids blocking entirely
## Prevention
- Never run `OPTIMIZE TABLE` or `ALTER TABLE` in production without RSU on Galera
- Add this check to DBA runbooks and monitoring alerts
- The Galera skill (`mariadb-galera-cluster-administration`) documents this pitfall
## Evidence
- Observed during Galera cluster administration sessions (Jul 2026)
- MEMORY.md entry: "OPTIMIZE TABLE=TOI DDL→Deadlock! RSU davor."
- Galera cluster: nodes 300/301/302, VIP .70:3306
+66
View File
@@ -0,0 +1,66 @@
---
pattern_id: PAT-009
title: "Stale NBD devices after RBD volume swap — delete VolumeAttachment + nbd teardown"
category: infrastructure
severity: high
status: active
first_observed: 2026-08-01
last_updated: 2026-08-30
related_systems: [rke2-kubernetes, ceph-cluster]
related_solution_docs:
- docs/solutions/architecture/2026-08-01-k8s-full-cluster-rebuild-procedure.md
related_skills: []
---
# Stale NBD devices after RBD volume swap — delete VolumeAttachment + nbd teardown
## Symptom
After swapping RBD images for PVCs (e.g. migrating to new Ceph pool or recreating
images), pods fail to mount with errors like:
```
MountVolume.MountDevice failed for volume "pvc-xxx" : rpc error: code = Internal
desc = rbd: map failed with error: /dev/nbd0 already in use
```
The NBD device is held by a stale mapping from the old RBD image, even though the
new image has the same name.
## Root Cause
When an RBD image is recreated (delete + create with same name), the Ceph CSI
driver's NBD mappings from the old image remain active. The Linux NBD layer
holds `/dev/nbdX` open, blocking new mounts to the same device path.
The Kubernetes VolumeAttachment object also references the old volume handle,
preventing the CSI driver from cleanly attaching the new volume.
## Mitigation
Three-step teardown procedure:
```bash
# 1. Delete the VolumeAttachment (allows CSI driver to release)
kubectl delete volumeattachment csi-cephfsplugin-<node>-<volume-handle>
# 2. Disconnect the stale NBD device on the target node
ssh <node> 'nbd-client -d /dev/nbd0' # or: qemu-nbd --disconnect /dev/nbd0
# 3. Restart the CSI node plugin to pick up clean state
kubectl delete pod -n kube-system csi-cephfsplugin-<node-id>
# (DaemonSet will respawn it)
```
After this, the pod can remount with the new RBD image.
## Prevention
- Before deleting RBD images, ensure all pods using them are scaled to 0
- Delete VolumeAttachments BEFORE deleting RBD images
- After RBD image recreation, restart CSI plugins on all nodes that had mounts
- Document this in the K8s disaster recovery runbook
## Evidence
- Observed during K8s cluster rebuild (Aug 2026) — 13 RBD volumes needed swapping
- Session: @session:default/20260731_113701_74fd2814 (250+ tool calls)
- Part of the full cluster rebuild procedure
+58
View File
@@ -0,0 +1,58 @@
---
pattern_id: PAT-006
title: "LinkedIn login via browser — nativeInputSetter required, session expires between nav"
category: integration
severity: medium
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: []
related_solution_docs:
- docs/solutions/bug-fixes/2026-07-10-ha-token-expiry-browser-session-loss.md
related_skills: [linkedin-personal-branding]
---
# LinkedIn login via browser — nativeInputSetter required, session expires between nav
## Symptom
Automated LinkedIn login via browser tools fails silently. Form fields appear filled
but LinkedIn doesn't recognize the input (login button stays disabled, or submission
fails with "invalid credentials").
Additionally, authenticated sessions expire between page navigations, requiring
re-login on almost every navigation step.
## Root Cause
1. **React-controlled inputs**: LinkedIn uses React, which manages input state internally.
Simply setting `.value` via DOM manipulation doesn't trigger React's onChange handler.
The `nativeInputSetter` approach is required:
```javascript
const setter = Object.getOwnPropertyDescriptor(
window.HTMLInputElement.prototype, 'value'
).set;
setter.call(inputElement, 'my-value');
inputElement.dispatchEvent(new Event('input', { bubbles: true }));
```
2. **Session expiry**: LinkedIn's SPA doesn't persist auth tokens reliably across
`browser_navigate` calls (each navigation may reset the JS context).
## Mitigation
- Always use `browser_console` with `nativeInputSetter` for form fills — NOT `browser_click`
followed by typing
- Perform login + desired action in a SINGLE browsing session (minimize navigations)
- UTF-8 files must be saved without BOM (LinkedIn rejects BOM in posted content)
## Prevention
- The `linkedin-personal-branding` skill documents this workflow
- Never use `browser_click` for LinkedIn form fields
- Minimize navigation steps after login
## Evidence
- Observed during LinkedIn branding sessions (Jul 2026)
- MEMORY.md entry: "Login via browser_console nativeInputSetter. NOT browser_click. Session expires between nav. UTF-8 no BOM."
+40
View File
@@ -0,0 +1,40 @@
# Skill Impact Tracker
> Inspired by WikiSkill (arXiv:2608.27454) — tracks skill modifications, their validation
> outcomes, and acceptance decisions. Prevents repeating failed modifications and provides
> an audit trail of skill evolution.
## How to Use
When a skill is patched, created, or deleted, append an entry to the table below.
The entry records WHAT changed, WHY, and WHETHER it helped (if validated).
### Entry Format
```
| Date | Skill | Change Type | Description | Validation | Outcome | Ref |
```
- **Change Type**: `created` | `patched` | `deleted` | `adopted`
- **Validation**: How was success measured? (`manual` | `tested` | `benchmark` | `n/a`)
- **Outcome**: `accepted` | `rejected` | `rolled-back` | `pending`
- **Ref**: Session ID or solution doc path
## Audit Trail
| Date | Skill | Change Type | Description | Validation | Outcome | Ref |
|------|-------|-------------|-------------|------------|---------|-----|
| 2026-08-30 | compound-learning | patched | Added WikiMaintainer phase + patterns/ guidance from WikiSkill paper analysis | tested | accepted | this session |
| 2026-08-30 | (retrospective) | created | 4 solution docs + 2 new patterns (PAT-008, PAT-009) from 30-day retrospective | tested | accepted | this session |
| 2026-08-13 | hindsight | patched | Expanded meta-memory automation table (decay scoring, drift detection, skill health) | tested | accepted | session 2026-08-13 |
| 2026-08-13 | (3 cron jobs) | created | Decay scoring, drift detection, skill health check — no_agent cron jobs | tested | accepted | session 2026-08-13 |
## Notes
- **Rejected proposals persist here** — they are NOT deleted. Future skill updates can
consult this log to avoid repeating failed approaches (the WikiSkill paper showed this
is critical for effective evolution).
- **Neutral changes** (no improvement, no regression) are recorded as `accepted` with
`validation: manual` — they may enable future improvements.
- **Rollbacks** are explicitly tracked — if a skill patch caused regressions and was
reverted, the entry stays with `outcome: rolled-back`.
+58
View File
@@ -0,0 +1,58 @@
---
pattern_id: PAT-002
title: "Traefik `reload` unreliable after conf.d edits — use `restart`"
category: infrastructure
severity: medium
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: [proxmox-cluster, rke2-kubernetes]
related_solution_docs:
- docs/solutions/bug-fixes/2026-07-26-paperless-ingressroute-hostname-fix.md
related_skills: []
---
# Traefik `reload` unreliable after conf.d edits — use `restart`
## Symptom
After modifying Traefik configuration files (especially dynamic config in `conf.d/`),
issuing a `reload` signal (SIGHUP) does not reliably pick up the changes. The old
configuration remains active, leading to stale ingress routes, incorrect routing,
or 404 errors.
## Root Cause
Traefik's hot-reload mechanism for file-based dynamic configuration can silently fail to
detect changes, especially when:
- Files are edited in-place (atomic rename not used)
- The file watcher misses events on certain filesystems (e.g., overlayfs, NFS)
- rsync is used without `--inplace` (creates temp file + rename, which the watcher may miss)
## Mitigation
Use `restart` instead of `reload` for Traefik container/service after conf.d edits:
```bash
# Instead of: docker kill -s HUP traefik (or systemctl reload traefik)
# Use:
docker compose restart traefik
# or: systemctl restart traefik
```
Additionally, when syncing config files via rsync, use `--inplace` to avoid
temp-file-rename patterns that confuse file watchers:
```bash
rsync --inplace -av ./conf.d/ /etc/traefik/conf.d/
```
## Prevention
- Always use `restart` (not `reload`) after Traefik config changes in the Traefik CT (99999)
- Use `rsync --inplace` when pushing config files to the Traefik host
- Document this in deployment runbooks
## Evidence
- Observed during Paperless v3 IngressRoute hostname fix (Jul 2026)
- Traefik runs in CT99999, conf.d directory
- MEMORY.md entry: "Traefik CT99999: `reload` unreliable after conf.d edits. Use `restart` or `rsync --inplace`."
+71
View File
@@ -0,0 +1,71 @@
---
pattern_id: PAT-004
title: "VFIO GPU passthrough fails — missing softdep + incomplete DRM blacklist"
category: infrastructure
severity: high
status: active
first_observed: 2026-07
last_updated: 2026-08-30
related_systems: [proxmox-cluster, rke2-kubernetes]
related_solution_docs:
- docs/solutions/bug-fixes/2026-07-24-vfio-pci-module-loading-race-condition.md
- docs/solutions/bug-fixes/2026-07-21-amd-gpu-passthrough-rombar.md
related_skills: []
---
# VFIO GPU passthrough fails — missing softdep + incomplete DRM blacklist
## Symptom
GPU passthrough to a VM fails intermittently or consistently. Symptoms include:
- `/dev/dri/renderD128` missing in the guest
- `amdgpu` driver not loading in guest
- Kernel BUG in host dmesg
- GPU device visible in `lspci` but not bound to `vfio-pci`
## Root Cause
Two intertwined issues:
1. **Module loading race condition**: Without `softdep amdgpu pre: vfio-pci`, the amdgpu
driver races to bind the GPU before vfio-pci can claim it. Normal bind time ~2.4s;
racing bind time ~302s (or never succeeds).
2. **Incomplete DRM blacklist**: If `drm` and `drm_kms_helper` are not blacklisted, the
kernel's DRM subsystem grabs the GPU before vfio-pci, preventing passthrough.
## Mitigation
Apply to the PVE host's modprobe config:
```bash
# /etc/modprobe.d/blacklist-drm.conf
blacklist drm
blacklist drm_kms_helper
# /etc/modprobe.d/amdgpu-vfio.conf
softdep amdgpu pre: vfio-pci
```
Then rebuild initramfs and reboot:
```bash
update-initramfs -u -k all
reboot
```
Also ensure `rombar=1` in the VM config for AMD GPUs (PCI ID 1002:xxxx):
```
# /etc/pve/qemu-server/<VMID>.conf
hostpci0: 0000:XX:YY.Z,pcie=1,rombar=1,x-vga=1
```
## Prevention
- Always configure `softdep` + DRM blacklist BEFORE attempting GPU passthrough on a new node
- Add to provisioning scripts (Ansible/Tofu) for any node with AMD GPU intended for passthrough
- ROCm 7.x: gfx1150 supported, gfx1036 NOT supported — verify GPU compute ISA before assignment
## Evidence
- Observed on ms-a2-1 (race condition, Jul 2026) and worker-05/VM 102 (kernel BUG, rombar)
- Solution docs: `2026-07-24-vfio-pci-module-loading-race-condition.md`,
`2026-07-21-amd-gpu-passthrough-rombar.md`
- MEMORY.md entry: "GPU:2 active—Immich(CT111/n5pro gfx1150)+Frigate(CT120/px7). ROCm7.14=gfx1150 supported,gfx1036 NOT."