Initial commit: Hermes Agent Skills collection

This commit is contained in:
Debian
2026-07-12 19:02:59 +00:00
commit e9cc106625
789 changed files with 233126 additions and 0 deletions
@@ -0,0 +1,392 @@
# Infinite Refill Scraper v4 (Pattern B)
## When to Use
- Target is >10,000 recipes (batch seeds run out)
- Continuous background operation without daily cronjob intervention
- Resilient against search exhaustion (alphabet/categories eventually cover everything)
## Architecture
```
Infinite Seed Generator URL Discovery (Playwright) Recipe Extraction (Requests)
│ │ │
▼ ▼ ▼
[Alphabet combos] ──► Search pages ──► new URLs ──► JSON-LD parse ──► real_recipes.jsonl
[Categories+rand] ──► (5 workers) ──► (queue) ──► (fast, no JS) ──► append-only
[Random 3-letter] ──► ──► dedup ──► ──► crash-safe
```
## Seed Generator (Python)
Cycles through 4 strategies infinitely, never running out:
```python
ALPHABET = "abcdefghijklmnopqrstuvwxyzäöü"
CATEGORIES = [
"Rezepte", "Hauptspeise", "Vorspeise", "Nachtisch", "Beilage",
"Fruehstueck", "Backen", "Grillen", "Salat", "Suppe", "Eintopf",
"Pasta", "Dessert", "Kuchen", "Brot", "Getraenk", "Fingerfood",
]
def _next_seeds(batch_size=200, cycle_counter):
seeds = []
while len(seeds) < batch_size:
cycle = cycle_counter % 4
if cycle == 0:
# Alphabetic 2-letter combos (e.g. "ab", "ac" ...):
a = ALPHABET[cycle_counter % len(ALPHABET)]
b = ALPHABET[(cycle_counter // len(ALPHABET)) % len(ALPHABET)]
term = a + b
elif cycle == 1:
# Category + random letter (e.g. "Pasta+a")
cat = CATEGORIES[cycle_counter % len(CATEGORIES)]
letter = random.choice(ALPHABET)
term = cat + "+" + letter
elif cycle == 2:
# Just category (e.g. "Pasta")
term = CATEGORIES[cycle_counter % len(CATEGORIES)]
else:
# Random 3-letter nonsense (e.g. "xyz")
term = "".join(random.choices(ALPHABET, k=3))
seeds.append(f"https://www.chefkoch.de/rs/s0/{term}/Rezepte.html")
cycle_counter += 1
return seeds, cycle_counter
```
**Coverage estimate:** 29 letters × 29 letters = 841 combos for 2-letter alone. Plus 17 categories × 29 = 493. Plus pure category = 17. Plus random 3-letter = infinite. Total: effectively unlimited.
## Two-Phase Extraction
### Phase 1: Link Discovery
```python
async def _fetch_links(page, url, timeout=30):
await page.goto(url, wait_until="domcontentloaded", timeout=timeout*1000)
await asyncio.sleep(1.5) # Let JS render
links = await page.eval_on_selector_all(
"a[href*='/rezepte/']",
"els => els.map(e => e.href)"
)
# Deduplicate and filter
seen = set()
result = []
for l in links:
if "chefkoch.de/rezepte/" in l and l not in seen:
seen.add(l)
result.append(l)
return result
```
### Phase 2: Recipe Extraction (NO Playwright)
```python
import urllib.request, re, json
async def _scrape_one(url, timeout=15):
req = urllib.request.Request(url, headers={
"User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 ...",
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": "de-DE,de;q=0.9,en;q=0.8",
})
with urllib.request.urlopen(req, timeout=timeout) as resp:
html = resp.read().decode("utf-8", errors="ignore")
blocks = re.findall(
r'<script type="application/ld\+json">(.*?)</script>',
html, re.DOTALL | re.IGNORECASE
)
for block in blocks:
data = json.loads(block.strip())
if isinstance(data, list):
for item in data:
if isinstance(item, dict) and item.get("@type") == "Recipe":
return _normalize_recipe(item, url)
elif isinstance(data, dict) and data.get("@type") == "Recipe":
return _normalize_recipe(data, url)
# Fallback: look for mainEntity
for block in blocks:
data = json.loads(block.strip())
if isinstance(data, dict) and "mainEntity" in data:
me = data["mainEntity"]
if isinstance(me, dict) and me.get("@type") == "Recipe":
return _normalize_recipe(me, url)
return None
```
**Why not use Playwright for recipe pages?** JSON-LD is embedded in raw HTML and extracts faster (no browser overhead). Playwright is only needed for search pages because Chefkoch serves a JS shell there.
## State Management
```python
STATE = {
"seed_cycle": 0, # Position in infinite sequence
"started_at": "", # ISO timestamp
"recipes_target": 25000
}
```
Save every batch. On restart, the script resumes from the last seed_cycle.
## Normalization Function
```python
def _normalize_recipe(data, url):
rid = (data.get("identifier", {}).get("value", "")
or url.split("/")[-1].split("-")[0])
title = data.get("name", "")
if not title:
return None
keywords = data.get("keywords", "")
tags = [t.strip().lower() for t in keywords.split(",") if t.strip()]
recipe_cat = data.get("recipeCategory", "")
cuisine = data.get("recipeCuisine", "")
if recipe_cat:
tags.append(recipe_cat.lower())
if cuisine:
tags.append(cuisine.lower())
tags = list(set(tags))
rating = data.get("aggregateRating", {})
nutrition = data.get("nutrition", {})
return {
"id": str(rid),
"title": title,
"description": data.get("description", ""),
"source_url": url,
"ingredients": data.get("recipeIngredient", []),
"instructions": data.get("recipeInstructions", ""),
"tags": tags,
"category": recipe_cat or (tags[0] if tags else "Sonstiges"),
"rating_value": rating.get("ratingValue"),
"rating_count": rating.get("ratingCount"),
"prep_time": data.get("prepTime", ""),
"cook_time": data.get("cookTime", ""),
"total_time": data.get("totalTime", ""),
"calories": str(nutrition.get("calories", "")),
"scraped_at": time.strftime("%Y-%m-%dT%H:%M:%S"),
}
```
## Execution Pattern
### DO NOT use LLM-based cronjob
**WRONG:** A cronjob that uses `hermes` LLM agent to run Python. Fails with:
- `Error 400: model is required` (if no model configured)
- `No Anthropic credentials` (if wrong provider configured)
- Unnecessary cost: paying LLM tokens just to run `wc -l` and `python script.py`
**CORRECT:** `terminal(background=True)` with `notify_on_complete`
```python
# Start directly
cd /profile/nutrition-coach && \
source .venv/bin/activate && \
python3 /path/to/chefkoch_scraper_v4.py
# In Hermes Agent:
# terminal(background=True, notify_on_complete=True)
```
### Why background process beats cron-checked scraper
| Aspect | Cronjob (daily check + start) | Background process |
|--------|-------------------------------|-------------------|
| Failure mode | LLM errors, credential mismatches | None (pure Python) |
| Latency | Up to 24h until next check | Continuous |
| State tracking | External (cron state DB) | Internal (scraper_state.json) |
| Backpressure | Manual | Automatic (sleep on queue overflow) |
| Monitoring | Cron output files | Log tail + notify_on_complete |
| Cost | LLM tokens per check | Zero |
## Performance Benchmarks
| Configuration | Recipes/min | 5,000 recipes | 25,000 recipes |
|--------------|-------------|---------------|----------------|
| 5 workers (batch, fixed seeds) | 6080 | ~6090 min | N/A (seeds exhausted) |
| Infinite v4 (single context) | 4060 | ~80120 min | ~710 hours |
| Infinite v4 (5 contexts, remote) | 80120 | ~4060 min | ~46 hours |
**Note:** v4 uses a single Playwright page context for discovery (lighter than 5 full workers). The extraction is HTTP requests which are much faster.*
## Known v4 Failure Modes
If the scraper is stuck at the exact same recipe count for multiple cron runs (check `scraper.log` — repeated `START run | current: N/N` with identical `N`), one of these three root causes is almost certainly present.
### Cause 1: `SEED_CYCLE` not declared global in `main()`
**Symptom:** `seed_cycle` in state JSON is `0` even after dozens of runs. The same exhausted 2-letter seeds are visited forever.
**Root cause:** In Python, a bare assignment `SEED_CYCLE = state["seed_cycle"]` inside `main()` creates a *local* variable if `global SEED_CYCLE` is not declared. The module-level global stays at `0`, so `_next_seeds()` always starts from `0`.
**Fix:** Add in `main()`:
```python
state.setdefault("seed_cycle", 0)
global SEED_CYCLE # ← required
SEED_CYCLE = state["seed_cycle"]
```
**Reproduce:** Open Python, import the module, inspect `chefkoch_scraper_v4.SEED_CYCLE` after `main()` runs — it remains `0` without the `global` line.
### Cause 2: Playwright EPIPE / `Executable doesn't exist`
**Symptom:** Log shows `Executable doesn't exist at ...chrome-linux/chrome` or Node `EPIPE` errors. The process exits immediately.
**Root cause:** Environment variable `PLAYWRIGHT_BROWSERS_PATH` is set in the caller's shell but not propagated to the background process spawned by `terminal(background=True)`. The agent's background launcher creates a new non-login bash that inherits `PATH` but not the caller's env overrides.
**Fix:** Hardcode the browser path inside the script itself:
```python
import os
os.environ.setdefault("PLAYWRIGHT_BROWSERS_PATH", "/home/.../.cache/ms-playwright")
```
Do **not** rely on shell `export` when launching via `terminal(background=True)`.
### Cause 3: Stale `completed_seeds` blockers in state
**Symptom:** Scraper runs but repeatedly visits the same search pages with no new URLs, or many search pages return zero links.
**Root cause:** `scraper_state.json` contains a `completed_seeds` dict that prevents revisiting certain terms. If these terms were exhausted naturally, the seeds never advance.
**Fix:** Wipe or prune the dict before a fresh run:
```python
import json
s = json.load(open("scraper_state.json"))
s.pop("completed_seeds", None)
s["seed_cycle"] = max(s.get("seed_cycle", 0), 1000)
json.dump(s, open("scraper_state.json", "w"), ensure_ascii=False)
```
### Quick diagnostic flow
```bash
# 1. Check recipe count isn't moving
wc -l ~/.hermes/profiles/nutrition-coach/real_recipes.jsonl
# 2. Check state
python3 -c "import json; s=json.load(open('scraper_state.json')); print('seed_cycle:', s.get('seed_cycle')); print('completed_seeds keys:', list(s.get('completed_seeds',{}).keys())[:5])"
# 3. Check Playwright browser
find /home/debian/.cache/ms-playwright -name chrome | head -1
# 4. Check if process is actually scraping (not just discovering)
ps auxf | grep -E 'chefkoch|chrome' | grep -v grep | wc -l
```
See `references/v4-known-issues-and-fixes.md` for the full patch snippets and state-reset script.
## File Locations (standard)
```
~/.hermes/profiles/nutrition-coach/
├── real_recipes.jsonl # Append-only output
├── scraper_state.json # Resume state
├── scraper.log # Human-readable log
└── scripts/
└── chefkoch_scraper_v4.py # Standalone script
```
## Monitoring
```bash
# Check progress
tail -n 5 ~/.hermes/profiles/nutrition-coach/scraper.log
wc -l ~/.hermes/profiles/nutrition-coach/real_recipes.jsonl
# Check if running
ps aux | grep chefkoch_scraper
# Restart if needed
cd ~/.hermes/profiles/nutrition-coach && \
source .venv/bin/activate && \
nohup python3 scripts/chefkoch_scraper_v4.py > scraper.log 2>&1 &
```
## Error Recovery
```
Error streak >= 100 → Script aborts automatically
→ Check Chefkoch availability: curl -I https://www.chefkoch.de
→ Verify Playwright install: playwright install chromium
→ Check disk space: df -h
→ Restart with fresh state (or resume from last saved seed_cycle)
```
## Production Deployment
### Browser path detection (common error)
If Playwright fails with `Executable doesn't exist at ...chrome-headless-shell`:
```bash
# Find existing browser installations
find /home/debian -name "chrome" -o -name "chromium" 2>/dev/null
# Usually found under /home/debian/.cache/ms-playwright/
# Set the environment variable before starting:
export PLAYWRIGHT_BROWSERS_PATH=/home/debian/.cache/ms-playwright
```
If the browser is not installed at all:
```bash
cd /profile/nutrition-coach && source .venv/bin/activate
playwright install chromium
```
**Do NOT rely on `playwright install` in the crontab — install once, reference via env var.**
### nohup wrapper for daemon mode
Create `~/.hermes/profiles/nutrition-coach/scripts/start_scraper_v4.sh`:
```bash
#!/bin/bash
LOG="/home/debian/.hermes/profiles/nutrition-coach/scraper_v4.out"
PIDFILE="/tmp/chefkoch_scraper_v4.pid"
if [ -f "$PIDFILE" ] && kill -0 "$(cat $PIDFILE)" 2>/dev/null; then
echo "Scraper already running (PID $(cat $PIDFILE))"; exit 0
fi
cd /home/debian/.hermes/profiles/nutrition-coach || exit 1
export PYTHONUNBUFFERED=1
export PLAYWRIGHT_BROWSERS_PATH=/home/debian/.cache/ms-playwright
nohup python3 scripts/chefkoch_scraper_v4.py --no-agent > "$LOG" 2>&1 &
echo $! > "$PIDFILE"
echo "Scraper started (PID $!)"
```
Then: `chmod +x scripts/start_scraper_v4.sh`
**Path note:** In this environment the script lives under `scripts/chefkoch_scraper_v4.py` inside the profile directory, not `~/.hermes/scripts/`. Always check `ls scripts/` before constructing the invocation path.
### Watchdog cronjob (no_agent)
```bash
cronjob action=create \
name=chefkoch-scraper-watchdog \
script=start_scraper_v4.sh \
schedule="*/10 * * * *" \
no_agent=true \
deliver=local
```
This restarts the scraper if it crashes, without involving any LLM.
### DO NOT use `terminal(background=true)` with `tee` or complex pipelines
`terminal(background=true, command="... | tee log.out")` — Output buffering causes `tee` to swallow data, and the `terminal` pseudo-tty layer conflicts with Playwright's subprocess management. Result: log file remains empty or stale, process output invisible.
✅ Start via explicit file redirect with `PYTHONUNBUFFERED=1`:
```python
# In Hermes Agent — background process with file redirect
terminal(
background=True,
command="cd /profile && PYTHONUNBUFFERED=1 python3 scripts/chefkoch_scraper_v4.py --no-agent >> scraper_v4.out 2>&1"
)
```
Or use the `start_scraper_v4.sh` wrapper script shown above. Always verify progress with `wc -l real_recipes.jsonl` and `tail scraper_v4.out` rather than relying solely on process output capture.