Files
hermes-skills/food-nutrition/meal-planning/references/multi-source-recipe-scraper/infinite-refill-scraper-v4.md
T

14 KiB
Raw Blame History

Infinite Refill Scraper v4 (Pattern B)

When to Use

  • Target is >10,000 recipes (batch seeds run out)
  • Continuous background operation without daily cronjob intervention
  • Resilient against search exhaustion (alphabet/categories eventually cover everything)

Architecture

Infinite Seed Generator     URL Discovery (Playwright)     Recipe Extraction (Requests)
     │                              │                                    │
     ▼                              ▼                                    ▼
[Alphabet combos]  ──►  Search pages  ──►  new URLs  ──►  JSON-LD parse  ──►  real_recipes.jsonl
[Categories+rand]  ──►  (5 workers)  ──►  (queue)    ──►  (fast, no JS)  ──►  append-only
[Random 3-letter]  ──►              ──►  dedup       ──►                ──►  crash-safe

Seed Generator (Python)

Cycles through 4 strategies infinitely, never running out:

ALPHABET = "abcdefghijklmnopqrstuvwxyzäöü"
CATEGORIES = [
    "Rezepte", "Hauptspeise", "Vorspeise", "Nachtisch", "Beilage",
    "Fruehstueck", "Backen", "Grillen", "Salat", "Suppe", "Eintopf",
    "Pasta", "Dessert", "Kuchen", "Brot", "Getraenk", "Fingerfood",
]

def _next_seeds(batch_size=200, cycle_counter):
    seeds = []
    while len(seeds) < batch_size:
        cycle = cycle_counter % 4
        if cycle == 0:
            # Alphabetic 2-letter combos (e.g. "ab", "ac" ...):
            a = ALPHABET[cycle_counter % len(ALPHABET)]
            b = ALPHABET[(cycle_counter // len(ALPHABET)) % len(ALPHABET)]
            term = a + b
        elif cycle == 1:
            # Category + random letter (e.g. "Pasta+a")
            cat = CATEGORIES[cycle_counter % len(CATEGORIES)]
            letter = random.choice(ALPHABET)
            term = cat + "+" + letter
        elif cycle == 2:
            # Just category (e.g. "Pasta")
            term = CATEGORIES[cycle_counter % len(CATEGORIES)]
        else:
            # Random 3-letter nonsense (e.g. "xyz")
            term = "".join(random.choices(ALPHABET, k=3))
        seeds.append(f"https://www.chefkoch.de/rs/s0/{term}/Rezepte.html")
        cycle_counter += 1
    return seeds, cycle_counter

Coverage estimate: 29 letters × 29 letters = 841 combos for 2-letter alone. Plus 17 categories × 29 = 493. Plus pure category = 17. Plus random 3-letter = infinite. Total: effectively unlimited.

Two-Phase Extraction

async def _fetch_links(page, url, timeout=30):
    await page.goto(url, wait_until="domcontentloaded", timeout=timeout*1000)
    await asyncio.sleep(1.5)  # Let JS render
    links = await page.eval_on_selector_all(
        "a[href*='/rezepte/']",
        "els => els.map(e => e.href)"
    )
    # Deduplicate and filter
    seen = set()
    result = []
    for l in links:
        if "chefkoch.de/rezepte/" in l and l not in seen:
            seen.add(l)
            result.append(l)
    return result

Phase 2: Recipe Extraction (NO Playwright)

import urllib.request, re, json

async def _scrape_one(url, timeout=15):
    req = urllib.request.Request(url, headers={
        "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 ...",
        "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
        "Accept-Language": "de-DE,de;q=0.9,en;q=0.8",
    })
    with urllib.request.urlopen(req, timeout=timeout) as resp:
        html = resp.read().decode("utf-8", errors="ignore")
    
    blocks = re.findall(
        r'<script type="application/ld\+json">(.*?)</script>',
        html, re.DOTALL | re.IGNORECASE
    )
    for block in blocks:
        data = json.loads(block.strip())
        if isinstance(data, list):
            for item in data:
                if isinstance(item, dict) and item.get("@type") == "Recipe":
                    return _normalize_recipe(item, url)
        elif isinstance(data, dict) and data.get("@type") == "Recipe":
            return _normalize_recipe(data, url)
    # Fallback: look for mainEntity
    for block in blocks:
        data = json.loads(block.strip())
        if isinstance(data, dict) and "mainEntity" in data:
            me = data["mainEntity"]
            if isinstance(me, dict) and me.get("@type") == "Recipe":
                return _normalize_recipe(me, url)
    return None

Why not use Playwright for recipe pages? JSON-LD is embedded in raw HTML and extracts faster (no browser overhead). Playwright is only needed for search pages because Chefkoch serves a JS shell there.

State Management

STATE = {
    "seed_cycle": 0,        # Position in infinite sequence
    "started_at": "",       # ISO timestamp
    "recipes_target": 25000
}

Save every batch. On restart, the script resumes from the last seed_cycle.

Normalization Function

def _normalize_recipe(data, url):
    rid = (data.get("identifier", {}).get("value", "") 
           or url.split("/")[-1].split("-")[0])
    title = data.get("name", "")
    if not title:
        return None
    
    keywords = data.get("keywords", "")
    tags = [t.strip().lower() for t in keywords.split(",") if t.strip()]
    recipe_cat = data.get("recipeCategory", "")
    cuisine = data.get("recipeCuisine", "")
    if recipe_cat:
        tags.append(recipe_cat.lower())
    if cuisine:
        tags.append(cuisine.lower())
    tags = list(set(tags))
    
    rating = data.get("aggregateRating", {})
    nutrition = data.get("nutrition", {})
    
    return {
        "id": str(rid),
        "title": title,
        "description": data.get("description", ""),
        "source_url": url,
        "ingredients": data.get("recipeIngredient", []),
        "instructions": data.get("recipeInstructions", ""),
        "tags": tags,
        "category": recipe_cat or (tags[0] if tags else "Sonstiges"),
        "rating_value": rating.get("ratingValue"),
        "rating_count": rating.get("ratingCount"),
        "prep_time": data.get("prepTime", ""),
        "cook_time": data.get("cookTime", ""),
        "total_time": data.get("totalTime", ""),
        "calories": str(nutrition.get("calories", "")),
        "scraped_at": time.strftime("%Y-%m-%dT%H:%M:%S"),
    }

Execution Pattern

DO NOT use LLM-based cronjob

WRONG: A cronjob that uses hermes LLM agent to run Python. Fails with:

  • Error 400: model is required (if no model configured)
  • No Anthropic credentials (if wrong provider configured)
  • Unnecessary cost: paying LLM tokens just to run wc -l and python script.py

CORRECT: terminal(background=True) with notify_on_complete

# Start directly
cd /profile/nutrition-coach && \
  source .venv/bin/activate && \
  python3 /path/to/chefkoch_scraper_v4.py

# In Hermes Agent:
# terminal(background=True, notify_on_complete=True)

Why background process beats cron-checked scraper

Aspect Cronjob (daily check + start) Background process
Failure mode LLM errors, credential mismatches None (pure Python)
Latency Up to 24h until next check Continuous
State tracking External (cron state DB) Internal (scraper_state.json)
Backpressure Manual Automatic (sleep on queue overflow)
Monitoring Cron output files Log tail + notify_on_complete
Cost LLM tokens per check Zero

Performance Benchmarks

Configuration Recipes/min 5,000 recipes 25,000 recipes
5 workers (batch, fixed seeds) 6080 ~6090 min N/A (seeds exhausted)
Infinite v4 (single context) 4060 ~80120 min ~710 hours
Infinite v4 (5 contexts, remote) 80120 ~4060 min ~46 hours

Note: v4 uses a single Playwright page context for discovery (lighter than 5 full workers). The extraction is HTTP requests which are much faster.*

Known v4 Failure Modes

If the scraper is stuck at the exact same recipe count for multiple cron runs (check scraper.log — repeated START run | current: N/N with identical N), one of these three root causes is almost certainly present.

Cause 1: SEED_CYCLE not declared global in main()

Symptom: seed_cycle in state JSON is 0 even after dozens of runs. The same exhausted 2-letter seeds are visited forever.

Root cause: In Python, a bare assignment SEED_CYCLE = state["seed_cycle"] inside main() creates a local variable if global SEED_CYCLE is not declared. The module-level global stays at 0, so _next_seeds() always starts from 0.

Fix: Add in main():

state.setdefault("seed_cycle", 0)
global SEED_CYCLE        # ← required
SEED_CYCLE = state["seed_cycle"]

Reproduce: Open Python, import the module, inspect chefkoch_scraper_v4.SEED_CYCLE after main() runs — it remains 0 without the global line.

Cause 2: Playwright EPIPE / Executable doesn't exist

Symptom: Log shows Executable doesn't exist at ...chrome-linux/chrome or Node EPIPE errors. The process exits immediately.

Root cause: Environment variable PLAYWRIGHT_BROWSERS_PATH is set in the caller's shell but not propagated to the background process spawned by terminal(background=True). The agent's background launcher creates a new non-login bash that inherits PATH but not the caller's env overrides.

Fix: Hardcode the browser path inside the script itself:

import os
os.environ.setdefault("PLAYWRIGHT_BROWSERS_PATH", "/home/.../.cache/ms-playwright")

Do not rely on shell export when launching via terminal(background=True).

Cause 3: Stale completed_seeds blockers in state

Symptom: Scraper runs but repeatedly visits the same search pages with no new URLs, or many search pages return zero links.

Root cause: scraper_state.json contains a completed_seeds dict that prevents revisiting certain terms. If these terms were exhausted naturally, the seeds never advance.

Fix: Wipe or prune the dict before a fresh run:

import json
s = json.load(open("scraper_state.json"))
s.pop("completed_seeds", None)
s["seed_cycle"] = max(s.get("seed_cycle", 0), 1000)
json.dump(s, open("scraper_state.json", "w"), ensure_ascii=False)

Quick diagnostic flow

# 1. Check recipe count isn't moving
wc -l ~/.hermes/profiles/nutrition-coach/real_recipes.jsonl

# 2. Check state
python3 -c "import json; s=json.load(open('scraper_state.json')); print('seed_cycle:', s.get('seed_cycle')); print('completed_seeds keys:', list(s.get('completed_seeds',{}).keys())[:5])"

# 3. Check Playwright browser
find /home/debian/.cache/ms-playwright -name chrome | head -1

# 4. Check if process is actually scraping (not just discovering)
ps auxf | grep -E 'chefkoch|chrome' | grep -v grep | wc -l

See references/v4-known-issues-and-fixes.md for the full patch snippets and state-reset script.

File Locations (standard)

~/.hermes/profiles/nutrition-coach/
├── real_recipes.jsonl          # Append-only output
├── scraper_state.json          # Resume state
├── scraper.log                 # Human-readable log
└── scripts/
    └── chefkoch_scraper_v4.py  # Standalone script

Monitoring

# Check progress
tail -n 5 ~/.hermes/profiles/nutrition-coach/scraper.log
wc -l ~/.hermes/profiles/nutrition-coach/real_recipes.jsonl

# Check if running
ps aux | grep chefkoch_scraper

# Restart if needed
cd ~/.hermes/profiles/nutrition-coach && \
  source .venv/bin/activate && \
  nohup python3 scripts/chefkoch_scraper_v4.py > scraper.log 2>&1 &

Error Recovery

Error streak >= 100 → Script aborts automatically
→ Check Chefkoch availability: curl -I https://www.chefkoch.de
→ Verify Playwright install: playwright install chromium
→ Check disk space: df -h
→ Restart with fresh state (or resume from last saved seed_cycle)

Production Deployment

Browser path detection (common error)

If Playwright fails with Executable doesn't exist at ...chrome-headless-shell:

# Find existing browser installations
find /home/debian -name "chrome" -o -name "chromium" 2>/dev/null

# Usually found under /home/debian/.cache/ms-playwright/
# Set the environment variable before starting:
export PLAYWRIGHT_BROWSERS_PATH=/home/debian/.cache/ms-playwright

If the browser is not installed at all:

cd /profile/nutrition-coach && source .venv/bin/activate
playwright install chromium

Do NOT rely on playwright install in the crontab — install once, reference via env var.

nohup wrapper for daemon mode

Create ~/.hermes/profiles/nutrition-coach/scripts/start_scraper_v4.sh:

#!/bin/bash
LOG="/home/debian/.hermes/profiles/nutrition-coach/scraper_v4.out"
PIDFILE="/tmp/chefkoch_scraper_v4.pid"
if [ -f "$PIDFILE" ] && kill -0 "$(cat $PIDFILE)" 2>/dev/null; then
    echo "Scraper already running (PID $(cat $PIDFILE))"; exit 0
fi
cd /home/debian/.hermes/profiles/nutrition-coach || exit 1
export PYTHONUNBUFFERED=1
export PLAYWRIGHT_BROWSERS_PATH=/home/debian/.cache/ms-playwright
nohup python3 scripts/chefkoch_scraper_v4.py --no-agent > "$LOG" 2>&1 &
echo $! > "$PIDFILE"
echo "Scraper started (PID $!)"

Then: chmod +x scripts/start_scraper_v4.sh

Path note: In this environment the script lives under scripts/chefkoch_scraper_v4.py inside the profile directory, not ~/.hermes/scripts/. Always check ls scripts/ before constructing the invocation path.

Watchdog cronjob (no_agent)

cronjob action=create \
  name=chefkoch-scraper-watchdog \
  script=start_scraper_v4.sh \
  schedule="*/10 * * * *" \
  no_agent=true \
  deliver=local

This restarts the scraper if it crashes, without involving any LLM.

DO NOT use terminal(background=true) with tee or complex pipelines

terminal(background=true, command="... | tee log.out") — Output buffering causes tee to swallow data, and the terminal pseudo-tty layer conflicts with Playwright's subprocess management. Result: log file remains empty or stale, process output invisible.

Start via explicit file redirect with PYTHONUNBUFFERED=1:

# In Hermes Agent — background process with file redirect
terminal(
    background=True,
    command="cd /profile && PYTHONUNBUFFERED=1 python3 scripts/chefkoch_scraper_v4.py --no-agent >> scraper_v4.out 2>&1"
)

Or use the start_scraper_v4.sh wrapper script shown above. Always verify progress with wc -l real_recipes.jsonl and tail scraper_v4.out rather than relying solely on process output capture.