8.1 KiB
v4 Known Issues & Fixes
Session-record of bugs found and fixed during production Chefkoch scraping runs.
This file is referenced from the main infinite-refill-scraper-v4.md reference.
Environment: Context for these fixes
- Host: Debian container, Python 3.11
- Playwright 1.60.0 installed in profile venv
- Browser cache:
/home/debian/.cache/ms-playwright/chromium-1223/ - Background process launcher: Hermes Agent
terminal(background=True) - Cron:
chefkoch-scraper-watchdogevery 10 minutes (no_agent=true)
Fix 1 — Playwright browser path not inherited
Failure signature
playwright._impl._api_types.Error: Executable doesn't exist at
/home/debian/.cache/ms-playwright/chromium-1071/chrome-linux/chrome
Note: The error references a different chromium version (-1071) than what is actually installed (-1223). This mismatch happens because Playwright falls back to its own lookup logic when PLAYWRIGHT_BROWSERS_PATH is absent.
Why it happens
Setting the env var in the Hermes Agent shell workspace does not propagate to the background bash launched by terminal(background=True). The background process inherits PATH and VIRTUAL_ENV but not custom export FOO=bar values.
Fix options (ranked)
-
Best: Hardcode inside
chefkoch_scraper_v4.pybefore any Playwright import:import os os.environ.setdefault("PLAYWRIGHT_BROWSERS_PATH", "/home/debian/.cache/ms-playwright") -
Acceptable: Shell wrapper with
execand explicitexport:#! /bin/bash export PLAYWRIGHT_BROWSERS_PATH=/home/debian/.cache/ms-playwright exec python3 scripts/chefkoch_scraper_v4.py --no-agent >> scraper_v4.out 2>&1 -
Do not: Rely on cron env or parent-shell export alone.
Fix 2 — global SEED_CYCLE missing
Failure signature
Scraper runs every 10 minutes (cron restart) but recipe count never increases. Log shows:
START run | current: 12736/25000
State file seed_cycle is always 0.
Why it happens
Simplified pseudocode:
SEED_CYCLE = 0 # module-level global
def _next_seeds(batch_size=50):
global SEED_CYCLE
for i in range(batch_size):
# ... compute seed based on SEED_CYCLE ...
SEED_CYCLE += 1
async def main():
state = _load_state()
SEED_CYCLE = state["seed_cycle"] # ← BUG: creates a local variable in main()
while ...:
seeds = _next_seeds(50) # reads/writes the GLOBAL SEED_CYCLE
# but the global is still 0 because main()'s assignment was local!
Because SEED_CYCLE is assigned inside main() without global SEED_CYCLE, Python creates a local variable shadowing the module global. _next_seeds() then reads the global (still 0), increments it, but on the next cron restart main() sets the local back to whatever is in state (0). The seeds are the same every single run.
Fix
async def main():
state = _load_state()
state.setdefault("seed_cycle", 0)
global SEED_CYCLE # ← required
SEED_CYCLE = state["seed_cycle"]
Verification
python3 -c "
import chefkoch_scraper_v4 as m
m.main() # or simulate the assignment
print('global SEED_CYCLE after main:', m.SEED_CYCLE)
"
Without global it prints 0. With global it prints whatever was in state.
Fix 3 — Stale completed_seeds locking the scraper
Failure signature
Scraper starts, opens browser, visits search pages, but new_urls is always empty because all returned URLs are already in existing. No new recipes are appended.
Why it happens
scraper_state.json accumulated a completed_seeds dict during a multi-worker batch run. The v4 infinite-refill script does not use completed_seeds, but the state file still carries it from an older script version. Certain seed terms are effectively blacklisted.
Fix
import json
s = json.load(open("scraper_state.json"))
s.pop("completed_seeds", None)
# Advance seed_cycle past known-exhausted 2-letter alphabet seeds
s["seed_cycle"] = max(s.get("seed_cycle", 0), 1000)
json.dump(s, open("scraper_state.json", "w"), ensure_ascii=False, indent=2)
Setting seed_cycle >= 1000 pushes the generator into the category+letter and category-only strategies, which yield far more fresh URLs than the aa, ab, ac alphabet exhaustion zone.
Fix 4 — Dual script / dual Python trap (watchdog silently fails)
Failure signature
Scraper cron log shows Scraper started (PID NNN) every 10 minutes. No errors ever appear. scraper_v4.out only contains Starting: X/Y line (the print() from main() before Playwright is used). Recipe count stays flat for days or weeks.
Why it happens
The watchdog script (~/.hermes/scripts/start_scraper_v4.sh) uses:
- Global
python3→ Playwright 1.59.0 (looks forchromium-1071, which does not exist in this environment). - Global script (
~/.hermes/scripts/chefkoch_scraper_v4.py) → missingglobal SEED_CYCLEinmain().
Both of these individually produce crashes (Fix 1 + Fix 2 above). Combined, the global python immediately crashes on browser launch, outputting nothing useful to the log. The watchdog sees the PID exit, and restarts it again 10 minutes later. This loop repeats forever with zero recipes.
Meanwhile, a working copy exists at:
~/.hermes/profiles/nutrition-coach/.venv/bin/python3(Playwright 1.60.0+, matches installed browsers)~/.hermes/profiles/nutrition-coach/scripts/chefkoch_scraper_v4.py(hasglobal SEED_CYCLE)
Both copies look very similar from a distance, making it easy to invoke the wrong pair.
Fix
Replace the watchdog shell script with the version that explicitly points to the profile-local python and script:
nohup /home/debian/.hermes/profiles/nutrition-coach/.venv/bin/python3 \
/home/debian/.hermes/profiles/nutrition-coach/scripts/chefkoch_scraper_v4.py \
--no-agent >> "$LOG" 2>&1 &
Also verify there is no stale PID file (otherwise the wrapper exits early assuming it's already running):
cat /tmp/chefkoch_scraper_v4.pid 2>/dev/null && (kill -0 "$(cat /tmp/chefkoch_scraper_v4.pid)" 2>/dev/null || rm -f /tmp/chefkoch_scraper_v4.pid)
Detection
# 1. Check if log ever got past the first line
wc -l ~/.hermes/profiles/nutrition-coach/scraper_v4.out
# Unhealthy: 1 line forever
# 2. Check which python the startup script actually calls
grep 'python3' ~/.hermes/scripts/start_scraper_v4.sh
# Unhealthy: just "python3 ..." (no venv path)
# 3. Check which script path it points to
grep 'chefkoch_scraper' ~/.hermes/scripts/start_scraper_v4.sh
# Unhealthy: ~/.hermes/scripts/chefkoch_scraper_v4.py
Recovery after fix
After updating the watchdog script:
# 1. Kill any stale scraper
pkill -f chefkoch_scraper_v4.py || true
rm -f /tmp/chefkoch_scraper_v4.pid
# 2. Advance state past exhaustion
python3 -c "
import json, os
p = os.path.expanduser('~/.hermes/profiles/nutrition-coach/scraper_state.json')
s = json.load(open(p))
s.pop('completed_seeds', None)
s['seed_cycle'] = max(s.get('seed_cycle', 0), 2000)
json.dump(s, open(p, 'w'), ensure_ascii=False, indent=2)
print('State reset — seed_cycle now', s['seed_cycle'])
"
# 3. Start with fixed script
bash ~/.hermes/scripts/start_scraper_v4.sh
# 4. Verify within 2 minutes
sleep 120 && wc -l ~/.hermes/profiles/nutrition-coach/real_recipes.jsonl
One-liner state reset (copy-paste)
python3 -c "
import json, os
p = os.path.expanduser('~/.hermes/profiles/nutrition-coach/scraper_state.json')
s = json.load(open(p))
s.pop('completed_seeds', None)
s['seed_cycle'] = max(s.get('seed_cycle',0), 1000)
json.dump(s, open(p,'w'), ensure_ascii=False, indent=2)
print('Reset OK — seed_cycle now', s['seed_cycle'])
"
Monitoring checklist (what to watch)
| Check | Command | Healthy / Unhealthy |
|---|---|---|
| Recipe count moving | wc -l real_recipes.jsonl |
Increases between checks |
| Process alive | ps auxf | grep chefkoch_scraper |
chromium + python present |
| State advancing | python3 -c "... seed_cycle ..." |
value increases over time |
| Log not looping | tail -n 5 scraper_v4.out |
NOT just "Starting: X/Y" only |
| Disk space | df -h . |
> 1 GB free |