# Chefkoch Scraper Evolution: v4 Playwright → v5.2 API Session learning from 2026-06-22: migrating the Chefkoch scraper from Playwright-based discovery to internal API discovery. This is a **10× speed increase** and removes all headless-browser fragility. ## Why v4 Stalled at ~25k | Problem | Symptom | Root cause | |---------|---------|-----------| | Search exhaustion | Same seeds return same recipes | 2-letter alphabet combos cover finite space | | Bot detection | Empty search pages, no links | Chefkoch serves JS shell to non-browser clients | | Playwright fragility | `Executable doesn't exist`, EPIPE | Browser binary mismatches, env propagation issues | | Seed cycle reset | `seed_cycle` stuck at 0 | Missing `global SEED_CYCLE` in `main()` | | Cookie walls | Content blocked | GDPR consent banners block link extraction | | Rate ~0.4/s | 100 recipes takes ~4 min | Browser overhead per search page | The fundamental issue: **Chefkoch search pages require a real browser**, and the browser-based discovery approach has natural throughput limits and maintenance overhead. ## The API Discovery Path (v5.2) ### How we found it During debugging, a test request to `api.chefkoch.de/v2/recipes?offset=0&limit=30` returned structured JSON with **384,252 recipes** listed. The endpoint is the same one Chefkoch's own frontend uses — no auth, no rate limiting observed. ### Architecture comparison | Aspect | v4/v5.1 (Playwright) | v5.2 (API) | |--------|----------------------|------------| | Discovery | Browser-rendered search pages | HTTP listing API (100 recipes/page) | | Detail fetch | JSON-LD regex from HTML | Structured JSON from API | | Browser needed | Yes (Chromium binary) | No | | Max reachable | ~25k (search exhaustion) | ~384k (full catalog) | | Ingredients | Raw strings from JSON-LD | Objects with amount/unit/name | | Avg speed | ~0.4 recipes/sec | ~5 recipes/sec | | Headless issues | Yes (cookies, JS timing, EPIPE) | None | | Maintenance | Browser updates, path management | Minimal (HTTP only) | ## v5.2 Architecture (query-rotation based) Because the listing API has an offset ceiling (~2000 for unfiltered, ~1000 for search queries), v5.2 uses **query rotation** instead of simple offset pagination: ```python QUERIES = [ "Hauptspeise", "Vorspeise", "Nachtisch", "Beilage", "Fruehstueck", "Backen", "Grillen", "Salat", "Suppe", "Eintopf", "Pasta", "Dessert", "Kuchen", "Brot", "Getraenk", "Fingerfood", "Curry", "Wok", "Auflauf", "Gratin", "Risotto", "Lasagne", "Lachs", "Haehnchen", "Rind", "Schwein", "Fisch", "Gemuese", "Kartoffeln", "Reis", "Linsen", "Vegan", "Vegetarisch", "Low-Carb", "Schnell", "Einfach", "Italienisch", "Asiatisch", "Tacos", "Bowl", "Wrap", "Nudeln", "Haehnchenbrust", "Hackfleisch", "Burger", "Pizza", "Pfannkuchen", "Quiche", "Muffins", "Smoothie", ] def _fetch_recipe_list(query, offset, limit=100): # Uses ?query=... not bare offset url = f"{API_BASE}?query={query}&offset={offset}&limit={limit}" ... ``` State tracks `query_cycle` (position in QUERIES list) not just `api_offset`. ## Migration notes **State field change:** v4/v5.1 uses `seed_cycle` (integer, position in infinite sequence). v5.2 uses `api_offset` (integer, offset in API listing). The state file can carry both: ```json { "seed_cycle": 9050, "api_offset": 25000, "source_counts": {"chefkoch": 25000, "hellofresh": 0}, "migrated_from_v4": true } ``` **Existing recipe IDs:** v4 IDs were `ck_347aaae514e25800` (16-char hex from URL). v5.2 IDs are `ck_4073161636112189` (numeric string from API). Both coexist in `real_recipes.jsonl` without conflict since dedup is per-ID. **When to migrate from v4/v5.1 to v5.2:** - Scraper is stuck at same count for multiple runs → migrate immediately - Target > 25,000 → v5.2 is the only viable path (search exhaustion) - Playwright/browser issues recurring → v5.2 eliminates browser dependency - Building from scratch → start with v5.2 ## Performance benchmark (observed) | Metric | v5.1 (Playwright) | v5.2 (API) | |--------|-------------------|------------| | Recipes/hour | ~1,500 | ~5,000–9,000 | | Recipes/min | ~25 | ~80–150 | | Failure rate | ~5-10% (timeouts, empty pages) | ~0.1% | | Memory footprint | High (Chromium process) | Low (pure Python + urllib) | | CPU usage | High (browser rendering) | Low (JSON parse only) | | Max reachable | ~25k (search exhaustion) | ~50k+ (query rotation) | ## Remaining work for 50k target At 25,020 recipes (current), v5.2 needs to fetch ~25,000 more. At ~800 recipes/minute with 16 workers: - **Time estimate:** ~35 minutes of continuous runtime - **API call count:** ~250 listing fetches + 25,000 detail fetches - **Query rotation:** ~10 unique queries required (each yields ~2,500 unique IDs) - **File size estimate:** ~25 MB of JSONL (assuming ~1 KB per recipe) **Note on offsets:** Do NOT attempt to reach offset 25,000 on a single query. The unfiltered listing exhausts at ~2,000 and search queries at ~1,000. Use query rotation to distribute the target across many small offset windows instead. **v5.3 refinements learned during deep scraping:** - `api_offset` must advance by `limit` (e.g. 100) per batch, not by 1. Advancing by 1 re-fetches the same 100-recipe window with a 1-recipe offset shift, producing zero new recipes while appearing to "progress" - 16 ThreadPool workers is the practical maximum; 32 workers showed no throughput improvement (~750 vs ~850 recipes/min) due to server-side latency becoming the dominant factor ## Pitfalls learned during migration 1. **Do not try `api.chefkoch.de/v2/recipes/{id}/ingredients`** — returns `resource_not_found`. Ingredients are inside the detail response under `ingredientGroups`. 2. **Tag filtering required** — API returns empty strings in `tags` array; always `strip()` and filter. 3. **Offset, not page** — The API uses `offset`, not `page`. Paginate with `offset += limit`, not `page += 1`. 4. **Count is approximate** — `count` field drifts; never use it as a hard loop bound. 5. **No `with=` query parameter works** — `with=ingredients,steps` is ignored; always fetch detail endpoint. 6. **Concurrent detail fetches** — ThreadPoolExecutor with 16 workers gives good throughput without hitting any observed limits.