Files
hermes-skills/food-nutrition/meal-planning/references/multi-source-recipe-scraper/chefkoch-scraper-v5-2-migration.md
T

115 lines
6.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chefkoch Scraper Evolution: v4 Playwright → v5.2 API
Session learning from 2026-06-22: migrating the Chefkoch scraper from Playwright-based discovery to internal API discovery. This is a **10× speed increase** and removes all headless-browser fragility.
## Why v4 Stalled at ~25k
| Problem | Symptom | Root cause |
|---------|---------|-----------|
| Search exhaustion | Same seeds return same recipes | 2-letter alphabet combos cover finite space |
| Bot detection | Empty search pages, no links | Chefkoch serves JS shell to non-browser clients |
| Playwright fragility | `Executable doesn't exist`, EPIPE | Browser binary mismatches, env propagation issues |
| Seed cycle reset | `seed_cycle` stuck at 0 | Missing `global SEED_CYCLE` in `main()` |
| Cookie walls | Content blocked | GDPR consent banners block link extraction |
| Rate ~0.4/s | 100 recipes takes ~4 min | Browser overhead per search page |
The fundamental issue: **Chefkoch search pages require a real browser**, and the browser-based discovery approach has natural throughput limits and maintenance overhead.
## The API Discovery Path (v5.2)
### How we found it
During debugging, a test request to `api.chefkoch.de/v2/recipes?offset=0&limit=30` returned structured JSON with **384,252 recipes** listed. The endpoint is the same one Chefkoch's own frontend uses — no auth, no rate limiting observed.
### Architecture comparison
| Aspect | v4/v5.1 (Playwright) | v5.2 (API) |
|--------|----------------------|------------|
| Discovery | Browser-rendered search pages | HTTP listing API (100 recipes/page) |
| Detail fetch | JSON-LD regex from HTML | Structured JSON from API |
| Browser needed | Yes (Chromium binary) | No |
| Max reachable | ~25k (search exhaustion) | ~384k (full catalog) |
| Ingredients | Raw strings from JSON-LD | Objects with amount/unit/name |
| Avg speed | ~0.4 recipes/sec | ~5 recipes/sec |
| Headless issues | Yes (cookies, JS timing, EPIPE) | None |
| Maintenance | Browser updates, path management | Minimal (HTTP only) |
## v5.2 Architecture (query-rotation based)
Because the listing API has an offset ceiling (~2000 for unfiltered, ~1000 for search queries), v5.2 uses **query rotation** instead of simple offset pagination:
```python
QUERIES = [
"Hauptspeise", "Vorspeise", "Nachtisch", "Beilage", "Fruehstueck",
"Backen", "Grillen", "Salat", "Suppe", "Eintopf", "Pasta", "Dessert",
"Kuchen", "Brot", "Getraenk", "Fingerfood", "Curry", "Wok", "Auflauf",
"Gratin", "Risotto", "Lasagne", "Lachs", "Haehnchen", "Rind", "Schwein",
"Fisch", "Gemuese", "Kartoffeln", "Reis", "Linsen", "Vegan", "Vegetarisch",
"Low-Carb", "Schnell", "Einfach", "Italienisch", "Asiatisch", "Tacos",
"Bowl", "Wrap", "Nudeln", "Haehnchenbrust", "Hackfleisch", "Burger",
"Pizza", "Pfannkuchen", "Quiche", "Muffins", "Smoothie",
]
def _fetch_recipe_list(query, offset, limit=100):
# Uses ?query=... not bare offset
url = f"{API_BASE}?query={query}&offset={offset}&limit={limit}"
...
```
State tracks `query_cycle` (position in QUERIES list) not just `api_offset`.
## Migration notes
**State field change:** v4/v5.1 uses `seed_cycle` (integer, position in infinite sequence). v5.2 uses `api_offset` (integer, offset in API listing). The state file can carry both:
```json
{
"seed_cycle": 9050,
"api_offset": 25000,
"source_counts": {"chefkoch": 25000, "hellofresh": 0},
"migrated_from_v4": true
}
```
**Existing recipe IDs:** v4 IDs were `ck_347aaae514e25800` (16-char hex from URL). v5.2 IDs are `ck_4073161636112189` (numeric string from API). Both coexist in `real_recipes.jsonl` without conflict since dedup is per-ID.
**When to migrate from v4/v5.1 to v5.2:**
- Scraper is stuck at same count for multiple runs → migrate immediately
- Target > 25,000 → v5.2 is the only viable path (search exhaustion)
- Playwright/browser issues recurring → v5.2 eliminates browser dependency
- Building from scratch → start with v5.2
## Performance benchmark (observed)
| Metric | v5.1 (Playwright) | v5.2 (API) |
|--------|-------------------|------------|
| Recipes/hour | ~1,500 | ~5,0009,000 |
| Recipes/min | ~25 | ~80150 |
| Failure rate | ~5-10% (timeouts, empty pages) | ~0.1% |
| Memory footprint | High (Chromium process) | Low (pure Python + urllib) |
| CPU usage | High (browser rendering) | Low (JSON parse only) |
| Max reachable | ~25k (search exhaustion) | ~50k+ (query rotation) |
## Remaining work for 50k target
At 25,020 recipes (current), v5.2 needs to fetch ~25,000 more. At ~800 recipes/minute with 16 workers:
- **Time estimate:** ~35 minutes of continuous runtime
- **API call count:** ~250 listing fetches + 25,000 detail fetches
- **Query rotation:** ~10 unique queries required (each yields ~2,500 unique IDs)
- **File size estimate:** ~25 MB of JSONL (assuming ~1 KB per recipe)
**Note on offsets:** Do NOT attempt to reach offset 25,000 on a single query. The unfiltered listing exhausts at ~2,000 and search queries at ~1,000. Use query rotation to distribute the target across many small offset windows instead.
**v5.3 refinements learned during deep scraping:**
- `api_offset` must advance by `limit` (e.g. 100) per batch, not by 1. Advancing by 1 re-fetches the same 100-recipe window with a 1-recipe offset shift, producing zero new recipes while appearing to "progress"
- 16 ThreadPool workers is the practical maximum; 32 workers showed no throughput improvement (~750 vs ~850 recipes/min) due to server-side latency becoming the dominant factor
## Pitfalls learned during migration
1. **Do not try `api.chefkoch.de/v2/recipes/{id}/ingredients`** — returns `resource_not_found`. Ingredients are inside the detail response under `ingredientGroups`.
2. **Tag filtering required** — API returns empty strings in `tags` array; always `strip()` and filter.
3. **Offset, not page** — The API uses `offset`, not `page`. Paginate with `offset += limit`, not `page += 1`.
4. **Count is approximate**`count` field drifts; never use it as a hard loop bound.
5. **No `with=` query parameter works**`with=ingredients,steps` is ignored; always fetch detail endpoint.
6. **Concurrent detail fetches** — ThreadPoolExecutor with 16 workers gives good throughput without hitting any observed limits.