# Chefkoch.de Internal API v2 — Discovery & Detail Discovered during v5.2 scraper development. This endpoint is **public and unauthenticated** (as of 2026-06) and returns structured JSON for 384,252+ recipes. ## Base URL ``` https://api.chefkoch.de/v2/recipes ``` No API key, no auth headers, no cookies required. ## Endpoints ### 1. Listing (Discovery) ``` GET /v2/recipes?offset={N}&limit={limit} ``` **Parameters** - `offset` — integer, start index (pagination) - `limit` — integer, max 100 per request (default 30) **Response shape** ```json { "count": 384252, "queryId": "", "results": [ { "recipe": { "id": "4073161636112189", "title": "Flammkuchen mit Kartoffeln und Schwarzwälder Schinken", "slug": "Flammkuchen-mit-Kartoffeln-und-Schwarzwaelder-Schinken", "siteUrl": "https://www.chefkoch.de/rezepte/4073161636112189/Flammkuchen-mit-Kartoffeln-und-Schwarzwaelder-Schinken.html", "difficulty": 1, "preparationTime": 15, "rating": {"rating": 4.75, "numVotes": 20}, "hasImage": true, "isPremium": true, "createdAt": "2021-11-05T13:36:51+01:00" }, "score": 384252 } ] } ``` **Discovery-only fields** (useful for filtering before detail fetch): - `id` — recipe ID (string) - `title` / `slug` / `siteUrl` — human metadata - `preparationTime` — minutes (int) - `difficulty` — 1=easy, 3=hard - `rating.rating` / `rating.numVotes` - `isPremium` / `isPlus` — premium flag ### 2. Detail (Full Recipe) ``` GET /v2/recipes/{id} ``` **Response shape** (key fields) ```json { "id": "4073161636112189", "title": "Flammkuchen…", "slug": "Flammkuchen-mit-Kartoffeln-und-Schwarzwaelder-Schinken", "siteUrl": "https://www.chefkoch.de/rezepte/...", "subtitle": "Originalrezept von Viki Fuchs", "additionalDescription": "", "instructions": "Den Ofen auf 220 °C Umluft vorheizen…", "ingredientGroups": [ { "header": "Für den Teig: ", "ingredients": [ { "id": "8891376", "name": "Öl", "unit": "EL", "unitId": "13", "amount": 4.0, "isBasic": false, "usageInfo": "", "foodId": "184" } ] } ], "tags": ["Schnell", "einfach", "Backen", "Frankreich"], "rating": {"rating": 4.75, "numVotes": 20}, "preparationTime": 15, "kCalories": 0, "servings": 2, "nutrition": null, "difficulty": 1 } ``` ## Structured Ingredients (API vs JSON-LD) | Aspect | JSON-LD (`recipeIngredient`) | API (`ingredientGroups`) | |--------|------------------------------|--------------------------| | Format | raw strings: `"4 EL Öl"` | objects: `{amount: 4, unit: "EL", name: "Öl"}` | | Groups | flat list | grouped by `header` (e.g. "Für den Teig:") | | Parsing | requires regex heuristics | already parsed | | Accuracy | ~95% (edge cases in parentheses) | ~99% | **Normalization recipe** (from API objects to flat strings): ```python ingredients = [] for group in data.get("ingredientGroups", []): header = group.get("header", "").strip() if header: ingredients.append(f"--- {header} ---") for ing in group.get("ingredients", []): parts = [str(p) for p in [ing.get("amount"), ing.get("unit"), ing.get("name")] if p] if parts: ingredients.append(" ".join(parts)) ``` ## Rate Limits & Politeness **Observed behavior:** - No explicit rate limiting (no 429 responses observed at 16 concurrent requests) - Server responds reliably with `application/json` - Total count fluctuates slightly (cache propagation) **Recommended politeness:** - 0.3s between listing batches - 16 concurrent detail fetches (ThreadPoolExecutor) - No backoff needed for normal operation ## Comparison: API vs Playwright | Metric | Playwright v4/v5 | API v5.2 | |--------|------------------|----------| | Discovery | Playwright scroll + DOM parse | 1 HTTP GET per 100 recipes | | Detail fetch | urllib.requests + JSON-LD regex | 1 HTTP GET per recipe | | Avg per recipe | ~3–5s | ~0.2s | | Browser dependency | Yes (chromium binary) | No | | Headless issues | Cookie banners, JS timing | None | | Max recipes reachable | ~25k (search exhaustion) | ~384k (full catalog) | | Bot detection risk | Moderate (requires stealth) | Very low (same API as site frontend) | ## Migration from Playwright to API ```python # old (v5.1) urls = await _fetch_links(page, seed_url) # Playwright recipe = await _scrape_one(page, url) # requests + JSON-LD # new (v5.2) recipes_meta = _fetch_recipe_list(offset, limit=100) # API listing recipe = _fetch_recipe_detail(meta["id"]) # API detail ``` ## Listing API: Offset Ceiling Reality The unfiltered listing endpoint (`/v2/recipes?offset=N&limit=100`) appears to return **the same global list** regardless of any `sort=`, `order=`, `with=`, or `difficulty=` parameters tested (2026-06-22). The only way to get diversity is to advance `offset`. **CRITICAL:** Offset has a practical ceiling. In testing: - `offset=0` returns 100 results ✅ - `offset=2000` returns 0 results ❌ (server silently drops) This means the unfiltered listing CANNOT be used to reach the full 384k catalog by offset alone. For deep scraping, use the **search query strategy** below. ## Search Queries via `?query=` The endpoint supports text search with real filtering: ``` GET /v2/recipes?query={term}&offset=0&limit=100 ``` Verified search queries (with approximate counts as of 2026-06): | Query | Count | Usable offsets | |-------|-------|----------------| | Hauptspeise | ~158,000 | 0–~1000 | | Kuchen | ~58,000 | 0–~1000 | | Vegan | ~28,000 | 0–~1000 | | Italienisch | ~7,000 | 0–~1000 | | (blank) | ~384,000 | 0–~2000 (barely) | **Search offset ceiling:** ~1000 results (10 pages of 100). Beyond that, the search returns empty. ## Deep Scraper Strategy for 25k+ Recipes (v5.2) Because no single query reaches beyond ~1000–2000 offsets, the scraper must **rotate queries**: 1. Maintain a list of 50+ search terms (categories, ingredients, cuisines) 2. Each run: pick a rotation window (e.g., 20 terms) 3. For each term: random offset between 0 and ~100, fetch listing 4. Filter known IDs, fetch details in parallel 5. Rotate to next term This prevents seed exhaustion while staying inside each query's usable window. ```python QUERIES = [ "Hauptspeise", "Vorspeise", "Nachtisch", "Beilage", "Fruehstueck", "Backen", "Grillen", "Salat", "Suppe", "Eintopf", "Pasta", "Dessert", "Kuchen", "Brot", "Getraenk", "Fingerfood", "Curry", "Wok", "Auflauf", "Gratin", "Risotto", "Lasagne", "Lachs", "Haehnchen", "Rind", "Schwein", "Fisch", "Gemuese", "Kartoffeln", "Reis", "Linsen", "Vegan", "Vegetarisch", "Low-Carb", "Schnell", "Einfach", "Italienisch", "Asiatisch", "Tacos", "Bowl", "Wrap", "Nudeln", "Haehnchenbrust", "Hackfleisch", "Burger", "Pizza", "Pfannkuchen", "Quiche", "Muffins", "Smoothie", ] def _fetch_list(query, offset, limit=100): url = f"https://api.chefkoch.de/v2/recipes?query={query}&offset={offset}&limit={limit}" ... ``` ## Performance observed (2026-06-22): **Initial projection under-estimated actual throughput significantly.** With 12 workers: ~10–15 recipes/sec (throughput ceiling is detail-fetch latency, not listing) With 16 workers: ~800–900 recipes/minute (~13–15 recipes/sec) With 32 workers: ~750 recipes/minute sustained (diminishing returns from context switching) **Key insight:** The listing fetch is negligible latency (~100ms). Detail fetch dominates (~200ms per recipe at 32 concurrent). 16 concurrent detail fetches saturates the network + server response pipe. More workers add no value. **Recommended configuration:** - 16 ThreadPool workers (sweet spot) - 0.3s delay between listing batches - limit=100 per listing request - Each run rotates to next query when hitting the offset ceiling ## State Management for API Scraper ### State file format ```json { "query_index": 0, "api_offset": 0, "source_counts": {"chefkoch": 25020}, "new_this_run": 0, "failed_listings": 0 } ``` **Field semantics:** - `query_index` — position in `QUERIES` list (rotates when exhausted) - `api_offset` — current offset within the active query's pagination window - `source_counts` — total recipes scraped per source - `new_this_run` — deduplicated count since last state write - `failed_listings` — consecutive failed listing requests (used for fast-exit) ### CRITICAL: Offset increment must be `+= limit`, not `+= 1` The scraper fetches **batches of `limit` recipes per listing call**. State must advance `api_offset` by `limit` (e.g., 100), not by 1. Incrementing by 1 causes the same 100-recipe batch to be re-fetched on every subsequent run, producing **zero new recipes** while appearing to "run successfully". ```python # CORRECT — advances to next batch def main(): state = _load_state() offset = state.get("api_offset", 0) limit = 100 # ... fetch listing at offset ... state["api_offset"] = offset + limit # 0 → 100 → 200 _save_state(state) # WRONG — re-fetches same batch every time def main(): state = _load_state() offset = state.get("api_offset", 0) # ... fetch listing at offset ... state["api_offset"] = offset + 1 # 0 → 1; next run: fetches offset 1 (same as offset 0 in effect) ``` **The same batch always returns the same 100 recipe IDs.** Deduplication drops everything, `new_this_run` stays at 0, but the state increments by 1 each run, crawling nowhere at ~1 recipe offset per launch. ## Pitfalls 1. **Do not use** `api.chefkoch.de/v2/recipes/{id}/ingredients` — returns `resource_not_found` (404) 2. **Ingredients live inside the detail response** under `ingredientGroups` — there is no separate endpoint 3. **Tag list may contain empty strings** — filter with `.strip()` before storage 4. **Count is approximate** — the `count` field drifts slightly; do not hardcode it as a loop bound 5. **Offset-based pagination with a hard ceiling** — unfiltered listing exhausts at ~2000; search queries at ~1000 6. **Query parameter filtering is NOT supported on the listing endpoint** — `sort=`, `difficulty=`, etc. are ignored 7. **Do not cache the `count` field as a loop bound** — use query rotation instead ## Headers Required Minimal headers suffice; full mimicry not required: ```python headers = { "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36", "Accept": "application/json", "Accept-Language": "de-DE,de;q=0.9,en;q=0.8", "Referer": "https://www.chefkoch.de/", } ```