Files

13 KiB

Recipe Data Sources (Tested 2026-06-28)

Active Sources (v7 Scraper — Direct MySQL)

Chefkoch.de — ~300k+ total, ~25k scraped

  • Access: REST API v2 (https://api.chefkoch.de/v2/recipes)
  • Method: Query rotation across 100+ German cooking terms, pagination via offset
  • Structured data: API returns JSON with title, rating, ingredients, tags, times
  • Anti-scraping: Rate-limited but no Cloudflare; 0.3-0.5s delay sufficient
  • Field mapping: id, title, category, tags, rating, prepTime, cookTime, totalTime, ingredients[].raw, siteUrl, nutrition.calories

Kuechengoetter.de — 30,000 recipes

  • Access: Sitemap (https://www.kuechengoetter.de/sitemaps/recipes-1.xml)
  • Method: Sitemap discovery → individual page fetch → JSON-LD extraction
  • Structured data: Schema.org Recipe in <script type="application/ld+json">
  • Fields extracted: name, aggregateRating.ratingValue, prepTime, cookTime, totalTime, recipeIngredient[], recipeCategory, recipeInstructions[], keywords→tags
  • Anti-scraping: None observed; polite delays (0.3-1.0s) work fine
  • Quality: Ratings present but sparse (many show rating=1, meaning single vote). Ingredient lists well-structured.

Kitchen Stories (kitchenstories.com/de) — 3,390 German recipes

  • Access: Sitemap (https://www.kitchenstories.com/sitemap.xml)
  • Filter: /de/rezepte/ (German recipe pages only; site has EN/DE bilingual)
  • Method: Sitemap discovery → individual page fetch → JSON-LD extraction
  • Structured data: Schema.org Recipe, 3 JSON-LD blocks per page (Recipe + WebSite + BreadcrumbList)
  • Fields: name, totalTime (ISO 8601 like PT20M), recipeIngredient[], recipeCategory, recipeInstructions[]
  • Anti-scraping: None observed
  • Quality: High-quality editorial recipes. No ratings in JSON-LD. Good metadata (times, categories).

HelloFresh.de — ~16,831 total, ~10,000 estimated unique, 555 already scraped

  • Access: Category pages via __NEXT_DATA__ (Next.js SSR hydration blob)
  • Listing URL pattern: https://www.hellofresh.de/recipes/{category}?skip={N}&take=50
  • Detail URL pattern: https://www.hellofresh.de/recipes/{slug}-{recipeId}
  • Method: Enumerate categories → paginate each → fetch detail per recipe → parse __NEXT_DATA__
  • Structured data: NOT JSON-LD. Next.js __NEXT_DATA__ script tag containing ssrPayload.dehydratedState.queries[]
    • Listing: find query with queryKey[0] == 'recipe.search'state.data.items[] (each has id, slug, name)
    • Detail: find query with queryKey[0] == 'foodContentHubRecipe.byId'state.data.recipe
  • Normalization: Custom _hf_normalize() — builds ingredients from yields[0].ingredients[] joined with ingredients[] lookup, instructions from steps[].instructionsMarkdown, calories from nutrition[] where name contains 'kcal'
  • Working categories (19/28): familien-rezepte (6,208), vegetarische-rezepte (4,625), einfache-rezepte (1,444), gesunde-rezepte (1,395), eintopf-rezepte (934), italienische-rezepte (481), deutsche-rezepte (463), asiatische-rezepte (185), amerikanische-rezepte (172), orientalische-rezepte (167), kartoffel-rezepte, mediterrane-rezepte, mexikanische-rezepte, nudel-rezepte, suppen, salate, burger, chinesische-rezepte, flammkuchen-rezepte
  • Dead categories (404): belibteste-rezepte, pancakes, pasta, pizza, reis-rezepte, sandwich, fleisch-rezepte, franzosische-rezepte, griechische-rezepte
  • Anti-scraping: None observed; Referer header recommended
  • Quality: HIGHEST of all sources. 99% have calories, 93% have ratings, all have instructions, description, difficulty, servings. Professionally curated with exact quantities.
  • ID prefix: hf_ (NOT hellofresh_) — must match v5.3 legacy IDs for deduplication
  • Headers: Requires Referer: https://www.hellofresh.de/ and standard browser UA

GuteKueche.de — 29,392 recipes (INTEGRATED in v6)

  • Access: Gzipped sitemap (https://cdn.gutekueche.de/sitemaps/recipe.xml.gz)
  • Discovery: robots.txt → Sitemap: directive led to CDN-hosted gzipped sitemap index → 4 sub-sitemaps, recipe.xml.gz has 29,392 recipe URLs
  • Method: Same as Kuechengoetter/Kitchen Stories — sitemap discovery → JSON-LD extraction. Scraper auto-detects gzip via magic bytes \x1f\x8b or config flag "gzipped": True.
  • Config: Added to SITEMAP_SOURCES dict with "gzipped": True flag. discover_sitemap_urls() enhanced with gzip support (urllib + gzip.decompress()).
  • URL pattern: https://www.gutekueche.de/<name>-rezept-<id>
  • Structured data: Schema.org Recipe in JSON-LD, 9/11 quality fields (highest of all sitemap sources)
  • Fields present: nutrition (calories ), recipeInstructions , aggregateRating (e.g. 4.5★), totalTime, recipeIngredient, recipeCategory, recipeYield, prepTime, cookTime
  • Missing: recipeCuisine, suitableForDiet
  • Anti-scraping: None; robots.txt allows all (only disallows /user/, /suche, /api/)
  • Quality: Strong — professional editorial recipes with calorie data and ratings. Second-best overall quality after HelloFresh.

Tested but Rejected

Source Reason
AllRecipes.com HTTP 403 Forbidden (Cloudflare)
Yummly.com HTTP 403 Forbidden
Tasty.co HTTP 404 on sitemap
Marmiton.org HTTP 404 on sitemap (French, URL structure changed)
EatThis.org SSL certificate verification failed
Lecker.de Sitemap XML files return 0 URLs (likely gzipped or JS-rendered)
Essen-und-trinken.de Has Recipe JSON-LD on recipe pages (6/11 fields, calories , ratings ), BUT robots.txt blocks ALL AI bots (ClaudeBot, ChatGPT, Claude-Web, CCBot, etc.). Respect robots.txt — do not scrape.
Maggi.de HTTP 403 Forbidden
Knorr.de HTTP 403 Forbidden
Serious Eats HTTP 403 Forbidden
Dr. Oetker No JSON-LD on recipe pages
Kochbar.de No JSON-LD
Marley Spoon No JSON-LD (JS-rendered)
Frauenzimmer.de JSON-LD present but no Recipe type on recipe pages
Bunte.de 404 on recipe page URLs (structure unclear)

Storage: MariaDB Galera Cluster

Recipe data migrated from JSONL to the existing Galera cluster (2026-06-28).

  • Why Galera over SQLite: User preference — reuse existing HA infrastructure (3-node cluster already running for HA Recorder) instead of adding new dependencies. Concurrent read/write via MaxScale readwritesplit.
  • Connection: 10.0.30.70:3306 (MaxScale VIP) → Galera cluster (db1=.71, db2=.72, db3=.73)
  • Credentials: DB=recipes, User=recipe_app@%, password in migration script
  • Table schema: 20 columns, InnoDB, utf8mb4_unicode_ci
  • Indexes: FULLTEXT ft_title_desc (title, description), FULLTEXT ft_ingredients (ingredients), BTREE idx_source, idx_category, idx_rating, idx_difficulty
  • Migration script: scripts/migrate_to_galera.py — INSERT IGNORE with batch=200, handles dict-valued category and nutrition fields, skips existing IDs
  • Pitfall: category field is a dict (not string) in 184 Chefkoch records, nutrition is a dict in 5004 records. Migration script must coerce these to strings/extract values before INSERT.
  • Pitfall: Galera replication overhead makes bulk inserts slower than local SQLite. Batch size 200 is the sweet spot — larger batches risk wsrep replication lag timeouts.
  • Architecture: Scraper writes directly to Galera via INSERT IGNORE (no JSONL intermediate). v7 eliminated the JSONL → migration → DB pipeline. JSONL (real_recipes.jsonl) is kept as historical archive only.
  • Pitfall: category field is a dict (not string) in 184 Chefkoch records, nutrition is a dict in 5004 records. recipe_to_row() in v7 coerces these to strings/extracts values before INSERT.
  • Pitfall: Galera replication overhead makes bulk inserts slower than local SQLite. Batch size 200 is the sweet spot — larger batches risk wsrep replication lag timeouts.
  • Pitfall: Per-row db_has_id() dedup is catastrophically slow through MaxScale — each check is a network roundtrip. v7 loads all IDs into a Python set() at startup via db_all_ids() (~0.3s for 28k), then all dedup is in-memory. Additionally, derive the likely recipe ID from the URL path and skip the HTTP fetch entirely if already known.

Potential Future Sources (not yet integrated)

Source Recipes Access Notes
Food.com Large JSON-LD confirmed on recipe pages English; sitemap index exists
BBC Good Food ~300/quarter Sitemap recipe files, JSON-LD confirmed English; premium content mixed in
HelloFresh.de 548 scrapednow active in v6 (see above) Was thought exhausted; actually has ~16,831 recipes

Recipe Source Scanning Methodology

When evaluating new recipe sources for the scraper:

  1. Don't judge by landing page alone. Landing/listing pages rarely embed Recipe JSON-LD — they have WebSite or ItemList schemas. Always fetch an individual recipe page and check for @type: "Recipe" in JSON-LD blocks.
  2. Check robots.txt first. Many sites declare their sitemap location there (Sitemap: <url>). Standard paths like /sitemap.xml frequently 404. Also check for AI-bot blocks — respect them.
  3. Sitemaps may be gzipped. Look for .gz extensions or Content-Encoding: gzip. Use gzip.decompress() in Python.
  4. Volume estimation from sitemap index. Fetch the index, count sub-sitemaps, sample 5-10 sub-sitemaps, extrapolate. Filter URLs by path patterns (/rezept-, /rezepte/, /recipes/).
  5. Quality scoring (11-field checklist): Count presence of: nutrition, recipeInstructions, aggregateRating, totalTime, recipeIngredient, recipeCategory, recipeYield, prepTime, cookTime, recipeCuisine, suitableForDiet. Score = present/11. HelloFresh ≈ 11/11 (custom parsing), GuteKueche ≈ 9/11, Kuechengoetter ≈ 6/11.
  6. Quality ≠ Volume. A source with 548 high-quality recipes (calories, instructions, ratings, difficulty) may be more valuable than a 30k source with sparse metadata. Evaluate both dimensions. Never dismiss a source solely for low volume without assessing data quality.

v7 Scraper Architecture (Direct MySQL — replaces v6)

v7 eliminates JSONL entirely. Writes directly to MariaDB Galera via PyMySQL.

recipe_scraper_v7.py
├── Phase 1: Sitemap sources (Kuechengoetter → GuteKueche → Kitchen Stories)
│   ├── discover_sitemap_urls() — fetch sitemap (auto-gzip), extract <loc> URLs
│   ├── scrape_sitemap_source() — visit each URL, extract JSON-LD
│   └── normalize_jsonld_recipe() — Schema.org → our format
├── Phase 2: HelloFresh (__NEXT_DATA__ extraction)
│   ├── _hf_fetch_category_page() — paginate categories via skip/take
│   ├── _hf_fetch_detail() — fetch recipe page, parse __NEXT_DATA__
│   ├── _hf_normalize() — convert HF nested structure to our format
│   └── Per-category skip tracking via state["hf_cat_offset"]
├── Phase 3: Chefkoch API (fills remaining gap to 100k)
│   ├── fetch_ck_page() — API query with rotation
│   └── fetch_ck_detail() — full recipe by ID
├── State: scraper_state.json (tracks tried URLs, query cycles, source counts, hf_cat_offset)
├── Dedup: db_all_ids() loads ALL IDs into Python set() at startup (~0.3s for 28k)
│   └── Pre-check: derives likely ID from URL path, skips HTTP fetch if already in set
├── Insert: INSERT IGNORE per recipe (idempotent, safe for re-runs)
├── Output: Direct to MariaDB Galera (10.0.30.70:3306, DB=recipes)
├── Cronjob: recipe-refill-v7 (daily 06:00, deliver: origin)
└── Target: 100,000 recipes

Key differences from v6:

  • No JSONL file, no migration step — single pipeline
  • In-memory dedup via db_all_ids() instead of per-row db_has_id() (avoids catastrophic Galera roundtrip latency)
  • URL pre-check skips HTTP fetch entirely for known recipes
  • Progress logging every 200 URLs
  • recipe_to_row() sanitizes dict/list fields (category, calories) before INSERT

Separate step — Embeddings: After scraping, run generate_embeddings.py to populate the embedding JSON column with 1024-dim harrier vectors. This is NOT done by the scraper.

Output format (unified across all sources, stored in recipes table):

{
  "id": "kuechengoetter_seelachs-mit-rucola-1",
  "name": "Seelachs mit Rucola",
  "title": "Seelachs mit Rucola",
  "category": "fish",
  "tags": ["fisch", "seelachs", "rucola"],
  "rating": 4.5,
  "rating_value": 4.5,
  "prep_time": "PT15M",
  "cook_time": "PT20M",
  "total_time": "PT35M",
  "ingredients": [{"raw": "300g Seelachsfilet"}, ...],
  "source_url": "https://...",
  "source": "kuechengoetter",
  "calories": "450 kcal"
}