13 KiB
13 KiB
Recipe Data Sources (Tested 2026-06-28)
Active Sources (v7 Scraper — Direct MySQL)
Chefkoch.de — ~300k+ total, ~25k scraped
- Access: REST API v2 (
https://api.chefkoch.de/v2/recipes) - Method: Query rotation across 100+ German cooking terms, pagination via offset
- Structured data: API returns JSON with title, rating, ingredients, tags, times
- Anti-scraping: Rate-limited but no Cloudflare; 0.3-0.5s delay sufficient
- Field mapping:
id,title,category,tags,rating,prepTime,cookTime,totalTime,ingredients[].raw,siteUrl,nutrition.calories
Kuechengoetter.de — 30,000 recipes
- Access: Sitemap (
https://www.kuechengoetter.de/sitemaps/recipes-1.xml) - Method: Sitemap discovery → individual page fetch → JSON-LD extraction
- Structured data: Schema.org Recipe in
<script type="application/ld+json"> - Fields extracted: name, aggregateRating.ratingValue, prepTime, cookTime, totalTime, recipeIngredient[], recipeCategory, recipeInstructions[], keywords→tags
- Anti-scraping: None observed; polite delays (0.3-1.0s) work fine
- Quality: Ratings present but sparse (many show rating=1, meaning single vote). Ingredient lists well-structured.
Kitchen Stories (kitchenstories.com/de) — 3,390 German recipes
- Access: Sitemap (
https://www.kitchenstories.com/sitemap.xml) - Filter:
/de/rezepte/(German recipe pages only; site has EN/DE bilingual) - Method: Sitemap discovery → individual page fetch → JSON-LD extraction
- Structured data: Schema.org Recipe, 3 JSON-LD blocks per page (Recipe + WebSite + BreadcrumbList)
- Fields: name, totalTime (ISO 8601 like PT20M), recipeIngredient[], recipeCategory, recipeInstructions[]
- Anti-scraping: None observed
- Quality: High-quality editorial recipes. No ratings in JSON-LD. Good metadata (times, categories).
HelloFresh.de — ~16,831 total, ~10,000 estimated unique, 555 already scraped
- Access: Category pages via
__NEXT_DATA__(Next.js SSR hydration blob) - Listing URL pattern:
https://www.hellofresh.de/recipes/{category}?skip={N}&take=50 - Detail URL pattern:
https://www.hellofresh.de/recipes/{slug}-{recipeId} - Method: Enumerate categories → paginate each → fetch detail per recipe → parse
__NEXT_DATA__ - Structured data: NOT JSON-LD. Next.js
__NEXT_DATA__script tag containingssrPayload.dehydratedState.queries[]- Listing: find query with
queryKey[0] == 'recipe.search'→state.data.items[](each hasid,slug,name) - Detail: find query with
queryKey[0] == 'foodContentHubRecipe.byId'→state.data.recipe
- Listing: find query with
- Normalization: Custom
_hf_normalize()— builds ingredients fromyields[0].ingredients[]joined withingredients[]lookup, instructions fromsteps[].instructionsMarkdown, calories fromnutrition[]where name contains 'kcal' - Working categories (19/28): familien-rezepte (6,208), vegetarische-rezepte (4,625), einfache-rezepte (1,444), gesunde-rezepte (1,395), eintopf-rezepte (934), italienische-rezepte (481), deutsche-rezepte (463), asiatische-rezepte (185), amerikanische-rezepte (172), orientalische-rezepte (167), kartoffel-rezepte, mediterrane-rezepte, mexikanische-rezepte, nudel-rezepte, suppen, salate, burger, chinesische-rezepte, flammkuchen-rezepte
- Dead categories (404): belibteste-rezepte, pancakes, pasta, pizza, reis-rezepte, sandwich, fleisch-rezepte, franzosische-rezepte, griechische-rezepte
- Anti-scraping: None observed; Referer header recommended
- Quality: HIGHEST of all sources. 99% have calories, 93% have ratings, all have
instructions,description,difficulty,servings. Professionally curated with exact quantities. - ID prefix:
hf_(NOThellofresh_) — must match v5.3 legacy IDs for deduplication - Headers: Requires
Referer: https://www.hellofresh.de/and standard browser UA
GuteKueche.de — 29,392 recipes (INTEGRATED in v6)
- Access: Gzipped sitemap (
https://cdn.gutekueche.de/sitemaps/recipe.xml.gz) - Discovery: robots.txt →
Sitemap:directive led to CDN-hosted gzipped sitemap index → 4 sub-sitemaps,recipe.xml.gzhas 29,392 recipe URLs - Method: Same as Kuechengoetter/Kitchen Stories — sitemap discovery → JSON-LD extraction. Scraper auto-detects gzip via magic bytes
\x1f\x8bor config flag"gzipped": True. - Config: Added to
SITEMAP_SOURCESdict with"gzipped": Trueflag.discover_sitemap_urls()enhanced with gzip support (urllib +gzip.decompress()). - URL pattern:
https://www.gutekueche.de/<name>-rezept-<id> - Structured data: Schema.org Recipe in JSON-LD, 9/11 quality fields (highest of all sitemap sources)
- Fields present: nutrition (calories ✅), recipeInstructions ✅, aggregateRating ✅ (e.g. 4.5★), totalTime, recipeIngredient, recipeCategory, recipeYield, prepTime, cookTime
- Missing: recipeCuisine, suitableForDiet
- Anti-scraping: None; robots.txt allows all (only disallows /user/, /suche, /api/)
- Quality: Strong — professional editorial recipes with calorie data and ratings. Second-best overall quality after HelloFresh.
Tested but Rejected
| Source | Reason |
|---|---|
| AllRecipes.com | HTTP 403 Forbidden (Cloudflare) |
| Yummly.com | HTTP 403 Forbidden |
| Tasty.co | HTTP 404 on sitemap |
| Marmiton.org | HTTP 404 on sitemap (French, URL structure changed) |
| EatThis.org | SSL certificate verification failed |
| Lecker.de | Sitemap XML files return 0 URLs (likely gzipped or JS-rendered) |
| Essen-und-trinken.de | Has Recipe JSON-LD on recipe pages (6/11 fields, calories ✅, ratings ✅), BUT robots.txt blocks ALL AI bots (ClaudeBot, ChatGPT, Claude-Web, CCBot, etc.). Respect robots.txt — do not scrape. |
| Maggi.de | HTTP 403 Forbidden |
| Knorr.de | HTTP 403 Forbidden |
| Serious Eats | HTTP 403 Forbidden |
| Dr. Oetker | No JSON-LD on recipe pages |
| Kochbar.de | No JSON-LD |
| Marley Spoon | No JSON-LD (JS-rendered) |
| Frauenzimmer.de | JSON-LD present but no Recipe type on recipe pages |
| Bunte.de | 404 on recipe page URLs (structure unclear) |
Storage: MariaDB Galera Cluster
Recipe data migrated from JSONL to the existing Galera cluster (2026-06-28).
- Why Galera over SQLite: User preference — reuse existing HA infrastructure (3-node cluster already running for HA Recorder) instead of adding new dependencies. Concurrent read/write via MaxScale readwritesplit.
- Connection:
10.0.30.70:3306(MaxScale VIP) → Galera cluster (db1=.71, db2=.72, db3=.73) - Credentials: DB=
recipes, User=recipe_app@%, password in migration script - Table schema: 20 columns, InnoDB, utf8mb4_unicode_ci
- Indexes: FULLTEXT ft_title_desc (title, description), FULLTEXT ft_ingredients (ingredients), BTREE idx_source, idx_category, idx_rating, idx_difficulty
- Migration script:
scripts/migrate_to_galera.py— INSERT IGNORE with batch=200, handles dict-valuedcategoryandnutritionfields, skips existing IDs - Pitfall:
categoryfield is a dict (not string) in 184 Chefkoch records,nutritionis a dict in 5004 records. Migration script must coerce these to strings/extract values before INSERT. - Pitfall: Galera replication overhead makes bulk inserts slower than local SQLite. Batch size 200 is the sweet spot — larger batches risk wsrep replication lag timeouts.
- Architecture: Scraper writes directly to Galera via
INSERT IGNORE(no JSONL intermediate). v7 eliminated the JSONL → migration → DB pipeline. JSONL (real_recipes.jsonl) is kept as historical archive only. - Pitfall:
categoryfield is a dict (not string) in 184 Chefkoch records,nutritionis a dict in 5004 records.recipe_to_row()in v7 coerces these to strings/extracts values before INSERT. - Pitfall: Galera replication overhead makes bulk inserts slower than local SQLite. Batch size 200 is the sweet spot — larger batches risk wsrep replication lag timeouts.
- Pitfall: Per-row
db_has_id()dedup is catastrophically slow through MaxScale — each check is a network roundtrip. v7 loads all IDs into a Pythonset()at startup viadb_all_ids()(~0.3s for 28k), then all dedup is in-memory. Additionally, derive the likely recipe ID from the URL path and skip the HTTP fetch entirely if already known.
Potential Future Sources (not yet integrated)
| Source | Recipes | Access | Notes |
|---|---|---|---|
| Food.com | Large | JSON-LD confirmed on recipe pages | English; sitemap index exists |
| BBC Good Food | ~300/quarter | Sitemap recipe files, JSON-LD confirmed | English; premium content mixed in |
| HelloFresh.de | Was thought exhausted; actually has ~16,831 recipes |
Recipe Source Scanning Methodology
When evaluating new recipe sources for the scraper:
- Don't judge by landing page alone. Landing/listing pages rarely embed Recipe JSON-LD — they have WebSite or ItemList schemas. Always fetch an individual recipe page and check for
@type: "Recipe"in JSON-LD blocks. - Check
robots.txtfirst. Many sites declare their sitemap location there (Sitemap: <url>). Standard paths like/sitemap.xmlfrequently 404. Also check for AI-bot blocks — respect them. - Sitemaps may be gzipped. Look for
.gzextensions orContent-Encoding: gzip. Usegzip.decompress()in Python. - Volume estimation from sitemap index. Fetch the index, count sub-sitemaps, sample 5-10 sub-sitemaps, extrapolate. Filter URLs by path patterns (
/rezept-,/rezepte/,/recipes/). - Quality scoring (11-field checklist): Count presence of:
nutrition,recipeInstructions,aggregateRating,totalTime,recipeIngredient,recipeCategory,recipeYield,prepTime,cookTime,recipeCuisine,suitableForDiet. Score = present/11. HelloFresh ≈ 11/11 (custom parsing), GuteKueche ≈ 9/11, Kuechengoetter ≈ 6/11. - Quality ≠ Volume. A source with 548 high-quality recipes (calories, instructions, ratings, difficulty) may be more valuable than a 30k source with sparse metadata. Evaluate both dimensions. Never dismiss a source solely for low volume without assessing data quality.
v7 Scraper Architecture (Direct MySQL — replaces v6)
v7 eliminates JSONL entirely. Writes directly to MariaDB Galera via PyMySQL.
recipe_scraper_v7.py
├── Phase 1: Sitemap sources (Kuechengoetter → GuteKueche → Kitchen Stories)
│ ├── discover_sitemap_urls() — fetch sitemap (auto-gzip), extract <loc> URLs
│ ├── scrape_sitemap_source() — visit each URL, extract JSON-LD
│ └── normalize_jsonld_recipe() — Schema.org → our format
├── Phase 2: HelloFresh (__NEXT_DATA__ extraction)
│ ├── _hf_fetch_category_page() — paginate categories via skip/take
│ ├── _hf_fetch_detail() — fetch recipe page, parse __NEXT_DATA__
│ ├── _hf_normalize() — convert HF nested structure to our format
│ └── Per-category skip tracking via state["hf_cat_offset"]
├── Phase 3: Chefkoch API (fills remaining gap to 100k)
│ ├── fetch_ck_page() — API query with rotation
│ └── fetch_ck_detail() — full recipe by ID
├── State: scraper_state.json (tracks tried URLs, query cycles, source counts, hf_cat_offset)
├── Dedup: db_all_ids() loads ALL IDs into Python set() at startup (~0.3s for 28k)
│ └── Pre-check: derives likely ID from URL path, skips HTTP fetch if already in set
├── Insert: INSERT IGNORE per recipe (idempotent, safe for re-runs)
├── Output: Direct to MariaDB Galera (10.0.30.70:3306, DB=recipes)
├── Cronjob: recipe-refill-v7 (daily 06:00, deliver: origin)
└── Target: 100,000 recipes
Key differences from v6:
- No JSONL file, no migration step — single pipeline
- In-memory dedup via
db_all_ids()instead of per-rowdb_has_id()(avoids catastrophic Galera roundtrip latency) - URL pre-check skips HTTP fetch entirely for known recipes
- Progress logging every 200 URLs
recipe_to_row()sanitizes dict/list fields (category, calories) before INSERT
Separate step — Embeddings:
After scraping, run generate_embeddings.py to populate the embedding JSON column with 1024-dim harrier vectors. This is NOT done by the scraper.
Output format (unified across all sources, stored in recipes table):
{
"id": "kuechengoetter_seelachs-mit-rucola-1",
"name": "Seelachs mit Rucola",
"title": "Seelachs mit Rucola",
"category": "fish",
"tags": ["fisch", "seelachs", "rucola"],
"rating": 4.5,
"rating_value": 4.5,
"prep_time": "PT15M",
"cook_time": "PT20M",
"total_time": "PT35M",
"ingredients": [{"raw": "300g Seelachsfilet"}, ...],
"source_url": "https://...",
"source": "kuechengoetter",
"calories": "450 kcal"
}