Files

13 KiB

International Recipe Sources (Researched 2026-06-28, Updated 2026-06-28)

Focus: Asian cuisine, light/healthy main courses, internationally diverse. English-language sites with JSON-LD for scraping + translation to German.

Tier 1 — Ready to Integrate (JSON-LD confirmed + sitemap available + no AI-bot blocks)

BBC Good Food (bbcgoodfood.com) — ~5,045 recipes TOP PICK

  • Sitemap: Quarterly recipe sitemaps at https://www.bbcgoodfood.com/sitemaps/YYYY-QN-recipe.xml (indexed from sitemap.xml)
  • Counts per quarter: Q2-2026: 149, Q1-2026: 156, Q4-2025: 210, Q3-2025: 179, Q2-2025: 182, Q1-2025: 214, Q4-2024: 292, Q3-2024: 179, Q2-2024: 160, Q1-2024: 223, Q4-2023: 290, Q3-2023: 262, Q2-2023: 132, Q1-2023: 203, Q4-2022: 296, Q3-2022: 231, Q2-2022: 272, Q1-2022: 207, Q4-2021: 250, Q3-2021: 206, Q2-2021: 154, Q1-2021: 191, Q4-2020: 278, Q3-2020: 129. Total: ~5,045
  • JSON-LD: Confirmed
  • Cuisine focus: Huge international coverage, including dedicated Asian/Indian/Thai sections
  • Robots.txt: Permissive — declares sitemaps, no AI-bot blocks
  • Quality: Professional editorial team, nutrition data, tested recipes
  • ID prefix: bgf_

RecipeTin Eats (recipetineats.com) — ~1,697 recipes

  • Sitemap: https://www.recipetineats.com/sitemap_index.xml → 4 post sitemaps (495+500+499+203 URLs)
  • JSON-LD: Confirmed
  • Cuisine focus: Pan-Asian (Japanese, Thai, Korean, Chinese), Australian, fusion. ~404 URLs match Asian keywords.
  • Robots.txt: Blocks anthropic-ai and Claude-Web specifically, but does NOT block generic crawlers. Standard UA with polite delays works.
  • Quality: High — detailed instructions, ratings, step photos. Many authentic Asian recipes.
  • ID prefix: rt_

Woks of Life (thewoksoflife.com) — ~1,555 recipes BEST CHINESE

  • Sitemap: https://thewoksoflife.com/sitemap_index.xml → post-sitemap.xml (997) + post-sitemap2.xml (558)
  • JSON-LD: Confirmed
  • Cuisine focus: Chinese (Cantonese, Sichuan, regional), some Asian-American
  • Robots.txt: Permissive (blocks wp-admin, cdn-cgi, wp-json only)
  • Quality: Family-authored, authentic Chinese techniques, detailed step photos
  • ID prefix: wol_

SkinnyTaste (skinnytaste.com) — ~2,354 recipes BEST LIGHT/HEALTHY

  • Sitemap: https://www.skinnytaste.com/sitemap_index.xml → 3 post sitemaps (685+862+807 URLs)
  • JSON-LD: Confirmed
  • Cuisine focus: Healthy, low-calorie, light mains, WW-friendly. Strong international coverage including Asian-inspired light dishes.
  • Robots.txt: Blocks Google-Extended, GPTBot, ChatGPT-User, PiplBot. Does NOT block generic crawlers.
  • Quality: Excellent — Gina Homolka provides nutrition data (calories, macros) for every recipe. Professionally tested.
  • ID prefix: st_

Rasa Malaysia (rasamalaysia.com) — ~1,390 recipes BEST SE-ASIAN

  • Sitemap: https://rasamalaysia.com/sitemap_index.xml → post-sitemap.xml (984) + post-sitemap2.xml (406)
  • JSON-LD: Confirmed
  • Cuisine focus: Malaysian, Singaporean, Thai, Chinese, pan-Southeast-Asian
  • Robots.txt: Permissive (blocks cdn-cgi, wp-admin only)
  • Quality: High — authentic Southeast Asian recipes with detailed cultural context
  • ID prefix: rm_

Pinch of Yum (pinchofyum.com) — ~1,483 recipes

  • Sitemap: https://pinchofyum.com/sitemap_index.xml → post-sitemap.xml (912) + post-sitemap2.xml (571)
  • JSON-LD: Confirmed
  • Cuisine focus: Comfort food, international, some Asian-inspired
  • Robots.txt: Permissive (blocks wp-admin, search only)
  • Quality: Good — standardized format, nutrition data sometimes included
  • ID prefix: poy_

Budget Bytes (budgetbytes.com) — ~1,981 recipes

  • Sitemap: https://www.budgetbytes.com/post-sitemap.xml (976) + post-sitemap2.xml (974) + post-sitemap3.xml (31)
  • JSON-LD: Confirmed
  • Cuisine focus: American comfort, international fusion, budget-friendly mains
  • Robots.txt: Permissive
  • ID prefix: bb_

Omnivore's Cookbook (omnivorescookbook.com) — ~772 recipes

  • Sitemap: https://omnivorescookbook.com/sitemap_index.xml → post-sitemap.xml (772 URLs)
  • JSON-LD: Confirmed
  • Cuisine focus: Modern Chinese, some Asian fusion
  • Robots.txt: Has Raptive ad-network block list but does NOT block generic crawlers
  • Quality: Good — authentic Chinese with modern adaptations
  • ID prefix: oc_

Hot Thai Kitchen (hot-thai-kitchen.com) — ~464 recipes MOST AUTHENTIC THAI

  • Sitemap: https://hot-thai-kitchen.com/sitemap_index.xml → post-sitemap.xml (464 URLs)
  • JSON-LD: Confirmed
  • Cuisine focus: Authentic Thai (not Westernized)
  • Robots.txt: Permissive (Yoast standard)
  • Quality: Exceptional — Pai (author) explains Thai cooking theory, ingredient substitution within Thai context
  • ID prefix: htk_

Spoon Fork Bacon (spoonforkbacon.com) — ~911 recipes

  • Sitemap: https://www.spoonforkbacon.com/sitemap_index.xml → post-sitemap.xml (911 URLs)
  • JSON-LD: Confirmed
  • Cuisine focus: Asian fusion, modern international
  • Robots.txt: Permissive (wp-admin, search only)
  • ID prefix: sfb_

Korean Bapsang (koreanbapsang.com) — ~275 recipes BEST KOREAN

  • Sitemap: https://www.koreanbapsang.com/sitemap_index.xml → post-sitemap.xml (275 URLs)
  • JSON-LD: Confirmed
  • Cuisine focus: Authentic Korean home cooking
  • Robots.txt: Permissive (Yoast standard)
  • Quality: Excellent — Susan (author) provides authentic Korean recipes with cultural context
  • ID prefix: kb_

Sanjeev Kapoor (sanjeevkapoor.com) — ~1,965 recipes BEST INDIAN

  • Sitemap: https://www.sanjeevkapoor.com/webcontent-sitemap.xml (1,965 URLs)
  • JSON-LD: ⚠️ Non-standard — uses @type: Recipe but in a different JSON structure (not standard schema.org/Recipe JSON-LD block). Has recipeInstructions and recipeIngredient fields. Needs custom parser.
  • Cuisine focus: Indian (all regions, vegetarian + non-veg)
  • Robots.txt: Declares sitemaps, no AI-bot blocks
  • Quality: Celebrity chef, authentic Indian recipes
  • ID prefix: sk_
  • Pitfall: JSON-LD is present but structured differently from WP recipe plugins. The @type appears as 'Recipe (note quote style) rather than "Recipe". May need regex extraction instead of standard JSON-LD parsing.

Minimalist Baker (minimalistbaker.com) — ~650 recipes

  • Sitemap: https://minimalistbaker.com/wp-sitemap.xmlpost-sitemap.xml (0 URLs — appears empty) + post-sitemap2.xml (650 URLs)
  • JSON-LD: Confirmed
  • Cuisine focus: Simple plant-based, ≤10 ingredients, GF options
  • Robots.txt: Disallows /r/, /ecourse/, /university/ paths. Blocks anthropic-ai.
  • Quality: High — Dana Schultz (author) provides consistently formatted recipes
  • ID prefix: mb_
  • Pitfall: post-sitemap.xml returns 0 URLs — use post-sitemap2.xml instead. This caused a misdiagnosis in the first research pass.

Tier 2 — JSON-LD Confirmed but Blocks AI Bots (use generic UA)

  • Sitemap: https://cookieandkate.com/post-sitemap.xml (920 URLs)
  • JSON-LD: Confirmed
  • Cuisine focus: Vegetarian, healthy, Mediterranean-inspired, some Asian fusion
  • Robots.txt: Blocks anthropic-ai, Claude-Web, bot "008". Does NOT block generic Mozilla UA.
  • ID prefix: ckate_

Healthy Nibbles (healthynibblesandbits.com) — ~640 recipes

  • Sitemap: https://healthynibblesandbits.com/post-sitemap.xml (640 URLs)
  • JSON-LD: Confirmed
  • Cuisine focus: Asian, healthy, light mains
  • Robots.txt: Blocks ChatGPT-User, GPTBot, anthropic-ai via Raptive. Generic UA works.
  • ID prefix: hn_

Tier 3 — Blocked or Not Feasible

Just One Cookbook (justonecookbook.com) — ~1,200+ recipes

  • Status: Cloudflare blocks ALL automated access. Returns Cloudflare challenge page for curl requests. Sitemap returns HTML error page, not XML.
  • JSON-LD: Exists (confirmed via browser in earlier session)
  • Note: Previously listed as "no sitemap" — incorrect. The site HAS a sitemap (sitemap_index.xml declared in robots.txt) but Cloudflare blocks programmatic access to it. Would need Playwright/browser automation to scrape.
  • Correction from earlier doc: Original reference said "No sitemap index or post-sitemap found" — this was wrong. The sitemap exists but is behind Cloudflare protection.

Maangchi (maangchi.com) — ~500+ recipes

  • Status: Returns HTTP 403 for all programmatic requests (even with browser-like UA)
  • JSON-LD: Could not verify (blocked)
  • Cuisine focus: Korean (most popular English-language Korean site)
  • Note: robots.txt is permissive (only blocks /post-comments/) but the server itself rejects non-browser requests.

AllRecipes (allrecipes.com)

  • Status: Sitemaps exist (sitemap_1.xml through sitemap_4.xml) but contain 0 URLs — appear to be empty or dynamically generated.
  • JSON-LD: Could not confirm on sampled recipe URLs
  • Note: Largest US recipe site but not scrapable via sitemap approach.

Archana's Kitchen (archanaskitchen.com) — 8,428 URLs

  • Status: Sitemap has 8,428 URLs but NO JSON-LD on recipe pages (0 blocks found). Would need custom HTML parsing.
  • Cuisine focus: Indian
  • Note: Large volume but no structured data — not worth the effort vs. Sanjeev Kapoor which has JSON-LD.

Bon Appétit (bonappetit.com) / Epicurious (epicurious.com)

  • Status: No JSON-LD on recipe pages. Both use Condé Nast's proprietary structured data format.
  • Note: High quality but not scrapable via JSON-LD approach.

Serious Eats (seriouseats.com)

  • Status: Sitemap exists but returned 0 recipe URLs. JSON-LD not confirmed on sampled URLs.
  • Note: May use WPRM (WordPress Recipe Maker) structured data instead of JSON-LD. Would need custom parser.

Pick Up Limes (pickuplimes.com) — ~240 recipes

  • Sitemap: https://www.pickuplimes.com/sitemap.xml — only 1 recipe URL found in sitemap (very small)
  • JSON-LD: Confirmed
  • Robots.txt: Returns HTML page, not a proper robots.txt
  • Note: Very small volume (~240). High quality vegan but not worth integration effort for the volume.
Priority Source Recipes Why
1 BBC Good Food ~5,045 Largest, international, permissive
2 SkinnyTaste ~2,354 Best for light/healthy gap
3 Woks of Life ~1,555 Best authentic Chinese
4 RecipeTin Eats ~1,697 Pan-Asian, high quality
5 Rasa Malaysia ~1,390 Best SE-Asian
6 Sanjeev Kapoor ~1,965 Best Indian (needs custom parser)
7 Budget Bytes ~1,981 Broad international, budget
8 Pinch of Yum ~1,483 Comfort food, international
9 Omnivore's Cookbook ~772 Modern Chinese
10 Hot Thai Kitchen ~464 Most authentic Thai
11 Spoon Fork Bacon ~911 Asian fusion
12 Korean Bapsang ~275 Authentic Korean
13 Minimalist Baker ~650 Simple vegan/GF
14 Cookie and Kate ~920 Vegetarian (use generic UA)
15 Healthy Nibbles ~640 Asian healthy (use generic UA)

Total potential: ~20,000+ new recipes, majority Asian or light/healthy.

Translation Strategy

Since these are English-language sources, recipes need translation before or after insertion:

  1. Pre-insertion translation (recommended): Translate title, description, ingredients, and instructions to German before DB insert. Use LLM batch translation (GLM-5.2 via noris, or a dedicated translation model).
  2. Store both languages: Keep original English text in a separate column (title_en, description_en) for reference.
  3. Ingredient mapping: English ingredients should map to the SAME canonical names in the ingredients table (e.g., garlic→knoblauch, chicken breast→hähnchenbrust). The CANONICAL_OVERRIDES dict in ingredient_normalizer.py already includes English→German mappings.
  4. Source prefix: Use the ID prefixes listed above for each source.

Scraping Pitfalls (International Sources)

  • Cloudflare blocks ≠ "no sitemap". Just One Cookbook was initially diagnosed as "no sitemap" — actually the sitemap exists but Cloudflare returns an HTML challenge page instead of XML. Always check HTTP status codes and content-type, not just whether XML parse succeeds.
  • AI-bot blocks in robots.txt are selective. Many sites (RecipeTin Eats, Cookie and Kate, Healthy Nibbles, SkinnyTaste) block anthropic-ai, Claude-Web, GPTBot, or ChatGPT-User specifically. They do NOT block generic Mozilla/5.0 user agents. Use a standard browser-like UA, not an AI-identifying one.
  • Empty first sitemap ≠ empty site. Minimalist Baker's post-sitemap.xml returns 0 URLs but post-sitemap2.xml has 650. Always check ALL sitemap files in the index.
  • Non-standard JSON-LD. Sanjeev Kapoor has recipe data but in a non-standard JSON structure (@type: 'Recipe with single quotes instead of "Recipe"). Standard JSON-LD parsers will miss it. May need regex-based extraction or custom parser.
  • Quarterly sitemap pattern. BBC Good Food organizes recipe sitemaps by quarter (YYYY-QN-recipe.xml). Must iterate through all quarters to get full coverage. Older quarters may not exist — check availability before assuming.