Files
hermes-skills/food-nutrition/meal-planning/references/semantic-ontology-classification.md
T

5.5 KiB

Semantic Ontology Classification — Implementation

Date: 2026-06-28 Status: Implemented, running in background

Overview

Rule-based semantic classifier that annotates every recipe across 11 ontology dimensions. Pure Python keyword matching — no LLM/API calls needed. Processes ~500 recipes/batch at ~50-500/s throughput.

Script

profiles/nutrition-coach/scripts/semantic_classifier.py

  • Incremental: LEFT JOIN recipe_semantics rs ON r.id = rs.recipe_id ... WHERE rs.recipe_id IS NULL
  • Batch insert: 500 recipes per executemany()
  • Fallback to one-by-one insert on batch error
  • Confidence: 0.85 (rule-based, not LLM-derived)
  • Runtime: ~10 min for 54k recipes

Database Table: recipe_semantics

CREATE TABLE recipe_semantics (
    recipe_id VARCHAR(255) PRIMARY KEY,  -- ⚠️ 255, not 128!
    diet_type ENUM('vegan','vegetarian','pescatarian','flexitarian','omnivore') NOT NULL,
    meal_course JSON NOT NULL,          -- breakfast,starter,soup,salad,main,side,dessert,snack,beverage,bread
    cuisine JSON NOT NULL,              -- italian,french,german,asian,mexican,thai,japanese,chinese,greek,spanish,turkish,middle_eastern,american,nordic,african,caribbean,austrian,swiss
    cooking_method JSON NOT NULL,       -- baking,frying,boiling,grilling,steaming,stewing,roasting,braising,raw,microwaving
    protein_source JSON NOT NULL,       -- beef,pork,poultry,fish,seafood,plant,dairy,egg,none
    allergens JSON NOT NULL,            -- gluten,eggs,dairy,nuts,soy,fish,shellfish,peanuts,sesame,celery,mustard,sulphites,lupin,molluscs
    flavor_profile JSON NOT NULL,       -- sweet,savory,salty,sour,bitter,umami,spicy,smoky,creamy,fresh
    complexity ENUM('quick','simple','moderate','elaborate') NOT NULL,
    season JSON NOT NULL,               -- spring,summer,autumn,winter,year_round
    caloric_density ENUM('light','balanced','hearty','indulgent') NOT NULL,
    preparation_style JSON NOT NULL,    -- one_pot,sheet_pan,pan_fried,oven_baked,grilled,no_cook,slow_cooked,meal_prep
    confidence FLOAT DEFAULT 0.85,
    analyzed_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4;

Ontology Dimensions

1. Diet Type (ENUM, single value)

Priority: explicit tags (vegan/vegetarisch) → meat detection → fish/seafood → dairy/egg → default omnivore.

2. Meal Course (JSON array)

Keyword matching on title+category+tags+ingredients+description. Fallback to category field if no match. Default: main.

3. Cuisine (JSON array)

18 cuisine types with German+English keywords. Also checks tags for country names. Default: ["german"] if deutsch in tags.

4. Cooking Method (JSON array)

10 methods. Default: ["boiling"] if no match.

5. Protein Source (JSON array)

Ordered priority: beef → pork → poultry → fish → seafood → egg → dairy → plant → none. Multiple can match.

6. Allergens (JSON array)

14 EU-regulated allergens. Keyword-based, broad coverage (German+English).

7. Flavor Profile (JSON array)

10 flavor dimensions. Desserts get sweet added if not already detected. Default: ["savory"].

8. Complexity (ENUM)

Based on max(prep_time, cook_time, total_time): ≤15min=quick, ≤30min=simple, ≤60min=moderate, >60min=elaborate. Falls back to tags (schnell/einfach).

9. Season (JSON array)

Keyword-based. Default: ["year_round"].

10. Caloric Density (ENUM)

From calories: <300=light, <500=balanced, <800=hearty, ≥800=indulgent. Falls back to keyword estimation (kalorienarm, deftig, dessert→indulgent).

11. Preparation Style (JSON array)

8 styles. Falls back to cooking method if no direct match.

Cronjob

semantic-classifier — daily at 07:00, delivers to telegram:223926918. Runs after scraper (06:00) and embeddings (06:30).

Pitfalls (Galera-Specific)

1. Foreign Key creation fails on collation mismatch

The recipes table uses utf8mb4_unicode_ci but new tables default to utf8mb4_general_ci. FK creation fails with errno 150. Fix: either omit the FK (preferred — no cascade needed) or specify COLLATE utf8mb4_unicode_ci on the new table.

2. JOINs between tables with different collations fail

LEFT JOIN recipe_semantics rs ON r.id = rs.recipe_id fails with Illegal mix of collations (utf8mb4_unicode_ci,IMPLICIT) and (utf8mb4_general_ci,IMPLICIT). Fix: add COLLATE utf8mb4_unicode_ci to the JOIN condition.

3. VARCHAR(128) is too short for some recipe IDs

Chefkoch IDs with encoded state fragments (e.g., ck_Kinderueberraschungsbrot.html#eyJvYm...) can be up to 248 chars. Use VARCHAR(255) for any column storing recipe IDs.

4. Time and calorie columns contain string values

Some prep_time, cook_time, total_time, and calories values are stored as strings (e.g., "PT20M" or "450 kcal"), not integers. Always wrap with a safe_int() helper that catches ValueError/TypeError before using in comparisons or arithmetic.

5. Concurrent writes cause Galera deadlocks

Running the semantic classifier and embedding generator simultaneously causes Deadlock found when trying to get lock (1213). Fix: add retry logic with rollback+backoff (safe_db_execute()), or run sequentially. The daily cronjob schedule (06:00 scraper → 06:30 embeddings → 07:00 classifier) ensures sequential execution.

6. Galera FK constraints are fragile

Even without collation issues, Galera's certification-based replication can reject DDL that works on standalone MySQL. Prefer logical relationships (application-level integrity checks) over physical FKs when working with Galera.