Files
hermes-skills/smart-home/smart-home-automation/references/ha-health-audit.md
T

9.2 KiB

HA Proactive Health Audit Methodology

Systematic health assessment of a Home Assistant instance — not tied to a specific outage, but used to find improvement opportunities, dead integrations, and hidden issues. Complements the reactive "Bulk Entity Audit" section in the SKILL.md (which diagnoses a known outage). This procedure finds problems the user doesn't know about yet.

Based on a 2026-07-12 audit session that reduced 28.5% unavailable entities to a prioritized action list.

Audit Procedure (6 Steps)

Step 1: Instance Metadata

curl -s "$HA_URL/api/config" -H "Authorization: Bearer $TOKEN" | python3 -c "
import sys,json
d=json.load(sys.stdin)
print(f'Version: {d.get(\"version\",\"?\")}')
print(f'State: {d.get(\"state\",\"?\")}')
comps=d.get('components',[])
print(f'Components: {len(comps)}')
"

Records HA version, running state, and component count. Useful for tracking growth over time and confirming the instance is responsive.

Step 2: Entity Health Statistics

import json
from collections import Counter

data = json.loads(open('/tmp/ha_states.json','rb').read().decode('utf-8','replace'))
total = len(data)
unavail = [e for e in data if e['state'] == 'unavailable']
unknown = [e for e in data if e['state'] == 'unknown']

print(f'Total entities: {total}')
print(f'Unavailable: {len(unavail)} ({len(unavail)/total*100:.1f}%)')
print(f'Unknown: {len(unknown)} ({len(unknown)/total*100:.1f}%)')

# Domain distribution
domains = Counter(e['entity_id'].split('.')[0] for e in data)
print(f'Domains (top 15): {domains.most_common(15)}')

Benchmark: <5% unavailable is healthy. 10-20% suggests stale integrations. >25% indicates systemic issues (dead integrations, network migration remnants, or a recent outage that didn't fully recover).

Step 3: Group Unavailable Entities by Source

This is the key diagnostic step — group unavailable entities by their integration/device prefix to identify which integrations are responsible for the bulk of unavailability.

groups = Counter()
for e in unavail:
    eid = e['entity_id']
    parts = eid.replace('.', '_').split('_')
    key = '_'.join(parts[:3])  # e.g. "sensor_spoolman_spool"
    groups[key] += 1

for g, c in groups.most_common(20):
    print(f'  {g}: {c}')

Interpretation: Groups with 50+ unavailable entities indicate an entire integration is down, not individual device failures. Common patterns:

  • sensor_spoolman_* → Spoolman server offline
  • *_hame_energy → HAME Energy integration dead (often replaced by Modbus)
  • select_x40_master / select_staubsauger_oben → Vacuum "feature-not-supported" entities (normal, but bloat)
  • *_double_take → Face recognition dependent on Frigate (cascade failure)

Step 4: Config Entry State Audit

curl -s "$HA_URL/api/config/config_entries/entry" -H "Authorization: Bearer $TOKEN" | python3 -c "
import sys,json
from collections import Counter
data=json.load(sys.stdin)
states=Counter(e.get('state','?') for e in data)
print(f'Config entry states: {dict(states)}')
for state in ['setup_retry','setup_error']:
    entries=[e for e in data if e.get('state')==state]
    for e in entries:
        print(f'  {e.get(\"domain\")} | {e.get(\"title\")} | {e.get(\"source\")}')
"

States:

State Meaning Action
loaded Working None
setup_retry Discovered but can't connect Check device/network, reconfigure
setup_error Configuration failed Check credentials, re-setup
not_loaded Disabled/ignored Clean up if no longer needed

Common findings:

  • Stale IPs from network migrations (e.g. 192.168.100.x titles when network is now 10.0.x)
  • not_loaded entries from abandoned integrations (Nest, DLNA, UPnP, HomeKit bridges)
  • setup_retry from VLAN-boundary mDNS failures (see SKILL.md "VLAN/Network Segmentation Diagnosis")

Step 5: Automation Status

autos = [e for e in data if e['entity_id'].startswith('automation.')]
on_a = sum(1 for a in autos if a['state'] == 'on')
off_a = sum(1 for a in autos if a['state'] == 'off')
print(f'Automations: {len(autos)} total, {on_a} on, {off_a} off')

# List disabled ones with friendly names
for a in autos:
    if a['state'] == 'off':
        name = a.get('attributes',{}).get('friendly_name', a['entity_id'])
        print(f'  {a["entity_id"]}: {name}')

Disabled automations are often forgotten after debugging. Check if they were disabled intentionally or left off after a fix was applied but never re-enabled. Especially check automations that were recently fixed — the fix may have been applied but the automation never reactivated.

Step 6: Staleness Detection

import datetime
now = datetime.datetime.now(datetime.timezone.utc)
stale = []
for e in data:
    lc = e.get('last_changed', '')
    if not lc:
        continue
    try:
        dt = datetime.datetime.fromisoformat(lc.replace('Z', '+00:00'))
        age = (now - dt).total_seconds() / 3600
        if age > 24 and e['state'] not in ('unavailable', 'unknown'):
            stale.append((e['entity_id'], e['state'], round(age)))
    except:
        pass
stale.sort(key=lambda x: x[2], reverse=True)
print(f'Stale entities (>24h, not unavail/unknown): {len(stale)}')

Large stale counts (hundreds) often correlate with update.* entities that haven't been polled since last restart, or sensors from integrations that silently stopped publishing. Cross-reference with Step 3 groups.

Synthesis: Prioritized Improvement List

Combine all 6 steps into a ranked table:

Priority Issue Entities Affected Effort Rationale
1 Dead integration removal (HAME, Spoolman, Kia) 194 S Largest reduction in unavailable count
2 Frigate recovery chain 13+ M Security cascade dependency
3 Reactivate disabled automations 7 S Immediate functional gain
4 Network migration cleanup 10+ S Stale IPs blocking reconnection
5 Entity bloat reduction 245+ M Disable unsupported vacuum entities

Scoring criteria (adapted from ideate skill):

  • Impact (30%): How many entities/users are affected?
  • Feasibility (25%): Can it be done without breaking production?
  • Strategic fit (20%): Does it align with the user's direction (e.g. Marstek replacing HAME)?
  • Effort (15% inversely weighted): S < 30 min, M < 2h, L > 2h
  • Compound potential (10%): Does fixing this unlock other improvements?

Audit Results Snapshot (2026-07-12)

Metric Value
HA Version 2026.7.1
Total entities 2,282
Unavailable 650 (28.5%)
Unknown 197 (8.6%)
Automations 35 (25 on, 7 off)
Config entries 152 loaded, 5 setup_retry, 13 not_loaded, 1 setup_error
Stale (>24h) 680

Top Unavailable Groups

Group Count Root Cause
sensor_spoolman_spool 101 Spoolman server down
select_x40_master 85 Vacuum feature-not-supported (normal bloat)
select_staubsauger_oben 48 Same
text/sensor/number/switch_hame_energy 74 HAME Energy completely dead (replaced by Marstek Modbus)
sensor_double_take 13 Frigate dependency (cascade)
light_tradfri_* 16+ Stale IP 192.168.100.60 (pre-migration)

Setup Retry Entries

Domain Title Source
lifx Kinderzimmer 1 dhcp
lifx Kinderzimmer 2 dhcp
tradfri 192.168.100.60 user
frigate 10.0.30.104:5000 user
ollama http://10.0.30.98:11434 user

Disabled Automations

Automation Friendly Name Likely Reason
automation.warmwasserbereitung_an DHW onetimecharge Disabled during tuning
automation.dhw_onetimecharge_when_dhw_charging_active DHW onetimecharge when DHW Charging active Related to above
automation.tibber_charge_battery Tibber Charge Battery Disabled during testing
automation.vogelabwehr_1/2/3 Vogelabwehr 1/2/3 Seasonal?
automation.emhass_batterie_kaskadensteuerung_sma_marstek EMHASS Kaskadensteuerung Fixed but not reactivated!

Pitfalls

  • Error log endpoint may 404 on some HA versions. /api/error_log returned 404 on HA 2026.7.1. Try /api/error/all or check via SSH instead. Don't treat the 404 as "no errors" — it means the endpoint isn't available, not that the log is empty.
  • Vacuum "feature-not-supported" entities inflate the unavailable count dramatically. Two Dreame vacuums contributed ~250 unavailable entities. These are normal (the integration creates entities for all possible features, unsupported ones stay unavailable). Don't try to "fix" them — consider disabling them in the entity registry to reduce noise.
  • update.* entities dominate the stale list. Most stale entities are update.* sensors that only change when an update is available. This is normal behavior, not a defect. Filter them out when assessing real staleness.
  • Config entry titles can be stale. A setup_retry TRADFRI entry titled "192.168.100.60" doesn't mean the integration is trying that IP — the title is cosmetic. Check the actual data field in .storage/core.config_entries for the real configured address.