9.2 KiB
HA Proactive Health Audit Methodology
Systematic health assessment of a Home Assistant instance — not tied to a specific outage, but used to find improvement opportunities, dead integrations, and hidden issues. Complements the reactive "Bulk Entity Audit" section in the SKILL.md (which diagnoses a known outage). This procedure finds problems the user doesn't know about yet.
Based on a 2026-07-12 audit session that reduced 28.5% unavailable entities to a prioritized action list.
Audit Procedure (6 Steps)
Step 1: Instance Metadata
curl -s "$HA_URL/api/config" -H "Authorization: Bearer $TOKEN" | python3 -c "
import sys,json
d=json.load(sys.stdin)
print(f'Version: {d.get(\"version\",\"?\")}')
print(f'State: {d.get(\"state\",\"?\")}')
comps=d.get('components',[])
print(f'Components: {len(comps)}')
"
Records HA version, running state, and component count. Useful for tracking growth over time and confirming the instance is responsive.
Step 2: Entity Health Statistics
import json
from collections import Counter
data = json.loads(open('/tmp/ha_states.json','rb').read().decode('utf-8','replace'))
total = len(data)
unavail = [e for e in data if e['state'] == 'unavailable']
unknown = [e for e in data if e['state'] == 'unknown']
print(f'Total entities: {total}')
print(f'Unavailable: {len(unavail)} ({len(unavail)/total*100:.1f}%)')
print(f'Unknown: {len(unknown)} ({len(unknown)/total*100:.1f}%)')
# Domain distribution
domains = Counter(e['entity_id'].split('.')[0] for e in data)
print(f'Domains (top 15): {domains.most_common(15)}')
Benchmark: <5% unavailable is healthy. 10-20% suggests stale integrations. >25% indicates systemic issues (dead integrations, network migration remnants, or a recent outage that didn't fully recover).
Step 3: Group Unavailable Entities by Source
This is the key diagnostic step — group unavailable entities by their integration/device prefix to identify which integrations are responsible for the bulk of unavailability.
groups = Counter()
for e in unavail:
eid = e['entity_id']
parts = eid.replace('.', '_').split('_')
key = '_'.join(parts[:3]) # e.g. "sensor_spoolman_spool"
groups[key] += 1
for g, c in groups.most_common(20):
print(f' {g}: {c}')
Interpretation: Groups with 50+ unavailable entities indicate an entire integration is down, not individual device failures. Common patterns:
sensor_spoolman_*→ Spoolman server offline*_hame_energy→ HAME Energy integration dead (often replaced by Modbus)select_x40_master/select_staubsauger_oben→ Vacuum "feature-not-supported" entities (normal, but bloat)*_double_take→ Face recognition dependent on Frigate (cascade failure)
Step 4: Config Entry State Audit
curl -s "$HA_URL/api/config/config_entries/entry" -H "Authorization: Bearer $TOKEN" | python3 -c "
import sys,json
from collections import Counter
data=json.load(sys.stdin)
states=Counter(e.get('state','?') for e in data)
print(f'Config entry states: {dict(states)}')
for state in ['setup_retry','setup_error']:
entries=[e for e in data if e.get('state')==state]
for e in entries:
print(f' {e.get(\"domain\")} | {e.get(\"title\")} | {e.get(\"source\")}')
"
States:
| State | Meaning | Action |
|---|---|---|
loaded |
Working | None |
setup_retry |
Discovered but can't connect | Check device/network, reconfigure |
setup_error |
Configuration failed | Check credentials, re-setup |
not_loaded |
Disabled/ignored | Clean up if no longer needed |
Common findings:
- Stale IPs from network migrations (e.g.
192.168.100.xtitles when network is now10.0.x) not_loadedentries from abandoned integrations (Nest, DLNA, UPnP, HomeKit bridges)setup_retryfrom VLAN-boundary mDNS failures (see SKILL.md "VLAN/Network Segmentation Diagnosis")
Step 5: Automation Status
autos = [e for e in data if e['entity_id'].startswith('automation.')]
on_a = sum(1 for a in autos if a['state'] == 'on')
off_a = sum(1 for a in autos if a['state'] == 'off')
print(f'Automations: {len(autos)} total, {on_a} on, {off_a} off')
# List disabled ones with friendly names
for a in autos:
if a['state'] == 'off':
name = a.get('attributes',{}).get('friendly_name', a['entity_id'])
print(f' {a["entity_id"]}: {name}')
Disabled automations are often forgotten after debugging. Check if they were disabled intentionally or left off after a fix was applied but never re-enabled. Especially check automations that were recently fixed — the fix may have been applied but the automation never reactivated.
Step 6: Staleness Detection
import datetime
now = datetime.datetime.now(datetime.timezone.utc)
stale = []
for e in data:
lc = e.get('last_changed', '')
if not lc:
continue
try:
dt = datetime.datetime.fromisoformat(lc.replace('Z', '+00:00'))
age = (now - dt).total_seconds() / 3600
if age > 24 and e['state'] not in ('unavailable', 'unknown'):
stale.append((e['entity_id'], e['state'], round(age)))
except:
pass
stale.sort(key=lambda x: x[2], reverse=True)
print(f'Stale entities (>24h, not unavail/unknown): {len(stale)}')
Large stale counts (hundreds) often correlate with update.* entities that haven't been polled since last restart, or sensors from integrations that silently stopped publishing. Cross-reference with Step 3 groups.
Synthesis: Prioritized Improvement List
Combine all 6 steps into a ranked table:
| Priority | Issue | Entities Affected | Effort | Rationale |
|---|---|---|---|---|
| 1 | Dead integration removal (HAME, Spoolman, Kia) | 194 | S | Largest reduction in unavailable count |
| 2 | Frigate recovery chain | 13+ | M | Security cascade dependency |
| 3 | Reactivate disabled automations | 7 | S | Immediate functional gain |
| 4 | Network migration cleanup | 10+ | S | Stale IPs blocking reconnection |
| 5 | Entity bloat reduction | 245+ | M | Disable unsupported vacuum entities |
Scoring criteria (adapted from ideate skill):
- Impact (30%): How many entities/users are affected?
- Feasibility (25%): Can it be done without breaking production?
- Strategic fit (20%): Does it align with the user's direction (e.g. Marstek replacing HAME)?
- Effort (15% inversely weighted): S < 30 min, M < 2h, L > 2h
- Compound potential (10%): Does fixing this unlock other improvements?
Audit Results Snapshot (2026-07-12)
| Metric | Value |
|---|---|
| HA Version | 2026.7.1 |
| Total entities | 2,282 |
| Unavailable | 650 (28.5%) |
| Unknown | 197 (8.6%) |
| Automations | 35 (25 on, 7 off) |
| Config entries | 152 loaded, 5 setup_retry, 13 not_loaded, 1 setup_error |
| Stale (>24h) | 680 |
Top Unavailable Groups
| Group | Count | Root Cause |
|---|---|---|
sensor_spoolman_spool |
101 | Spoolman server down |
select_x40_master |
85 | Vacuum feature-not-supported (normal bloat) |
select_staubsauger_oben |
48 | Same |
text/sensor/number/switch_hame_energy |
74 | HAME Energy completely dead (replaced by Marstek Modbus) |
sensor_double_take |
13 | Frigate dependency (cascade) |
light_tradfri_* |
16+ | Stale IP 192.168.100.60 (pre-migration) |
Setup Retry Entries
| Domain | Title | Source |
|---|---|---|
| lifx | Kinderzimmer 1 | dhcp |
| lifx | Kinderzimmer 2 | dhcp |
| tradfri | 192.168.100.60 | user |
| frigate | 10.0.30.104:5000 | user |
| ollama | http://10.0.30.98:11434 | user |
Disabled Automations
| Automation | Friendly Name | Likely Reason |
|---|---|---|
automation.warmwasserbereitung_an |
DHW onetimecharge | Disabled during tuning |
automation.dhw_onetimecharge_when_dhw_charging_active |
DHW onetimecharge when DHW Charging active | Related to above |
automation.tibber_charge_battery |
Tibber Charge Battery | Disabled during testing |
automation.vogelabwehr_1/2/3 |
Vogelabwehr 1/2/3 | Seasonal? |
automation.emhass_batterie_kaskadensteuerung_sma_marstek |
EMHASS Kaskadensteuerung | Fixed but not reactivated! |
Pitfalls
- Error log endpoint may 404 on some HA versions.
/api/error_logreturned 404 on HA 2026.7.1. Try/api/error/allor check via SSH instead. Don't treat the 404 as "no errors" — it means the endpoint isn't available, not that the log is empty. - Vacuum "feature-not-supported" entities inflate the unavailable count dramatically. Two Dreame vacuums contributed ~250 unavailable entities. These are normal (the integration creates entities for all possible features, unsupported ones stay unavailable). Don't try to "fix" them — consider disabling them in the entity registry to reduce noise.
update.*entities dominate the stale list. Most stale entities areupdate.*sensors that only change when an update is available. This is normal behavior, not a defect. Filter them out when assessing real staleness.- Config entry titles can be stale. A
setup_retryTRADFRI entry titled "192.168.100.60" doesn't mean the integration is trying that IP — the title is cosmetic. Check the actualdatafield in.storage/core.config_entriesfor the real configured address.