Table of Contents
- Ground truth before reading a single error
- Finding 1: the cameras caused 4,524 of the 5,077 warnings
- Finding 2: a cloud integration that reported loaded and delivered nothing
- Finding 3: 123 errors from a REST sensor that could not handle null
- Finding 4: eleven KNX addresses that would never answer
- Finding 5: 100 unavailable entities, and why that number means little
- Finding 6: the garage door that was reported open
- Finding 7: what the restart itself surfaced
- Repair order and the restart at the end
- Frequently asked questions
My production Home Assistant has been running as a Docker container on a small cloud VM since 2024, on version 2024.11.1, with about 570 entities, 73 automations and a KNX bus behind it. I had not read its log in months. In September I made myself do it properly: one full day of log, every error and warning counted and grouped before touching anything, and only then a repair list in order of priority. This post is that list. It includes two findings that needed no repair and one template that briefly reported the garage door as open.
The numbers for the 24 hours from 12 September, 15:25, to 13 September, 15:25: 10,490 log lines, 775 ERROR entries, 5,077 WARNING entries. Several entries share a cause, so before reading any of them I grouped them by the logger name in square brackets, which is the fifth field in this log format:
# one day of log, counted before reading anything
docker logs --since 24h homeassistant 2>&1 | wc -l # 10490
docker logs --since 24h homeassistant 2>&1 | grep -c ' ERROR ' # 775
docker logs --since 24h homeassistant 2>&1 | grep -c ' WARNING ' # 5077
# group by logger name (field 5), most frequent first
docker logs --since 24h homeassistant 2>&1 \
| grep -E ' (ERROR|WARNING) ' \
| awk '{print $5}' | sort | uniq -c | sort -rn | head -20Ground truth before reading a single error
Before interpreting anything I checked the things that would make the log meaningless if they were off. Container running since 27 August, one historical restart, not OOM-killed. Host with 3.8 GiB RAM, 2.3 GiB free, the container at about 800 MiB. Swap 4 GiB with 1 MiB used. Disk 63 percent full. The recorder database at 60 MiB with a 4.3 MiB write-ahead log, and a read-only SQLite quick_check returning ok. A config check inside the container with no errors. The four configuration files on the VM matched the repository by SHA-256. I also ran a static check of all 73 automations against the entity registry, with one rule for reading the result: an entity that is not registered is not automatically missing.
That took twenty minutes and ruled out the boring explanations. Memory and disk were sufficient, the database check passed, and the config on the machine matched the repository.
Finding 1: the cameras caused 4,524 of the 5,077 warnings
The cameras accounted for most of the warnings. 2,886 were warnings about slow updates from the two Tapo cameras. Another 1,638 came from two automations that set the motion detection parameters on those cameras. Their trigger was a time pattern on second 10, which fires once a minute, and they wrote the same settings every time. Both run in mode single. One call to the camera integration hung, so from that moment on every further run was refused with an already running warning, 1,440 times for one camera and 198 times for the other. The trace of the last accepted run of the first automation was from 6 September, so it had not run for a week and had logged a warning once a minute during that time.
I repaired this the same afternoon in two steps. Reloading the config entries of the two cameras released the hung calls, without restarting Home Assistant or the cameras. In the automations, a setting is now only written when the entity does not already report the target value, which was 24 places in the file. Trigger and mode stayed as they were. Switching to mode parallel or silencing the warning would have left the hung call in place and removed the only sign of it. I did not add a timeout around the camera calls, so a call that hangs in the library can block these automations again, and the already running warning is how I would notice. The example below shows the condition.
# before: fires every minute at second 10, mode single,
# and writes the same detection settings every time
triggers:
- trigger: time_pattern
seconds: "10"
mode: single
# after: only write when the setting actually differs
conditions:
- condition: template
value_template: >
{{ not is_state('switch.camera_motion_detection', 'on') }}Finding 2: a cloud integration that reported loaded and delivered nothing
The FusionSolar integration polls Huawei's cloud kiosk for daily, monthly and yearly energy. 144 failed polls in 24 hours, one every ten minutes, each followed by three errors about missing energy fields, 576 log entries in total. All four of its sensors stood at unknown. The integration itself reported loaded, which is what the integrations page shows, so nothing looked wrong from the UI.
The local Modbus sensors on the same inverter still reported a daily yield of 22.63 kWh at the same time, so I could read the yield locally while the cloud sensors stood at unknown. Before fixing the kiosk reference I checked whether anything used those four sensors: no automation, no dashboard, not the energy configuration. I disabled the integration instead of repairing it. Home Assistant answered with require_restart and failed_unload, which is its way of saying the disable only completes after a restart. What I take from this is that loaded only describes the integration's setup. Whether its data arrives has to be checked on the sensors. My values have come from the local Modbus registers for a year, and the cloud integration was still set up although nothing used its sensors.
Finding 3: 123 errors from a REST sensor that could not handle null
The my-PV heating rod exposes its state over a local HTTP endpoint, read by the REST sensor pattern. 123 update errors in the day, 110 of them on the battery charge sensor: the API returned null for a field, the template passed it through, and Home Assistant refused the string None as a value with unit W. Eight further timeouts on the same endpoint. The fix treats missing data as unavailable instead of raising an error:
# rest.yaml, the sensor that failed 110 times in a day
- name: "Battery charge"
unit_of_measurement: W
# the API returns null while the device reboots; "None" is not a number
value_template: >
{{ value_json.m2sum | default(0, true) | float(0) }}
availability: >
{{ value_json.m2sum is number }}One detail cost an extra round: my first version fell back to the string unknown, which produced the same conversion error, because Home Assistant 2024.11 writes the native value before it re-evaluates availability. The version above, a numeric fallback plus a separate availability template, went through rest.reload without a restart, and the sensor showed 0 W a minute later. On this core version a transient numeric fallback can still be written during the transition. I have only tested this on 2024.11.
Finding 4: eleven KNX addresses that would never answer
Roughly once an hour the KNX integration logged read timeouts for eleven group addresses, the operating mode state addresses of eleven thermostats. The errors predate everything else in the log, so this was not new. I looked the addresses up in the ETS project: they belong to a forced-mode object whose read flag is disabled. Home Assistant was sending a read request to an object that is configured never to answer. There is no readable mode feedback on these devices.
I removed the eleven state addresses from the YAML instead of guessing replacements or disabling state synchronisation globally. The mode commands, temperature and setpoint feedback are untouched. Home Assistant now has no independent confirmation of which mode a thermostat is in. For that, a proper feedback object would have to be provided in ETS. It is the same problem I wrote about when the thermostats dropped to standby after every restart.
Finding 5: 100 unavailable entities, and why that number means little
100 of 570 state objects were unavailable, 38 of them restored from the registry with no live source. Grouped: 44 Tapo entities from a camera that is deliberately switched off, 19 from the mobile app on phones, 13 TP-Link, 11 automations that are disabled leftovers, and a handful of templates and helpers. Two of those needed action: a smart plug has been in setup_in_progress since 28 August and is not reachable on its ports, and HACS has been asking for re-authentication since 8 September. The other 90 or so were expected.
I mention this because bulk-deleting unavailable entities is a popular cleanup step, and here it would have removed entities that automations reference. I checked each group for use before removing anything.
Finding 6: the garage door that was reported open
This one came a few days later, during the restart for the KNX changes, and it is the one I consider most serious. The raw door contact went unavailable for a moment while the KNX integration came up. A template sensor derived the door state from it with the expression not is_state(contact, 'on'). Unavailable is not on, so for that moment the template reported the garage door as open, and the dashboard showed it.
I checked the recorder data from shortly before and after the restart. I found no switch command for the door, no knx.send to its address and no triggered automation. That is no proof that Home Assistant sent nothing, because I have not verified that all of this would have been recorded. Whether the door moved I can neither prove nor rule out, because the bus telegrams were not recorded. The template now has an availability condition and becomes unavailable together with its source instead of inventing a state. That change went live through a template reload, no restart needed.
# before: unavailable counted as "open"
- binary_sensor:
- name: "Garage door open"
state: "{{ not is_state('binary_sensor.garage_contact', 'on') }}"
# after: the template becomes unavailable together with its source
- binary_sensor:
- name: "Garage door open"
state: "{{ not is_state('binary_sensor.garage_contact', 'on') }}"
availability: >
{{ states('binary_sensor.garage_contact') not in ['unknown', 'unavailable'] }}Finding 7: what the restart itself surfaced
The restart on 16 September, done only after the changes above were prepared and checked, brought up two more. The old Shelly custom integration threw a TypeError on shutdown, made blocking calls in the event loop and left WebSocket errors behind. The boiler temperature sensor had two problems. Its template used float without a fallback for an unavailable source, and it published 701 where the precise sensor showed 70.1 degrees, a scaling error. Three disabled automations compare that sensor against 50 and 55 degrees, so it had to be fixed before any of them is enabled again.
The boiler template now checks that its source is available and numeric, uses float(0) and has the scaling corrected by a factor of ten. After the reload the sensor showed 66.1 instead of 661 degrees. For the Shelly integration I wrote a small compatibility module that moves the library import and shutdown into the executor, serialises the connection checks and closes the sockets with a bounded wait, plus four regression tests. It is a local patch on top of a custom component and has to be re-checked after every HACS update of that integration. The second restart came up with zero ERROR lines, the boiler at 66.2 degrees and the garage door closed.
Repair order and the restart at the end
The repair order was: unblock the camera automations, disable the dead cloud integration and fix the plug and the HACS login, fix the null handling, remove the KNX addresses that cannot answer, and only then clean up the disabled leftovers by actual use. Two things were off the table as diagnostic steps: a blanket restart, and a blanket deletion of unavailable entities. The restart came last, after the changes, and it was the step that found finding 6 and 7.
Three of the seven findings affected operation: the hung camera calls, the null handling, and the template that turned unavailable into an open door. The KNX read timeouts and the errors from the restart did not affect operation, but they had causes I could remove. The cloud integration was unused, and most of the 100 unavailable entities were expected. The camera automation had been logging its warning since 6 September, and I only saw it when I read the log. My update workflow has me read the log for a few minutes after each update. Between updates I had not been reading it.
Frequently asked questions
Is 775 errors a day a lot for Home Assistant?
Here it was mostly four causes repeating. I judge the number by the causes behind it and by what stopped working, and that was three things.
Should I restart Home Assistant when the log fills up?
Not as a first step. I would save the log and the automation traces first and look at the repeating entries, because a restart changes the state I want to examine.
Why does a template report a state when its source is unavailable?
Because not is_state(x, 'on') is also true for unknown and unavailable. An availability template that excludes those two source states fixes it.



