Chapter 19

Health, Network Map and diagnostics

Use Health & Diagnostics to read Home Assistant, the bridge, the fabric, sessions, subscriptions and the network as separate layers, then use Network Map, live events, the system log, metrics and System Information to build diagnostic evidence that leaks nothing about your environment.

One “offline” can have six different causes

When a controller shows No Response, all it means is that the controller cannot reliably use some Matter endpoint right now. The cause may be a dropped Home Assistant WebSocket, a bridge that is stopped or failed, a fabric that has not been commissioned yet, a controller session that is not active, a subscription that has disappeared, mDNS advertising on the wrong interface, or a mapping that failed on the endpoint itself. Health is worth using because it shows all of those layers at the same time, so you locate the problem before you act instead of commissioning again and again.

Stable 2.0.55 has reachable /health and /network-map routes. HealthPage combines the Health Dashboard, the Live Event Log, Network Diagnostics, System Information and the Translation Editor. Network Map visualizes the controller vendor groups merged by rootVendorId, together with the Hub, Bridge, Device and Failed nodes; it is not one node per fabric. These pages are Stable operations features, but they describe the state as HAMH knows it, not a guarantee about how a controller's UI behaves.

Diagnostic data is sensitive data: an exported diagnostic can contain entity names, bridge names, network interfaces, error text and Matter identity information. Share only the smallest fragment, through a trusted channel, after a person has read it and redacted it. This guide neither supplies nor asks for any real value from your environment.
LayerHealthy signalUnhealthy signalNext check
Home Assistantconnecteddisconnected, or ready failsThe HA URL, permissions, the WebSocket and latency
Bridgerunning, no failed entitiesstopped or failed, plus the status reasonThe bridge log, its settings and host resources
FabricThe controller fabrics you expect are presentZero fabrics, or records you did not expectCommissioning history and the controller-side state
Sessionpeer active, with recent activityActivity stalled for a long timeNetwork reachability and the controller hub
SubscriptionA wildcard or endpoint-specific subscription existsZero subscriptions, or rebuilt over and overLive events and session health
mDNS / interfaceBound to a reachable LAN interfaceWarnings about Docker, Thread or extra interfacesThe network settings in Chapter 20

How to read health status and connection signals

The backend's basic Health rolls the HA connection and the bridge counts into healthy, degraded or unhealthy. It is healthy only when HA is connected and no bridge is stopped or failed. It is degraded when HA is connected but a bridge is stopped or failed. It is unhealthy when HA is not connected. This is a service-level summary: it does not mean every entity is fine, and it does not mean every controller supports everything.

Detailed Health gives you, for every bridge, its status and reason, port, priority, device count, fabric count, failed entity count, controller warnings and entity diagnostics, plus a session and subscription summary. The session activity fields let you separate traffic that is still moving at the transport layer from Interaction Model commands that have stalled. Subscriptions are split further into whole-node wildcard, endpoint-specific and unknown scope. Seeing that a session exists is still not proof that control works.

A fabric is a trust relationship. A session is a secure connection over a period of time. A subscription is the controller asking the bridge to report certain data on its own. A fabric that exists with zero sessions may mean the controller hub is temporarily offline or the network cannot reach it. A session with zero subscriptions may mean you can read and write but get no continuous state reporting. Sessions and subscriptions both present while one endpoint has failed points at the mapping and the source entity first.

Liveness and readiness: Stable 2.0.55 exposes health live and ready endpoints. live only says the process can answer; ready is decided by whether Home Assistant is connected. They are made for process monitoring, and must not be read as every bridge, fabric and device being healthy.

DiagnosticsPage exists, but v2.0.55 has no route to it

An important version fact: DiagnosticsPage.tsx is present in the v2.0.55 code, and HealthPage reuses the LiveEventLog from it; but routes.tsx registers no separate DiagnosticsPage route. So in a Stable 2.0.55 tutorial you cannot claim there is a separate Diagnostics page in the sidebar, and you cannot hand out an invented URL. The entry point you can actually use is the live event area embedded in the Health & Diagnostics page, together with Export Diagnostic.

FeatureRelease channelMaturityController support
The Health and Network Map routesStable 2.0.55Stable operations UINot tied to a particular controller; how complete the data is depends on the live connections
The Live Event Log inside HealthStable 2.0.55Stable diagnostic componentThe event types are what HAMH observes; no guarantee a controller exposes the same detail
A separate DiagnosticsPagePresent in the codeNo route registeredMust not be written up as a reachable Stable page
Translation EditorEmbedded in Health in Stable 2.0.55Local UI override toolDoes not affect the controller's language
Server Mode and experimental endpoint healthObservable within StableExperimental (experimental-in-Stable)Varies by controller type

Network Map draws its graph from the bridge and device APIs. It is not a packet capture, and it is not the physical topology your router sees. The controller vendor groups on the graph are built and de-duplicated from the rootVendorId of the commissioned fabric records; several fabrics from the same vendor merge into one node, so the node and edge counts are not an exact fabric count and not a one-to-one topology. For exact fabric records, read Detailed Health. The lines do not show live packet direction or signal strength either. A Failed node is an endpoint HAMH failed to build, which is not the same as a fault a controller declares.

From the summary to a root cause you can verify

  1. Open Health & Diagnostics and read the global summary first

    Check the version, uptime, Home Assistant connected, and the running / total bridge ratio. Do not press Restart first. If HA is disconnected, deal with the shared upstream problem. If only one bridge has failed, narrow down to that bridge.

  2. Expand bridge and fabric health one bridge at a time

    Compare the device count, fabric count, failed entity count and controller warnings. Then read the session and subscription summary for each fabric. Record only "present or absent, recent or stalled, and the trend in the counts"; never copy fabric or node identity values into a record you will share.

  3. Build a timeline with the Live Event Log

    Keep the event types you need: state update, command received, entity error / warning, session opened / closed, subscription changed, bridge started / stopped. Use the filter chips to isolate the symptom, and reproduce one low-risk action first. Clear Events only clears the list you are watching in the frontend; it is not a fix.

  4. Run Network Diagnostics

    Read the pass / warn / fail results and the recommendations, then expand the interface table. Confirm that the bound interface is a reachable LAN interface, that the IPv4 switch matches your deployment, and that IPv6 is present and does not rely only on a link-local address that cannot cross subnets. Do not paste the real addresses on screen into a ticket.

  5. Switch to Network Map and cross-check

    Confirm that the Hub-to-bridge links, the controller vendor groups built and de-duplicated by rootVendorId, and the Device nodes all match what you expect; for individual fabric records, go back to Detailed Health. Press Refresh Data to reload the data. Dragging a node only changes the layout held in the browser's local storage; Undo takes back one move, Reset Layout clears the saved positions, and Fullscreen only changes the view.

  6. Export a diagnostic or change settings only at the end

    If you still cannot pin it down, press Export Diagnostic, then read the file offline and redact every identity, network address, entity name and secret fragment inside the error text. When you need to change mDNS, the firewall or a VLAN, go to Chapter 20. When you need to restart or reset, go back to the backup and restore procedures in Chapter 18 and Chapter 22.

Network Map, logs, metrics and the translation tool

Network Map interactions

Network Map supports zoom, pan, the MiniMap, Controls, node dragging, single-step Undo, Reset Layout, Refresh Data and Fullscreen. The positions you drag are saved in the current browser's local storage; they do not survive a different browser, clearing site data, or Reset Layout. Refresh Data fetches the bridges again and then loads the devices for each bridge. It does not perform a Matter factory reset, and it does not repair failed entities.

The system log and Live Event

The system log is where you read startup, mDNS, the HA connection, bridge failures, plugin problems and resource warnings. Live Event leans toward the runtime ordering of state, commands, sessions and subscriptions. Line the two up: mark the time of the symptom first, then read the neighboring events and log lines. Protocol debug can include per-packet information, so turn it on only temporarily, protect it strictly, and put the level back when you are done.

Metrics and System Information

The metrics JSON contains uptime, heap and RSS, bridge total / running / stopped / failed, device and fabric totals, HA connected, and the registry counts; the Prometheus format also carries a status and device count label for each bridge. System Information shows the version and the runtime environment, which is what you cross-check the architecture, the Node version and the resources against. Watch the trend when you monitor; do not read a single spike as a leak.

Translation Editor

The Translation Editor lets you pick a language, search keys, filter by missing or edited, edit local overrides, reset one key at a time, reset all, copy or export JSON, import JSON, and create and remove a custom language. Overrides and custom languages are restored from the browser's local storage. They affect the HAMH UI in this browser: they do not change Home Assistant's translations, and they are never pushed to a controller.

Before you import a translation, confirm that the JSON holds only string keys and values, with no description of your environment and no secrets. Reset on a key returns it to the built-in string; Reset All clears the local edits for that language. If you delete a custom language, that browser no longer offers it. To move one between browsers, Export JSON first, Import it on the target, and keep a copy you can roll back to.

ToolQuestion it answersWhat it will not do
Health & DiagnosticsWhether HA, the bridges, the fabrics, the sessions and the subscriptions are healthyProve that every controller UI supports something
Live EventWhich events happened around the symptomReplace the server log long term
Network DiagnosticsWhether the interface and mDNS configuration look suspectChange your router or firewall for you
Network MapThe logical relationships HAMH sees right nowShow packet paths or Wi-Fi signal
MetricsTrends in resources and countsBring its own access control once you expose it
Translation EditorUI text overrides in this browserTranslate a controller or change backend behavior

Three diagnostic paths

Every controller fails at once: start with HA connected and with whether every bridge went degraded together; if HA is disconnected, fix the shared connection first. If the bridges are all running, read the bound interface and the mDNS warnings in Network Diagnostics. That hits the shared root cause more often than commissioning each bridge again.

Only one controller fails: the fabric is still there, but its session or subscription has stalled while the other controllers' sessions are fine. Check that controller's hub and the VLAN and mDNS path first; do not factory reset the whole bridge. A difference in controller support can also make a single device type go missing.

Only one entity fails: the bridge and the sessions are healthy, and Network Map shows a failed device or the Health entity diagnostic gives a reason. Go back to the bridge detail page and fix the HA state, the filter or the mapping; Refresh Data in Network Map is there to confirm the result, not to repair anything.

Minimal reproduction: pick a low-risk endpoint that is not a lock, an alarm or a high-power appliance, send one controller command and make one HA state change, and watch the command, the state update and the subscription. Do not use a safety-critical device for a diagnostic test.

Common pitfalls and safe fallback points

  1. Health shows unhealthy

    Check: HA connected and ready; unhealthy reflects an HA that is not connected before anything else. Safe fix: confirm the HA service, URL, permissions and WebSocket, and do not reset the fabric first; look at bridge recovery once the connection is back.

  2. Health is healthy, but the controller still shows No Response

    Check: session activity and subscriptions on the matching fabric, Network Diagnostics, and the controller hub. Safe fix: fix mDNS, IPv6 or multicast, or the controller hub, and restart a single bridge if you have to; healthy is not a guarantee that an endpoint is supported.

  3. Network Map keeps loading

    Check: whether the Bridges API and the devices call for every bridge both finish, the browser console and WebSocket, and whether one bridge is very large. Safe fix: go back to Bridges to find the one that failed and press Refresh Data once; do not press Reset Layout over and over, because it only clears positions.

  4. Live Event shows Offline

    Check: whether the reverse proxy forwards the WebSocket upgrade, and whether the base path matches. Safe fix: fix the proxy, then reload Health; do not conclude from a dropped frontend socket that a Matter fabric is gone.

  5. Network Diagnostics is bound to an extra interface

    Check: what each interface is for, container and Thread interfaces in particular. Safe fix: bind the real LAN interface in the start options, restart the service and run the diagnostics again; take the interface name from your own read-only screen.

  6. Translation overrides disappear in another browser

    Check: whether they are in the original browser's local storage, and whether you ever exported them. Safe fix: import from a JSON file you trust; if the data is already cleared and there is no export, go back to the built-in translations, and do not mistake a translation file for a system backup.

FAQ

Can I open DiagnosticsPage directly in Stable 2.0.55?
Nothing here lets you claim that you can. The component file exists, but routes.tsx registers no separate route; use the Live Event, Network Diagnostics, System Information and Export Diagnostic embedded in Health.
Does a healthy Health mean every device can be controlled?
No. The global healthy state mainly looks at HA and at bridges that are stopped or failed; you still have to check failed entities, sessions, subscriptions, the mapping and controller support.
Is Network Map a live packet diagram of the network?
No. It builds a logical graph from the bridge and device data and lets you rearrange the layout; it does not show physical routing, latency or signal strength.
Can I expose the metrics endpoint straight to a monitoring platform?
You should not expose it directly. It carries the version, resource figures, bridge labels and counts, which is still operational information; keep it on a controlled network, behind proxy authentication and least privilege.
Does clearing Live Event fix a subscription?
No. Clear Events only clears the events the frontend has collected for display; session and subscription problems are fixed from the network, the controller and bridge health.
Does the Translation Editor change what other users see?
Local overrides are saved in the current browser; no other browser and no controller picks them up automatically. Moving them anywhere else takes a manual export and import.

Pinned-version sources