Chapter 22

Backup, restore, migration and full troubleshooting

Separate what a bridge config backup covers from what a full identity snapshot covers, then rehearse preview, selective restore, migration and rollback; finally work through commissioning, No Response, mDNS, failed entities, low-resource hosts and disaster resets with "symptom → check → safe fix".

Downloading a file is not the same as being able to restore

The recoverable state of a Matter bridge covers at least the bridge config, the entity mappings, the Matter identity and fabric credentials, the plugin enabled / disabled state of each bridge, and some administrative assets. If the identity is missing you can rebuild the bridge settings, but the controllers still have to commission again; and starting the same identity on two hosts at once produces duplicate services and session and mDNS behavior you cannot predict.

Backup & Restore in Stable 2.0.55 Settings offers all of these together: downloadable config and full backups, a restore preview, bridge selection, overwrite, include mappings, restore identity, internal snapshots, auto backup and retention. A restore can overwrite your current settings and ask the application to restart, so before every restore you also download a backup of the current state and keep an explicit rollback point.

A full backup is a secret: an archive that contains the identity holds Matter key material and fabric credentials. That is what keeps the existing commissioning alive, and it is also why you must never share it, commit it to version control or put it in public cloud storage. Keep it on encrypted, controlled offline storage with revocable access, and verify file integrity on a schedule.
QuestionEvidence neededSafe fallback point
Just a mis-edited settingA bridge export or config backup, and a diffA selective restore that leaves identity alone
The host or the volume failedA full backup, its version and its checksumStop the old instance, then restore in full
A controller shows No ResponseHealth, sessions, mDNS, the logsFix the network or restart that one bridge; do not reset first
The identity is confirmed unusableProof the backup will not restore, plus a controller cleanup planA factory or disaster reset, as the last resort

Four kinds of export and exactly what each covers

Bridge export JSON (Chapter 18) mainly holds bridge definitions, and you use it to move filters and basic settings. A config backup ZIP holds the Bridges and entityMappings inside backup.json and marks itself explicitly as containing no identity; its archive carries only the per-bridge plugin enabled / disabled flags that exist, it carries no bridge icon, and after a restore you still have to commission again.

A full backup ZIP asks for the identity to be included, but it only packs one in where that bridge's identity directory actually exists, and it adds the matching bridge icon; on the plugin side you still get nothing but the per-bridge enabled / disabled flags. includesIdentity: true is metadata about the request as a whole, not proof that every bridge was packed successfully. A stored snapshot uses the same conditional identity and icon scope; it can be manual or automatic, and retention deletes the older files. After a restore, compare identitiesRestored against the number of bridges you selected before you decide whether the identities survived intact.

The code does not pack every storage asset into these archives: installed plugin packages, installed-plugins.json, per-plugin config, storage and secrets, Lock Credentials and device-images are all outside the backup scope; a standalone mapping profile export is not an extra asset that gets collected for you either. App settings such as Basic Auth, auto recovery and the backup preference do not appear in the backup.json restore scope. Record the non-secret settings somewhere else and keep a trusted original of each asset; secrets go only into a dedicated secret manager and an encrypted backup, never into your operations notes.

What Mapping means here: entity mappings go into both the config backup and the full backup, and a restore lets you choose whether to apply them; a mapping profile is a portable rule file on a separate path. Neither one replaces the Matter identity.

Version, maturity and support limits

CapabilityRelease channelProduct maturityController support
Backup, preview, restore and snapshotsStable 2.0.55A Stable operations capability; rehearse the destructive partsA correct identity can keep the fabric, but controller reconnection still depends on the network
Auto RecoveryStable 2.0.55A Stable settingRetries failed bridges only; it does not touch a running bridge
Update CheckerStable 2.0.55An informational featureShows different update guidance for add-on, Docker and npm installs
Restoring Server Mode and Camera / SecurityTheir state can be backed up within StableExperimental (experimental-in-Stable)Restoring the data is not the same as full controller support

Update Checker compares current against latest from the release information and shows the release notes and the runtime environment it detected; it does not install an update for you. The update method depends on whether you run the add-on, Docker or npm, so do not turn "a new version exists" into "we have upgraded safely". Read the release notes, take a backup, confirm which image or package you can roll back to, and only then update inside a maintenance window.

Auto Backup is enabled by default with retention set to 5, and an automatic snapshot is attempted at graceful shutdown rather than on a periodic schedule. Auto Recovery is enabled by default with an interval of 60 seconds; the UI range is 10–3600 seconds. It only restarts failed bridges periodically, and it also triggers recovery after HA reconnects; a running bridge is left alone. If the root cause is a port conflict, a bad setting or not enough memory, Recovery only produces retry records — it will not repair the configuration.

Restoring across versions: the archive carries version metadata, but the sources promise nothing about moving freely between any two versions in either direction. The safest route is to restore successfully on the same Stable version first and then upgrade by the release notes; before you cross versions, keep the original archive and an environment you can still boot on the old version.

Create, verify, restore and roll back

  1. Build a backup inventory

    In Settings → Backup & Restore, create a manual snapshot, and download a config backup and a full backup as well; record the HAMH version, the date, the number of bridges, the number of mappings, and whether identity and icons are included. Never copy an identity or a secret value into the inventory.

  2. Move it off the host and verify it

    Copy the archive to an encrypted offline location, compute a checksum, and test that the ZIP opens and that you can read the README and the backup metadata. Do not unpack the identity into a shared folder, and clear the temporary files once the test is done.

  3. Create a rollback point before any change

    Before an upgrade, a plugin install, a migration or a restore, take one more full snapshot of the current state. Record the image or package version you are on, the storage mount and the non-secret settings you need; make sure the old image is still available.

  4. Run Restore Preview first

    Upload a trusted archive and check the version, the creation time, includes identity, whether each bridge already exists, and the mapping count. Select only the bridges you want to restore; leave overwrite off as it comes, and set include mappings and restore identity to match your plan.

  5. Stop the conflicting source, then restore

    When you migrate, stop the old HAMH first so that one identity has exactly one active instance. Run Restore and read the restored, skipped and errors counts; if the UI asks for a restart, use a graceful restart and wait for HA and the bridges to settle.

  6. Verify identity stability layer by layer

    Confirm the version, HA connected, bridges running, the device and fabric counts, failed entities, sessions and subscriptions, and the mDNS interface; then use a low-risk endpoint to test both directions, controller to HA and HA to controller. "The page loads" is not a complete success.

  7. If it fails, roll back instead of overwriting again

    Stop the new instance, keep the redacted logs, and either restore the pre-change snapshot or start the old instance again; keep exactly one active instance at a time. If the archive reported errors, do not go on to overwrite more bridges — work out the version, the storage permissions and the missing files first.

Migration, identity stability and low-resource operation

Migration and identity stability

The safe migration order is "full backup → verify the archive → stop the old instance → restore the new instance with the same storage and the complete identity → start one new instance → verify". Only when the bridge ID, the identity directory and the endpoint identity are all kept together do you have a chance of the controllers not needing to commission again. Copying the bridge config alone, or letting the storage volume come up as an empty directory, will not get you there.

The stable identity feature and persistent entity identity reduce the chance that endpoints are renumbered after a restart, but a large change to the filters, the mappings or a plugin's device set can still change what the controller sees. Do not rename things, rewrite filters, switch to Server Mode or jump several versions while you migrate; prove the identity works unchanged first, then change one item at a time.

The settings and assets inventory

Keep a separate checklist of the non-secret settings: whether Basic Auth comes from the environment or from stored settings, Auto Recovery enabled and interval, backup auto and retention, the mDNS start options, the base path and the log level. For a secret value, record only that "a secret manager supplies it", never the value itself. Bridge icons are in the full backup; for device images, keep the original files ready; a local translation override has to be exported separately from the browser that holds it.

Low-resource operation

The official Low-Resource document at the same version says that HAMH loads the Matter cluster definitions, the HA registry and V8 overhead at startup, and that memory grows with the endpoint count. When resources are tight, shrink the filters first, cut the endpoint count, turn off auto composed where you do not need it, move non-essential large add-ons elsewhere, and watch the heap and RSS trends in metrics along with host OOM signals.

Force Sync may be skipped under heap pressure; a process that restarts with no stack trace, where the last thing you see is Killed or a container exit, points to OOM and is the classic clue. On plain Docker or npm you can tune the Node heap by the official guidance; in the add-on the entrypoint sets it dynamically, so do not invent a UI option that is not there. Swap is a buffer, not an answer to an unbounded endpoint count.

Resource-pressure signalDo firstAvoid
The heap sits near its limit for a long timeShrink the entity or bridge set, check the pluginsRunning Force Sync again and again
An exit code, or a host OOMCheck the host events, add usable RAM or reduce the loadOnly turning on debug logging, which adds load
A large HA request times outCheck the HA load and the message timeout settingResetting the fabric straight away
A start-up spike from several bridges at onceAdjust the startup priorityPressing Restart All over and over

Updates, host migration and disaster recovery

Routine updates: Update Checker only tells you a version exists. Read the release notes, take a full snapshot, keep the image you are on, update a single environment, and verify it layer by layer. If a schema or controller regression appears, stop the new version and roll back to the old version and the original snapshot; do not keep resetting things inside the broken environment.

Moving to a new host: prepare host networking, IPv6, mDNS and persistent storage on the new instance, but do not start it on the same identity. Stop the old instance, then restore the full backup, keeping the bridge config and the network identity stable; only once that works do you close the rollback window on the old instance.

Lost storage: if you have a verified full backup, restore it on the same Stable version. If all you have is a config backup, accept that you will commission again, and clear the old relationships on the controller side before you pair. Only when there is no usable backup at all do you move to a disaster reset, and then treat rebuilding the controllers, HAMH and the automations as a project in its own right.

A disaster reset is not a troubleshooting shortcut: it destroys the existing fabric relationships and the tidying you did on the controller side. Run it only when the identity is already lost or damaged, the full backup will not restore, and you have ruled out the network and session problems.

Symptom → check → safe fix

  1. Commissioning fails, or the bridge is not found

    Check: the bridge is running, it is not yet commissioned, the phone and the hub are on the same segment, IPv6, the mDNS bound interface, multicast and the operational firewall. Safe fix: go back to a plain single subnet first, correct the interface and the firewall, and open the commissioning window again; do not run factory resets one after another.

  2. No Response after commissioning succeeded

    Check: HA connected, whether the fabric still exists, sessions and subscriptions, whether mDNS is advertising the wrong interface, and the controller hub. Safe fix: repair the shared network; if you need more, confirm autoForceSync first and then restart or Force Sync that one bridge; when the other controllers are fine, look at that hub and its support first.

  3. mDNS drops in and out, or shows duplicate records

    Check: AP multicast and IGMP, the mDNS reflector, multiple interfaces, whether there has been an unclean power loss, and the real fabric count. Safe fix: bind the LAN interface, fix multicast, do a graceful restart and wait out the cache TTL; follow the commissioning cleanup procedure only for fabrics that really are surplus.

  4. A bridge shows Failed

    Check: the status reason, HA, the port, storage permissions, memory and the plugins. Safe fix: fix the root cause, then let Auto Recovery or a single-bridge restart retry it; if the recovery history keeps failing, turn the retries off and handle it by hand.

  5. Only some entities are failed

    Check: the failed reason, HA unavailable, the filter, the device class, and the mapping and composed links. Safe fix: fix the source or the mapping and restart that one bridge; do not reset the whole bridge.

  6. Restore Preview looks fine but Restore reports errors

    Check: the error on each bridge, exists and overwrite, the version, archive integrity and storage permissions. Safe fix: stop overwriting anything else and restore the current snapshot; reproduce it on an isolated copy, and once you are sure, retry only the items that failed.

  7. After a migration the controller finds two services

    Check: whether the old host or the old container is still active, and the mDNS cache. Safe fix: get back to a single active instance at once, stop the other end cleanly, and wait for or clear the controller cache; do not reset both ends together.

  8. A low-resource host restarts without warning

    Check: metrics, host OOM, container exits, the endpoint count and large plugins. Safe fix: cut the entity count, stagger bridge startup, and add RAM or set a suitable heap; do not tighten the Recovery interval and manufacture a restart loop.

  9. Last resort: a disaster reset

    Check: that you have proved the full backup will not restore, that the identity is unusable, and that network and controller problems are ruled out. Start the target bridge into Running first; on v2.0.55 a Stopped or Failed bridge may be returned and started without the reset running at all. Safe fix: keep an archive of the current state and redacted logs, remove the old bridge on each controller one at a time, then run Factory Reset against the Running HAMH bridge, cross-check the fabric and commissioning status after the action, and rebuild and commission again one bridge at a time; rebuild rooms and automations last. Pairing data is only ever shown in the local UI — do not copy it into a document.

FAQ

Can a config backup keep the controller commissioning?
No. It explicitly contains no Matter identity, so the bridges will need to be commissioned again. To keep the existing fabric, use a protected full backup or snapshot.
Does a full backup include plugins, Lock Credentials, device images and every setting?
Not that whole scope. It includes the bridges, the entity mappings, the identity where it exists, the matching icons and the per-bridge plugin enabled / disabled flags; it does not include plugin packages, the registry, per-plugin config, secrets, Lock Credentials, device-images or general app settings.
Does Auto Recovery restart healthy bridges?
No. Both the hint on the setting and the BridgeService behavior deal with failed bridges only; a running bridge is not disturbed.
Does Update Checker upgrade automatically?
No. It compares versions and shows guidance for your deployment; the backup, the update and the rollback are still yours to run.
During a move, can I run the old and the new host together to test?
Not on the same Matter identity. Stop the old instance before you start the new one, and if you need to roll back, stop the new one first.
When is a factory reset the right move?
Only when the commissioning identity genuinely has to be cleared, when a full backup will not restore, or when you have decided to commission again — and you already have a controller cleanup and rebuild plan. On v2.0.55, confirm the bridge is Running first and cross-check the fabric and commissioning after the action; a Stopped or Failed bridge may be started without being reset. An ordinary No Response, an mDNS problem or a failed entity does not need a reset first.

Pinned-version sources