Simplified Disaster Recovery: Unplanned Failover
Your enterprise must define procedures for determining that the original primary site is unrecoverable (and does not come back and interact with client applications).
When you have determined that the original primary site is unrecoverable, you can execute the following procedure to migrate client applications to the DR site. Note that, once you activate the DR site, a tail end of application data may be lost due to the abrupt failure of the original primary site. See Replication for Disaster Recovery for details.
Generally, you must not make modifications to the FTL server yaml configuration files. Some updates to the FTL realm configuration are required; FTL can perform those updates automatically. However, if you want to update the FTL realm configuration manually, see dr-cluster-sample-dr-activate.json in the samples/yaml/dr-simple directory.
-
Determine that the primary site is unrecoverable according to the standards defined by your enterprise. Generally, all FTL servers at the primary site must be down (or totally unreachable). If the primary site comes back and interacts with client applications that have not yet failed over to the DR site, split-brain occurs.
-
Activate the DR site. Run the
activate_drcommand at the DR site. Theauto_config_updateflag is optional.Example:
tibftladmin -ftls <url> --activate_dr --auto_config_update -
Ensure the DR site has activated (is now primary). Do this by checking FTL server's operating mode at the DR site. See Monitoring for Disaster Recovery.
-
Modify the FTL realm definition so that all persistence clusters used by client applications at the DR site are activated.
-
Case 1. You specified the
auto_config_updateflag when activating the DR site. No further action is needed. FTL updated the realm definition automatically. The primary set of all persistence clusters must now contain the persistence services running at the DR site. -
Case 2. You did not specify the
auto_config_updateflag when activating the DR site. You must update the FTL realm definition. For each DR-enabled persistence cluster used by client applications at the DR site, update theprimary_setfield of each persistence cluster. The value should be the name of the server set that is running at the DR site. See Persistence Configuration for Disaster Recovery.
-
-
Ensure that persistence services at the DR site are now active (that is, the persistence cluster has a leader and is not in the DR role). See Monitoring for Disaster Recovery.
-
Direct client applications to the DR site. Generally, this is done by restarting the clients with new URLs, or remapping DNS in your enterprise.