Simplified Disaster Recovery: Planned Failover
Use this procedure if you want to migrate client applications to the DR site without losing application data. For example, you might be intending to designate the DR site as the permanent primary site, or you might want to test your enterprise's DR procedures. If you want to fail back to the original primary site, you can perform the Simplified Disaster Recovery: Failback After Planned Failover procedure after performing this procedure.
However, consider whether your goal is to simulate a DR event. If so, you might want to shutdown the original primary site instead of running the deactivate_dr command (see the steps below). After performing this procedure, you would then perform the Simplified Disaster Recovery: Failback After Unplanned Failover procedure instead.
Generally, you must not make modifications to the FTL server yaml configuration files. Some updates to the FTL realm configuration are required; FTL can perform those updates automatically. However, if you want to update the FTL realm configuration manually, see dr-cluster-sample-dr-activate.json in the samples/yaml/dr-simple directory.
Perform the following steps.
-
Verify the state of persistence services at the DR site. They must be acting as standby services for the original primary site. Ensure that pending message counts are similar between the two sites. See Monitoring for Disaster Recovery.
-
Prepare the system for planned failover to the DR site by issuing the
prepare_planned_dr_failovercommand at the original primary site.On running this command, FTL stops accepting client application data (for example, messages and acknowledgments) at the original primary site. FTL ensures all application data at the original primary site is replicated to the DR site before indicating success to the admin tool (or web API).
Note: If something goes wrong before you can activate the DR site, you can restart the servers at the original primary site to restore service at the original primary site. Enabling disk persistence for your persistence clusters is recommended in this case.Example:
tibftladmin -ftls <urls> --prepare_planned_dr_failover -
Double-check the state of the system. For example, all persistence services at the DR site must have suspended status (indicating that replication of data is complete), and message counts must be identical across the two sites. See Monitoring for Disaster Recovery.
-
Perform final deactivation of the original primary site.
-
Case 1. Run the
deactivate_drcommand at the original primary site.Do this only when you are sure that the DR site is ready for client connections. When you run this command, FTL servers at the original primary site is out of service (no client connections) until the DR site has been activated.
Note: If something goes wrong before you can activate the DR site, you can run theactivate_drcommand to restore service at the original primary site. For this to work, you must have enabled disk persistence (or taken a backup) for all persistence clusters. Otherwise, application data is not preserved.Example:
tibftladmin -ftls <urls> --deactivate_dr -
Case 2. Your goal is to simulate an unplanned DR event. Instead of issuing the
deactivate_drcommand, shutdown the original primary site instead. In this case, you would perform the Simplified Disaster Recovery: Failback After Unplanned Failover procedure after finishing this procedure.
-
-
Activate the DR site. Run the
activate_drcommand at the DR site. Theauto_config_updateflag is optional.Example:
tibftladmin -ftls <url> --activate_dr --auto_config_update -
Ensure that the DR site has activated (is now primary). Do this by checking FTL server's operating mode at the DR site. See Monitoring for Disaster Recovery.
-
Modify the FTL realm definition so that all persistence clusters used by client applications at the DR site are activated.
-
Case 1. You specified the
auto_config_updateflag when activating the DR site. No further action is needed. FTL updated the realm definition automatically. The primary set of all persistence clusters must now contain the persistence services running at the DR site. -
Case 2. You did not specify the
auto_config_updateflag when activating the DR site. You must update the FTL realm definition. For each DR-enabled persistence cluster used by client applications at the DR site, update the primary_set field of each persistence cluster. The value must be the name of the server set that is running at the DR site. See Persistence Configuration for Disaster Recovery.
-
-
Ensure that persistence services at the original primary site are now active (that is, the persistence cluster has a leader and is not in the DR role). See Monitoring for Disaster Recovery.
-
Direct client applications to the DR site. Generally, this is done by restarting the clients with new URLs, or remapping DNS in your enterprise.
-
Verify the operating mode of FTL servers at the original primary site (must be DR). Verify the state of persistence services at the original primary site (must be standby). See Monitoring for Disaster Recovery.
If you instead shutdown the original primary site to simulate an unplanned DR event, skip this step.