Simplified Disaster Recovery: Failback After Planned Failover

Use this procedure when you want to fail back to the original primary site after performing a planned failover to the DR site.

This procedure is similar to Simplified Disaster Recovery: Failback After Unplanned Failover. The difference is that you explicitly deactivated the original primary site through the deactivate_dr command, so the presence of auto.init.primary.on.no.contact does not affect the procedure. No special action is needed to avoid split-brain.

Some updates to the FTL realm configuration are required; FTL can perform those updates automatically. However, if you want to update the FTL realm configuration manually, see dr-cluster-sample.json in the samples/yaml/dr-simple directory.

  1. Verify the state of persistence services at the original primary site. They must now be acting as standby services for the (activated) DR site. Ensure that pending message counts are similar between the two sites. See Monitoring for Disaster Recovery.

  2. Prepare the system for planned failover to the original primary site by running the prepare_planned_dr_failover command at the (activated) DR site.

    On running this command, FTL stops accepting client application data (for example, messages and acknowledgments) at the DR site. FTL ensures all application data at the DR site is replicated to the original primary site before indicating success to the admin tool (or web API).

    Note: If something goes wrong before you can activate the original primary site, you can restart the servers at the DR site to restore service at the DR site. Enabling disk persistence for your persistence clusters is recommended in this case.

    Example: tibftladmin -ftls <urls> --prepare_planned_dr_failover

  3. Double-check the state of the system. For example, all persistence services at the original primary site must have "suspended" status (indicating that replication of data is complete), and message counts must be identical across the two sites. See Monitoring for Disaster Recovery.

  4. Perform final deactivation of the DR site by running the deactivate_dr command at the DR site.

    Do this only when you are sure the original primary site is ready for client connections. When you run this command, FTL servers at the DR site are out of service (no client connections) until the original primary site has been activated.

    Note: If something goes wrong before you can activate the original primary site, you can run the activate_dr command to restore service at the DR site. For this to work, you must have enabled disk persistence (or taken a backup) for all persistence clusters. Otherwise, application data is not preserved.

    Example: tibftladmin -ftls <urls> --deactivate_dr

  5. Activate the original primary site. Run the activate_dr command at the original primary site. The auto_config_update flag is optional.

    Example: tibftladmin -ftls <url> --activate_dr --auto_config_update

  6. Ensure the original primary site has activated (is now primary). Do this by checking FTL server's operating mode at the original primary site. See Monitoring for Disaster Recovery.

  7. Modify the FTL realm definition so that all persistence clusters used by client applications at the original primary site are activated.

    • Case 1. You specified the auto_config_update flag when activating the original primary site. No further action is needed. FTL updated the realm definition automatically. The primary set of all persistence clusters must now contain the persistence services running at the original primary site.

    • Case 2. You did not specify the auto_config_update flag when activating the original primary site. You must update the FTL realm definition. For each DR-enabled persistence cluster used by client applications at the original primary site, update the primary_set field of each persistence cluster. The value must be the name of the server set that is running at the original primary site. See Persistence Configuration for Disaster Recovery.

  8. Ensure that persistence services at the original primary site are now active (that is, the persistence cluster has a leader and is not in the DR role). See Monitoring for Disaster Recovery.

  9. Direct client applications to the original primary site. Generally, this is done by restarting the clients with new URLs, or remapping DNS in your enterprise.

  10. Verify the operating mode of FTL servers at the DR site (must be DR). Verify the state of persistence services at the DR site (must be standby). See Monitoring for Disaster Recovery.