Simplified Disaster Recovery: Failback After Unplanned Failover

If you have executed an unplanned failover, as in Simplified Disaster Recovery: Unplanned Failover, and are now ready to migrate client applications back to the original primary site, perform the steps below.

Some updates to the FTL realm configuration are required; FTL can perform those updates automatically. However, if you want to update the FTL realm configuration manually, see dr-cluster-sample.json in the samples/yaml/dr-simple directory.

Conceptually, this is a planned failover to the original primary site. However, since the original failover was unplanned, you did not have an opportunity to deactivate the original primary site. If you did not specify auto.init.primary.on.no.contact during Simplified Disaster Recovery: Setup, no special action is needed, because the original primary site automatically deactivates itself when it contacts the (activated) DR site. In this case, this procedure is similar to Simplified Disaster Recovery: Failback After Planned Failover.

However, if you did specify auto.init.primary.on.no.contact (because you wanted to ensure availability of the primary site), you must have a plan for avoiding "split-brain" if the original primary site happens to restart while the link to the (activated) DR site is down. Opt whichever strategy meets the standards of your enterprise and is convenient to implement.

  • Split-brain option 1. Delete all data directories at the original primary site before restarting FTL servers. Also remove auto.init.primary.on.first.startup from the FTL server yaml configuration file at the primary site. This prevents FTL servers at the original primary site from initializing until they contact the DR site.

  • Split-brain option 2. Temporarily remove auto.init.primary.on.no.contact from the FTL server yaml configuration file at the original primary site. This prevents FTL servers at the original primary site from initializing until they contact the DR site. Remember to add it back later.

  • Split-brain option 3. Specify --skip.auto.init.primary.on.no.contact <timestamp> on the command line to FTL servers at the original primary site. This acts as a one-time suppression of auto.init.primary.on.no.contact. It prevents FTL servers at the original primary site from initializing until they contact the DR site. When the DR site is contacted, --skip.auto.init.primary.on.no.contact is ignored until <timestamp> increases, so you can leave it on the command line of FTL server. Also see FTL Server Executable.

  • Split-brain option 4. Take no special action, because your enterprise has some external method for handling "split-brain". For example, your enterprise might be able to determine that no clients remain at the original primary site, or your enterprise might have high confidence in the stability of the link to the DR site. Note that, in this scenario, if the original primary site happens to initialize as primary, you can run the deactivate_dr command at the original primary site, just like you would for a planned failover.

  1. Restart FTL servers at the original primary site.

    • Case 1. You did not specify auto.init.primary.on.no.contact. No special action needed. The FTL servers at the original primary site automatically deactivate themselves on contacting the (activated) DR site.

    • Case 2. You have chosen "split-brain option 1", above. Delete all realm and persistence data directories prior to restarting each FTL server. The FTL servers at the original primary site copy realm configuration and persistence data on contacting the (activated) DR site.

    • Case 3. You have opted "split-brain option 2", above. Remove auto.init.primary.on.no.contact from the FTL server yaml configuration file at the original primary site. The FTL servers at the original primary site automatically deactivate themselves on contacting the (activated) DR site, at which point you can add auto.init.primary.on.no.contact back to the yaml file.

    • Case 4. You have opted "split-brain option 3", above. Pass --skip.auto.init.primary.on.no.contact <timestamp> on the command line to FTL servers at the original primary site. The timestamp must be the number of seconds since the epoch. The timestamp must be the same for all FTL servers. Precision is not important; the important thing is that all FTL servers have the same timestamp and that the timestamp does not change (until the next unplanned DR failover).

      The FTL servers at the original primary site automatically deactivates themselves on contacting the (activated) DR site. At this point --skip.auto.init.primary.on.no.contact is ignored (unless its value increases in the future). You can leave it on the command line.

    • Case 5. You have opted "split-brain option 4" above. No special action needed. The FTL servers at the original primary site automatically deactivate themselves on contacting the (activated) DR site. However, if the network link to the DR site is down, be prepared to run the deactivate_dr command at the original primary site.

  2. Verify the operating mode of FTL servers at the original primary site (must be DR). See Monitoring for Disaster Recovery.

    If you have opted "split-brain option 4", and you see an operating mode of primary at the original primary site, run the deactivate_dr command at the original primary site.

  3. Verify the state of persistence services at the original primary site. They should now be acting as standby services for the (activated) DR site. Ensure that pending message counts are similar between the two sites. See Monitoring for Disaster Recovery.

  4. Prepare the system for planned failover to the original primary site by running the prepare_planned_dr_failover command at the (activated) DR site.

    On running this command, FTL stops accepting client application data (for example, messages and acknowledgments) at the DR site. FTL ensures all application data at the DR site is replicated to the original primary site before indicating success to the admin tool (or web API).

    Note: If something goes wrong before you can activate the original primary site, you can restart the servers at the DR site to restore service at the DR site. Enabling disk persistence for your persistence clusters is recommended in this case.

    Example: tibftladmin -ftls <urls> --prepare_planned_dr_failover

  5. Double-check the state of the system. For example, all persistence services at the original primary site should have "suspended" status (indicating that replication of data is complete), and message counts should be identical across the two sites. See Monitoring for Disaster Recovery.

  6. Perform final deactivation of the DR site by issuing the deactivate_dr command at the DR site.

    Do this only when you are sure the original primary site is ready for client connections. When you run this command, FTL servers at the DR site are out of service (no client connections) until the original primary site has been activated.

    Note: If something goes wrong before you can activate the original primary site, you can run the activate_dr command to restore service at the DR site. For this to work, you must have enabled disk persistence (or taken a backup) for all persistence clusters. Otherwise, application data is not preserved.

    Example: tibftladmin -ftls <urls> --deactivate_dr

  7. Activate the original primary site. Run the activate_dr command at the original primary site. The auto_config_update flag is optional.

    Example: tibftladmin -ftls <url> --activate_dr --auto_config_update

  8. Ensure the original primary site has activated (is now primary). Do this by checking FTL server's operating mode at the original primary site. See Monitoring for Disaster Recovery.

  9. Modify the FTL realm definition so that all persistence clusters used by client applications at the original primary site are activated.

    • Case 1. You specified the auto_config_update flag when activating the original primary site. No further action is needed. FTL updated the realm definition automatically. The primary set of all persistence clusters should now contain the persistence services running at the original primary site.

    • Case 2. You did not specify the auto_config_update flag when activating the original primary site. You must update the FTL realm definition. For each DR-enabled persistence cluster used by client applications at the original primary site, update the primary_set field of each persistence cluster. The value must be the name of the server set that is running at the original primary site. See Persistence Configuration for Disaster Recovery.

  10. Ensure that persistence services at the original primary site are now active (that is, the persistence cluster has a leader and is not in the DR role). See Monitoring for Disaster Recovery.

  11. Direct client applications to the original primary site. Generally, this is done by restarting the clients with new URLs, or remapping DNS in your enterprise.

  12. Verify the operating mode of FTL servers at the DR site (must be DR). Verify the state of persistence services at the DR site (must be standby). See Monitoring for Disaster Recovery.