Selecting a Disaster Recovery Model

The choice of the classic or simplified DR model determines the specifics of what goes into the FTL server yaml configuration file (for parameter reference, see FTL Server Configuration Parameters). It also determines which administrative steps you perform to carry out DR failover and failback (for command reference, see FTL Administration Utility).

  • In the "classic" DR model, you specify drto in the FTL server yaml configuration file at the primary site, and you specify drfor in the FTL server yaml configuration file at the DR site. The value of each parameter is the URLs of the other site.

  • In the simplified DR model, you specify dr in the FTL server yaml configuration file at both sites. At the primary site, the value of this parameter is the URLs of the DR site; at the DR site, the value is the URLs of the primary site. There are some optional parameters for your convenience, including auto.init.primary.on.first.startup and auto.init.primary.on.no.contact.

Realm configuration and monitoring for DR is largely the same between the two models. See Persistence Configuration for Disaster Recovery and Monitoring for Disaster Recovery.

The choice of DR model largely comes down to convenience for the administrator. Both models support planned and unplanned failover and failback. Both models support DR for satellite sites (that is routes). New deployments should generally use the simplified model, which requires fewer actions from the administrator. Pre-existing deployments can continue to use the classic model or migrate to the simplified model.

A more detailed comparison follows. Many of the differences revolve around how to avoid split-brain, an undesirable situation where both sites are active and potentially interacting with clients. In a DR setup, clients should connect to one site or the other, not both.

The classic DR model is designed to ensure availability of the primary site. In the classic model, the primary site always continues to run, even if the DR site is down. However, to avoid split-brain, this requires manual steps from the administrator at various points in the process, including:

  • Manual edits to FTL server yaml configuration files, to designate a site as primary or dr.

  • Manual cleanup of data directories when preparing a deactivated site. (Specifically, when preparing the original primary site for failback, or resetting the DR site after successful failback.)

The simplified DR model is designed for ease of use. Wherever possible, administrative actions are carried out as FTL server web API calls or invocations of the FTL administration tool (tibftladmin). You do not explicitly designate a site as primary or DR in FTL server yaml configuration files.

By default, the simplified DR model is configured to avoid split-brain, but does not guarantee availability of the primary site in certain corner cases. Specifically, if the DR site is down or unreachable, and the primary site suffers a complete shutdown of all servers, then on restart the primary site waits to contact the DR site before accepting client connections.

Note: A rolling restart or upgrade of the primary site does not affect the availability of the primary site, even if the DR site is down. The only issue is a complete restart of the primary site. In this case, if the network to the DR site is down, the primary site does not know if the DR site has been activated or not.

You can configure the simplified DR model to achieve the same availability guarantees as the classic model (and you are encouraged to do so). However, special care is required when bringing back the original primary site after an unplanned DR failover in order to prevent the original primary site from activating and introducing split-brain. Specifically, this involves updating the command line of FTL server at the original primary site (note that other options are available as well). As long as you have clearly defined procedures for what to do after an unplanned failover, the simplified DR model is the right choice.

See Classic Disaster Recovery or Simplified Disaster Recovery. If you are enabling DR for persistence clusters at satellite sites, also see Disaster Recovery for Routes.