Replication for Disaster Recovery
Replication of application data to the DR site is asynchronous. As there is typically significant latency between the primary and DR sites, synchronous replication would often be prohibitively expensive.
When the persistence services at the primary site acknowledge a client's API call (for example, send or acknowledge a message), the primary site leader has confirmed that the data is stored by the persistence services at the primary site, as usual. If the link to the DR site is up, and the DR persistence services are up-to-date, the primary site leader immediately forwards the data to the DR site. However, the primary site leader does not confirm that the DR site has received the data before acknowledging the client's API call.
This means that if an unplanned DR failover occurs, a tail end of application data could be lost, as the DR site has not confirmed receipt of the most recent records.
Only the primary site leader replicates data to the DR site. The primary site leader replicates data only to the DR site leader. So, during normal operation, application data is transmitted once to the DR site. The DR site leader forwards application data to other persistence services running at the DR site.
Replication of persistence data to a disaster recovery site requires sufficient WAN capacity to transfer data in a timely manner. You must ensure sufficient WAN capacity for expected peak data volume.
A small data gap always exists because replication to the disaster recovery site is asynchronous. The data gap can grow in the following situations:
-
Peak message activity exceeds the WAN capacity. The data gap grows until the message volume decreases. When message volume decreases, replication at the disaster recovery site can catch up to the data state at the main site, and the data gap shrinks.
-
The WAN link malfunctions. Replication completely stops, unless you have redundant WAN lines for fault tolerance, and the data gap grows rapidly. Even with redundant lines, a data gap can grow if the backup capacity is less than the message volume. The data gap can also grow if failover introduces a significant delay.
For information on how to monitor the status of replication, see Monitoring for Disaster Recovery.