UPC Off-Site Disaster Recovery & Fault Tolerance Solutions

2026-06-05

Foreword

With the deepening of enterprise digital transformation, core business systems (such as ERP, core databases, production management systems, etc.) have placed higher requirements on the continuity of the underlying infrastructure. To effectively prevent the devastating blows that regional disasters such as ransomware, fires, earthquakes, regional power outages, or network interruptions can bring to the production center, this solution builds a remote asynchronous disaster recovery solution based on distributed architecture and automated disaster recovery orchestration technology, ensuring that enterprises have rapid business recovery capabilities when extreme disasters occur.


I. Construction Background and Challenges

Traditional disaster recovery solutions based on centralized storage often face the following challenges:

High Construction Costs:

Storage hardware at both the production and disaster recovery ends must be of the same brand and model, resulting in severe hardware lock-in, and requiring additional purchases of expensive disaster recovery gateways or replication software licenses.

Poor Infrastructure Environment Compatibility:

Production environments are often composed of multiple architectures such as physical machines and virtualization. Traditional solutions struggle to achieve unified planning and lack the failover capability for disaster recovery.

Difficult Operations, Maintenance, and Drills:

The disaster recovery switchover process is complex, involving multiple disconnected stages such as storage cascading, network switching, and application startup; daily disaster recovery drills carry high risks and are time-consuming, leading to disaster recovery systems being "built but not dared to be used."

II. Solution Overview

This solution adopts software-defined infrastructure (hyper-converged) to implement modern disaster recovery concepts, aiming to build an elastic, intelligent, and highly available remote disaster recovery system for enterprises. Combined with intelligent disaster recovery software, it achieves data-level and application-level protection for the production center, achieving the goals of cost reduction and efficiency improvement, simplified operations and maintenance, and security and controllability.

The disaster recovery data center deploys a hyper-converged cluster, pooling computing, storage, and network resources as the "base camp" for remote takeover resources. Based on replication and recovery functions, it establishes data replication, failover, and disaster recovery between sites, and establishes encrypted data transmission channels through dedicated line networks. On weekdays, it provides a backup data storage pool, and when a disaster occurs, it immediately transforms into computing resources. Once a disaster occurs, business can be quickly switched to the disaster recovery cluster, and business operations can be rapidly restored.

III. Disaster Recovery Failover

Replica object data synchronization, failover, and disaster recovery are performed between multiple clusters at the same site or different sites, achieving cluster or site-level disaster recovery protection.

Emergency Failover:

When the original object can no longer continue to provide services, historical recovery points are used to start the replica object, achieving rapid business recovery and minimizing the downtime caused by disasters. It is typically applied to unplanned failover scenarios such as sudden disasters.

Planned Failover:

Combining replication and failover, the data of the original object is synchronized to the replica object before transferring the business to the replica object, achieving zero-data-loss business switchover and providing reliable assurance for business continuity. It is typically applied to planned failover scenarios such as disaster recovery drills, disaster warnings, or migration of replication objects, focusing on ensuring data recoverability (RPO=0).

Disaster Recovery & Failback:

When the original object environment of the production center is repaired and regains operating conditions, the disaster recovery system will initiate a smooth failback process. It will automatically identify and extract the incremental data generated during the disaster recovery takeover period and reverse-synchronize it to the production center, ensuring absolute data consistency and zero loss. The original object continues to provide services and resumes periodic automatic replication.

Permanent Failover:

When the original object cannot be recovered, the original object is permanently switched to the replica object, and the original object is removed from the replication plan, with the replica object converted to a normal resource object that continues to provide services.

Disaster Recovery Drills:

There is no need to build a separate drill environment. Disaster recovery drills can be conducted by replicating production environment virtual machines in the existing environment without affecting the normal business operations of the production environment.