Jianbo Ding, Yuxuan Qin
Consensus protocols form the foundational building blocks of fault-tolerant distributed data systems, enabling nodes to agree on a consistent shared state despite failures, network partitions, and heterogeneous communication delays. The Raft algorithm, designed explicitly for understandability and practical deployability, has become the de facto foundation for a wide class of modern replicated storage systems. However, as organizations require data infrastructure to span geographically distributed (GD) environments, vanilla Raft exhibits significant performance limitations rooted in its single-leader architecture and majority quorum constraints, causing write latency to scale with cross-region wide-area network (WAN) round-trip times (RTTs). This survey provides a comprehensive review of consensus mechanisms from foundational Raft semantics to advanced variants engineered for GD deployments, covering state machine replication (SMR), Multi-Raft architectures, flexible quorum designs, and Byzantine fault tolerance (BFT). We analyze prominent WAN consensus protocols including WPaxos, EPaxos, and HotStuff, examining their theoretical guarantees and practical trade-offs in latency, throughput, and operational complexity. The survey further examines how production systems such as CockroachDB, TiKV, and YugabyteDB integrate and extend these protocols to achieve global-scale consistency. By synthesizing recent advances in adaptive leader election, hierarchical consensus, hybrid protocol design, and BFT convergence with crash-fault-tolerant (CFT) alternatives, this paper provides a structured reference for researchers and engineers designing the next generation of GD data infrastructure.