Raft Consensus Algorithm
The Raft-Consensus-Algorithm is a protocol designed to manage a replicated log in the context of Distributed-Systems. It was developed by Diego-Ongaro and John-Ousterhout at Stanford-University as an alternative to the notoriously complex Paxos algorithm. The primary design goal of Raft is understandability, making it easier for engineers to build reliable systems like Etcd and HashiCorp-Consul.
Operational Overview
The algorithm functions by electing a distinguished leader which has complete responsibility for managing the replicated log. The leader accepts log entries from clients, replicates them on other servers, and tells the servers when it is safe to apply log entries to their state machines. To maintain Fault-Tolerance, Raft decomposes the consensus problem into three relatively independent subproblems: Leader Election, Log Replication, and Safety.
In a typical Cluster, nodes can be in one of three states: Follower, Candidate, or Leader. Under normal operation, there is exactly one leader, and all other nodes are followers. If a follower receives no communication over a period called the election timeout, it becomes a candidate and initiates a new election to maintain Availability.
Detailed specifications and proofs of correctness can be found in the original paper "In Search of an Understandable Consensus Algorithm" or on the Official Raft Website.