
System Design: CAP Theorem
CAP Theorem
The CAP Theorem states that in a distributed system, it is impossible to simultaneously guarantee consistency, availability, and partition tolerance. Assuming partition tolerance, when a network partition occurs, the system must choose between consistency and availability.
In practice, the CAP Theorem describes constraints, not system categories. Every distributed system must tolerate network partitions because partitions are unavoidable, others you would be forced to host everything in a single data center which is not realistic.
Thus, once a partition occurs, engineers must decide whether the system prioritizes consistency or availability for that specific scenario.
CAP is a lens for reasoning about system behavior under failure, rather than a rule that forces systems into rigid boxes.
Dealing With Network Partitions
In practice, partitions occur due to network congestion, hardware issues, cloud infrastructure failures or anything that causes nodes to not be able to communicate reliably.
At that point, the system must choose whether to reject requests to preserve consistency or serve requests with potentially stale data to preserve availability
| System Behavior During Partition | Resulting Tradeoff |
|---|---|
| Rejects requests to avoid stale data | Favors consistency |
| Serves requests despite stale data requests to avoid stale data | Favors availability |
| Pauses some operations selectively | Hybrid approach |
Consistency (C), Availability (A), Partition Tolerance (P)
Consistency (C)
Consistency means that all clients/nodes see that same data at the same time. After a successful write, every subsequent read should return the updated value, regardless of which client/node processes the request
Availability (A)
Availability means that every request receives a response, even if the response does not contain the most recent data. The system should remain responsive as long as it is operational
Partition Tolerance (P)
Partition tolerance means that the system continues operating despite network failures that split nodes into isolated groups. Messages between partitions may be delayed or dropped entirely.
Partition tolerance is not operational in real world distributed systems. Networks fail unpredictably, and systems must be designed with this reality in mind. Because partition tolerance is mandatory, leading to the real trade off in CAP occurs between consistency and availability.
Consistency / Partition Tolerant System (Banking Transactions)
Systems that require consistency across nodes and partition tolerance for partial network failures
Availability / Partition Tolerant System (Social Media News Feed)
Systems that require immediate availability and partition tolerance for partial network failures
Consistency / Availability (Rare)
Systems that usually run in a single data center that require immediate availability and consistency across nodes (itself)
(Consistency / Availability) / Partition Tolerant System Hybrid (Online Shopping Cart)
Systems that require tolerance for partial network failures and switch between requiring availability and requiring consistency to handle different stages of the user journey effectively