Jc-alt logo
jc
System Design: CAP Theorem

System Design: CAP Theorem

··
3 min read
·system design

CAP Theorem

The CAP Theorem states that in a distributed system, it is impossible to simultaneously guarantee consistency, availability, and partition tolerance. Assuming partition tolerance, when a network partition occurs, the system must choose between consistency and availability.

In practice, the CAP Theorem describes constraints, not system categories. Every distributed system must tolerate network partitions because partitions are unavoidable, others you would be forced to host everything in a single data center which is not realistic.

Thus, once a partition occurs, engineers must decide whether the system prioritizes consistency or availability for that specific scenario.

CAP is a lens for reasoning about system behavior under failure, rather than a rule that forces systems into rigid boxes.

Dealing With Network Partitions

In practice, partitions occur due to network congestion, hardware issues, cloud infrastructure failures or anything that causes nodes to not be able to communicate reliably.

At that point, the system must choose whether to reject requests to preserve consistency or serve requests with potentially stale data to preserve availability

System Behavior During PartitionResulting Tradeoff
Rejects requests to avoid stale dataFavors consistency
Serves requests despite stale data requests to avoid stale dataFavors availability
Pauses some operations selectivelyHybrid approach

Consistency (C), Availability (A), Partition Tolerance (P)

Consistency (C)

Consistency means that all clients/nodes see that same data at the same time. After a successful write, every subsequent read should return the updated value, regardless of which client/node processes the request

Availability (A)

Availability means that every request receives a response, even if the response does not contain the most recent data. The system should remain responsive as long as it is operational

Partition Tolerance (P)

Partition tolerance means that the system continues operating despite network failures that split nodes into isolated groups. Messages between partitions may be delayed or dropped entirely.

Partition tolerance is not operational in real world distributed systems. Networks fail unpredictably, and systems must be designed with this reality in mind. Because partition tolerance is mandatory, leading to the real trade off in CAP occurs between consistency and availability.

Consistency / Partition Tolerant System (Banking Transactions)

Systems that require consistency across nodes and partition tolerance for partial network failures

Availability / Partition Tolerant System (Social Media News Feed)

Systems that require immediate availability and partition tolerance for partial network failures

Consistency / Availability (Rare)

Systems that usually run in a single data center that require immediate availability and consistency across nodes (itself)

(Consistency / Availability) / Partition Tolerant System Hybrid (Online Shopping Cart)

Systems that require tolerance for partial network failures and switch between requiring availability and requiring consistency to handle different stages of the user journey effectively