What Is Clock Skew and Clock Drift?
Clock skew is the difference between the readings of two clocks at one instant. Clock drift is the rate at which one clock gains or loses time against a reference. Skew is measured in units of time, for example 45 milliseconds. Drift is measured as a rate, usually in parts per million (ppm). An uncorrected quartz clock in a server drifts by about 10 to 50 ppm, which is 1 to 4 seconds a day. Drift is the cause, and skew is the effect that builds up when nothing corrects it.
Clock drift: the rate of error
Every computer keeps time with a quartz crystal oscillator, a small part that vibrates at a rated frequency. No crystal runs at exactly its rated frequency. A crystal with a 20 ppm error gains or loses up to 20 microseconds every second. That is about 1.7 seconds per day and close to a minute per month.
Temperature is the largest cause. The frequency of a crystal changes with heat, so a server in a warm rack drifts faster than one in a cool rack. Age matters as well, because a crystal changes slowly over years of use. Virtual machines add a third cause: the guest clock loses ticks when the hypervisor takes CPU time away from it.
Drift is why two machines set to the same time in the morning disagree by the evening. Drift never stops, so correction must be continuous rather than a one-time adjustment.
Clock skew: the gap at one moment
Skew is a snapshot. If node A reads 10:00:00.000 and node B reads 10:00:00.045 at the same instant, the skew between them is 45 milliseconds. Skew has three sources: different starting values when the clocks were set, drift accumulated since the last correction, and network delay in the synchronization messages themselves.
Skew is what application code experiences. A service never sees another machine's drift rate directly. It sees a timestamp that is 45 milliseconds different from its own, and it must decide what that means.
Why skew breaks distributed systems
Ordering by timestamp. Many databases resolve concurrent writes with last-writer-wins, a rule that keeps the write with the newest timestamp. With 50 milliseconds of skew, a write that happened second can have an older timestamp and be discarded.
Leases and timeouts. A lease is a lock that expires after a fixed time, for example 10 seconds. A node whose clock runs 2 seconds fast believes the lease expired early and may act while another node still holds it.
Security checks. Kerberos rejects an authentication ticket when the skew between client and server exceeds 5 minutes by default. AWS rejects a signed API request when its timestamp differs from server time by more than 15 minutes. A TLS certificate is invalid on a machine whose clock is outside the certificate's validity window.
Logs and traces. A request that crosses four services with 200 milliseconds of skew can appear to receive its response before it was sent. Debugging from such logs is unreliable.
How NTP and PTP correct drift
The Network Time Protocol (NTP) is the standard way to keep clocks aligned. A client exchanges four timestamps with a server, estimates the round-trip delay, and computes its offset from the server. Over the public internet, NTP keeps a machine within tens of milliseconds of true time. On a local network with a nearby server, it keeps a machine within about 1 millisecond.
Drift correction has two forms. Slewing changes the clock rate slightly so that time catches up or slows down without a jump. Stepping sets the clock to a new value at once. The reference NTP daemon slews when the offset is under 128 milliseconds and steps when the offset is larger. Stepping backward is dangerous because a program can observe time moving in reverse, so durations should be measured with a monotonic clock that only increases.
The Precision Time Protocol (PTP) uses hardware timestamps in the network card to remove most software delay. On a local network with PTP-capable switches, it reaches accuracy below 1 microsecond. Trading systems and telecom networks use it. Most web systems do not need it.
When to stop trusting wall clocks
Some systems avoid the question. A logical clock is a counter that records the order of events rather than the time of day. Lamport timestamps and vector clocks are two forms of logical clock, and they order events correctly regardless of skew. Google Spanner takes the other route: its TrueTime service reports a time interval rather than a single value, typically under 7 milliseconds wide, and a transaction waits out that interval before it commits.
Clock skew vs clock drift
| Aspect | Clock skew | Clock drift |
|---|---|---|
| Definition | Difference between two clocks at one instant | Rate at which one clock diverges from a reference |
| Unit | Time, for example 45 ms | Rate, for example 20 ppm |
| Typical size | Tens of ms over the internet, about 1 ms on a LAN with NTP | 1 to 4 seconds a day for uncorrected quartz |
| Main causes | Set-time differences, accumulated drift, network delay | Crystal tolerance, temperature, age, virtualization |
| Effect | Wrong ordering, early lease expiry, rejected requests | Growing skew between corrections |
| Correction | Synchronization with NTP or PTP | Continuous slewing by the NTP daemon |
Key Takeaways
- Skew is a distance, drift is a speed. Skew is the gap between two clocks at one instant. Drift is how fast that gap grows.
- Quartz drifts by 10 to 50 ppm. That is 1 to 4 seconds a day, which is why every server runs an NTP client.
- NTP keeps skew near 1 millisecond on a LAN and within tens of milliseconds over the internet. PTP reaches below 1 microsecond with hardware support.
- Never order writes by wall clock alone. Use logical clocks, version numbers, or an uncertainty bound such as TrueTime.
Time synchronization is one of the core topics in Grokking System Design Fundamentals.
The consistency and replication chapters of Grokking the System Design Interview show where clock assumptions appear in real designs.
For the patterns that replace wall clocks, such as Lamport clocks and leases, see System Design Patterns.

GET YOUR FREE
Coding Questions Catalog

$99

$197

$72