Twitter's Software Architecture: Timelines, Fan-Out, and the JVM Backend
Twitter's architecture has three defining parts: home timelines built ahead of time by fan-out on write into Redis, a self-built database called Manhattan for tweets and user data, and a backend of Scala and Java services. Those services replaced the original Ruby on Rails application. Fan-out on write means that when a user posts, the tweet's id is pushed into every follower's cached timeline at that moment. Reading a timeline is then a single cache lookup.
The design is shaped by one ratio. In a public talk, a Twitter engineer put reads at about 300,000 requests per second against about 6,000 tweets written per second. Twitter is a reading system, so the work is moved to write time.
From Ruby on Rails to JVM services
Twitter began as one Ruby on Rails application on MySQL. As traffic grew, the single application could not keep up, and outages were common enough that the error page, the Fail Whale, became famous.
The fix was to move the backend to the Java virtual machine. Twitter rewrote its services in Scala and Java, and built Finagle, an open-source library for writing RPC services, so that every service handled connections, retries, and timeouts the same way. Engineers reported more than a tenfold increase in requests served per machine after the move. Rails remained only at the edge for a time and was then retired from the request path.
The result is a service-oriented backend. Tweetypie serves tweet objects, Gizmoduck serves user objects, a social graph service serves follower lists, and a timeline service assembles what you see.
The timeline fan-out design, step by step
- A tweet arrives at the write API. It is assigned an id by Snowflake, Twitter's id generator, which encodes the time so ids sort in order.
- The tweet is stored. The text and metadata go to the tweet store, and any media goes to Blobstore, Twitter's object store.
- The fan-out service reads the follower list. It asks the social graph service for everyone who follows the author.
- The tweet id is inserted into each follower's home timeline. Each timeline is a list in a Redis cluster, replicated three times, and capped at 800 entries. Only the id is stored, not the tweet.
- A follower opens the app. The timeline service reads the id list from Redis in one call.
- The ids are turned into tweets. The service asks Tweetypie for the tweets and Gizmoduck for the authors, in parallel, and returns the page.
The talk gave a delivery target of five seconds from post to follower timeline, with a median of about 3.5 seconds. Accounts with millions of followers cannot meet that target by fan-out alone, so their tweets are merged in at read time instead. That mix is the answer interviewers want when they ask about celebrities.
Fan-out on write compared with fan-out on read
| Question | Fan-out on write (Twitter's default) | Fan-out on read |
|---|---|---|
| Work happens | When the tweet is posted | When a follower opens the timeline |
| Read cost | One cache lookup, then hydrate the ids | Query every followed account, then merge and sort |
| Write cost | One insert per follower | One write |
| Suits | Most users, who read far more than they post | Accounts with millions of followers |
| Storage | A cached list per active user | Little extra storage |
| Freshness | A few seconds behind | Immediate |
Twitter uses both. Ordinary accounts fan out on write. The highest-follower accounts are excluded from fan-out and merged in when a timeline is read. An interview answer that names this hybrid, and says where the cutoff should be, is complete.
What database does Twitter use?
Twitter's main store is Manhattan, a distributed key-value database it built after outgrowing Cassandra. Twitter has said Manhattan holds tweets, direct messages, and account data, and serves tens of millions of queries per second across its clusters. Before Manhattan, tweets lived in sharded MySQL and then in Cassandra.
| Data | Store |
|---|---|
| Tweets, direct messages, accounts | Manhattan |
| Home timelines | Redis cluster, ids only |
| Images and video | Blobstore, served through CDNs |
| Search index | Earlybird, a modified Lucene index held in memory |
| Follower graph | A dedicated graph service, originally FlockDB on MySQL |
| Hot objects | Memcached and Redis caches in front of the services |
Manhattan is multi-tenant, meaning many teams share one cluster with their own key spaces and their own consistency settings. That is why a single system can hold both tweets, which tolerate slight delay, and account data, which cannot.
The Twitter tech stack in one table
| Layer | Technology |
|---|---|
| Service languages | Scala and Java on the JVM, with Finagle for RPC |
| Original application | Ruby on Rails, now retired from serving |
| Id generation | Snowflake |
| Timeline cache | Redis |
| Primary database | Manhattan |
| Search | Earlybird (Lucene-based) |
| Cluster management | Mesos and Aurora, with Kubernetes adopted later |
| Analytics | Hadoop, Scalding (Scala on Hadoop), Heron for streaming |
How to present Twitter in a system design interview
Start with the read-to-write ratio and say that it justifies doing work at write time. Draw the write path through fan-out into Redis, then the read path as a lookup plus hydration. Name the celebrity problem before you are asked, and explain the hybrid. Finish with storage: a timeline cache that holds ids, a durable store for tweets, and an object store for media.
For the patterns behind this design, read what design pattern Twitter uses. For the interview loop itself, read what a Twitter interview is like.
How to Prepare
- Practice the news feed design until you can draw it in ten minutes. Grokking the System Design Interview covers Twitter, Instagram, and Facebook news feed designs, including the fan-out trade-off.
- Learn the building blocks separately. Caching, sharding, and id generation each appear in this design. Grokking System Design Fundamentals explains them one at a time.
- Study the patterns by name. Fan-out, cache-aside, and the hybrid push-pull feed are covered in System Design Patterns.
- Check the hiring bar before you apply. Read what the Twitter hiring requirements are.

GET YOUR FREE
Coding Questions Catalog

$99

$197

$72