A home timeline looks like one list, but at Twitter it has long been two different systems wearing the same user interface. The Following timeline is a reverse-chronological list of posts from accounts you follow, and its hard problem is delivery: getting each new post into millions of lists quickly and cheaply. The For You timeline is a recommendation feed, and its hard problem is choice: finding a few dozen posts worth showing out of a vast pool, in well under a second.
This article builds both from first principles. The generic trade-off between pushing posts at write time and pulling them at read time is covered in news feed architecture; here we go deeper into what Twitter actually built: the shape of a timeline cache entry, which users get fan-out at all, how reads merge, hydrate and filter, and how the recommendation pipeline sources and ranks candidates. The For You description follows the design Twitter published with its open-sourced recommendation code in 2023. In January 2026 X published a newer system that replaces the heavy ranker with a Grok-based transformer, so treat that section as a well-documented reference design, not today's production.
Two timelines, two different problems
Following promises completeness and order: every post from every followed account, newest first. For You promises relevance: the posts most likely to be worth your time, from anyone. Nobody can tell whether it missed a post, but everyone notices when it is dull or repetitive.
Those promises lead to opposite designs. Following is a data delivery problem, solved by precomputing each user's list so that reading it is a cheap lookup. For You is a search and ranking problem, solved by computing the list at request time because the ranking depends on the moment: what you just engaged with, what is trending, what you have already seen. Both share the lower layers: a tweet store with a tweet cache in front, a social graph service, and visibility filtering. The diagram shows how they fit together.
The Following timeline cache
The timeline cache stores, for each user, a list of compact entries rather than posts. An entry needs the post id, the author id (so filters can drop it without loading the post) and a few flags, such as whether it is a repost and of which original. Post ids are time-ordered (Twitter's Snowflake ids put a millisecond timestamp in the high bits), so sorting by id is sorting by time and cursors can simply be ids. Keeping entries to a couple of dozen bytes means a list of several hundred entries costs a few kilobytes per user, which is what makes keeping it in memory affordable.
The list is capped. Nobody scrolls back through thousands of posts in the Following view, and an uncapped list turns every insertion into a slow operation on a growing structure. When a user scrolls past the cap, the server falls back to a slower path that queries the posts of followed accounts directly. The cache is also allowed to lose entries: it is a derived view that can always be rebuilt from the social graph and the author timelines, and that single property drives most of the operational design.
# Timeline cache on a Redis-like store, sharded by user_id.
ENTRY = struct.Struct("<QQB") # post_id, author_id, flags (17 bytes)
CAP = 800 # illustrative cap
def insert(cache, user_id, post_id, author_id, flags):
key = f"tl:{user_id}"
if not cache.exists(key): # inactive or evicted user: do not recreate here
return
cache.lpush(key, ENTRY.pack(post_id, author_id, flags))
cache.ltrim(key, 0, CAP - 1)
def read_page(cache, user_id, max_id=None, count=40):
raw = cache.lrange(f"tl:{user_id}", 0, CAP - 1)
if not raw:
return None # miss: caller rebuilds (see below)
entries = [ENTRY.unpack(r) for r in raw]
if max_id is not None:
entries = [e for e in entries if e[0] < max_id]
return entries[:count]Insertion never creates a missing list, because a missing list means the user is inactive or was evicted. A read miss returns nothing rather than an empty timeline, so the caller rebuilds instead of showing a blank screen.
Fan-out workers: bounding the multiplication
The cost to control is deliveries, not posts. With illustrative figures (not Twitter's) of 4,600 posts per second and 200 followers per author on average, that is about 920,000 timeline insertions per second, and one post from an account with 50 million followers is 50 million insertions on its own. When a post is created, the tweet service stores it and emits an event. Fan-out workers consume the event, page through the author's follower list in chunks of a few thousand, and issue batched inserts to the cache shards that own those followers. Chunking matters: a single task for 50 million followers would take minutes and would be retried from the start on any failure, while chunked tasks run in parallel and retry independently.
Three rules keep the work bounded. First, only active users get fan-out: users who have not opened the app in, say, 30 days have no cached list, and the insert above is skipped for them. Second, accounts above a follower threshold are not fanned out at all; their posts are merged in at read time, which is the hybrid described in the news-feed article. Third, deliveries have priority classes: fan-out to users who are online right now goes to a fast queue, everyone else to a bulk queue that can fall behind during spikes without anyone noticing.
Deletes and privacy changes use the same machinery in reverse, but correctness does not depend on them. A deleted post may linger in cached lists for a while; the read path drops it because the tweet cache returns a tombstone. That is the general principle: fan-out is an optimisation for speed, and the read path is where correctness is enforced.
The Following read path: merge, hydrate, filter
Serving a Following page takes four steps. Fetch the cached entries (or rebuild on a miss). Fetch recent posts from any followed accounts that are above the fan-out threshold and merge them by id. Hydrate the surviving ids into full posts with one batched multi-get against the tweet cache. Then filter: drop posts from blocked or muted authors, deleted posts, posts from protected accounts the viewer may no longer see, and posts withheld in the viewer's country.
def following_timeline(user, max_id=None, count=40):
entries = read_page(cache, user.id, max_id, count * 2)
if entries is None:
entries = rebuild(user) # bounded; see failure modes
big = graph.followed_large_accounts(user.id) # usually a short list
extra = [(pid, a, 0) for a in big
for pid in author_timeline(a, max_id, limit=count)]
merged = heapq.nlargest(count * 2, entries + extra, key=lambda e: e[0])
posts = tweet_cache.multi_get([e[0] for e in merged]) # one round trip
blocked = graph.blocked_and_muted(user.id) # cached per user
out = [p for p in posts
if p and not p.deleted and p.author_id not in blocked
and visibility.allowed(user, p)]
out = inject_own_recent_posts(user, out) # read-your-writes
return out[:count]Filtering at read time is deliberate: blocks, mutes and deletions change after fan-out, and undoing them in millions of cached lists would never quite complete. The last line handles a subtle expectation: a user who has just posted expects to see the post immediately, even if fan-out to their own list is still queued, so the server injects their own recent posts directly.
For You, stage one: candidate sources
The For You pipeline in Twitter's published code has three stages: candidate sourcing, ranking, and heuristics and filters. It is orchestrated by a service called Home Mixer, built on a framework called Product Mixer that composes candidate pipelines, feature hydrators, scorers and selectors.
Candidate sources each propose posts cheaply, trading precision for speed. The in-network source is the search index (Earlybird), which finds recent posts from accounts you follow and ranks them with a small light-ranker model; the repository's README says about half of the posts come from this source. An engagement-prediction model called Real Graph, which estimates how likely you are to interact with each account, decides whose posts are worth considering. Out-of-network sources include UTEG, an in-memory graph of user-to-post interactions built on the GraphJet framework, which answers questions like 'what have the people I engage with recently liked?'. There are also embedding-based sources: SimClusters represents users and posts as sparse memberships in detected communities, and TwHIN provides dense embeddings from a knowledge graph of users, posts and interactions, so nearest-neighbour search finds posts close to your interests. A coordination service, tweet-mixer, fetches out-of-network candidates across these sources.
The principle is recall first: sources may be noisy, because ranking sorts the pool, but the good posts must be in it within a strict time budget.
For You, stage two and three: ranking, heuristics and mixing
Ranking scores every candidate with a neural network, the heavy ranker. It does not predict a single 'relevance' number. It predicts the probabilities of several different engagements (like, reply, repost, watching a video, opening the author's profile, and negative feedback such as 'show less'), and the final score is a weighted sum of those probabilities. The weights are product decisions: raising the weight on replies favours conversation. Features come from hydrators that attach author, post and viewer-author relationship signals before scoring, and models are served by a Rust service called Navi.
A pure score order would produce a bad feed, so heuristics follow. Author diversity stops one prolific account from filling the screen. Content balance keeps a mix of in-network and out-of-network posts. Feedback fatigue lowers the score of authors or topics the user recently dismissed. Visibility filters remove content for legal, safety and user-preference reasons. Finally the mixer interleaves posts with ads, follow suggestions and conversation modules, and records what was served so the next request does not repeat it.
def for_you(user, budget_ms=900):
deadline = now_ms() + budget_ms
sources = [in_network_index, uteg_graph, simclusters_ann, twhin_ann, frs]
pools = run_parallel(sources, user, timeout=deadline - 500) # partial ok
cands = dedupe(flatten(pools)) - already_served(user)
cands = hydrate_features(cands, user, timeout=deadline - 250)
try:
probs = ranker.predict(cands, timeout=deadline - 60)
for c, p in zip(cands, probs):
c.score = sum(WEIGHTS[k] * p[k] for k in WEIGHTS)
except Timeout:
for c in cands:
c.score = c.post_id # recency fallback, never an empty feed
ranked = sorted(cands, key=lambda c: c.score, reverse=True)
ranked = author_diversity(ranked, max_per_author=2)
ranked = balance_in_out_network(ranked)
ranked = [c for c in ranked if visibility.allowed(user, c.post)]
return mix_in_ads_and_modules(ranked[:50], user)
Worked example: one request, end to end
Take a user who follows 400 accounts, one of which has 50 million followers, and opens the app after a few hours away. For Following, the cache holds entries delivered by fan-out while the user was away. The large account was never fanned out, so its last few posts are fetched from its author timeline and merged in. Hydration then drops two posts from an account muted after fan-out, and one deleted post. Forty posts go back in a few tens of milliseconds, mostly the hydration round trip.
For You is more expensive; the illustrative budget below shows where the time goes. The sources return perhaps a thousand candidates, one embedding source times out and is skipped, the ranker scores the rest, and heuristics trim the list to about fifty. The request still finishes on time because every stage has a deadline and a way to degrade.
| Stage | Illustrative budget | Degrades to |
|---|---|---|
| Candidate sources (parallel) | 300-400 ms | skip slow sources |
| Feature hydration | 150-250 ms | default feature values |
| Heavy ranking | 150-200 ms | recency order |
| Heuristics, filters, mixing | 20-50 ms | no degradation; must run |
Failure modes and operations
| Failure | What users see | Mitigation |
|---|---|---|
| Fan-out backlog after a spike | posts appear late in Following | priority queues for online users; alert on delivery lag p99, not queue depth |
| Cache shard lost | misses for millions of users at once | rate-limit rebuilds per shard; serve a degraded timeline from recent author posts while rebuilding |
| Hot post from a huge account | tweet cache shard overloaded | replicate hot keys into in-process caches; see hot key mitigation |
| Ranker slow or down | For You less relevant | recency fallback at a deadline; never return an empty page |
| Filter data stale | muted or blocked content shows | cache block lists with short TTLs; invalidate on change events |
Rebuild storms deserve special attention: when a cache shard restarts empty, every user on it misses at once, and unbounded rebuilds overload the stores they read from. Cap rebuild concurrency and apply backpressure from the rebuild queue rather than letting callers retry freely. The celebrity problem appears on the read side too: a post from a very large account is requested by millions of timelines within seconds, which is a textbook case for hot key mitigation. Shard the timeline cache by user id, as described in sharding strategies, so that a hot user's list never lands on the same shard as a hot post.
Trade-offs worth arguing about
- Store ids, not bodies. Small lists and immediately visible deletes, paid for with a hydration round trip per read.
- Fan-out threshold. Lower means less write amplification and more read-time merging. Tune it by measuring read latency against delivery cost.
- Precompute For You or not. Precomputing saves compute but serves stale rankings to users who may never open the app.
What to do next
- Write down your own load model: posts per second, average and maximum follower counts, and timeline reads per second. Compute insertions per second; that number decides whether fan-out on write is affordable.
- Define the timeline entry format and a cap, and estimate cache memory as active users times cap times entry size.
- Implement fan-out that skips inactive users and accounts above a threshold, and merge the skipped authors at read time.
- Move every correctness rule (deletes, blocks, mutes, visibility) to the read path and test it by blocking an account after fan-out.
- Give every For You stage a deadline and a fallback, then run a game day where one candidate source and the ranker are killed.
- Add a rate-limited rebuild path, then restart an empty cache shard in staging and measure the load on the stores behind it.