Pattern: Real-Time Updates
Many products need to show something to a user the moment it happens: a new message, a stock price, a driver's position, a collaborator's edit, a notification count. HTTP was built for the client to ask and the server to answer, so pushing updates to clients needs deliberate design. This chapter covers the options for delivering updates, the architecture for fanning them out to many clients, and the failure cases that interviewers probe.
1. Decide how real-time "real-time" must be
Before choosing a technology, pin down the requirement:
- Latency. Within a second? Within thirty seconds? A dashboard that refreshes every ten seconds does not need a persistent connection.
- Direction. Server to client only, or both ways?
- Frequency and volume. One update a minute or a hundred a second per client?
- Audience. One recipient, a small group, or millions subscribed to the same stream?
- Loss tolerance. Is a missed update harmless (a stock tick, because the next one comes soon) or damaging (a chat message)?
The answers decide whether you need the simplest approach or a full push infrastructure. Over-engineering real-time is a common mistake.
2. Techniques for getting updates to a client
Short polling
The client asks the server for changes on a timer, say every five seconds.
- Pros: the simplest. Works everywhere, stateless servers, easy to cache.
- Cons: wasteful, since most polls return nothing. Latency is up to the poll interval. Cost grows with clients and frequency.
- Use when: updates are infrequent, a delay of seconds is fine, and the audience is moderate.
Long polling
The client sends a request, and the server holds it open until there is something to send or a timeout expires. The client then immediately sends another.
- Pros: near-immediate delivery over ordinary HTTP, works through most proxies.
- Cons: every message costs a new request, and servers hold many open requests.
- Use when: you need push-like behaviour without WebSockets, or as a fallback.
Server-sent events
A single long-lived HTTP response over which the server streams events to the client. The browser reconnects automatically and can resume from the last event identifier.
- Pros: simple, works over plain HTTP, automatic reconnection and resume.
- Cons: one direction only, from server to client.
- Use when: feeds, notifications, live scores, progress updates.
WebSockets
A persistent, two-way connection upgraded from HTTP. Either side can send at any time with low overhead.
- Pros: low latency and overhead, bidirectional.
- Cons: stateful servers, harder load balancing and scaling, proxies and firewalls occasionally interfere, and you must handle reconnection, heartbeats and authentication yourself.
- Use when: chat, collaborative editing, multiplayer games, and anything chatty in both directions.
Push notifications
For apps in the background, the operating system's push channel wakes the device with a message. It is best-effort and rate limited, and delivery is not guaranteed, so use it to say "something happened", and let the app fetch the data.
A selection guide
| Need | Choose |
|---|---|
| Seconds of delay is fine, low volume | Polling |
| Server to client updates, simple | Server-sent events |
| Two-way, low latency | WebSockets |
| Behind restrictive networks, simple push | Long polling |
| App is closed | Mobile push notification |
Say that you would start with the simplest technique that meets the latency requirement, and that production systems often combine one of these with a polling fallback.
3. The architecture: connections, a broker and fan-out
The hard part is not one connection. It is delivering one event to the right set of connected clients, across many servers.
Connection servers. A fleet of servers holds the persistent connections. They are stateful but thin: authenticate, track which user and channels each connection cares about, and forward messages. A well-tuned server can hold on the order of tens of thousands or more idle connections, so a million connected clients needs tens to a hundred or more servers. Treat the figure as something to measure.
The routing problem. A publisher creates an event for user 42, but user 42's connection lives on one of many servers. How does the event find it? Two common answers:
- A pub-sub broker. Each connection server subscribes to the channels its clients care about. The publisher sends the event once to the broker, and the broker delivers it to every server subscribed to that channel. Each server then writes it to the right sockets.
- A session registry. Keep a map of user to connection server. The sender looks up the server and sends the message directly. Cheaper per message for one-to-one traffic, but you must keep the map accurate.
Choosing channels. A channel can be per user (notifications), per conversation (chat), per document (collaboration), or per topic (a stock symbol). The number and size of channels determine broker load. A channel with ten million subscribers is a different problem from ten million channels with one subscriber each.
Fan-out cost. Delivering one event to subscribers costs socket writes. For small groups that is trivial. For a broadcast to millions, you need hierarchical fan-out, in which the broker sends to servers, and each server writes to its own thousands of sockets, and you accept that delivery takes seconds, not milliseconds.
4. Reliability: what happens when connections drop
Persistent connections break constantly on mobile networks. A design that treats the live connection as the only way to get data will lose updates.
Principle: the live push is an optimisation, and durable state is the truth.
- Sequence numbers or event identifiers. Every update has an increasing number per channel. The client remembers the last one it received.
- Catch-up on reconnect. After reconnecting, the client says "give me everything after number n", served from durable storage, such as a log or the database. A missed push is recovered without special cases.
- Heartbeats. Both sides send periodic pings so that dead connections are detected, and idle connections are not silently dropped by intermediaries.
- Reconnect with backoff and jitter. After a server restart or outage, clients reconnecting all at once cause a thundering herd. Randomised exponential backoff spreads the load.
- At-least-once delivery with deduplication. The server may resend after an ambiguous failure, and the client discards events it has already seen, by their numbers.
- Acknowledgements where loss matters, so the server knows what to retry.
For data where only the latest value matters (a position, a price), send the latest state instead of every change, so a client that missed several updates is brought current by one message.
5. Scaling and operations
- Connection limits and load balancing. Balance new connections across servers by least connections. Plan for draining: when you deploy, close connections gradually so that clients reconnect over time, not all together.
- Per-connection memory. Buffers for slow clients accumulate. Apply a limit and drop or disconnect clients that cannot keep up, instead of letting one slow consumer exhaust a server.
- Backpressure. If a client reads slowly, do not queue unlimited updates. Coalesce updates (keep only the latest per key) or disconnect.
- Authentication and authorisation. Authenticate at connection time, and authorise every subscription: a user must not be able to subscribe to another user's private channel.
- Rate limits. On both inbound messages and subscriptions per connection.
- Multi-region. Place connection servers near users. Events published in one region must be replicated to the regions where subscribers are. Say where the broker lives and how cross-region delivery works, and accept added latency for far-away subscribers.
- Observability. Track connected clients, connect and disconnect rates, message delivery latency, dropped messages and the lag of the catch-up path.
6. Special cases
Presence (who is online). Heartbeats with short expiry, and updates sent only to those watching. Do not broadcast every change to every contact.
Collaborative editing. Concurrent edits to one document must merge without conflict. The usual approaches are operational transformation, which rewrites operations against each other, and conflict-free replicated data types, which are designed so that all replicas converge regardless of the order in which operations arrive. In an interview, naming these and explaining that a central server orders operations, or that the data type guarantees convergence, is sufficient.
Live location. Send the latest position at a modest rate, and let the client interpolate. Do not store every point for the live view.
Live counters and dashboards. Aggregate on the server and push the totals at a fixed rate, such as once a second, instead of pushing every event.
Very large audiences (live streams, sports). Fan out through a hierarchy and a CDN, accept a few seconds of delay, and send aggregated updates rather than every event.
7. A worked example
Problem. Show a delivery customer the courier's live position and status changes, for 2 million active deliveries.
Reasoning.
- Direction: server to client only, so server-sent events or WebSockets, with the simpler SSE if it suffices. Latency of a few seconds is acceptable.
- The courier app sends its position every few seconds to a location service, which updates the latest-position store and publishes to a channel named for the delivery.
- A connection server holds each customer's connection and subscribes to its delivery channel on the broker. When a position arrives, the broker delivers it to the server, which writes it to the customer's socket.
- Only the latest position matters, so a slow or reconnecting customer receives the current state, not a backlog.
- Status changes (picked up, arriving, delivered) are important, so they go into a durable log with sequence numbers, and a reconnecting client catches up from its last number. They also trigger a push notification, in case the app is closed.
- Capacity: 2 million connections at 50,000 per server needs about 40 connection servers plus headroom. The position update rate is millions of updates over a few seconds, which the broker and a sharded latest-position store handle.
What I would say about the trade-off. "Positions can be a few seconds stale or skipped, since the next one supersedes it. Status changes must not be lost, so they are durable and replayable."
8. Interview questions and model answers
Q: How would you push updates to a web client? It depends on direction and latency. For server-to-client updates, server-sent events are the simplest. For two-way, low latency traffic like chat or collaboration, WebSockets. For infrequent updates where seconds of delay are fine, plain polling is enough. I use the simplest one that meets the requirement.
Q: How do you deliver an event to the right connection among thousands of servers? Connection servers subscribe to the channels their clients care about on a pub-sub broker. The publisher sends once to the broker, which delivers to the servers that have subscribers. Alternatively, a session registry maps each user to a server for direct delivery.
Q: What if the connection drops and a message is missed? Every update has a sequence number, the client remembers the last one, and on reconnect it asks for everything after that from durable storage. The live push is an optimisation, and the stored data is the truth.
Q: How do you handle a million clients reconnecting after a restart? Clients use exponential backoff with jitter, servers limit new connections per second while recovering, and deployments drain connections gradually.
Q: What about a slow client? Bound the per-connection buffer, coalesce updates to the latest value where possible, and disconnect clients that cannot keep up so they cannot exhaust server memory.
9. Common mistakes
- Using WebSockets for something polling would handle.
- Treating the live connection as the only source of data, with no catch-up path.
- No heartbeat, so dead connections linger.
- Reconnecting without backoff, causing a thundering herd.
- Broadcasting presence or typing events to everyone.
- Authorising at connect time but not per subscription.
- Unbounded per-client buffers.