Design a Cloud Drive: File Storage, Sync and Video Delivery
Designing a cloud drive, a photo service or a video platform is mostly the problem of storing very large objects cheaply and durably, and moving them to and from clients efficiently. The engineering that earns marks lives in chunking, de-duplication, the separation of metadata from content, and conflict handling across devices.
The chapter builds the design in order, then goes deep on uploads, durability, sync and video. Each deep dive compares a weak, a solid and an excellent answer.
1. Understanding the problem
Users store files in the cloud, access them from several devices, and share them with others. The service must feel like a local folder that happens to be everywhere.
Functional requirements
Core:
- Upload and download files, including very large ones.
- Files sync automatically across a user's devices.
- Share a file or folder with other people.
Confirm in or out: version history, preview and thumbnails, search, offline access, collaborative editing. A sensible opening: "I will design upload, download, sync and sharing, with versions, and leave real-time co-editing out."
Non-functional requirements
- Durability above everything. Losing a customer's file is the worst failure, so the system must never lose acknowledged data.
- Reliable uploads over poor networks. A dropped connection must not restart a 4 GB upload from zero.
- Efficient use of bandwidth and storage.
- Metadata consistency. Two devices must agree on what a folder contains.
- Availability for reads, with moderate latency for large transfers.
Estimation
Assume 100 million users with 10 GB each, and 10 million files uploaded a day averaging 1 MB.
| Quantity | Calculation | Result |
|---|---|---|
| Total content | about 1 exabyte before replication | |
| Upload rate | about 116 MB per second average | |
| Files | about 100 per user | files |
| Metadata | 1 KB per file | about 10 TB |
What the numbers say. An exabyte cannot live in a database, so content goes to object storage. Metadata is a thousand times smaller and far hotter, so it lives in a sharded database. Content is huge and cold, metadata is small and hot, and the design separates them.
2. The set up
Core entities
- File (and folder): name, parent, owner, size, permissions, current version.
- File version: an ordered manifest listing the chunks that make up the content.
- Chunk: an immutable piece of data identified by the hash of its content.
- Device, and the change log per user.
API
POST /v1/files/prepare { path, size, chunk_hashes: [...] } -> { missing: [...], upload_urls: [...] }
PUT <signed upload url> (raw chunk bytes, sent straight to object storage)
POST /v1/files/commit { path, chunk_hashes: [...], base_version } -> { version }
GET /v1/files/{id}/download -> manifest + signed download URLs
GET /v1/changes?since=41 -> ordered list of changes after sequence 41
Large bytes never pass through application servers if you can avoid it. The server issues a pre-signed URL: a time-limited link scoped to one object, and the client talks to object storage directly. That keeps application capacity independent of bytes transferred.
3. High-level design
Two stores with different jobs:
- Object storage holds chunks, immutable, addressed by content hash, replicated or erasure-coded for durability.
- A metadata database, sharded by user, holds the file tree, versions, manifests and permissions.
A metadata service sits in front of the database, issues signed URLs, and runs the commit step. Clients upload and download chunks directly. A garbage collector removes chunks that no manifest references.
<!--fig:hld-->How an upload works
- The client splits the file into chunks, computes a hash for each and calls prepare with the list.
- The metadata service replies with the chunks it does not already have and signed URLs for those.
- The client uploads the missing chunks in parallel directly to object storage, retrying any that fail.
- The client calls commit. The metadata service checks that every chunk exists, then atomically points the file at the new manifest as a new version.
The commit is the only moment the file becomes visible, and it is atomic, so no device ever sees a half-uploaded file.
4. Potential deep dives
Deep dive 1: How do you upload very large files reliably?
The challenge. A 4 GB upload on a mobile connection will be interrupted. Restarting from zero is unacceptable.
Weak: a single streaming upload through an application server. Any interruption loses everything, and every byte flows through your servers, so application capacity scales with bandwidth.
Solid: chunked uploads to object storage with resume. Split the file into chunks of a few megabytes. Upload each independently with a signed URL. After an interruption, the client asks which chunks are already present and uploads only the rest. Chunks can go in parallel, which also speeds up the transfer.
Excellent: content-addressed chunks with de-duplication and delta sync. Name each chunk by the hash of its content. Before uploading, the client sends the hashes, and the server says which it lacks. Identical data is stored once, so uploading a file that already exists, in your account or someone else's, transfers almost nothing. When a large file changes slightly, only the changed chunks are new.
Two refinements:
- Fixed-size versus content-defined chunking. With fixed-size chunks, inserting one byte near the start shifts every later boundary, so all chunks change. Content-defined chunking picks boundaries from the data itself using a rolling hash, so a small edit changes only nearby chunks. It costs more CPU and gives much better delta sync.
- Privacy of de-duplication. If the server confirms "I already have this chunk", one user can learn that someone else holds a particular file. Many services scope de-duplication to an account, or use extra proof that the client really holds the data.
Deep dive 2: How do you make storage durable?
The challenge. Disks fail, racks lose power, data silently corrupts, and regions go offline. Acknowledged data must survive.
Weak: one copy on one disk, with periodic backups. A single failure loses data written since the last backup, and silent corruption goes unnoticed until someone opens the file.
Solid: replicate across failure domains. Keep three or more copies on different disks, racks and availability zones. Reads are fast and recovery is simple, at three times the storage.
Excellent: tiered replication and erasure coding, with integrity checks. Replicate recent, hot data for speed, and move colder data to erasure coding: split it into data fragments plus parity fragments so any of the rebuild the data. A scheme of 10 data plus 4 parity survives the loss of any four fragments and uses 1.4 times the storage instead of three times. The cost is CPU and slower reconstruction.
Layer on integrity: store a checksum with every chunk, verify on read, and run background scrubbing that reads data and repairs corruption before it spreads. Replicate across regions for disaster recovery. State that durability figures such as "eleven nines" are design targets built from replication and fast repair, not promises to quote.
Deep dive 3: How do devices stay in sync?
The challenge. A change on one device must reach the others, including ones that were offline, without losing changes or applying them twice.
Weak: each device polls the whole tree periodically and compares. Expensive, slow to notice changes, and it does not scale to millions of devices.
Solid: a notification triggers a fetch of the changes. When a file changes, the server notifies other devices, which then ask for the changed items. A persistent connection or a push message carries the notification.
Excellent: an ordered change log per user. Every committed change appends an entry to the user's change log with an increasing sequence number. A device remembers the last sequence number it has applied, and asks for "changes since n". Notifications are only a hint to ask sooner. A device that was off for a week asks with its old number and receives exactly what it missed, in order. This is the same pattern as message sync in a chat system: a durable ordered log is the truth and live notification is an accelerator.
<!--fig:sync-->Conflicts. Two devices edit the same file offline and both commit against the same base version. The second commit finds the base has moved. For opaque files, silently picking a winner is data loss, so keep both: store the later one as a "conflicted copy" and let the user decide. For structured documents, a merge may be possible. Say which applies.
Detecting local changes. Devices watch the file system for events and also scan periodically, because event delivery is not perfectly reliable.
Deep dive 4: Sharing and permissions
The challenge. Control who can read or edit each file, including through shareable links, without a slow check on every request.
Weak: put permissions on each blob in object storage. Blobs know nothing about users, and permissions per chunk would be duplicated across every file that shares the chunk.
Solid: access control entries in the metadata service. Store permissions on files and folders as entries linking users or groups to roles, with inheritance down the tree. Check them on every metadata request and when issuing signed URLs. A shared link is a high-entropy token mapped to a file, a permission level and an optional expiry or password.
Excellent: the same, with careful propagation and short-lived access. Make signed URLs expire quickly so a leaked link does little harm. For folder sharing, avoid rewriting permissions on every descendant, and resolve inheritance at read time against a cached path. Log access to sensitive files. The key point to state: access is enforced in the metadata layer and at URL issuance, never in the blob store.
5. Video delivery
A video platform adds a media pipeline and heavy use of a CDN.
Upload and processing. Store the original as a blob. A job queue then triggers transcoding into several resolutions and bitrates, and into formats suited to different devices. Transcoding is CPU-heavy and parallel, so split the video into segments and process them on many workers. Generate thumbnails, extract metadata and run safety or copyright checks before publishing.
Streaming. Use adaptive bitrate streaming. Cut the video into segments of a few seconds at each quality level and describe them in a manifest. The player downloads the manifest, picks a quality that fits the current bandwidth, and switches up or down as conditions change. Common protocols are HLS and MPEG-DASH, which run over ordinary HTTP, so standard CDNs can cache the segments.
CDN strategy. A small share of titles gets most views, so cache popular segments at the edge and leave the long tail at the origin. For live events, the edge fetches each new segment once and serves it to many viewers.
Cost. Storage and delivery dominate. Store fewer renditions for rarely watched videos, move cold content to cheaper storage, and watch the CDN cache hit rate.
6. What is expected at each level
Mid-level. You separate metadata from content, store files in object storage and propose chunking so large uploads can resume. You describe a basic sync mechanism.
Senior. You explain pre-signed URLs, content-addressed chunks with de-duplication and a commit step that makes a version visible atomically. You choose the change log for sync, handle conflicts without data loss, and discuss durability with replication or erasure coding.
Staff. You discuss cost and operations at exabyte scale: tiering, garbage collection safety, scrubbing, cross-region replication and its cost, abuse and malware scanning, and privacy trade-offs of de-duplication. You can also argue what not to build, such as real-time co-editing in the first version.
7. Interview questions and model answers
Q: Why not store files in the database? Large blobs bloat the database, hurt backups and replication, and waste its strengths. I store content in object storage and keep a small metadata record, which lets each scale on its own terms.
Q: How do you make uploads reliable for a very large file? Split it into chunks, upload them in parallel through pre-signed URLs, retry individual chunks, and commit the manifest once all arrive. A dropped connection resumes from the last good chunk.
Q: How does de-duplication work, and what is the catch? Chunks are addressed by content hash, and the client asks which hashes the server lacks. Only those are uploaded, so identical data is stored once. The catch is a privacy leak if one user can learn that another has a file, so I scope de-duplication carefully.
Q: How do you keep devices in sync? An ordered change log per user. A device asks for changes since its last sequence number, so a missed notification is harmless. Conflicts from offline edits keep both versions.
Q: How do you make storage durable? Replicate across failure domains, or erasure-code colder data, checksum every chunk, scrub in the background, and replicate across regions for disaster recovery.
Q: How do you delete data safely? Deleting a file removes the manifest reference, not the chunks. A garbage collector later removes chunks that no manifest references, after a grace period that also protects in-flight uploads.
8. Common mistakes
- Streaming uploads through application servers.
- Storing blobs in the metadata database.
- One-piece uploads that restart from zero on failure.
- Making a file visible before all its chunks have landed.
- Silently overwriting on a sync conflict.
- Treating durability as a vendor promise instead of designing it.
- Enforcing permissions in the blob store.