Design a Ticket Booking System
Ticket booking looks like a standard e-commerce flow, and that is the trap. The hard part is contention: when tickets for a popular show go on sale, hundreds of thousands of people compete for a few thousand seats, and the system must never sell the same seat twice. It is the cleanest interview problem for concurrency control, temporary reservations and surviving a traffic stampede.
The chapter follows the usual shape: understand the problem, set up the interface, build the high-level design, then go deep on the questions interviewers use to separate levels.
1. Understanding the problem
Users browse events, pick seats, and pay. Venues, events and seats are finite, so every booking competes with others for the same inventory.
Functional requirements
Core:
- Users can search for and view events and their available seats.
- Users can book a ticket for a specific seat, or for general admission.
- A seat must never be sold to two people.
Confirm in or out: seat maps with choice of seat versus general admission, payments inside the system, refunds and cancellations, a waiting list, and resale. A sensible opening: "I will design browsing, search and booking with seat selection, with the guarantee that a seat is never double-sold, and treat resale and refunds as extensions."
Non-functional requirements
- Strong consistency for booking. Two buyers cannot both succeed for one seat. Here correctness beats availability: it is better to reject a request than to oversell.
- High availability and low latency for browsing. Reading event pages is far more frequent than booking and can tolerate slightly stale data.
- Handle massive spikes. A popular on-sale can bring orders of magnitude more traffic than normal.
- Fairness. People who arrive first should have a fair chance.
- Scale. Assume 10 million daily users in normal operation, and a flash sale of 50,000 seats attracting 500,000 simultaneous buyers.
Estimation
| Quantity | Calculation | Result |
|---|---|---|
| Normal browsing | 10 million users 20 page views / 86,400 s | about 2,300 reads per second |
| Normal bookings | 1 million tickets a day | about 12 per second |
| Flash sale | 500,000 users requesting seat maps within a minute | about 8,000 requests per second, with bursts far higher |
| Seat inventory | 50,000 seats in one event | tiny: kilobytes of state |
What the numbers say. The data is small, and normal booking volume is modest. What is hard is a concentrated burst on a small amount of contended state. You cannot solve this by adding generic capacity. You must control concurrency and traffic admission.
2. The set up
Core entities
- Event: name, venue, start time.
- Venue with a seat map.
- Seat (per event): section, row, number, price, status.
- Booking: user, seats, state, payment reference.
- User.
API
GET /v1/events?query=...&city=... -> search results
GET /v1/events/{id}/seats -> seat map with availability
POST /v1/events/{id}/reserve { seat_ids } -> { reservation_id, expires_at }
POST /v1/bookings { reservation_id, payment_token, idempotency_key }
-> { booking_id }
Booking is a two-step flow: reserve, then pay. The reservation holds the seats for a few minutes while the user pays. Without a hold, a user could pick a seat, fill in payment details, and then find it gone, and the system would have to keep seats unreserved during payment, which is where double sales happen.
3. High-level design
Split the system by access pattern:
- Event and search service. Read-heavy, cacheable, tolerant of slight staleness. Back it with a cache and CDN for event pages, and a search index for queries.
- Booking service. Low volume but correctness-critical. It owns the seat state and the reservation logic.
- Seat holds, a fast store with expiry, and a bookings database with transactions as the system of record.
- A payment provider reached through the booking service.
Separating reads from the contended write path lets you scale and protect them independently. Event pages can be served from a CDN to millions of people without touching the booking database at all.
A seat is a state machine
Each seat moves through a small set of states. This diagram is worth drawing in the interview, because every correctness question reduces to guarding its transitions.
<!--fig:states-->- Available to Held: when a user selects the seat. This transition must be atomic, so only one user can win it.
- Held to Booked: when payment is confirmed within the hold time.
- Held to Available: when the hold expires or the user cancels.
4. Potential deep dives
Deep dive 1: How do you prevent double booking?
The challenge. Two people click the same seat at the same moment. Exactly one must win.
Weak: check, then write. The service reads the seat, sees it is available, and then marks it sold. Two requests can both read "available" before either writes, and both succeed. This is the classic race condition, and describing it correctly is half the answer.
Solid: make the check and the write one atomic database operation. Use a transaction with a conditional update, so the database itself arbitrates:
UPDATE seats
SET status = 'held', held_by = :user, held_until = now() + interval '10 minutes'
WHERE event_id = :e AND seat_id = :s AND status = 'available';
-- affected rows = 1 means you won, 0 means someone else did
Only one concurrent update can change a row from available to held, so one request sees one affected row and the others see zero. Alternatives that achieve the same: a row lock taken before the check (pessimistic locking), or a version column with compare-and-set (optimistic locking), where a failed update means the data changed under you. Optimistic control suits low contention. For a hot seat, many failed retries waste work, so a pessimistic lock or a single writer is better there.
Excellent: a hold with an expiry, held in a fast store, with the database as the final authority. Use a fast store with time-to-live support for the hold: setting a key for the seat only if it does not already exist, with an expiry, is an atomic, sub-millisecond operation that handles the high request rate. If the user does not pay in time, the key expires and the seat becomes available again without any cleanup job. When payment succeeds, write the booking to the transactional database and mark the seat booked, treating that write as the source of truth. The database still enforces uniqueness on (event, seat), so even a bug or a lost hold cannot create two bookings for one seat. Layer the protections: the fast store absorbs the contention, and a database constraint is the safety net.
State the failure to design for: the hold store loses data. If it does, you must not oversell, so the final commit checks the database constraint, and at worst a user is told their reservation lapsed.
Deep dive 2: How do you survive the flash sale?
The challenge. 500,000 people arrive within a minute for 50,000 seats. Traffic is concentrated on a few hot rows and a handful of endpoints.
Weak: scale out the booking service and hope. More servers do not help when the bottleneck is contention on the same rows, and an overloaded database fails for everyone, including people who would have succeeded.
Solid: rate limit and cache aggressively. Cache event pages and static seat-map assets on a CDN. Rate limit per user. Make the seat availability read cheap, served from the cache with a short TTL, and let only the reserve call touch the contended path.
Excellent: a virtual waiting room that controls admission. Put users who arrive during a spike into a queue, show them a position and estimated wait, and admit them at a rate the booking path can handle. Admission control watches signals such as error rate and database load, and adjusts the rate. Each admitted user gets a signed pass that the booking service verifies.
<!--fig:waiting-->This converts an uncontrolled stampede into a steady stream, protects the database, and gives a fairness story: roughly first come, first served. Add protections around it: bot detection and per-account limits so that one person cannot hold many places, and a random draw within the first seconds if exact arrival order is not meaningful. Say that the waiting room is a product decision as much as a technical one.
Deep dive 3: How do users see current availability?
The challenge. Thousands of people are looking at the seat map while seats change state.
Weak: query the booking database for every page view. It adds heavy read load to the database that must stay responsive for bookings.
Solid: serve availability from a cache that is updated on state changes. Keep the seat map in a cache keyed by event. Update it when a seat is held, booked or released, and use a short TTL as a safety net. Slight staleness is acceptable here because the real check happens atomically at reserve time.
Excellent: push updates and design for stale reads. Send availability changes to connected clients through a pub-sub channel, so the seat map updates live without polling. Be explicit that the display is a hint, not a guarantee: a seat shown as free may be taken by the time the user clicks, so the reserve call returns a clear "seat unavailable" response, and the UI suggests alternatives. For large venues, return availability by section first and load seat-level detail on demand to keep payloads small.
Deep dive 4: Search and the event catalogue
The challenge. Users search by artist, city, date and category, with filters and typo tolerance.
Weak: SQL LIKE queries on the events table.
They do not rank, tolerate typos or scale to complex filters.
Solid: a search index fed from the database. Index events in a dedicated search engine. Keep it in sync by publishing changes from the database through a queue or change data capture. Search is a read model: it can lag by seconds, and the database remains the source of truth.
Excellent: a search index plus geography and popularity signals. Add geo queries for "events near me", boost by popularity and recency, and cache the most common queries. Keep inventory counts out of the index or update them lazily, since they change constantly. A search result that shows "sold out" for a few seconds longer than it should is harmless.
Deep dive 5: Payment and failure handling
The challenge. The user pays, but the network fails at the wrong moment, or the payment provider is slow.
Weak: charge the card and then try to mark the seat as booked. If the booking write fails after the charge, you have taken money for a ticket that does not exist.
Solid: use an idempotency key and a booking state. Create the booking in a pending state with an idempotency key, charge the provider with the same key, and move to confirmed only after the provider confirms. A retry with the same key returns the original result and does not charge twice.
Excellent: treat payment as an asynchronous workflow with reconciliation. The hold must outlast the payment attempt, so extend or pin the hold while a payment is in progress. If a payment response is lost, a background process queries the provider for the true status and resolves the booking either way, refunding automatically if the seat can no longer be given. Webhooks from the provider are verified and processed idempotently. Expire pending bookings that never complete and release their seats. Describe a compensation path for every step, so no failure leaves a charged customer without a ticket or a ticket without a charge.
5. What is expected at each level
Mid-level. You identify the double-booking race and prevent it with a transaction or lock. You propose a reservation with a time limit and a cache for event pages.
Senior. You explain atomic conditional updates, optimistic versus pessimistic control and when each fits, and a hold with expiry. You separate read and write paths, propose a waiting room for spikes, and handle payment failure with idempotency.
Staff. You reason about the whole trade-off space: fairness, bots, the cost of overselling versus rejecting, multi-region placement of inventory, operational runbooks for an on-sale, and how you would load test with realistic contention. You also discuss which guarantees you provide to the user and how you communicate failures.
6. Interview questions and model answers
Q: How do you stop two people buying the same seat? I make the seat transition from available to held a single atomic operation, either a conditional update in the database or an atomic set-if-absent with expiry in a fast store. Only one request can win it. A uniqueness constraint on event and seat in the bookings table is the final safety net.
Q: Why hold seats instead of booking immediately? Payment takes time. A hold with an expiry reserves the seat during that time without permanently removing it from sale, and it releases the seat automatically if the user abandons the purchase.
Q: What happens when 500,000 people arrive at once? A virtual waiting room admits users at a rate the booking path can sustain, driven by load signals, and hands each a signed pass. Event pages and seat maps come from a CDN, so only the reserve call touches the contended path.
Q: Optimistic or pessimistic locking? Optimistic when conflicts are rare, because it avoids holding locks. For a hot seat in a flash sale, many optimistic retries waste work, so I use an atomic conditional update or a short pessimistic lock on that row.
Q: What if the payment succeeds but the response is lost? The booking is pending with an idempotency key. A reconciliation job asks the provider for the real status and finalises or refunds. The user is never charged without either a ticket or a refund.
Q: Where would you accept stale data? Event pages, search results and the seat-map display. Never the reserve and book steps, which are checked atomically against the source of truth.
7. Common mistakes
- Checking availability and then writing in two separate steps.
- Booking at payment time with no hold, so seats are lost mid-payment.
- Using a cache as the only record of a booking.
- Trying to scale the contended path by adding generic servers.
- Charging before the booking is safely recorded, with no reconciliation.
- Forgetting to release expired holds.
- No plan for bots or one user grabbing many places.