Design a Payment System
Payments are the problem where "mostly correct" is not acceptable. Money must never be created, lost or moved twice, and the system depends on an external provider that is slow, occasionally wrong and sometimes silent. A good answer is organised around three ideas: idempotency, an honest state machine, and a ledger that can be audited.
This chapter is written for the common interview version of the problem: an application that accepts customer payments through an external payment provider and records them correctly. It does not cover building a card network or a bank. It also does not give legal or compliance advice. Where regulations matter, such as card data security rules, say that you would involve the relevant specialists and confirm current requirements.
1. Understanding the problem
A customer pays a merchant, through a checkout in a website or app. The system takes the payment instruction, calls a payment provider to move the money, and records the outcome. Later it may refund, and it must always be able to explain where every unit of money went.
Functional requirements
Core:
- A customer can pay for an order, and the system tells them whether it succeeded.
- A payment is processed exactly once, even if the customer or the network retries.
- The system keeps an accurate record of every payment, refund and balance, and can reconcile it against the provider's records.
Confirm in or out: multiple payment methods, subscriptions and recurring billing, multi-currency, payouts to merchants, disputes and chargebacks, and fraud checks. A sensible opening: "I will design one-off payments through an external provider with refunds, a ledger and reconciliation, and treat subscriptions and payouts as extensions."
Non-functional requirements
- Correctness above all. No double charges, no lost payments, no money without a record.
- Durability. A recorded payment must survive failures.
- Reliability under partial failure. The provider call can time out or fail in ambiguous ways, and the system must still converge to the right state.
- Auditability. Every change is traceable.
- Reasonable latency. A payment completes within a few seconds, and availability matters, but correctness wins when the two conflict.
- Security. Sensitive payment data is protected and kept out of your own systems where possible.
Estimation
Assume 1,000 payments per second at peak, 100 million per day at the extreme.
| Quantity | Calculation | Result |
|---|---|---|
| Payments per day | stated | |
| Average rate | about 1,160 per second | |
| Rows per payment | payment, state history, ledger entries, events | about 10 |
| Storage per year | payments about 2 KB of records | about 70 TB |
What the numbers say. The throughput is modest and the volume is easy for a sharded relational database. The difficulty is not scale. It is correctness under failure: every one of those payments involves a network call to an unreliable dependency, and every failure mode must resolve correctly.
2. The set up
Core entities
- Payment: identifier, amount, currency, customer, merchant, order reference, state, idempotency key, provider reference.
- Ledger entry: account, direction (debit or credit), amount, payment reference, time.
- Account: customer, merchant and platform accounts the ledger tracks.
- Refund: linked to a payment, with its own state.
API
POST /v1/payments
headers: Idempotency-Key: 6f1d...
body: { "order_id": "o-77", "amount": 4999, "currency": "INR", "payment_token": "tok_..." }
returns: 201 { "payment_id": "p-9", "state": "processing" | "succeeded" | "failed" }
GET /v1/payments/{id} -> current state
POST /v1/payments/{id}/refunds { "amount": 1000 }
Two details matter immediately.
Amounts are integers in the smallest currency unit, such as paise or cents, never floating point, because floating point cannot represent decimal money exactly.
The client sends a payment token, not raw card details. The sensitive data is collected by the provider's hosted fields or SDK and exchanged for a token, so your servers never see or store full card numbers. That keeps most of the security and compliance burden with the provider. Say that you would confirm the current compliance requirements with specialists.
3. High-level design
A payment service owns the lifecycle. It stores each payment in a payments database, writes ledger entries in a ledger, and calls an external payment provider to move the money. The provider also sends asynchronous results through webhooks, and a reconciliation job compares your records with the provider's settlement files.
<!--fig:hld-->The key design principle: the provider is not a reliable participant. Requests can be lost, responses can be lost, and the provider can tell you the result later. The payment service is designed so that each of those cases leads to a correct final state.
The payment as a state machine
Every payment is in exactly one state, and transitions are the only way it changes.
<!--fig:states-->The state you must not omit is unknown, sometimes called pending. When a call to the provider times out, you do not know whether the money moved. Pretending it failed risks a double charge if you retry. Pretending it succeeded risks a free order. The honest answer is "unknown", followed by a resolution step.
4. Potential deep dives
Deep dive 1: How do you make sure a payment is processed exactly once?
The challenge. The customer double-clicks. The mobile app retries after a timeout. A load balancer replays a request. Each could cause a second charge.
Weak: trust the client not to retry. Retries are inevitable, and a design that depends on them not happening will double charge.
Solid: an idempotency key. The client generates a unique key per logical payment attempt and sends it on every retry. The server stores the key together with the payment, and when it sees the key again it returns the stored result instead of starting a new payment. A unique constraint on the key makes this safe against two simultaneous requests.
Excellent: idempotency end to end, in one transaction, with the provider too. Three details raise the answer:
- Store the key and the new payment in the same database transaction. If they were separate, a crash between them would create a payment with no key, or a key with no payment, and a retry would behave wrongly.
- Pass the same idempotency key to the provider, so that if your retry reaches the provider twice, the provider also deduplicates. Providers commonly support idempotency keys for this reason.
- Reject a reused key with a different request body. The same key with different parameters is a client bug, and returning an error is safer than guessing.
Phrase the guarantee precisely: the network provides at-least-once delivery, and idempotency converts it into exactly-once effects.
Deep dive 2: What do you do when the provider call times out?
The challenge. You sent a charge and got no answer. Did the customer pay?
Weak: treat the timeout as a failure and tell the customer to try again. If the charge actually went through, the retry charges them twice, or you have taken money for an order you reported as failed.
Solid: mark the payment unknown, then ask the provider. Record the payment as unknown. A background process queries the provider for the status of that idempotency key or reference, and moves the payment to succeeded or failed once the provider answers. Meanwhile show the customer "processing" rather than a final answer.
Excellent: resolve unknowns with several signals, and never guess. Combine the webhook the provider will eventually send, polling by reference at increasing intervals, and the next day's reconciliation as a last line of defence. Whichever arrives first resolves the payment, and the others are ignored because the transition is idempotent. If the customer retries during the unknown window, the same idempotency key returns "processing" instead of charging again. Alert on payments stuck in unknown for longer than an expected threshold, since those are exactly the cases a human needs to look at.
Deep dive 3: How do you record money so it can be trusted?
The challenge. Balances must be correct and auditable. Updating a "balance" column in place makes errors invisible and history unrecoverable.
Weak: store a balance per account and update it. Concurrent updates race, a bug silently changes money, and there is no trail showing how a balance came to be.
Solid: record every movement as an immutable transaction record. Append a record for each payment and refund and never modify it. Balances are computed from the records. This gives history and auditability, and corrections are made by adding new records, not editing old ones.
Excellent: a double-entry ledger. Every movement of money is recorded as a pair of entries, a debit in one account and a credit in another, and the entries for one event always sum to zero. For a 100.00 payment that includes a 3.00 platform fee, the ledger records the customer's debit of 100.00, a credit to the merchant of 97.00 and a credit to platform fees of 3.00, all written in one transaction.
<!--fig:ledger-->Why this is the standard answer:
- Money is never created or destroyed in the books, because debits equal credits. A violation is a detectable bug.
- Balances are derived from entries, so they can always be recomputed and checked.
- The ledger is append-only. A refund is a new set of entries that reverse the originals, not a deletion.
- It supports audit and reconciliation naturally.
Keep a balance as a cached, derived value for speed, and verify it against the ledger periodically.
Deep dive 4: How do you reconcile with the provider?
The challenge. Your records and the provider's records can differ because of lost responses, bugs, late settlements, fees and disputes.
Weak: assume they match. Differences accumulate unnoticed until finance finds them months later.
Solid: a daily reconciliation job. Download the provider's settlement report, match each line to a payment by provider reference, and flag mismatches: payments you recorded but the provider did not, the reverse, and differences in amount or state.
Excellent: automated matching, categorised exceptions and fix workflows. Match automatically, classify the exceptions (missing in your system, missing at the provider, amount mismatch, duplicate), and route each class to an automated fix or a human queue with the evidence attached. Make the fixes themselves ledger entries with a reason, never silent edits. Track the reconciliation match rate as a health metric, because a falling rate signals a bug before customers notice.
Deep dive 5: Webhooks and asynchronous results
The challenge. The provider tells you the result later through an HTTP callback, and you must handle it safely.
Weak: accept the webhook and update the payment. Anyone who learns the URL can forge a webhook and mark an order paid, and retries by the provider would apply the update more than once.
Solid: verify and deduplicate. Verify the webhook's signature using the shared secret so that only the provider can create valid events. Record each event identifier and ignore ones already processed. Return success quickly, and do the processing asynchronously.
Excellent: verify, deduplicate, order-check and trust the provider as the source of truth for the outcome. Check that the transition is valid for the payment's current state, since webhooks can arrive out of order or twice, and ignore stale ones. For important events, fetch the current state from the provider rather than trusting the payload alone. Make the handler idempotent, and store raw events for later investigation. Respond with a success status only after the event is durably recorded, so the provider's retries are your safety net.
Deep dive 6: Refunds, failures and consistency across services
Refunds are payments in reverse and follow the same discipline: an idempotency key, a state machine, ledger entries that reverse the original ones, and a provider call that may be ambiguous. Limit the total refunded to the original amount.
Cross-service consistency. The order service must learn that a payment succeeded. Use the outbox pattern: write the payment event to an outbox table in the same transaction as the state change, and publish it from there, so the order is updated reliably without a distributed transaction. Handle the reverse failure with compensation: if the order cannot be fulfilled after payment, issue a refund.
Concurrency. Two operations on one payment, such as a refund arriving while a webhook updates the state, must not corrupt it. Use optimistic locking with a version column or a row lock, and validate every transition against the state machine.
Deep dive 7: Scale, availability and security
Scale. Shard the payments and ledger by a key such as the merchant or customer so that one payment's records live together, and no distributed transaction is needed for the main flow. Cache read-only data, but never cache the authority for state.
Availability. If the primary provider is down, a second provider can serve as a fallback, with care: a payment whose first attempt is unknown must be resolved before it is attempted elsewhere, or the customer may be charged twice.
Security. Keep card data out of your systems by using tokens, encrypt sensitive data at rest, restrict who and what can create refunds, log every action, and use strong authentication for internal tools. Confirm the applicable regulations and standards with compliance specialists rather than relying on memory.
5. What is expected at each level
Mid-level. You describe the payment flow, store payments in a database with a state, and recognise that retries cause double charges. You propose an idempotency key when prompted.
Senior. You design idempotency in one transaction and through to the provider, include the unknown state and a resolution process, and explain a double-entry ledger. You handle webhooks safely and discuss reconciliation.
Staff. You reason about the whole operation: provider failover without double charging, disputes and chargebacks, multi-currency and rounding, regulatory boundaries, the reconciliation operation and its metrics, and how you would prove the ledger is correct. You describe how finance, risk and engineering share the data.
6. Interview questions and model answers
Q: How do you avoid double charging? An idempotency key from the client, stored with the payment in the same transaction and passed to the provider too. A repeat of the key returns the stored result. The network gives at-least-once delivery and idempotency makes the effect exactly-once.
Q: The provider call timed out. What now? Mark the payment unknown and resolve it by querying the provider, listening for the webhook and, finally, reconciliation. Never assume failure and retry as a new payment.
Q: Why a ledger instead of a balance column? A double-entry ledger is append-only and every event's entries sum to zero, so money cannot silently appear or vanish, history is complete and balances can be recomputed and audited.
Q: How do you trust a webhook? Verify its signature, deduplicate by event identifier, check the state transition is valid, and, for critical events, fetch the authoritative state from the provider.
Q: Why store amounts as integers? Floating point cannot represent decimal fractions exactly, so rounding errors creep into money. Integers in the smallest currency unit are exact.
Q: How does the order system learn the payment succeeded? Through an event written to an outbox in the same transaction as the payment's state change, then published reliably. A failed fulfilment triggers a compensating refund.
7. Common mistakes
- No idempotency, or idempotency that is not atomic with the payment record.
- Treating a timeout as a failure and retrying as a new payment.
- Updating a balance column instead of keeping a ledger.
- Floating point for money.
- Trusting a webhook without verifying it.
- Skipping reconciliation, so differences are found months later.
- Storing raw card data instead of tokens.