Mid to Senior Engineer

System Design Interview Prep

A structured path from the interview framework through core concepts, key technologies and patterns to eighteen full problem breakdowns, each with diagrams and weak, solid and excellent answers to every deep dive.

Chapter 17 of 36Patterns · Multi-Step Processes and Workflows

Pattern: Multi-Step Processes and Workflows

Many real operations are not one database transaction but a sequence across services: place an order, charge the card, reserve stock, create a shipment, notify the customer. Each step can succeed, fail or hang, and a failure part-way leaves the system in a half-finished state. This pattern is about how to make such processes correct, observable and recoverable without a transaction that spans everything.

It builds on the sagas and outbox ideas in the consistency chapter. This chapter turns them into a design choice you can defend.

1. Why a normal transaction does not work

Inside one database, a transaction makes several changes all-or-nothing. Once the steps live in different services, each with its own database or an external provider, there is no shared transaction. A distributed transaction protocol such as two-phase commit exists, but it holds locks across services while waiting, blocks if the coordinator fails, and is unsupported by most external systems. You cannot enrol a payment provider or an email service in your transaction.

So you design for a different guarantee: eventual completion with a defined outcome for every failure. Each step either succeeds or is compensated, and the process always ends in a known state.

2. What goes wrong in a multi-step process

Make this list in the interview, because it is what the pattern exists to solve:

  • A step fails after earlier steps succeeded. The earlier effects must be undone, or the process must retry until it succeeds.
  • A step times out and you do not know whether it happened.
  • The process crashes half-way. After restart it must continue, not start again or forget.
  • A message is delivered twice, so a step runs twice.
  • Steps arrive out of order, or a later event overtakes an earlier one.
  • A step is slow, and the process must wait for hours or days (an approval, a delivery) without holding resources.
  • The business cancels while the process is running.

3. The building blocks

Idempotent steps. Every step must be safe to run twice, because retries and redeliveries are certain. Use an idempotency key derived from the process and the step, so that the second call returns the first result.

Durable state. The progress of the process (which steps are done, with what results) must be stored durably, so a crash does not lose it. The state is the source of truth, not the messages in flight.

Retries with backoff. Transient failures are retried automatically, with exponential backoff and jitter and a limit.

Timeouts. Every step and the process as a whole have a deadline, after which a defined action occurs (fail, compensate or escalate).

Compensating actions. For each step that changed something, define the action that undoes it semantically: refund a charge, release a reservation, cancel a shipment. A compensation is not a rollback. It is a new action, which can itself fail and must be retried until done. Some steps cannot be undone (an email sent), so order steps with the hardest-to-reverse last.

The outbox. A service that changes its state and must publish an event writes both in one local transaction, and a relay publishes from the outbox. This prevents the "database updated but message lost" failure.

4. Two ways to coordinate: orchestration and choreography

Choreography: services react to events

There is no central controller. Each service listens for events and does its part, then publishes its own event. The order service publishes OrderPlaced. Payment reacts, charges, and publishes PaymentCompleted. Shipping reacts to that, and so on. If payment fails, it publishes PaymentFailed, and the order service reacts by cancelling.

  • Pros: loose coupling, no single point of control, easy to add a new reactor (analytics listens to OrderPlaced without anyone changing).
  • Cons: the process is implicit. No single place shows the whole flow, so it is hard to see the current state, hard to debug why an order is stuck, and hard to change the order of steps. Risk of cyclic dependencies between events. Handling timeouts and overall deadlines is awkward.

Orchestration: a coordinator owns the flow

A central component, the orchestrator or workflow engine, defines the sequence, calls each service, records the result, decides what to do on failure, and runs compensations. The services do not know about each other.

<!--fig:workflow-->
ORCHESTRATION CHOREOGRAPHY 1 call 2 call 3 call OrderPlaced to payment paid Workflow engine state + retries + timers Payment Inventory Shipping Event bus Order Payment Shipping One place shows the whole flow and itsstate. Easy to reason about, to retry andto add timeouts. The engine must be reliable. No central owner, loose coupling. The flowis implicit across services, so it is harder tosee, debug and change. Figure 1. Orchestration (a coordinator owns the flow) versus choreography (services react to each other's events).
  • Pros: the flow is explicit and visible in one place, with its current state queryable. Retries, timeouts and compensation are handled centrally. Easy to change the sequence and to add conditions and waits.
  • Cons: the orchestrator is a critical component that must be reliable and scalable, and services become coupled to the orchestrator's calls.

Choosing

SituationPrefer
A few steps, simple, loosely related reactionsChoreography
Many steps, branches, waits, a business process you must monitorOrchestration
Need to answer "where is order 123 and why is it stuck?"Orchestration
Several independent consumers react to one eventChoreography (events)

A common, balanced answer: orchestrate the core business flow, and use events for side effects such as notifications and analytics.

5. Sagas in practice

A saga is the set of local transactions plus compensations. With orchestration:

  1. Reserve stock (compensation: release stock).
  2. Charge payment (compensation: refund).
  3. Create shipment (compensation: cancel shipment).

If step 3 fails permanently, the orchestrator runs compensations for 2 and then 1, in reverse order, retrying each until it succeeds. The order ends in a cancelled state, a valid terminal state.

Points to make:

  • Intermediate states are visible. Between steps 1 and 3, stock is reserved but the order is not complete. Design the model so that a pending state is explicit (reserved, payment_pending), and other flows respect it.
  • Isolation is weak. Unlike a database transaction, other operations can see and act on partial results. Counter-measures include semantic locks (marking an item as pending), commutative operations, and ordering steps so that the riskiest happen first.
  • A pivot step. In a sequence, the point after which the process must complete rather than compensate. Steps before it are compensatable, steps after it are retried until they succeed. Identify it.
  • Compensation must be reliable. A refund that fails must be retried, alerted on, and eventually handled by a person.

6. Durable execution and workflow engines

Hand-building an orchestrator on a queue and a table works for small flows. It becomes complicated as you add timers, retries, waits for human input, versioning and visibility. Workflow engines, which implement durable execution, solve this class of problem.

The idea: you write the process as ordinary code, and the engine records every step's result in durable history. If a worker crashes, another resumes the workflow by replaying the history, skipping steps already done, and continuing at the right place. The engine provides:

  • Automatic retries with backoff per step.
  • Durable timers, so a workflow can sleep for a week without holding a thread.
  • Signals, to deliver external events (an approval) to a waiting workflow.
  • Visibility, where you can query the state and history of any workflow instance.
  • Versioning support, for changing a workflow while instances are still running.

A discipline required by replay: the workflow code itself must be deterministic, and all side effects (calls, random numbers, current time) must go through the engine's step mechanism so they are recorded and not repeated on replay.

In an interview, you do not need a product name. Say that you would use a workflow engine with durable execution for any process with several steps, waits or compensation, and describe what it gives you. Mention the cost: another system to operate, and the determinism rule for workflow code.

7. State machines for the entity

Independently of the engine, model the business entity (the order) as an explicit state machine: created, payment pending, paid, shipped, delivered, cancelled, refunded. Define the allowed transitions, and reject any transition that is invalid from the current state. Persist the state with a version, so concurrent events cannot corrupt it. This makes late and duplicate events harmless: an event that arrives for a state already passed is ignored.

8. Observability and operations

Because processes are long-lived and distributed, you need to see them.

  • A process identifier (correlation id) carried in every call and log line, so you can trace one order across services.
  • Queryable state and history for each instance: where it is, what it has done, why it is waiting.
  • Metrics: counts by state, time in each state, failure and compensation rates.
  • Alerts for processes stuck in a state beyond an expected time, and for compensation failures.
  • Manual intervention tools: retry a step, skip a step, cancel a process, with an audit trail. Real systems need these.
  • Dead-letter handling for steps that exhausted retries.

9. A worked example

Problem. Design the order flow for an online store: reserve stock, charge the customer, create a shipment, and send a confirmation, with refunds on failure.

Reasoning.

  1. This has several steps across services, external dependencies (a payment provider), compensation and a need to answer "what is the status of order 123?". That favours orchestration.
  2. The order service starts a workflow instance per order, identified by the order id. The workflow engine persists its progress.
  3. Step 1, reserve stock, with a reservation that expires if the process stalls. Compensation: release.
  4. Step 2, charge payment, with an idempotency key from the order id. A timeout leaves the payment in an unknown state, so the workflow queries the provider and waits for the webhook before deciding. Compensation: refund.
  5. Step 3, create the shipment, retried until it succeeds, since the order has been paid. This is after the pivot: if shipping fails permanently, a person intervenes or the order is refunded.
  6. Step 4, notifications, are side effects published as events through an outbox. Failure to email does not fail the order.
  7. If the customer cancels while it is running, a signal reaches the workflow, which stops and runs the compensations for completed steps.
  8. The order entity is a state machine, so a duplicate PaymentCompleted event arriving later is ignored.
  9. Dashboards show orders by state and age, and alert on any order stuck in payment_pending or compensating beyond a threshold.

What I would say about the trade-off. "I use an orchestrator for the core flow, because I need visibility and controlled compensation, and events for side effects. The cost is operating a workflow engine, and the benefit is that every failure ends in a defined state."

10. Interview questions and model answers

Q: How do you keep data consistent across services without a distributed transaction? I use a saga: a sequence of local transactions with compensating actions, coordinated by an orchestrator or by events. Steps are idempotent, state is stored durably, and every failure ends in a defined state, either completed or compensated.

Q: Orchestration or choreography? For a business process with several steps, conditions and a need for visibility, orchestration, because the flow is explicit and failures are handled in one place. For loosely related reactions to an event, choreography. I often orchestrate the core flow and use events for side effects.

Q: What if a compensation fails? It is retried with backoff until it succeeds, since a compensation must be reliable. If it still fails, it alerts a person and the process stays in a visible compensating state.

Q: How do you handle a step that times out? I treat the outcome as unknown, make the step idempotent with a key, and then query the downstream system or wait for its callback before deciding. I never assume failure and run it again as a new operation.

Q: How do you make a service update its database and publish an event reliably? The outbox pattern: write the state change and the event to an outbox table in one local transaction, and a relay publishes from the outbox, with consumers deduplicating.

Q: What does durable execution give you? Workflow code that survives crashes: progress is recorded, so a process resumes where it stopped, with built-in retries, durable timers, signals and queryable history. The code has to be deterministic, with side effects routed through the engine.

11. Common mistakes

  • Trying to use a distributed transaction across services and external providers.
  • Steps that are not idempotent, so retries cause duplicates.
  • Keeping process state only in messages or memory.
  • No compensation for a step, or compensation that is not retried.
  • Treating a timeout as a failure and starting over.
  • Choreography for a complex flow that nobody can then trace.
  • No way to see, retry or cancel a stuck process.
Header Logo