Testing, Deployment and CI/CD for Backend Services
Shipping reliably is part of the job. Interviewers ask how you test a service, how you release without downtime, how you change a database safely, and how you roll back. This chapter covers the testing strategy for backends, the deployment strategies (rolling, blue-green, canary), feature flags, database migrations, and the CI/CD pipeline that ties them together, with runnable examples of the logic behind each.
1. Testing a backend service
| Level | What it exercises | Notes |
|---|---|---|
| Unit | pure logic: pricing, validation, state machines | fast, many; no I/O |
| Integration | your code with a real database, queue or cache | the highest-value tests for backends: SQL, transactions, serialisation, migrations |
| Contract | the agreement between a consumer and a provider | catches breaking API changes without running everything |
| Component / service | the whole service in isolation with dependencies stubbed | checks wiring, config, HTTP surface |
| End to end | several real services together | few, slow, for critical journeys |
| Load and performance | throughput, latency, resource use under load | before launches and in CI for regressions |
| Security | dependency scanning, static analysis, authorisation tests | automate in the pipeline |
| Chaos / resilience | behaviour under faults | targeted, in staging first |
A healthy suite has many fast unit tests, a solid layer of integration tests, and few end-to-end tests. The balance for backends tilts toward integration tests, because most bugs live at boundaries (queries, serialisation, transactions) that mocks hide.
Principles
- Test behaviour through the public interface (HTTP, a service method), not private functions.
- Deterministic: control time, randomness, IDs and ordering; no test depends on another.
- Fast feedback: the suite runs on every commit; slow suites get skipped.
- Realistic dependencies where it matters: use a real database (a container started by the test run, or an ephemeral instance) rather than mocking SQL.
- Test failures, not only success: timeouts, duplicates, malformed input, partial failures, concurrency.
- Arrange, act, assert, with clear names such as "rejects a transfer that would overdraw the account".
Use real databases in integration tests
Mocking the data layer proves your mock works. A real engine catches wrong SQL, missing indexes, constraint violations and transaction behaviour. Containers (Testcontainers) make this practical.
import sqlite3
def make_db():
db = sqlite3.connect(":memory:") # a fresh database per test: isolation and speed
db.executescript("""
CREATE TABLE accounts (id INTEGER PRIMARY KEY, balance INTEGER NOT NULL CHECK (balance >= 0));
INSERT INTO accounts VALUES (1, 100), (2, 0);
""")
return db
def transfer(db, src, dst, amount):
try:
with db:
db.execute("UPDATE accounts SET balance = balance - ? WHERE id = ?", (amount, src))
db.execute("UPDATE accounts SET balance = balance + ? WHERE id = ?", (amount, dst))
return True
except sqlite3.IntegrityError:
return False
def test_transfer_moves_money():
db = make_db()
assert transfer(db, 1, 2, 40) is True
assert db.execute("SELECT id, balance FROM accounts ORDER BY id").fetchall() == [(1, 60), (2, 40)]
def test_overdraft_is_rejected_and_nothing_changes():
db = make_db()
assert transfer(db, 1, 2, 500) is False
assert db.execute("SELECT balance FROM accounts ORDER BY id").fetchall() == [(100,), (0,)] # the constraint and the rollback protect the data
for t in (test_transfer_moves_money, test_overdraft_is_rejected_and_nothing_changes):
t()
Test doubles at the edges
Stub or fake things you do not own or that are slow and nondeterministic: payment gateways, email, the clock, other teams' services. Prefer fakes (an in-memory implementation) and contract tests over elaborate mocks, and run a small number of tests against the real third-party sandbox.
class FakePaymentGateway:
def __init__(self, fail_on=None):
self.charges, self.fail_on = [], fail_on
def charge(self, token, amount):
if token == self.fail_on:
raise ConnectionError("gateway down")
self.charges.append((token, amount))
return f"ch_{len(self.charges)}"
def checkout(gateway, token, amount, order_log):
try:
charge_id = gateway.charge(token, amount)
except ConnectionError:
order_log.append("payment_failed")
return None
order_log.append("paid")
return charge_id
log = []
assert checkout(FakePaymentGateway(), "tok", 500, log) == "ch_1" and log == ["paid"]
log2 = []
assert checkout(FakePaymentGateway(fail_on="bad"), "bad", 500, log2) is None and log2 == ["payment_failed"] # the failure path is tested too
Contract testing
In a service ecosystem, the consumer defines what it needs from the provider (fields and types it reads); the provider's build verifies it still satisfies every consumer's contract. This catches a renamed field before deployment, without deploying both services in a shared environment.
consumer_contract = {"id": int, "email": str, "plan": str} # the fields this consumer actually reads
def provider_satisfies(response, contract):
return all(k in response and isinstance(response[k], t) for k, t in contract.items())
assert provider_satisfies({"id": 1, "email": "a@x", "plan": "pro", "extra": True}, consumer_contract) # extra fields are fine
assert not provider_satisfies({"id": 1, "email_address": "a@x", "plan": "pro"}, consumer_contract) # a renamed field breaks the consumer
Property-based and mutation testing (names worth knowing)
Property-based testing (Hypothesis, jqwik, fast-check) generates many random inputs and checks invariants ("sorting then sorting again changes nothing", "decode(encode(x)) == x"). Mutation testing changes your code in small ways to check that tests actually fail; it measures test quality better than line coverage.
2. The CI/CD pipeline
Continuous integration: every change is merged frequently and automatically built and tested. Continuous delivery: every passing change is deployable at any time. Continuous deployment: every passing change goes to production automatically.
A typical pipeline:
- Commit and review (small pull requests, code owners).
- Build (compile, resolve dependencies from a lockfile, build the container image once).
- Static checks: formatting, lint, type checks, dependency and secret scanning, licence checks.
- Tests: unit, integration, contract.
- Package and publish an immutable, versioned artefact (the same image goes to every environment; configuration is injected, not rebuilt).
- Deploy to staging, run smoke and end-to-end checks.
- Deploy to production progressively, with automated health checks and rollback.
- Post-deploy verification and monitoring of the new version's metrics.
Principles: build once, deploy many; trunk-based development with short-lived branches; fast pipelines (minutes, not hours); environment parity; infrastructure as code (Terraform, CloudFormation, Pulumi) and GitOps (the desired state lives in version control and a controller reconciles); everything reproducible from a commit.
3. Deployment strategies
| Strategy | How | Rollback | Trade-off |
|---|---|---|---|
| Recreate | stop the old, start the new | redeploy the old | downtime |
| Rolling update | replace instances gradually | roll back gradually or reverse | old and new versions run side by side, so APIs and data must be compatible; needs readiness checks |
| Blue-green | run a full new environment (green) beside the old (blue), then switch traffic | flip back instantly | double the capacity during the switch; databases still need compatibility |
| Canary | send a small share of traffic to the new version, watch metrics, increase gradually | route traffic back | needs good metrics and automated analysis; the safest for risky changes |
| Shadow (mirroring) | copy live traffic to the new version without affecting users | n/a | validates behaviour and performance without risk; careful with side effects |
| Feature flag rollout | deploy dark, enable per user or percentage | turn the flag off | decouples deploy from release; flags need cleanup |
A canary decision, in code
A canary compares the new version's error rate against the baseline using enough traffic to be meaningful, then promotes or rolls back automatically.
import math
def canary_decision(base_errors, base_total, canary_errors, canary_total, min_requests=500, z_crit=2.33):
if canary_total < min_requests:
return "wait" # not enough data to decide
p1, p2 = base_errors / base_total, canary_errors / canary_total
pooled = (base_errors + canary_errors) / (base_total + canary_total)
se = math.sqrt(pooled * (1 - pooled) * (1 / base_total + 1 / canary_total)) or 1e-9
z = (p2 - p1) / se
return "rollback" if z > z_crit else "promote" # only a significantly worse error rate fails the canary
assert canary_decision(100, 50000, 3, 400) == "wait"
assert canary_decision(100, 50000, 2, 1000) == "promote" # 0.2 % versus 0.2 %
assert canary_decision(100, 50000, 30, 1000) == "rollback" # 3 % versus 0.2 %
Zero-downtime requirements
- Graceful shutdown: on
SIGTERMstop accepting new requests, finish in-flight ones (within a deadline), close connections, then exit. - Readiness probes so traffic only reaches instances that are ready.
- Connection draining at the load balancer.
- Backward and forward compatibility between adjacent versions of services, APIs, messages and database schemas, since old and new run together.
4. Feature flags
A feature flag gates behaviour at runtime. Uses: dark launches, gradual rollouts, A/B tests, kill switches, per-customer entitlements, trunk-based development of unfinished features. Hygiene matters: name and own each flag, set expiry, remove it after rollout, and test both states. A percentage rollout must be sticky (the same user always gets the same answer) which a hash of the user ID provides.
import hashlib
def flag_enabled(flag, user_id, percent):
bucket = int(hashlib.sha256(f"{flag}:{user_id}".encode()).hexdigest(), 16) % 100
return bucket < percent
users = [f"user-{i}" for i in range(10000)]
enabled = [u for u in users if flag_enabled("new-checkout", u, 20)]
assert 1700 < len(enabled) < 2300 # about 20 % of users
assert all(flag_enabled("new-checkout", u, 20) for u in enabled) # sticky: asking again gives the same answer
assert all(flag_enabled("new-checkout", u, 50) for u in enabled) # raising the percentage only adds users, never removes
5. Database migrations in a live system
Schema changes must be compatible with both the old and the new code, because during a rolling deploy both run at once. Use the expand and contract (parallel change) pattern:
- Expand: add the new column or table (nullable or with a default). Old code ignores it.
- Deploy code that writes both old and new (and reads the old).
- Backfill existing rows in small batches, throttled.
- Switch reads to the new column; verify.
- Stop writing the old column.
- Contract: drop the old column in a later release.
Never rename or drop a column in the same release that stops using it. Keep migrations small, versioned, reviewed and idempotent; test them on production-sized data; avoid long locks (create indexes concurrently, add constraints as not-validated first, then validate). Have a rollback plan: forward-fix is often safer than reversing a destructive migration, and backups must be tested.
db = sqlite3.connect(":memory:")
db.executescript("CREATE TABLE users (id INTEGER PRIMARY KEY, full_name TEXT); INSERT INTO users VALUES (1,'Asha Rao'),(2,'Ravi Iyer'),(3,'Meera');")
# 1. expand
db.execute("ALTER TABLE users ADD COLUMN first_name TEXT")
# 3. backfill in batches
def backfill(batch=2):
rows = db.execute("SELECT id, full_name FROM users WHERE first_name IS NULL LIMIT ?", (batch,)).fetchall()
for uid, full in rows:
db.execute("UPDATE users SET first_name = ? WHERE id = ?", (full.split()[0], uid))
return len(rows)
batches = 0
while backfill():
batches += 1
assert batches == 2 # three rows processed two at a time
assert db.execute("SELECT id, first_name FROM users ORDER BY id").fetchall() == [(1, "Asha"), (2, "Ravi"), (3, "Meera")]
assert db.execute("SELECT COUNT(*) FROM users WHERE full_name IS NOT NULL").fetchone()[0] == 3 # the old column is untouched until the contract step
6. Configuration, secrets and environments
- Twelve-factor app habits: configuration in the environment, one codebase, stateless processes, logs as streams, disposable processes with fast start and graceful shutdown.
- Secrets in a secrets manager (Vault, cloud KMS and secret stores), never in code or images; rotate them; grant least privilege through workload identity.
- Environments: development, staging (as close to production as practical), production; avoid snowflake servers; use the same artefact everywhere.
- Immutable infrastructure: replace instances rather than patching them.
7. Containers and orchestration (the essentials)
A container image packages the application and its dependencies; keep images small (multi-stage builds), pinned and scanned, and run as a non-root user. Kubernetes schedules containers: a Deployment manages replicas and rolling updates; a Service gives a stable address; liveness and readiness probes drive restarts and traffic; resource requests and limits control placement and protect neighbours; a Horizontal Pod Autoscaler scales on metrics; ConfigMaps and Secrets inject configuration. You do not need to be a Kubernetes expert for most backend interviews, but you should explain these concepts.
8. Rollback and incident readiness
- Make rollback boring: a one-command, practised path back to the previous version; keep the last known good artefact; ensure data changes are backward compatible so rollback is possible.
- Deploy small and often: small changes are easier to diagnose and revert.
- Deploy outside peak hours and Fridays if rollback is risky; prefer automation that makes timing irrelevant.
- Runbooks, on-call rotations and postmortems (blameless, with action items).
9. Common mistakes
- Mocking the database and shipping broken SQL.
- Slow, flaky pipelines that people bypass.
- Rebuilding artefacts per environment so production runs code that was never tested.
- Breaking changes in a single step (dropping a column the old version still reads).
- Deploying without readiness checks or graceful shutdown, dropping requests.
- Feature flags that live forever and interact unpredictably.
- No automated rollback or health-based promotion.
- Secrets in repositories or images.
- Testing only the happy path.
10. Practice questions
- How would you test a service that talks to a database, a queue and a payment gateway?
- What is a contract test and what problem does it solve?
- Compare rolling, blue-green and canary deployments. When would you use each?
- How do you rename a database column with zero downtime?
- What is trunk-based development and why does it need feature flags?
- Describe a CI/CD pipeline you would build for a new service.
- How do you make sure a deploy does not drop in-flight requests?
- A deploy increased errors. What is your process, from detection to resolution?