Intermediate to senior

Backend Interview Prep

Fourteen chapters on HTTP and API design, SQL, indexing and transactions, NoSQL, authentication, caching, concurrency, messaging, resilience, deployment and observability, with tested SQL and Python.

Chapter 6 of 14Security and performance · Authentication and Authorisation

Authentication and Authorisation

Authentication answers "who are you?"; authorisation answers "what may you do?". They are separate concerns that are constantly confused, and getting either wrong is how breaches happen. Interviewers ask how you store passwords, how sessions and tokens differ, what a JWT is and is not, how OAuth works, and how you prevent users from reaching each other's data. This chapter covers each with code you can run.

1. Storing passwords

Never store passwords in plain text, and never use a fast hash (MD5, SHA-1, plain SHA-256) alone. Attackers who steal the database can test billions of guesses per second against fast hashes.

Use a slow, memory-hard, salted password hash: Argon2id (preferred), scrypt, bcrypt or PBKDF2 with a high iteration count.

  • Salt: a unique random value per password, stored alongside the hash. It defeats precomputed rainbow tables and makes identical passwords hash differently.
  • Work factor: tuneable cost, raised over time as hardware improves.
  • Pepper (optional): a secret kept outside the database (an HSM or secret manager), mixed into the hash.
  • Constant-time comparison when comparing hashes, to avoid timing leaks.
  • Re-hash on login when parameters are upgraded.
import hashlib, hmac, os

def hash_password(password, iterations=200_000):
    salt = os.urandom(16)
    digest = hashlib.pbkdf2_hmac("sha256", password.encode(), salt, iterations)
    return f"pbkdf2_sha256${iterations}${salt.hex()}${digest.hex()}"

def verify_password(password, stored):
    algo, iterations, salt_hex, digest_hex = stored.split("$")
    candidate = hashlib.pbkdf2_hmac("sha256", password.encode(), bytes.fromhex(salt_hex), int(iterations))
    return hmac.compare_digest(candidate.hex(), digest_hex)           # constant-time comparison

stored = hash_password("correct horse battery staple")
assert verify_password("correct horse battery staple", stored)
assert not verify_password("wrong password", stored)
assert hash_password("same") != hash_password("same")                  # unique salts: identical passwords look different
assert stored.count("$") == 3 and len(stored.split("$")[2]) == 32      # the salt is stored with the hash

Also: rate-limit and lock out repeated failures, check new passwords against breached-password lists, support multi-factor authentication, never log passwords, and make reset flows safe (single-use, short-lived tokens, no account enumeration in messages).

2. Sessions versus tokens

Server-side sessions

After login the server creates a random session ID, stores the session data (user ID, expiry) in a store such as Redis or a database, and sends the ID in a cookie. Each request looks the session up.

  • Immediate revocation (delete the session), easy logout-everywhere, small cookie.
  • Needs a shared session store when you have several servers, and a lookup per request.

Stateless tokens (JWT)

The server signs a token containing claims; any server with the key can verify it with no lookup.

  • Scales horizontally with no shared store; works well across services.
  • Hard to revoke before expiry; contents are visible to anyone holding the token; larger than a session ID.
Session IDJWT
Stateserver-sidein the token
Revocationinstantneeds a denylist or short expiry
Sizesmalllarger
Lookup per requestyesno (signature check)
Cross-service useneeds shared storenatural

Neither is universally better. A common pattern: short-lived access token (minutes) plus a long-lived refresh token that is stored server-side so it can be revoked and rotated.

3. JWT in detail

A JWT has three base64url parts separated by dots: header (algorithm, type), payload (claims) and signature. The signature covers header.payload. The payload is encoded, not encrypted: anyone can read it. Never put secrets in it.

Standard claims: iss (issuer), sub (subject, the user), aud (audience), exp (expiry), iat (issued at), nbf (not before), jti (unique ID).

import base64, json, hmac, hashlib, time

def b64(data):
    return base64.urlsafe_b64encode(data).rstrip(b"=").decode()

def b64decode(s):
    return base64.urlsafe_b64decode(s + "=" * (-len(s) % 4))

def sign_jwt(claims, secret):
    header = b64(json.dumps({"alg": "HS256", "typ": "JWT"}, separators=(",", ":")).encode())
    payload = b64(json.dumps(claims, separators=(",", ":")).encode())
    sig = hmac.new(secret, f"{header}.{payload}".encode(), hashlib.sha256).digest()
    return f"{header}.{payload}.{b64(sig)}"

def verify_jwt(token, secret, now=None, audience=None):
    header_b64, payload_b64, sig_b64 = token.split(".")
    header = json.loads(b64decode(header_b64))
    if header.get("alg") != "HS256":                                   # never trust the token's own algorithm choice (the "alg: none" attack)
        raise ValueError("unexpected algorithm")
    expected = hmac.new(secret, f"{header_b64}.{payload_b64}".encode(), hashlib.sha256).digest()
    if not hmac.compare_digest(expected, b64decode(sig_b64)):
        raise ValueError("bad signature")
    claims = json.loads(b64decode(payload_b64))
    if claims["exp"] <= (now if now is not None else time.time()):
        raise ValueError("expired")
    if audience and claims.get("aud") != audience:
        raise ValueError("wrong audience")
    return claims

secret = b"server-side-secret-key-32-bytes!!!"
token = sign_jwt({"sub": "user-42", "aud": "api", "exp": 2_000_000_000}, secret)
assert verify_jwt(token, secret, now=1_900_000_000, audience="api")["sub"] == "user-42"

def must_fail(fn):
    try:
        fn()
    except ValueError:
        return True
    return False

assert must_fail(lambda: verify_jwt(token, b"another-secret-key-another-secret!", now=1_900_000_000))   # wrong key
assert must_fail(lambda: verify_jwt(token, secret, now=2_000_000_001))                                      # expired
assert must_fail(lambda: verify_jwt(token, secret, now=1_900_000_000, audience="admin"))                    # wrong audience

# tampering with the payload breaks the signature
h, p, s = token.split(".")
forged = b64(json.dumps({"sub": "admin", "aud": "api", "exp": 2_000_000_000}, separators=(",", ":")).encode())
assert must_fail(lambda: verify_jwt(f"{h}.{forged}.{s}", secret, now=1_900_000_000))

# anyone can READ the payload without the key: it is encoded, not encrypted
assert json.loads(b64decode(token.split(".")[1]))["sub"] == "user-42"

JWT pitfalls:

  • Accepting alg: none, or letting the token choose between symmetric and asymmetric algorithms (algorithm confusion). Fix the expected algorithm on the server.
  • Not validating exp, iss and aud.
  • Long expiry with no revocation story.
  • Storing sensitive data in the payload.
  • Storing tokens in localStorage where XSS can read them (see the frontend security chapter).
  • Using asymmetric signing (RS256, ES256) when many services only need to verify: they hold the public key, and only the issuer holds the private key.

4. OAuth 2.0 and OpenID Connect

OAuth 2.0 is a framework for delegated authorisation: a user lets an application access their data on another service without sharing their password. OpenID Connect (OIDC) adds an identity layer on top (an ID token saying who the user is), which is what "Sign in with Google" uses.

Roles: resource owner (the user), client (the app), authorisation server (issues tokens), resource server (the API).

Authorisation code flow with PKCE (the standard for web, mobile and SPAs)

  1. The client generates a random code verifier and sends its hash (code challenge) when redirecting the user to the authorisation server.
  2. The user authenticates and consents; the server redirects back with a short-lived authorisation code.
  3. The client exchanges the code and the original code verifier for tokens at the token endpoint.
  4. The server checks that the verifier matches the earlier challenge. An attacker who steals the code cannot redeem it without the verifier.
import base64, hashlib, os

verifier = base64.urlsafe_b64encode(os.urandom(32)).rstrip(b"=").decode()
challenge = base64.urlsafe_b64encode(hashlib.sha256(verifier.encode()).digest()).rstrip(b"=").decode()

def server_checks(code_challenge, presented_verifier):
    expected = base64.urlsafe_b64encode(hashlib.sha256(presented_verifier.encode()).digest()).rstrip(b"=").decode()
    return expected == code_challenge

assert server_checks(challenge, verifier)
assert not server_checks(challenge, "an-attacker-guess")        # a stolen code without the verifier is useless
assert len(verifier) >= 43                                       # the RFC asks for 43 to 128 characters

Other essentials: the state parameter defends against CSRF on the redirect, exact redirect URI matching, scopes to limit access, short access tokens plus refresh tokens (rotate them, store securely), and the client credentials flow for service-to-service calls. The old implicit flow and password grant are discouraged.

5. Authorisation models

ModelIdeaFits
RBAC (role-based)users have roles; roles grant permissionsmost business apps: admin, editor, viewer
ABAC (attribute-based)policies over attributes of user, resource and contextfine-grained rules: "managers can approve expenses in their department under 50,000"
ReBAC (relationship-based)permissions follow relationships in a graphdocuments shared with people and groups (Google Docs style)
ACLsa list per resource of who may do whatfile systems, per-object sharing

Always enforce authorisation on the server, on every request, close to the data. Hiding a button in the UI is not a control.

ROLE_PERMS = {"viewer": {"read"}, "editor": {"read", "write"}, "admin": {"read", "write", "delete"}}

def can(user, action, resource):
    if action not in ROLE_PERMS.get(user["role"], set()):
        return False                                             # RBAC: the role must allow the action
    if resource["tenant_id"] != user["tenant_id"]:
        return False                                             # tenant isolation: never cross tenants
    if action == "write" and user["role"] == "editor" and resource["owner_id"] != user["id"]:
        return False                                             # ABAC rule: editors may edit only their own items
    return True

doc = {"owner_id": 1, "tenant_id": "t1"}
assert can({"id": 1, "role": "editor", "tenant_id": "t1"}, "write", doc)
assert not can({"id": 2, "role": "editor", "tenant_id": "t1"}, "write", doc)      # someone else's document
assert not can({"id": 3, "role": "admin", "tenant_id": "t2"}, "read", doc)        # another tenant, even an admin
assert not can({"id": 4, "role": "viewer", "tenant_id": "t1"}, "write", doc)
assert can({"id": 5, "role": "admin", "tenant_id": "t1"}, "delete", doc)

The most common authorisation bug: IDOR

Insecure direct object reference (also called broken object-level authorisation, the top item on the OWASP API Security list): the API checks that you are logged in but not that this object is yours. Changing /orders/101 to /orders/102 returns someone else's order.

Fix: load objects through the authenticated user's scope (WHERE id = ? AND user_id = ?), or check ownership explicitly, and return 404 (not 403) when you do not want to reveal existence.

orders = {101: {"id": 101, "user_id": 1}, 102: {"id": 102, "user_id": 2}}

def get_order_vulnerable(order_id, current_user_id):
    return orders.get(order_id)                                  # authenticated, but not authorised

def get_order_safe(order_id, current_user_id):
    o = orders.get(order_id)
    return o if o and o["user_id"] == current_user_id else None  # scoped to the caller; the route maps None to 404

assert get_order_vulnerable(102, 1) is not None                  # user 1 can read user 2's order
assert get_order_safe(102, 1) is None
assert get_order_safe(101, 1)["id"] == 101

Related: mass assignment (a client sets role: "admin" in a JSON body that is bound straight onto a model). Use explicit allow-lists of writable fields.

6. Other authentication topics

  • API keys: simple identifiers for programmatic access. Treat them as passwords: hash at rest, show once, support rotation and scoping, and send in headers, never URLs.
  • Mutual TLS for service identity; service accounts and short-lived credentials (workload identity) instead of long-lived static secrets.
  • Single sign-on (SSO): SAML or OIDC with an identity provider, typical for enterprises.
  • Multi-factor authentication: TOTP apps, WebAuthn/passkeys (phishing-resistant), SMS as the weakest option.
  • Session security: rotate the session ID on login (prevents session fixation), HttpOnly; Secure; SameSite cookies, idle and absolute timeouts.
  • Account recovery is an attack surface: use single-use, expiring tokens and notify the user.
  • Secrets management: keep secrets out of code and images; use a secrets manager; rotate; scan repositories.
  • Webhook signatures: verify an HMAC of the payload with a shared secret, and check a timestamp to prevent replays.
import hmac, hashlib, time

def sign_webhook(body, secret, ts):
    return hmac.new(secret, f"{ts}.".encode() + body, hashlib.sha256).hexdigest()

def verify_webhook(body, secret, ts, signature, now, tolerance=300):
    if abs(now - ts) > tolerance:
        return False                                              # too old or too new: a replay
    return hmac.compare_digest(sign_webhook(body, secret, ts), signature)

body, wh_secret = b'{"event":"paid","id":"p1"}', b"whsec_test"
sig = sign_webhook(body, wh_secret, 1000)
assert verify_webhook(body, wh_secret, 1000, sig, now=1100)
assert not verify_webhook(b'{"event":"paid","id":"p2"}', wh_secret, 1000, sig, now=1100)   # altered payload
assert not verify_webhook(body, wh_secret, 1000, sig, now=5000)                            # replayed much later

7. A secure login checklist

  1. Transport over HTTPS only (with HSTS).
  2. Slow salted password hashing; breached-password checks.
  3. Rate limiting and lockout or progressive delays; bot protection.
  4. MFA available, required for sensitive roles.
  5. Generic error messages ("invalid credentials") to avoid account enumeration (and similar response time).
  6. New session ID after login; secure cookie flags; short-lived access tokens.
  7. Logging of authentication events and alerts on anomalies.
  8. Authorisation checks on every request, at the object level.

8. Common mistakes

  • Fast hashes (SHA-256, MD5) for passwords, or no salt.
  • Confusing authentication and authorisation; checking login but not ownership (IDOR).
  • Trusting the client: role or user ID in a request body or an unsigned cookie.
  • JWTs as sessions with no revocation and long lifetimes; accepting alg: none.
  • Secrets in the payload or in source control.
  • Roll-your-own cryptography or token schemes; use vetted libraries and standards.
  • Reusing one signing key everywhere with no rotation plan.
  • Different error messages for "no such user" and "wrong password".
  • Authorisation only in the front end.

9. Practice questions

  1. How should passwords be stored? Why is SHA-256 not enough?
  2. Compare server-side sessions and JWTs. How do you revoke a JWT?
  3. What are the parts of a JWT, and what can an attacker do with one they steal?
  4. Explain the OAuth authorisation code flow and what PKCE adds.
  5. What is the difference between OAuth and OpenID Connect?
  6. Describe RBAC and ABAC with an example where RBAC is not enough.
  7. What is an IDOR and how do you prevent it systematically?
  8. How would you secure service-to-service calls inside a cluster?
Header Logo