Skip to content

Designing a URL shortener

The classic interview system, built small and measured: where the reads go, what the cache is for, and the one decision that made everything else easy.

, 16 min read

Scale target
100M redirects / day, 1M new links / day
Consistency
Strong for creation, eventual for analytics
Core tech
PostgreSQL · Redis · Kafka
Status
Built and load-tested locally
On this page

The problem

Turn long URLs into short codes and redirect quickly. The interesting part is the ratio: reads outnumber writes by roughly a hundred to one, so the design is really a read-path design with a small write path attached.

Requirements

Functional

  • Create a short code for a URL, optionally with a custom alias.
  • Redirect a code to its URL.
  • Count clicks per code.

Non-functional

  • p99 redirect under 50 ms.
  • A created link must redirect immediately.
  • Analytics may lag by minutes.

Out of scope

  • Accounts, link editing and expiry policies.

Assumptions

  • 100M redirects/day ≈ 1,160/s average, ~5,000/s peak.
  • 1M new links/day; 7-character base62 codes give 3.5 trillion combinations.
  • 500 bytes per link → about 180 GB a year before indexes.

Architecture

Fig. 1 — Read path

A client calls the API gateway, which forwards to the shortener service. The service looks the code up in a Redis cache first and reads PostgreSQL only on a miss. Every redirect also publishes a click event to Kafka asynchronously; an analytics worker consumes those events.
Redirects never wait on analytics: the click event is published after the response is decided. Open full size
Diagram description

A client calls the API gateway, which forwards to the shortener service. The service looks the code up in a Redis cache first and reads PostgreSQL only on a miss. Every redirect also publishes a click event to Kafka asynchronously; an analytics worker consumes those events.

A stateless service in front of PostgreSQL, with Redis absorbing the read load. Click tracking is pushed off the request path entirely.

APIs

POST/links
Create a link. Body: url, optional alias. Returns the code.
GET/{code}
302 redirect to the stored URL.
GET/links/{code}/stats
Click counts by day (eventually consistent).

Data model

schema.sql
CREATE TABLE links (
    code        VARCHAR(10) PRIMARY KEY,
    url         TEXT        NOT NULL,
    created_at  TIMESTAMPTZ NOT NULL DEFAULT now()
);

Caching

What is cached, and for how long
WhatWhereTTLInvalidation
code → urlRedis24 h, refreshed on hitNone needed: links are immutable
Missing codesRedis5 minExpires; protects PostgreSQL from scans

Links never change, which is the decision that made the cache trivial: no invalidation, only expiry.

Failure scenarios

Failure scenarios
ScenarioImpactDetectionMitigation
Redis unavailable All reads hit PostgreSQL; latency rises.Cache error rate, p99 latency alert.Read replicas sized for full read load for a short window.
Kafka unavailable Click events lost or delayed.Producer error metrics.Local buffer with bounded size; redirects unaffected.
Hot link (viral) One Redis key takes most traffic.Per-key request rate.Short in-process cache in front of Redis for top keys.

Observability

  • Redirect latency (p50/p99) and status codes by route.
  • Cache hit ratio — the single most useful number here.
  • Kafka producer errors and consumer lag.

Trade-offs

Generating codes
OptionProsConsVerdict
Hash the URL Stateless; same URL, same code.Collisions need handling; custom aliases fit awkwardly.Rejected
Counter + base62 ChosenNo collisions; short codes.Sequential codes are guessable; needs a coordinated counter.Chosen, with ranges handed out per instance
Random codes Unguessable; no coordination.Collision checks on every write.Kept for custom aliases only

What I would change

I spent too long on the counter service. Handing each instance a range of 10,000 codes at start-up removed the coordination problem in an afternoon, and I would start there next time. I would also measure the cache hit ratio before adding the in-process layer; 99.2% turned out to be plenty.

Filed under

  • Caching
  • Redis
  • PostgreSQL
  • Kafka

Related

  1. System design, , 9 min read

    Idempotent consumers, without the ceremony

    What actually happens when the same message arrives twice, and the smallest set of changes that makes a consumer safe to retry.

    • Messaging
    • Kafka
    • PostgreSQL
  2. System design, , 10 min read

    Redis caching: the invalidation I got wrong first

    Cache-aside looked simple until two writers raced. What I changed, and what I would measure before adding a cache next time.

    • Caching
    • Redis
  3. Journal, , 8 min read

    Rebuilding the billing flow as events

    Why the second version of an event-driven design is usually the simpler one, and the retry question that decided the whole shape.

    • Building
    • Kafka

Next

Bring me a structure that stopped fitting

If a system has outgrown its own diagram, that is the interesting part. Describe it and I will sketch the trade-offs with you.

Start a conversation