Designing a URL shortener
The classic interview system, built small and measured: where the reads go, what the cache is for, and the one decision that made everything else easy.
- 100M redirects / day, 1M new links / day
- Strong for creation, eventual for analytics
- PostgreSQL · Redis · Kafka
- Built and load-tested locally
On this page
The problem
Turn long URLs into short codes and redirect quickly. The interesting part is the ratio: reads outnumber writes by roughly a hundred to one, so the design is really a read-path design with a small write path attached.
Requirements
Functional
- Create a short code for a URL, optionally with a custom alias.
- Redirect a code to its URL.
- Count clicks per code.
Non-functional
- p99 redirect under 50 ms.
- A created link must redirect immediately.
- Analytics may lag by minutes.
Out of scope
- Accounts, link editing and expiry policies.
Assumptions
- 100M redirects/day ≈ 1,160/s average, ~5,000/s peak.
- 1M new links/day; 7-character base62 codes give 3.5 trillion combinations.
- 500 bytes per link → about 180 GB a year before indexes.
Architecture
Fig. 1 — Read path
Diagram description
A client calls the API gateway, which forwards to the shortener service. The service looks the code up in a Redis cache first and reads PostgreSQL only on a miss. Every redirect also publishes a click event to Kafka asynchronously; an analytics worker consumes those events.
A stateless service in front of PostgreSQL, with Redis absorbing the read load. Click tracking is pushed off the request path entirely.
APIs
- POST
/links - Create a link. Body: url, optional alias. Returns the code.
- GET
/{code} - 302 redirect to the stored URL.
- GET
/links/{code}/stats - Click counts by day (eventually consistent).
Data model
CREATE TABLE links (
code VARCHAR(10) PRIMARY KEY,
url TEXT NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);Caching
| What | Where | TTL | Invalidation |
|---|---|---|---|
| code → url | Redis | 24 h, refreshed on hit | None needed: links are immutable |
| Missing codes | Redis | 5 min | Expires; protects PostgreSQL from scans |
Links never change, which is the decision that made the cache trivial: no invalidation, only expiry.
Failure scenarios
| Scenario | Impact | Detection | Mitigation |
|---|---|---|---|
| Redis unavailable | All reads hit PostgreSQL; latency rises. | Cache error rate, p99 latency alert. | Read replicas sized for full read load for a short window. |
| Kafka unavailable | Click events lost or delayed. | Producer error metrics. | Local buffer with bounded size; redirects unaffected. |
| Hot link (viral) | One Redis key takes most traffic. | Per-key request rate. | Short in-process cache in front of Redis for top keys. |
Observability
- Redirect latency (p50/p99) and status codes by route.
- Cache hit ratio — the single most useful number here.
- Kafka producer errors and consumer lag.
Trade-offs
| Option | Pros | Cons | Verdict |
|---|---|---|---|
| Hash the URL | Stateless; same URL, same code. | Collisions need handling; custom aliases fit awkwardly. | Rejected |
| Counter + base62 | No collisions; short codes. | Sequential codes are guessable; needs a coordinated counter. | Chosen, with ranges handed out per instance |
| Random codes | Unguessable; no coordination. | Collision checks on every write. | Kept for custom aliases only |
What I would change
I spent too long on the counter service. Handing each instance a range of 10,000 codes at start-up removed the coordination problem in an afternoon, and I would start there next time. I would also measure the cache hit ratio before adding the in-process layer; 99.2% turned out to be plenty.