System Design Twitter Course

System Design Twitter Course

Lesson 28: The Backpressure Problem — Handling AI Lag During Traffic Spikes

sysdesign101's avatar
sysdesign101
Jul 16, 2026
∙ Paid

What We’re Building Today

A bounded work queue with configurable depth sitting between Redpanda ingestion and Ollama classification, preventing unbounded memory growth during traffic bursts

  • A priority lane that guarantees verified-account tweets are classified within 200ms P99 even when the queue is saturated at 500 items

  • A circuit breaker that falls back to a deterministic rule-based classifier when queue depth crosses threshold, keeping the pipeline alive under 10× load spikes


Why This Matters

In 2015, Slack’s message delivery system ran into what their engineering team called the “thundering herd after reconnect” problem. When a network partition healed, every disconnected client reconnected simultaneously and flooded the fanout service with pending message requests. The service had no mechanism to signal to upstream producers that it was at capacity — it accepted everything, queued everything, and then fell over trying to process everything at once. Recovery took longer than the original outage because the backlog compounded.

NEXUS faces the same physics. Ollama classifies at roughly 50 tweets/sec on commodity hardware. Redpanda can ingest 2,000 tweets/sec. Without a pressure-relief valve between them, a viral moment — a breaking news event, a celebrity post — creates a queue that grows at 1,950 items/sec. At that rate, you have 15 seconds before you’ve accumulated 30,000 pending items. At 20 seconds, the Node.js process is paging. At 30 seconds, OOM.


Core Concepts

User's avatar

Continue reading this post for free, courtesy of sysdesign101.

Or purchase a paid subscription.
© 2026 SystemDR · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture