Lesson 28: The Backpressure Problem — Handling AI Lag During Traffic Spikes
What We’re Building Today
A bounded work queue with configurable depth sitting between Redpanda ingestion and Ollama classification, preventing unbounded memory growth during traffic bursts
A priority lane that guarantees verified-account tweets are classified within 200ms P99 even when the queue is saturated at 500 items
A circuit breaker that falls back to a deterministic rule-based classifier when queue depth crosses threshold, keeping the pipeline alive under 10× load spikes
Why This Matters
In 2015, Slack’s message delivery system ran into what their engineering team called the “thundering herd after reconnect” problem. When a network partition healed, every disconnected client reconnected simultaneously and flooded the fanout service with pending message requests. The service had no mechanism to signal to upstream producers that it was at capacity — it accepted everything, queued everything, and then fell over trying to process everything at once. Recovery took longer than the original outage because the backlog compounded.
NEXUS faces the same physics. Ollama classifies at roughly 50 tweets/sec on commodity hardware. Redpanda can ingest 2,000 tweets/sec. Without a pressure-relief valve between them, a viral moment — a breaking news event, a celebrity post — creates a queue that grows at 1,950 items/sec. At that rate, you have 15 seconds before you’ve accumulated 30,000 pending items. At 20 seconds, the Node.js process is paging. At 30 seconds, OOM.



