A database failover in eu-north did not complete cleanly and left read replicas serving stale connections. p95 latency on the REST API rose from 90 ms to roughly 1.4 s. Writes were unaffected throughout.
The stuck replicas were cycled at 09:51 and latency returned to normal within ten minutes. We have since shortened the replica health-check timeout from 30 s to 8 s so the load balancer drops a stale replica before clients notice.
A malformed payload from a single customer integration caused the delivery worker to retry in a tight loop, consuming the shared queue budget. Deliveries to other customers were delayed by up to 40 minutes. No events were lost.
Per-tenant rate limiting has been added to the delivery worker.
Object storage was migrated to the new cluster. The API remained available; uploads were queued for approximately four minutes during the cut-over.