Kubernetes AI inference latency spike during peak hours
Problem – Latency Spike in Kubernetes‑Hosted AI Inference Service During peak traffic windows the inference endpoint that serves ~100 ms predictions suddenly starts responding in 1 s +. The spike is repeatable, lasts for the duration of the load burst, and then returns to baseline once traffic subsides. Key observations: CPU and memory usage on GPU‑accelerated pods stay … Read more