Resolved
Operations back to normal
Monitoring
Restoring the required compute capacity resolved the degradation of inference performance. The system is recovering, and our team is still actively monitoring.
Investigating
Degraded Performance: Elevated Latency and Reduced Throughput
We are currently experiencing degraded performance affecting inference requests.
Customers may experience higher time-to-first-token (TTFT), increased inter-token latency, and reduced throughput compared with normal service levels.
Our engineering team is actively working to restore the affected capacity and rebalance traffic. We are monitoring the system closely and will provide further updates as recovery progresses.
Impact: Elevated latency and reduced throughput for affected US inference traffic