On 17 August, between 17:24 and 18:05 UTC, a small percentage of SIP call transfers failed in our US East region during a routine configuration rollout. An issue in the automation that manages our SIP signaling servers prevented outgoing servers from being taken out of service safely, and those servers shut down while calls were still active on them. Transfers have completed normally since 18:05 UTC. Calls themselves stayed connected, and no other region was affected.
The configuration change restarts the servers that handle SIP signaling one at a time. Before a server is shut down, it is removed from service so that no new calls reach it, and it is then given some time to finish the calls it is already handling.
In this case that removal did not complete expectedly and the outgoing servers kept receiving new calls for the entire drain duration, then shut down on schedule with calls still active on them.
Only SIP call transfers (TransferSIPParticipant) in our US East region were affected, between 17:24 and 18:05 UTC. This represented 0.04% of all active calls in that window, and 1.1% of the calls that attempted a transfer. Customers who were impacted would have seen affected transfer requests return a 408; the underlying call stayed connected and only the transfer failed. Inbound and outbound calling were unaffected, as were calls that did not attempt a transfer, and no other region was affected.