Investigating reports of failed SIP Transfer calls in US East

Incident Report for LiveKit

Postmortem

Summary

On 17 August, between 17:24 and 18:05 UTC, a small percentage of SIP call transfers failed in our US East region during a routine configuration rollout. An issue in the automation that manages our SIP signaling servers prevented outgoing servers from being taken out of service safely, and those servers shut down while calls were still active on them. Transfers have completed normally since 18:05 UTC. Calls themselves stayed connected, and no other region was affected.

Root Cause

The configuration change restarts the servers that handle SIP signaling one at a time. Before a server is shut down, it is removed from service so that no new calls reach it, and it is then given some time to finish the calls it is already handling.

In this case that removal did not complete expectedly and the outgoing servers kept receiving new calls for the entire drain duration, then shut down on schedule with calls still active on them.

Timeline (UTC)

  • 16:52 - A routine configuration rollout begins in US East.
  • 17:24 - SIP transfer requests begin to fail for a subset of active calls.
  • 17:36 - Automated monitoring detects the elevated failure rate and our team begins investigating.
  • 17:40 - A further subset of active calls is affected.
  • 17:41 - Replacement capacity comes online.
  • 18:05 - Last of the impacted calls attempts a transfer and record a failure.

Scope of Impact

Only SIP call transfers (TransferSIPParticipant) in our US East region were affected, between 17:24 and 18:05 UTC. This represented 0.04% of all active calls in that window, and 1.1% of the calls that attempted a transfer. Customers who were impacted would have seen affected transfer requests return a 408; the underlying call stayed connected and only the transfer failed. Inbound and outbound calling were unaffected, as were calls that did not attempt a transfer, and no other region was affected.

Mitigations and Follow-ups

  • We have deployed an alert for calls that end unexpectedly when a server shuts down.
  • We are changing our rollout process so that a server which cannot be removed from service safely halts the rollout.
  • We are preventing new calls from being routed to servers that are shutting down.
  • We are improving monitoring of the automation that manages SIP server rotation.
  • We are returning a more specific error when a transfer request cannot be delivered.
Posted Sep 01, 2026 - 04:45 PDT

Resolved

We have identified the root cause of both spikes and have ensured safeguards to prevent a recurrence. We've been monitoring since 18:05 UTC and have seen no further transfer failures, and have confirmed that SIP transfers are operating normally. We will follow up with a detailed post-mortem.
Posted Aug 17, 2026 - 12:22 PDT

Update

No transfer failures has been observed since 18:05 UTC and SIP transfers are currently completing normally. Impact was limited to a small percentage of calls between 15:48-16:08 UTC and 17:24-18:05 UTC. We're actively monitoring while we investigate the cause and put safeguards in place to prevent a recurrence.
Posted Aug 17, 2026 - 11:22 PDT

Investigating

We're investigating an elevated rate of failures when transferring active SIP calls in the US East region. SIP calls themselves remain connected, and inbound and outbound calling are otherwise operating normally.
Posted Aug 17, 2026 - 11:10 PDT
This incident affected: Regional SIP (US East - SIP).