Earlier today, Doozy experienced an outage that impacted end-user connectivity. The root cause was a major global incident at one of our third-party CDN providers. Below is a quick recap of what happened, what was affected, and what we are doing next.
Timeline & Impact
11:30 AM: We started seeing elevated errors as the provider’s global connectivity issues began.
1:00 PM: The issue escalated into a full outage, blocking inbound requests from reaching Doozy. Anyone trying to use the app inside Slack would have seen an exclamation mark showing the request failed.
2:40 PM: Service was fully restored.
What Actually Happened
The outage affected traffic coming into Doozy. Our backend stayed up the whole time. Scheduled content, automations, and outbound messages to Slack continued to run normally. The failure was isolated to the inbound network path controlled by the upstream provider.
Because the network failure also affected the infrastructure we use for our status page, that page was temporarily unavailable.
Prevention & Next Steps
We know this caused disruption, and we are moving quickly to make sure it does not happen again.
Building Network Redundancy:
We are implementing fallback paths so that if a CDN provider has an issue, it will not take down connectivity to Doozy. This will help us avoid single-provider points of failure.
Decoupling the Status Page:
We are restructuring the status page so it loads independently of our main infrastructure. Even if the core application is impacted, the status page will stay reachable and continue providing real-time updates.
Resolved
Earlier today, Doozy experienced an outage that impacted end-user connectivity. The root cause was a major global incident at one of our third-party CDN providers. Below is a quick recap of what happened, what was affected, and what we are doing next.
Timeline & Impact
11:30 AM: We started seeing elevated errors as the provider’s global connectivity issues began.
1:00 PM: The issue escalated into a full outage, blocking inbound requests from reaching Doozy. Anyone trying to use the app inside Slack would have seen an exclamation mark showing the request failed.
2:40 PM: Service was fully restored.
What Actually Happened
The outage affected traffic coming into Doozy. Our backend stayed up the whole time. Scheduled content, automations, and outbound messages to Slack continued to run normally. The failure was isolated to the inbound network path controlled by the upstream provider.
Because the network failure also affected the infrastructure we use for our status page, that page was temporarily unavailable.
Prevention & Next Steps
We know this caused disruption, and we are moving quickly to make sure it does not happen again.
Building Network Redundancy:
We are implementing fallback paths so that if a CDN provider has an issue, it will not take down connectivity to Doozy. This will help us avoid single-provider points of failure.
Decoupling the Status Page:
We are restructuring the status page so it loads independently of our main infrastructure. Even if the core application is impacted, the status page will stay reachable and continue providing real-time updates.
Monitoring
The Slack app and Web app have recovered from this outage. We're continuing to monitor the situation.
We'll be sharing a full update later today.
Monitoring
We're continuing to see intermittent issues connecting to Doozy. Our 3rd party CDN provider is rolling out a fix and we hope to see access to Doozy recovering shortly.
Monitoring
It looks like things are beginning to improve. We're continuing to monitor the situation closely.
Identified
There is a global outage impacting our 3rd party CDN provider. You may see elevated errors when trying to make requests to Doozy through the web and the Slack app.
We're monitoring the situation and looking to put workarounds in place.