Security Patch

Scheduled Maintenance Report for Apwide

Postmortem

Summary

On 2026-07-19, our production services experienced an extended outage following a planned database security maintenance announced by OVH. While the maintenance was initially expected to last no more than 30 minutes, an infrastructure failure at the OVH datacenter significantly prolonged the interruption.

After assessing the situation and the available recovery options, we activated our Business Continuity Plan (BCP) by restoring the production database from the latest available backup and upgrading it to PostgreSQL 18. All services were fully restored on 2026-07-20 at 11:15 UTC.

Customer Impact

During the incident, Golive and Time Squad were unavailable for approximately 20 hours, preventing customers from accessing the applications and performing normal operations.

As part of the recovery process, we restored the database from the latest available backup. The most recent backup had been completed 25 minutes before the outage began, resulting in a potential data loss window of up to 25 minutes.

A review of the database audit logs showed that, due to the weekend period, only a limited number of write operations occurred during this timeframe. These updates were primarily related to deployment and monitoring activities, significantly reducing the effective customer impact.

Root Cause

The root cause of the incident was a hardware failure affecting a storage device in the OVH datacenter during a planned security maintenance operation on the managed PostgreSQL infrastructure operated by Aiven on behalf of OVH.

Initially, the incident was communicated as an extension of the planned maintenance, delaying visibility into the actual nature and severity of the problem. Several hours later, OVH published a dedicated incident confirming that the outage was caused by a hardware failure rather than the maintenance activity itself.

Incident Timeline

Context

The incident occurred during a planned maintenance window announced by OVH on 2026-07-16. OVH informed customers that they would apply a security patch to our production database over the weekend. The expected service interruption was no longer than 30 minutes.

2026-07-19

14:39 UTC
Our monitoring systems and availability probes detected an outage affecting our production services. As this matched the scheduled maintenance window, our support engineers considered the interruption expected.

16:00 UTC
After allowing an additional one-hour buffer beyond the planned maintenance window, our monitoring continued to report that the database was unavailable. Our support engineers checked the OVH status page, which still showed the maintenance as ongoing with no additional information. The OVH control panel also reported the database status as "Updating". At this stage, the extended downtime was still considered part of the maintenance.

17:15 UTC
The OVH status page remained unchanged. To gather additional information, our engineers consulted the public OVH Discord community, where other customers reported experiencing the same issue: databases remained in the "Updating" state while waiting for the security patch to complete.

21:20 UTC
With the announced maintenance window approaching its end, the database still unavailable, and business hours about to begin in the Australia/Oceania region, our support engineers opened a Priority 1 (P1) support ticket with OVH.

21:25 UTC
OVH acknowledged receipt of the incident and confirmed that the ticket had been assigned.

21:41 UTC
OVH confirmed an incident involving Aiven, the service responsible for managing the cloud database platform. They indicated that the issue was related to the planned maintenance operation and that the outage was taking significantly longer than the initially announced 30 minutes. They committed to providing further updates.
https://public-cloud.status-ovhcloud.com/incidents/f02sjn0qw27f

At this point, we updated the Golive and Time Squad status pages to report a service outage.

2026-07-20

00:46 UTC
After more than ten hours of downtime without any meaningful update, our support engineers requested additional information from OVH to determine whether our Business Continuity Plan (BCP) should be activated. During this investigation, we also discovered that the OVH user interface did not allow us to create a database fork from an existing backup.

04:55 UTC
As no progress had been communicated publicly, our support engineers contacted OVH support again and also reached out to the on-call support team.

05:30 UTC
OVH support confirmed that the planned maintenance was still in progress but could not provide further details regarding the broader incident. We expressed our concerns about the lack of communication and requested additional information to support our operational decision-making. We also requested that OVH create a dedicated public incident rather than continuing to associate the outage solely with the maintenance operation.

07:20 UTC
Our engineers discovered that creating a database fork from a backup using the OVH REST API was possible, even though the option was unavailable in the web interface.

The most recent backup had been created at 14:15 UTC, while the outage started at 14:40 UTC, resulting in a potential data loss window of approximately 25 minutes.

After reviewing the database audit logs, our engineers confirmed that only a limited number of write operations had occurred during this period, primarily related to deployment and status monitoring activities due to the weekend. Based on this assessment, management decided to activate the Business Continuity Plan.

07:48 UTC
OVH published a dedicated incident, confirming that the outage was caused by a hardware failure rather than simply an extended maintenance window.
https://public-cloud.status-ovhcloud.com/incidents/bn54j3232tlv

08:00 UTC
Because our PostgreSQL version did not support database forking through the OVH web interface, we took advantage of the BCP activation to upgrade the database to PostgreSQL 18, a version that had already been successfully validated in our development and staging environments.

08:50 UTC
The restored database became available on PostgreSQL 18. Database integrity and sanity checks completed successfully, and Time Squad was restarted.

08:55 UTC
Time Squad was fully operational. Our status page was updated, and Golive services were restarted.

09:05 UTC
Despite several restart attempts, some Golive components failed to start correctly. Our engineers began investigating the root cause.

09:30 UTC
The investigation identified the issue as a backlog of asynchronous Atlassian webhook events (issue creation, updates, etc.) that had accumulated during the outage. Once connectivity was restored, Atlassian attempted to replay all missed events simultaneously, creating a significant load on our infrastructure before caches had been rebuilt.

09:55 UTC
A temporary mitigation was implemented by blocking incoming webhook traffic. This allowed Golive to complete its startup sequence and resume serving normal transactional traffic. The status page was updated accordingly.

10:13 UTC
A performance improvement that had already been validated in our development and staging environments and scheduled for an upcoming release was deployed to production. Webhook traffic was then re-enabled, and the combination of rebuilt caches and the performance improvement proved effective.

11:15 UTC
After more than one hour of stable event processing and normal system behavior, the incident was declared resolved. The status page was updated to indicate that all systems were fully operational.

Actions Taken

During the incident, we initially could not activate our Business Continuity Plan because the OVH web interface did not allow us to create a database fork from a backup. This limitation was caused by the legacy PostgreSQL version running on our production environment.

As part of the recovery process, we upgraded our production database to PostgreSQL 18, the latest version supported by OVH and already validated in our non-production environments. This upgrade removes this limitation and enables faster recovery should a similar incident occur in the future.

Follow-up Actions

  • Work with OVH to improve communication during major incidents, particularly regarding impact assessment and incident classification.
  • Request earlier publication of dedicated incident reports instead of extending planned maintenance notifications when an unexpected failure occurs.
  • Continue reviewing and improving our Business Continuity Plan activation criteria to enable faster decision-making when infrastructure providers experience prolonged outages.
  • Follow up with OVH on the restoration of the original database service, which remains unavailable at the time of writing this post-mortem.
Posted Jul 20, 2026 - 15:27 CEST

Completed

We have deployed a patch to reduce the number of asynchronous events that need to be processed by our system.

Combined with the application caches now being fully warmed up, this has resulted in stable system performance for more than an hour. Our systems are fully operational and handling the current load without issue.

We will continue to monitor the situation closely, but the incident can now be considered resolved.

A detailed post-mortem will be shared to provide a complete timeline of the incident, its root cause, and the actions taken to resolve it.

Thank you for your understanding and patience throughout this incident.
Posted Jul 20, 2026 - 13:42 CEST

Update

Our hosting provider has published a public incident report regarding the root cause of the database failures affecting Golive and Time Squad:
https://public-cloud.status-ovhcloud.com/incidents/bn54j3232tlv

As the resolution on his side side has taken longer than expected, we activated our business continuity plan and restored a new database from one of our backups.

Time Squad is now fully operational.

The main Golive features are also back online and operational. However, notifications, automations, and scheduling conflict checks are currently experiencing performance issues due to the length of the outage and the recovery of asynchronous events. We are actively working to fully resolve these issues, but you may still experience occasional disruptions over the coming hours.

Once again, we sincerely apologize for the impact of this incident. Although the root cause was outside of our control, we took every possible step to minimize its impact by implementing our business continuity plan.
Posted Jul 20, 2026 - 12:11 CEST

Update

The incident is still ongoing, and our hosting provider is actively working to resolve it.

We will provide another update as soon as we have additional information.

We sincerely apologize for the extended service disruption and appreciate your patience.
Posted Jul 20, 2026 - 07:46 CEST

Update

Following an incident during a maintenance operation performed by our hosting provider, our database service is currently unavailable.

We are in contact with them, and they are actively working to resolve the issue.

We will keep you updated as soon as we have more information.

We apologize for the inconvenience this has caused.
Posted Jul 20, 2026 - 00:25 CEST

In progress

Scheduled maintenance is currently in progress. We will provide updates as necessary.
Posted Jul 17, 2026 - 20:30 CEST

Scheduled

Our hosting provider is currently rolling out a security patch across several infrastructure components.

This weekend, they will apply the patch to the machines hosting our production database.

The scheduled maintenance window is from 2026-07-17 18:30 UTC to 2026-07-20 07:00 UTC.

During this maintenance period, an outage of up to 30 minutes may occur.

We apologize for any inconvenience this may cause. This maintenance is being carried out by our hosting provider and is beyond our control.

Thank you for your understanding.
Posted Jul 17, 2026 - 15:00 CEST
This scheduled maintenance affected: Golive Cloud (Golive Cloud - App, Golive Cloud - API, Golive Cloud - Email Notifications, Golive Cloud - Automations & Webhooks) and Time Squad Cloud.