Polaris DC01 - RCA – July 19, 2026

Introduction

This document serves as a Root Cause Analysis for the service interruption experienced by Polaris customers–hosted in DC01 (Seattle Data Center).

The goal of this document is to share our findings regarding the event, specify the root cause analysis, outline actions to be taken to solve the downtime event, as well as preventive measures Clarivate is taking to avoid similar cases in future.

Event Timeline

Service interruption was experienced across multiple Polaris environments hosted in DC01 during the following timeframe:

July 19, 2026, from 01:25 AM until 11:14 AM PDT.

During the event, impacted customers were unable to access their Polaris environments until recovery and migration activities were completed and services were validated. The impact was limited to workloads on a single affected server and did not constitute a data center-wide outage.

Root Cause Analysis

Clarivate Engineers investigated this event to determine the root cause analysis with the following results:

Our engineering teams identified that a critical hardware component within one of the physical servers hosting these environments had failed, rendering that server inaccessible and unable to support the workloads running on it. In addition, we recognize that a gap in our escalation workflow and procedures meant the issue was not addressed as swiftly as our standards require, which extended the time to full recovery.

We take full ownership of this and have already put measures in place to ensure a faster, more decisive response going forward. To restore service, our engineers migrated all affected workloads from the failed server and redistributed them across multiple healthy servers; once migration was complete, all monitoring alerts cleared and Polaris services were fully restored and validated.

Technical Action Items and Preventive Measures

Clarivate has taken the following actions and preventive measures to avoid such an occurrence in future:

  • We are strengthening our monitoring and on-call escalation so that any single-host disruption immediately and automatically triggers engineering engagement, enabling faster assessment and response.

  • We are formalizing a comprehensive host-failure recovery runbook with clearly defined team responsibilities and escalation paths to shorten the path from detection to restoration.

  • We are improving recovery automation to accelerate the migration of affected workloads onto healthy infrastructure and speed the return to service.

  • We are enhancing the resilience of our hosting architecture by distributing customer workloads more evenly across our infrastructure, reducing the potential impact of any single hardware failure.

Customer Communication

Innovative is committed to providing customers with prompt and ongoing updates during Cloud events. Ongoing and prompt updates on service interruptions appear in the system status portal at this address:http://status.exlibrisgroup.com/

These updates are automatically sent as emails to registered customers