---
title: "Handling Power Outages and Network Isolation in a HA Cluster"
canonical: "https://safekit.eviden.com/best-practises/power-outage-and-network-isolation-in-a-cluster/"
description: "Learn how to manage network isolation and power outages in a high availability cluster. SafeKit proposes a SANless architecture with integrated split-brain checkers and automatic resynchronization to ensure data integrity and 24/7 uptime without a shared disk."
category: "best-practises"
lang: "en"
topics: "What are the different scenarios in case of network isolation in a cluster?, What are the different scenarios in case of power outage in a cluster?"
---

# Handling Power Outages and Network Isolation in a HA Cluster

## What are the different scenarios in case of network isolation in a cluster? {#isolation}

### A single network

When there is a network isolation, the default behavior is:

  * as heartbeats are lost for each node, each node goes to ALONE and runs the application with its virtual IP address (double execution of the application modifying its local data),
  * when the isolation is repaired, one ALONE node is forced to stop and to resynchronize its data from the other node,
  * at the end the cluster is PRIM-SECOND (or SECOND-PRIM according the duplicate virtual IP address detection made by Windows).

### Two networks with a dedicated replication network

When there is a network isolation, the behavior with a dedicated replication network is:

  * a dedicated replication network is implemented on a private network,
  * heartbeats on the production network are lost (isolated network),
  * heartbeats on the replication network are working (not isolated network),
  * the cluster stays in PRIM/SECOND state.

### A single network and a splitbrain checker

When there is a network isolation, the behavior with a split-brain checker is:

  * a split-brain checker has been configured with the IP address of a witness (typically a router),
  * the split-brain checker operates when a server goes from PRIM to ALONE or from SECOND to ALONE,
  * in case of network isolation, before going to ALONE, both nodes test the IP address,
  * the node which can access the IP address goes to ALONE, the other one goes to WAIT,
  * when the isolation is repaired, the WAIT node resynchronizes its data and becomes SECOND.

Note: If the witness is down or disconnected, both nodes go to WAIT and the application is no more running. That's why you must choose a robust witness like a router.

 

## What are the different scenarios in case of power outage in a cluster?

### Primary node power outage

When a power outage stops only the primary node:

  * there is an automatic failover on the secondary node, which becomes ALONE and restarts the application,
  * when node 1 is rebooted, it becomes SEDOND after resynchronization of replicated data,
  * the roles of primary and secondary can be swapped by an adminsitrator if needed.

### Secondary node power outage

When a power outage stops only the secondary node:

  * there is no failover, the primary becomes ALONE and the application continues its execution on node 1,
  * when node 2 is rebooted, it becomes SEDOND after resynchronization of replicated data.

### General power outage - case 1

When a power outage stops both nodes, the default behavior is:

  * both nodes goes to STOP,
  * when node 1 is rebooted, it does not go into ALONE state and does not restart the application because it doesn't know if it has the up-to-date data. So it goes to the WAIT state waiting for the restart of the other node,
  * when node 2 is rebooted, both nodes return to their previous PRIM/SECOND states.

### General power outage - case 2

When a power outage stops both nodes, the behavior with syncdelta is :

  * syncdelta is set for example to 10 minutes in the configuration meaning that start of a node is accepted even if its data is 10 minutes behind the last sync,
  * when node 1 is rebooted, it goes to ALONE and restarts the application assuming that the restart is done within 10 minutes after the power failure,
  * when node 2 is rebooted, it becomes SECOND after resynchronization of replicated data,
  * Note: if node 2 is rebooted the first, then it becomes ALONE and node 1 will become SECOND at its start.

 

