Search <book_title>...

Cluster Server 7.4.1 Administrator's Guide - Linux

Last Published: 2019-10-17

Product(s): InfoScale & Storage Foundation (7.4.1)

Platform: Linux

Section I. Clustering concepts and terminology
Section II. Administration - Putting VCS to work
Section III. VCS communication and operations
Section IV. Administration - Beyond the basics
Section V. Veritas High Availability Configuration wizard
1. Introducing the Veritas High Availability Configuration wizard
2. Administering application monitoring from the Veritas High Availability view
  1. Administering application monitoring from the Veritas High Availability view
  2. Administering application monitoring settings
Section VI. Cluster configurations for disaster recovery
Section VII. Troubleshooting and performance
1. VCS performance considerations
2. Troubleshooting and recovery for VCS
Section VIII. Appendixes

VCS behavior when a resource fails to come online

In the following example, the agent framework invokes the Online function for an offline resource. The resource state changes to WAITING TO ONLINE.

VCS goes through the following steps when a resource fails to come online:

If the Online function times out, VCS examines the value of the ManageFaults attribute first at the resource level and then at the service group level. If ManageFaults is defined at the resource level, VCS overrides the corresponding values that are specified at the service group level. VCS takes action based on the ManageFaults values that are specified at the resource level.
If ManageFaults is set to IGNORE at the resource level, the resource state changes to OFFLINE|ADMIN_WAIT. The resource-level value overrides the service group-level value.
If ManageFaults is set to ACT at the resource level, VCS calls the Clean function with the CleanReason set to Online Hung.
If resource-level ManageFaults is set to "" or blank, VCS checks the corresponding service group-level value, and proceeds as follows:
- If ManageFaults is set to NONE, the resource state changes to OFFLINE|ADMIN_WAIT.
- If ManageFaults is set to ALL, VCS calls the Clean function with the CleanReason set to Online Hung.
If ManageFaults is set to NONE, the resource state changes to OFFLINE|ADMIN_WAIT.
If ManageFaults is set to ALL, VCS calls the Clean function with the CleanReason set to Online Hung.
If the Online function does not time out, VCS invokes the Monitor function. The Monitor routine returns an exit code of 110 if the resource is online. Otherwise, the Monitor routine returns an exit code of 100.
VCS examines the value of the OnlineWaitLimit (OWL) attribute. This attribute defines how many monitor cycles can return an offline status before the agent framework declares the resource faulted. Each successive Monitor cycle increments the OnlineWaitCount (OWC) attribute. When OWL= OWC (or if OWL= 0), VCS determines the resource has faulted.
VCS then examines the value of the ManageFaults attribute. If the ManageFaults is set to NONE, the resource state changes to OFFLINE|ADMIN_WAIT.
If the ManageFaults is set to ALL, VCS calls the Clean function with the CleanReason set to Online Ineffective.
If the Clean function is not successful (exit code = 1), the agent monitors the resource. It determines the resource is offline, and calls the Clean function with the Clean Reason set to Online Ineffective. This cycle continues till the Clean function is successful, after which VCS resets the OnlineWaitCount value.
If the OnlineRetryLimit (ORL) is set to a non-zero value, VCS increments the OnlineRetryCount (ORC) and invokes the Online function. This starts the cycle all over again. If ORL = ORC, or if ORL = 0, VCS assumes that the Online operation has failed and declares the resource as faulted.