Search <book_title>...

NetBackup™ Deployment Guide for Kubernetes Clusters

Last Published: 2024-10-04

Product(s): NetBackup (10.5)

Elastic media server

All the replicas for the media server are always up and running which incurs unnecessary cost to customers. The basic media server pod power management (Elastic media server) feature provides Auto scaling of media server replicas based on the CPU and memory usage as well as the jobs queued due to maximum jobs per media server settings to reduce the cost.

Enabling/disabling the auto scaling feature

For enabling/disabling the auto scaling feature, following media server CR inputs are required:

replicas: Describes the maximum number of replicas that the media server can scale up to.
minimumReplicas: Describes the minimum number of replicas of the media server running. This is an optional field. If not specified, the value for minimumReplicas field will be set to the default value of 1.
From version 10.5, along with CPU and memory usage, the media server scaleout is also seen if the jobs are found in queued state due to the maximum job per media server settings. To configure this setting, refer to the configuration parameter bpsetconfig below:

To enable the elasticity of media server, the value of replicas must be more than value of minimumReplicas.

To disable the autoscaling feature of media server, ensure that the value of replicas is equal to the value of minimumReplicas.

Note:

The values of minimumReplicas and replicas must be greater than 0 to enable the elasticity of media server.

Status attributes of elastic media server CR

Following table describes the ElasticityAttributes that describes the attributes associated with the media server autoscaler. These attributes are only applicable if autoscaler is running.

Fields	Description
ExpectedReplicas	Indicates the ideal number of replicas computed by media server autoscaler that must be running. Default value is 0. It will be 0 if media server autoscaler is disabled. Note: If autoscaler is enabled (after that autscaler is tuned off) the value would be set to the value of minimumReplicas. It will be minimumReplicas even if media server autoscaler is disabled.
ActiveReplicas	Indicates the actual number of replicas that must be running to complete the ongoing operations on the media servers. Default value is 0. It will be 0 if media server autoscaler is disabled. Note: If autoscaler is enabled (after that autscaler is tuned off) the value would be set to the value of minimumReplicas. It will be minimumReplicas even if media server autoscaler is disabled.
NextIterationTime	Indicates the next iteration time of the media server autoscaler that is, the media server autoscaler will run after NextIterationTime only. Default value is empty.
NextCertificateRenewalTime	Next time to scale up all registered media servers for certificate renewal.

Configuration parameters

ConfigMap

A new ConfigMap with name nbu-media-autoscaler-configmap is created during deployment and the key-value pairs would be consumed for tuning the media server autoscaler. This ConfigMap is common to all the media server CR objects and supports the following keys:

Parameters	Description
memory-low-watermark-in-percent	Low watermark for memory usage.
memory-high-watermark-in-percent	High watermark for memory usage.
cpu-low-watermark-in-percent	Low watermark for CPU usage.
cpu-high-watermark-in-percent	High watermark for CPU usage.
scaling-interval-in-seconds	Interval after which media server autoscaler should run.
stabilitywindow-time-in-seconds	CPU and memory usage is calculated between two time intervals. This key indicates the time interval to be considered for collecting usage.
stability-count	CPU and memory usages are calculated by averaging out on multiple readings. This key indicates the number of readings to be considered.
graceful-shutdown-interval-in-seconds	The time interval after which the media server autoscaler should run incase it is not able to scale in due to running jobs on media server pods.
delayed-scalein-notifications-interval-in-minutes	The time interval between two successive notifications in the event that a scale in does not occur.

Note:

If you are upgrading to latest version, change the default values of the following parameters: scaling-interval-in-seconds : "90" stability-window-time-in-seconds : "10" stability-count : "3" graceful-shutdown-interval-in-seconds : "120" cpu-high-watermark-in-percent: "80"

bpsetconfig

A new entry has been added in the primary server bp.conf that is consumed by media server autoscaler. This value applies to all the Cloud Scale Technology managed media servers.

Parameters	Description
MAX_JOBS_PER_K8SCLUSTER_MEDIA_SERVER	Maximum number of jobs that can run on each media server. This value can be set using bpsetconfig CLI. Veritas NetBackup™ Commands Reference Guide

Parameters

Description

MAX_JOBS_PER_K8SCLUSTER_MEDIA_SERVER

Maximum number of jobs that can run on each media server. This value can be set using bpsetconfig CLI.

Veritas NetBackup™ Commands Reference Guide

Media server scaling

Parameters	Description
Scale-out	If all the active media servers managed by the Cloud Scale Technology are at their capacity due to the maximum jobs per media server settings and if there are more jobs in queue, scale-out is performed and multiple replicas may get scaled out due to the media server settings. Additionally, if there are no jobs in queue due to this settings and if the CPU or memory consumption in the specified value provided in configMap but if any existing media server is idle that is, no jobs are running on it, then scale-out will not be performed. If all the existing media servers which are ready have jobs running on them, media server autoscaler will scale out a media server pod.
Scale-in	If the CPU and memory consumption is below the specified values provided in configMap, media server autoscaler will scale in the media server pods. Ensure that the running jobs are completed. Note: The scale-in does not happen until there are jobs in the queue due to the maximum job per media server settings.
Note: The media server autoscaler scales out a single pod at a time in case a scale-out happens due to CPU and memory usage. It may exit from the multiple pods in case the scale-out happens due to the throttled jobs. The media server autoscaler can scale-in multiple pods at a time.

Parameters

Description

Scale-out

If all the active media servers managed by the Cloud Scale Technology are at their capacity due to the maximum jobs per media server settings and if there are more jobs in queue, scale-out is performed and multiple replicas may get scaled out due to the media server settings.

Additionally, if there are no jobs in queue due to this settings and if the CPU or memory consumption in the specified value provided in configMap but if any existing media server is idle that is, no jobs are running on it, then scale-out will not be performed. If all the existing media servers which are ready have jobs running on them, media server autoscaler will scale out a media server pod.

Scale-in

If the CPU and memory consumption is below the specified values provided in configMap, media server autoscaler will scale in the media server pods. Ensure that the running jobs are completed.

Note:

The scale-in does not happen until there are jobs in the queue due to the maximum job per media server settings.

Note:

The media server autoscaler scales out a single pod at a time in case a scale-out happens due to CPU and memory usage. It may exit from the multiple pods in case the scale-out happens due to the throttled jobs. The media server autoscaler can scale-in multiple pods at a time.

Note:

If the scale-in does not happen due to background processes running on the media server, a notification would be sent on NetBackup Web UI after regular time interval as configured in the autoscaler ConfigMap. For more details, see the following section:

The time taken for media server scale depends on the value of scaling-interval-in-seconds configuration parameter. During this interval, the jobs would be served by existing media server replicas based on NetBackup throttling parameters. For example, Maximum concurrent jobs in storage unit, Number of jobs per client, and so on.

Cluster's native autoscaler takes some time as per scale-down-unneeded-time attribute, which decides on the time a node should be unneeded before it is eligible to be scaled down. By default this is 10 minutes. To change this parameter, edit the cluster-autoscaler's current deployment settings using the following commands and then edit the existing value:

AKS: az aks update --resource-group $RESOURCE_GROUP_NAME --name $CLUSTER_NAME --cluster-autoscaler-profile scale-down-unneeded-time=5m
EKS: kubectl -n kube-system edit deployment cluster-autoscaler

Note the following:

For scaled in media servers, certain resources and configurations are retained to avoid reconfiguration during subsequent scale out.
- Kubernetes services, persistent volume claims and persistent volumes are not deleted for scaled in media servers.
- Host entries for scaled in media servers are not removed from NetBackup primary server. Hence scaled in media server entries will be displayed on Web UI / API.
For scaled down media servers, the deleted media servers are also displayed on Web UI/API during the credential validation for database servers.

Certificate renewal scheduling

In order to ensure that the certificates are renewed on the media server replicas for which the pods were scaled in by media server autoscaler for longer duration, the NetBackup operator would scale out all the media server replicas once in a month for which certificate was issued. The scaled-out media server replicas would be then scaled in as per media server autoscaler.

Handling of sudden incoming jobs

Based on the configured schedules, if a large number of jobs are expected to run at certain time, the maximum number of jobs per media server should be configured to ensure that the required number of media server pods are scaled out and the jobs are properly distributed.

For configuration, please refer - <Link to the new configuration parameter>

More Information

Troubleshooting AKS and EKS issues