8.5 Alarms & vCenter Appliance Health Monitoring
Key Takeaways
vCenter ships predefined alarm definitions, such as host connection state, host CPU and memory usage, datastore usage on disk, and certificate status, defined at the vCenter root so every object inherits them.
A custom alarm is created by choosing a target type, one or more rules whose trigger is a condition or an event, a Warning or Critical severity, and actions such as email, SNMP traps, or running a script.
Acknowledging a triggered alarm stops its actions from repeating while the alarm stays visible; Reset to Green clears it manually, which event-based alarms often need.
The vCenter Server Management Interface (port 5480) shows appliance health for CPU, memory, database, storage, and swap, plus disk usage per partition, such as /storage/seat and /storage/log.
Email alarm actions need vCenter's mail settings, and SNMP trap actions need SNMP receivers, both configured under vCenter > Configure > Settings > General.
8.5 Alarms & vCenter Appliance Health Monitoring
Predefined Alarms (7.10)
vCenter includes a large set of predefined alarm definitions, mostly defined on the vCenter Server root object, so they apply to every matching object below it. Examples:
| Predefined Alarm | Watches |
|---|---|
| Host connection and power state | Hosts that disconnect or become not responding |
| Host CPU usage / Host memory usage | Sustained high utilization |
| Virtual machine CPU usage / Virtual machine memory usage | Sustained high VM utilization |
| Datastore usage on disk | Datastore capacity: Warning at 75% and Critical at 85% by default |
| Cannot connect to storage | Loss of storage paths or connectivity |
| vSphere HA alarms | Failover in progress, host status, or insufficient failover resources |
| Certificate status | Certificates close to expiring (Section 6.5) |
| vSAN health alarms | Skyline Health findings for vSAN clusters |
| Host hardware alarms | Sensor-based health (fans, power supplies, temperature) |
Working with Triggered Alarms
- Triggered alarms appear on the object and in Monitor > Issues and Alarms > Triggered Alarms.
- Acknowledge: records that someone is handling the alarm and stops its actions from repeating. The alarm stays in the triggered list until the condition clears.
- Reset to Green: manually clears an alarm. You often need this for event-based alarms, which have no condition that clears on its own.
- Disable or edit a predefined definition where it is defined (Configure > Alarm Definitions on the vCenter object). It is usually better to create a custom alarm than to change defaults.
Creating a Custom Alarm (7.11)
Go to the object where the alarm should apply (for example a datacenter, cluster, or datastore folder), then open Configure > Alarm Definitions > Add. The wizard has these steps:
- Name and Targets: name the alarm and pick the target type, such as virtual machines, hosts, clusters, datastores, or networks. The alarm applies to objects of that type at or below where it is defined.
- Alarm Rule: define IF and THEN:
- IF: a condition or state (for example Datastore Disk Usage is above 90% for 5 minutes) or an event (for example Host connection lost, or an SSO event such as
com.vmware.sso.PrincipalManagement). - THEN: set the severity to Warning or Critical, and add actions:
- Send email notifications: requires vCenter's mail settings (SMTP server and sender) under vCenter > Configure > Settings > General.
- Send SNMP traps: requires SNMP receivers in the same settings.
- Run script: runs a command or script on the vCenter appliance.
- For VM or host targets, some rules also offer VM or host actions, such as power operations or entering maintenance mode.
- Choose how often actions repeat while the alarm stays triggered.
- Add more rules if needed. Each rule can have its own severity and actions.
- IF: a condition or state (for example Datastore Disk Usage is above 90% for 5 minutes) or an event (for example Host connection lost, or an SSO event such as
- Reset Rule: decide when the alarm returns to normal, for example when usage drops below the threshold, or on a clearing event.
- Review and Enable.
Tips:
- Define alarms as high in the inventory as makes sense, so new objects inherit them.
- Use a duration ("for 5 minutes") on metric conditions to avoid alerts from brief spikes.
- Creating alarms needs the Alarms privileges on the object.
Monitoring the vCenter Server Appliance (5.2)
vCenter is itself a workload to watch. The vCenter Server Management Interface (VAMI) at https://<vcenter-fqdn>:5480 shows:
| VAMI Area | What to Check |
|---|---|
| Summary > Health Status | Overall health plus CPU, Memory, Database, Storage, and Swap health, shown green, yellow, orange, or red |
| Monitor > CPU & Memory | Appliance utilization trends |
| Monitor > Disks | Usage per partition, such as /storage/log (logs), /storage/seat (statistics, events, alarms, tasks), /storage/db, and /storage/core |
| Monitor > Network | Throughput and errors |
| Monitor > Database | vPostgres usage, including growth of SEAT data |
| Services | Status of each vCenter service, with start, stop, and restart |
Common findings:
/storage/seatfilling up from high statistics levels or long retention. Lower the statistics level (Section 8.1), shorten retention, or grow the disk./storage/logfilling up after verbose logging. Revert the log levels and clean up.- Low memory warnings ("Appliance is running low on memory") on appliances sized too small, such as Tiny in production. Resize the appliance to the correct deployment size.
From the appliance shell, service-control --status --all lists the state of every vCenter service. In the vSphere Client, Monitor > Skyline Health runs health checks for vCenter and vSAN.
Scenario: Designing Useful Alarm Coverage
A new cluster goes live, and the operations team wants to know about problems without drowning in email. A sensible plan:
- Keep the predefined alarms, but check their thresholds. On large datastores, 75% and 85% may be too early, so a copy of the alarm with a different threshold on the right folder may be better than editing the default.
- Add custom alarms for business-specific risks, such as a datastore folder for a critical application, or VMs with snapshots that could fill storage.
- Route actions: send email for Critical alarms and SNMP traps to the monitoring platform. Set the repeat interval so the same alarm doesn't send new messages every few minutes.
- Use event-based alarms for security-relevant changes, for example SSO principal management events, and reset them to green after review.
- Watch vCenter itself: check the VAMI health, especially
/storage/seatand/storage/logusage, and the vCenter service status after every patch.
A datastore alarm keeps sending email every 5 minutes while the storage team expands the LUN. The team wants the emails to stop but the alarm to remain visible until the capacity problem is fixed. What should the administrator do?
Acknowledge the triggered alarm
Reset the alarm to green
Delete the predefined Datastore usage on disk definition
Disable Storage DRS
An administrator must be emailed when any datastore in the Finance folder exceeds 90% usage for more than 10 minutes. Which configuration meets this requirement?
Create an alarm on the Finance VM folder with Virtual Machines as the target and a CPU usage condition
Create an alarm on the Finance datastore folder with Datastores as the target, a Disk Usage above 90% for 10 minutes condition, and a Send email notifications action, with vCenter mail settings configured
Edit each VM's advanced settings to send SNMP traps when disk usage exceeds 90%
Enable Storage I/O Control with a 90% congestion threshold
The vCenter Server Management Interface shows the Storage health as red, and Monitor > Disks shows /storage/seat almost full after the statistics level was raised to Level 4. What is the most appropriate response?
Restart the vpxd service to clear the partition
Delete files from /storage/seat through the appliance shell
Lower the statistics level or retention and, if needed, increase the appliance disk that holds /storage/seat
Move the vCenter VM to a larger datastore using Storage vMotion
Sections you finish are checked off in the contents.
You've completed this section
Continue exploring other exams