14.1 Orchestrator Monitoring, Trigger Alerts & Job Health
Key Takeaways
Monitoring covers Machines, Processes, Queues, and Queues SLA with general and individual views, for up to the last 30 days in the real-time tab.
Enhanced monitoring and the Historical tab (up to two years, refreshed every 20 minutes) are powered by Insights.
Each monitoring page needs View on Monitoring plus View on the component, such as Queues or Machines.
Trigger settings can stop or kill stuck jobs, alert on jobs stuck Pending or Resumed or running too long, and disable a trigger after consecutive failures.
Combine monitoring with job logs, recordings, and webhooks for a complete view of production health.
14.1 Orchestrator Monitoring, Trigger Alerts & Job Health
Core Concept: The exam description lists "Orchestrator monitoring". Orchestrator's Monitoring pages show the health of machines, processes, queues, and queue SLAs for up to the last 30 days, while trigger and job settings raise alerts and end stuck jobs automatically.
The Monitoring Pages
The Monitoring section has four components:
| Component | What it shows |
|---|---|
| Machines | Status of machines and their runtimes |
| Processes | Jobs across statuses, errors, and durations |
| Queues | Queue item statuses, processing times, and errors |
| Queues SLA | Items at risk of missing their SLA, based on SLA predictions |
Each component has two views:
- General view: all resources of that type together, for example every queue in the folder.
- Individual view: one resource, opened by selecting it in the general view.
Filters include Interval, which shows the last day or the last hour, and Include Subfolders, which is No by default.
Real-time and historical tabs
- Real time shows continuously updated information for up to one month. The Enhanced monitoring toggle switches to Insights-powered real-time views for machines, processes, and queues, with more filters.
- Historical is powered by the Insights service, shows data up to two years back, refreshes every 20 minutes, and is available only when Insights is enabled. Its dashboards can be filtered, downloaded, and scheduled for email delivery.
Permissions
Monitoring needs two grants per component: View on Monitoring plus View on the component.
| Permissions | Page shown |
|---|---|
| Monitoring View + Machines View | Monitoring > Machines |
| Monitoring View + Queues View | Monitoring > Queues |
| Monitoring View + Jobs View | Monitoring > Jobs (processes) |
| Monitoring Edit + Queues View or Jobs View | Allows dismissing errors from the Error Feed widget |
| Folders View | Filtering by subfolders |
Without both View permissions, the page is not displayed at all.
Alerts on Triggers and Jobs
Triggers (time, queue, event, and API) offer settings that keep jobs from silently hanging:
| Setting | Effect |
|---|---|
| Schedule ending of job execution | Stop (graceful) or Kill (forced) a job stuck in Pending or Running after a set time, from 1 minute up to 10 days, 23 hours, and 59 minutes. With Stop plus automatic Kill, the kill timer starts only after the job is in Stopping. |
| Generate an alert if the job is stuck (in pending or resumed status) | Raises an Error-severity alert when jobs stay Pending or Resumed longer than the set duration (1 minute to 11 days) |
| Generate an alert if the job started and has not completed | Raises an Error-severity alert when a running job exceeds the set duration |
| Set execution-based trigger disabling | Disables the trigger after a number of consecutive failed executions (0–100, 0 means never), with an optional grace period of 0–30 days; stopped jobs are not counted |
| Schedule automatic trigger disabling | Disables the trigger at a set date and time |
The schedule-ending timer counts from the scheduled time even if the job waited in the queue. A job scheduled for 1:00 PM with "stop after 20 minutes" stops at 1:20 PM even if it only started at 1:15 PM.
Users choose which alerts they receive, in the product or by email, through their notification settings in Automation Cloud. The Alerts permission governs access to alerts at the tenant level.
Other Monitoring Signals
- Job details and logs: the Jobs page shows state, source (manual, time trigger, queue trigger, API trigger, Assistant, and more), host identity, and logs. Faulted jobs show the error.
- Healing Agent column: shows Issues Detected when Healing Agent found problems during execution.
- Recordings: faulted unattended jobs can include screenshots or video when recording is enabled.
- Webhooks: send job, queue, robot, and trigger events to external monitoring tools in real time.
- Queue item details: each item's history, exceptions, and output.
A Monitoring Routine for Production
- Daily: open Monitoring > Processes for the production folder and review faulted jobs in the Error Feed.
- Queues: check Monitoring > Queues for rising application exceptions, which often mean a target system changed.
- SLA: review Queues SLA for items at risk and add robots or change priorities before deadlines pass.
- Machines: confirm that machines are connected and runtimes are not all busy, which would leave jobs Pending.
- Triggers: make sure critical triggers have stuck-job alerts and schedule-ending rules.
Scenario
A queue-triggered process sometimes hangs on a modal dialog and stays Running all night, blocking the only runtime. Fixes:
- Set Generate an alert if the job started and has not completed to 45 minutes, so operators hear about the hang.
- Set Schedule ending of job execution to Stop, then Kill if the job does not stop, so the runtime is freed.
- In the workflow, use Should Stop in the transaction loop so a Stop request ends the job cleanly.
- Enable recording for failed jobs to capture the dialog that caused the hang.
A support analyst has View on Queues but still cannot see Monitoring > Queues. What is missing?
Edit on Queues.
View on the Monitoring permission set.
Create on Execution Media.
The Orchestrator Administrator tenant role.
A time trigger is set to stop jobs 20 minutes after their scheduled time. A job scheduled for 1:00 PM waits in Pending until 1:15 PM before it starts. When is it stopped?
At 1:35 PM, 20 minutes after it started.
Never, because the timer only applies to Running jobs.
At 1:20 PM, because the time counts from the schedule even while the job is queued.
At 2:00 PM, at the next trigger run.
Which trigger setting raises an Error-severity alert when jobs remain Pending or Resumed longer than expected?
Set execution-based trigger disabling.
Schedule automatic trigger disabling.
Keep Account/Machine allocation on job resumption.
Generate an alert if the job is stuck (in pending or resumed status).
Sections you finish are checked off in the contents.