9.3 Performance Monitoring and Management of Change
Key Takeaways
- After a recent calibration, PV-versus-gauge disagreement is more often a plugged impulse line, drained wet leg, or isolation valve than a CPU bug.
- A sticky valve shows controller output moving while position feedback and PV lag or jump; that is not solved by raising PID gain first.
- A setpoint or output clamp with a documented process basis (for example NPSH) is management of change, not an HMI preference.
- Interlock bypasses need time limits, compensating measures, and a log; alarm-setpoint changes that delay hazard recognition are MOC.
- Controller firmware, even a 'small' patch, is MOC: backup, regression test, window, and confirmation that SIS stays segregated.
SAT closes a contract. The plant then starts teaching you what FAT could not. Knowledge area 2.F continues into performance monitoring, diagnostics, and management of change (MOC) after startup. The PE exam tests whether you will retune a healthy controller, rewrite firmware, or walk to the impulse line.
Performance monitoring and diagnostics
Once the unit is running, watch the system as a machine with a health trend, not only as a collection of faceplates.
- Loop KPIs: time in saturation, oscillation, valve travel, and error relative to a simple benchmark. A loop that hunted after a rate change may be a process-gain change (fouling, level, composition)—tuning drift—not a failed CPU.
- Valve signatures: plot controller output against position feedback and PV. A sticky valve holds, then jumps; output may wind up. Raising gain on a sticky valve makes the jump worse.
- Instrument diagnostics: HART or fieldbus status, impulse-line plugging signatures, thermocouple burnout bits. These are often more honest than an operator's 'the DCS is wrong.'
- Infrastructure: CPU loading versus the FAT baseline, network error counters, UPS battery tests, cabinet temperature, alarm flood rate. A creeping CPU load after a historian change is a capacity problem returning from the previous section.
'The DCS is wrong' is a diagnosis of last resort. After a recent successful calibration, a control-room PV that disagrees with a local gauge is more often a plugged impulse line, a drained or overfilled wet leg, a partly closed isolation valve, a heat-trace failure, or a leaking fill fluid than a processor that suddenly forgot engineering units. The field check is cheap. Re-imaging firmware is not. The same humility applies to analytical instruments: a sample-system lag looks like a slow controller.
Sticky valves get blamed on the DCS because the faceplate still 'works'—you can move setpoint and the output number changes. If position feedback barely moves until a large kick, then jumps, the stem, packing, positioner, or air supply is the process, not the scan rate. A bump test that moves output while PV and position lag is the diagnostic. Replacing the CPU first is how you spend the shutdown on the wrong object.
What belongs in MOC
MOC is not paperwork for its own sake. It is the review that asks whether a change can affect safety, environment, product, or the operator's ability to see a trip. After startup, the changes that look small on a screen are the ones that fail this test.
- Graphics: moving a trip indicator off the overview, inverting running/stopped colors, or adding a navigation path that buries the ESD status is an HMI-standard issue and usually MOC, because it changes human response time. Cosmetic tag-font changes with no status meaning are lighter—but still documented so two engineers do not 'fix' the same screen differently.
- Alarm setpoints: if the new setpoint delays recognition of a hazard or floods the operator, it is MOC and should pass through alarm rationalization, not a night-shift handshake. A nuisance analog chatter filter on a healthy transmitter is a different review than raising a high-high that is the last warning before a relief valve.
- Interlock bypasses: formal MOC, a time limit, compensating measures (person at the pump, extra patrol, SIS still in force if this is only a DCS perm), and a log. Permanent bypasses are design changes, not bypasses.
- Controller firmware: scan-time changes, communication stacks, and vendor patches. Treat firmware as MOC even when the bulletin is cyber-related: backup, restore test, offline regression of the affected sequences, a maintenance window, and confirmation that the SIS is not using the same engineering laptop as the BPCS without a procedure. 'It is only a patch' is how two redundant CPUs get different builds.
Tuning a PID with the same strategy (gain, reset, derivative, same PV, same valve) is usually a controls-engineer plus operations awareness item with a log—not a full PHA—unless the new tuning is being used to hide a sticky valve or to fight a constraint that should be an interlock. Changing strategy (adding override, changing fail direction, sharing a measurement with SIS) is MOC.
Worked example — operator asks to raise a setpoint clamp
A charge-pump flow controller has a setpoint high clamp at 80%. Operations wants 90% to make rate. This is the exam's favorite 'is it tuning or MOC?' fork. Do not start by moving the clamp. Recover the basis.
If the control narrative, pump curve, or PHA documents 80% because of NPSH margin, downstream hydraulic limit, or vessel overflow, the clamp is a process/equipment limit wearing an HMI shirt. Raising it is MOC: check the equipment, check whether SIS uses the same measurement, update the narrative and graphics, train the console, and do not discover the limit by cavitating the pump. If the 80% was an arbitrary start-up conservative number, and the P&ID and narrative already allow 100% with no safety or equipment basis, you still document the change and confirm the basis in writing. That confirmation is the review. Skipping it because 'it is only a clamp' is how unofficial interlocks disappear.
If the clamp is doing the job of an interlock—the only thing preventing a high-high—the correct path is not a quiet clamp edit. The correct path is MOC toward a real independent protection layer, or a documented operating limit with alarms that match the philosophy. Tuning the PID will not replace that discussion.
| Change type | Required review |
|---|---|
| PID gain/reset, same strategy and fail direction | Controls engineer plus operations awareness; log the before/after; not usually a full PHA |
| Alarm setpoint that changes time to recognize a hazard | Alarm rationalization and MOC |
| Graphic that can hide a trip or invert run/stop status | HMI standard review under MOC |
| Interlock bypass | Formal MOC, time limit, compensating measures, bypass log |
| Setpoint or output clamp with a process/equipment basis | MOC against the documented limit; check SIS sharing |
| Controller firmware or communication stack | MOC: backup/restore, regression, window, BPCS/SIS segregation |
| I/O range or engineering-unit change | MOC: scaling, alarms, any SIS or interlock using the same measurement |
| Night-shift 'temporary' force left in place | Treat as an undocumented bypass—clear or convert to MOC |
Exam traps: raising PID gain on a sticky valve; re-imaging a controller because a wet leg drained; treating a documented NPSH clamp as a faceplate preference; patching firmware on both redundant CPUs without a restore point.
An operator reports that the DCS is wrong because the control-room PV is 2 psi below a local gauge and the loop is cycling. Calibration last week was successful. The best first field hypothesis is:
The control narrative documents an 80% setpoint high clamp on a charge-pump flow loop because of NPSH margin at high rate. Operations asks to raise the clamp to 90% to make rate. This request is:
A flow loop's PV is sluggish and the operator says the controller is dead, but the output faceplate moves when you change setpoint. Position feedback barely moves until a large kick, then jumps. This pattern most strongly indicates: