H-CLAIR: Physical Safety Under a Fully Compromised Automation Plane

Research Project 2026
Md Nazmul Kabir Sikder

Overview

Industrial control systems increasingly use learned anomaly detectors to gate automated controllers: normal-looking telemetry goes to the primary controller, and an alarm hands control to a fallback. This project asks whether that design actually keeps the plant safe. In a closed-loop water-distribution network, the detector catches every ordinary fault, never fires on a clean day, and defers on almost every attacked sample, yet most attacked runs still drive tanks outside their safe operating range.

The gap is structural rather than a matter of detector calibration. Anomaly detection asks whether telemetry looks abnormal; physical safety depends on whether the command about to reach a pump is safe for the current plant state. A correct deferral to "close the pumps" can prevent overflow but cause underflow when the network is already draining, and a compromised controller can skip the detector entirely and issue a harmful command directly.

Approach

H-CLAIR moves the safety decision to actuation time. It treats the whole automation plane, including telemetry processing, the detector, the learned controller, orchestration logic, and the fallback, as fully compromised, and keeps a small, separate physical-safety trusted computing base:

  • Authenticated state. Plant state comes from a separate, minimal sensor path. Each record carries an endpoint identity, sequence number, timestamp, and HMAC tag, and only records that verify and are fresh may update the shield's state.
  • Set-valued tracking. The shield maintains the set of plant states consistent with authenticated readings and the commands that actually reached the actuator, never the commands that compromised software requested.
  • Complete mediation. Every request, from the primary branch or the fallback, is checked against the set of commands whose successors stay inside a certified safe region for every possible state and disturbance. Safe requests pass unchanged; unsafe ones are replaced with a certified restoring action.
  • Explicit assurance loss. When trusted observations are delayed or dropped, the state set widens. If no command remains safe for every consistent state, H-CLAIR reports loss of assurance and fails closed instead of silently reusing stale state.

Guarantees

  • Request-channel noninterference. Changing telemetry, the detector decision, the agent, or the fallback cannot enlarge the set of commands the shield will admit.
  • Invariant preservation. Under a truthful authenticated endpoint, a sound reachability model, a declared disturbance envelope, and complete mediation, the plant stays inside its safe region for every adaptive sequence of requests.
  • Assurance horizon. A state- and model-dependent bound on how long the plant can run without fresh trusted state before no command is safe for every consistent state, replacing a fixed freshness timeout.

Detector replacement does not require recertifying the safety logic; changes to the reachability model, trusted endpoint, timing envelope, or actuator proxy do.

Evaluation

  • Detection does not enforce safety: adaptive white-box evasion on SWaT, and closed-loop consequence studies on the BATADAL C-Town water-distribution network.
  • Containment under complete request compromise: an end-to-end Modbus/TCP tank loop where the unguarded controller leaves the safe band on up to half of its cycles while the shield records no excursions, plus a held-out nonlinear plant whose dynamics the shield never sees.
  • Authenticated channel: forged, modified, and replayed frames are all rejected, and the assurance horizon is validated under delay, burst loss, and total denial of trusted observations.
  • Cost: the verifier, set propagation, and guard run in microseconds; the price of stale state appears as a higher intervention rate rather than as safety violations.