Security for a plant that doesn’t have a security team.
A few staff. A tight budget. A contracted operator and an integrator you call when something breaks. Most industrial-security guidance is written for a company with a SOC, an Active Directory forest and a DMZ — and then it is handed to a water system serving four thousand people.
This is the version for the second one. It is free, it is open, and the parts that don’t apply to you are labelled instead of implied.
Read this before the phases. This is a mitigation plan, not a product and not a conformance claim. Real rollout depends on your specific device inventory and each device’s capability, and this page cannot know either.
The single most important sequencing fact: the cryptographic fix — per-device certificates — depends on the device supporting it, and a lot of legacy equipment does not. So this is built so that you get most of the risk reduction from Phases 0–2 regardless of your hardware, and reach the real fix in Phase 3 only if the equipment allows. If it doesn’t, you are not left with “nothing until you buy new PLCs.”
Five phases, cheapest and highest-impact first
Phases 0–2 are hardware-independent. Phase 3 is hardware-gated. Phase 4 never ends, which is the point.
Phase 0 — Stop the bleeding
hours to days · low / no costThe highest-impact, cheapest actions. Do these before anything else, and before reading further if you have to pick one thing.
- Get the controller off direct internet exposure. This is CISA’s number-one recommendation and it defeats the remote-attacker path entirely for most victims. Remove port-forwards and NAT to the controller; require a VPN for any remote engineering access.
- Check your own exposure. Search your public IP ranges (Shodan or equivalent) and confirm nothing is internet-reachable on the native industrial ports.
- Inventory every device — model, firmware version, and whether it is or ever was reachable. You cannot protect what you have not listed, and an operator plus an integrator can finish the list in an afternoon.
- Change default credentials; disable unused services and ports on the controllers and on the engineering workstation.
Phase 1 — Segment
days to weeks · one switch + one firewall- Put the operational network on its own VLAN behind a firewall — a small-utility version of the Purdue model, not a full multi-zone DMZ.
- Allowlist only the flows actually required between the HMI / engineering station and the controller, and deny everything else — especially anything crossing from the business side or the internet.
- Lock down the engineering workstation. It is the highest-value pivot on the network; treat it as sensitive equipment, not as an office PC.
Phase 2 — Compensating controls
weeks · modest costFor the residual risk segmentation doesn’t cover — and as the durable, permanent control if your hardware can’t do Phase 3.
- Passive monitoring on a mirror/SPAN port. A lightweight IDS, or even a small alerter that flags unexpected writes to control tags. The attack seen in the wild is an unauthorised write to a controller, so alerting on anomalous writes is high-signal.
- Strict change control on who may connect to the controller, and when.
Phase 3 — Per-device certificates: the real fix
weeks to months · hardware-permittingOnly available if the hardware and firmware support it — verify per device. This is the layer our open proof of concept demonstrates the principle of.
- Time synchronisation first, and it is easy to miss. Certificate validity windows and “the current revocation list” both depend on devices agreeing what time it is. Industrial equipment often has drifted clocks and no discipline around it — confirm this before anything else, or the whole PKI fails in ways that look like unrelated network trouble.
- Stand up a lightweight internal certificate authority, issue per-device certificates, and enable mutual TLS with identity binding — so the controller authenticates which engineering station, not merely that a certificate is signed by the right authority.
- Issuance is the easy part. Revocation and rotation are not. Neither is a few lines of
opensslthe way issuance is; both need a real workflow or a purpose-built tool. Budget for the full lifecycle, not just standing up the CA.
Phase 4 — Operate and maintain
ongoing · the durable needWhen a flaw cannot be patched away, the safe state has to be maintained rather than achieved.
- A revocation process: who revokes a credential, and when — engineer offboarding, device decommission, suspected compromise. This is the operational half of Phase 3’s cryptography.
- A rotation schedule with dates on a real calendar, well before expiry rather than after.
- Monitoring and a one-page incident runbook sized for the utility — not a SOC. Build the real one with your primacy agency and an integrator.
- Repeat the Phase 0 exposure check on a schedule. Exposure comes back; someone adds a port-forward for a good reason on a bad day.
Questions to hand your integrator
You do not need to understand cryptography to use this page. You need someone who supports your system to give a straight yes / no / unknown for each line.
Before anything else
- [ ]Is our controller reachable from the internet right now? This should be “no,” unconditionally, regardless of anything below.
- [ ]Do we have an inventory of every device, its firmware version, and whether it has ever been internet-reachable?
Does our hardware support the real fix?
- [ ]Does our specific firmware version support per-device certificate authentication? This is hardware-gated and it is a yes/no your integrator or vendor rep can answer directly.
- [ ]If yes: does the controller reject a connection using another device’s certificate — not just any certificate signed by the right authority? Ask to be shown a rejection, not told that it works.
- [ ]Is revocation configured, so a compromised or offboarded credential can actually be recalled — and what happens if the revocation source can’t be reached? Someone should be able to answer that on purpose.
- [ ]What happens to the write-monitoring from Phase 2 once encryption is on? It goes blind. Confirm what replaces the visibility rather than assuming it still works.
- [ ]When an engineer leaves, is there a process to issue a new credential and separately retire the old one? Re-issuing alone revokes nothing. Ask to see both steps, not just the first.
If our hardware can’t do it
- [ ]Is the controller on its own segment, behind a firewall, allowlisting only the traffic it needs?
- [ ]Is there any monitoring that would flag an unexpected write to a control tag? This doesn’t require new hardware to have an answer — ask what is already watching.
“Unknown” is a fine answer. It just means that is the next thing to find out — not a failure, and not something to be embarrassed about in front of an integrator.
The first sixty minutes
Not a complete incident-response plan, not a substitute for your primacy agency’s requirements, and not professional IR support. An orientation for the first hour.
Before everything below
Safety of the water system and the people who depend on it comes before anything on this page. If continuing safe operation requires an action that conflicts with any step here, do the safe thing.
Evidence preservation is always secondary to public health and safety. Nothing here should ever delay an action needed to keep the system operating safely.
Call — don’t wait until you’re sure
- →CISA, 24/7 — report@cisa.gov · (888) 282-0870. Report anomalous activity even before you’re certain. That is what the line is for.
- →Ask about mandatory reporting on that first call. There is a federal regime (CIRCIA) with a 72-hour clock for a covered entity’s substantial incidents and 24 hours for a ransom payment. Whether and exactly when it binds your utility turns on rulemaking status and your covered-entity determination, and this page deliberately will not answer that for you. Make it your first question — if a clock is running, it started when the incident did, not when you finish investigating.
- →Your state drinking-water primacy agency. Many states have their own reporting requirement on top of the federal one. This varies by state and no single page can state it accurately for all fifty.
Before you touch anything, if it’s safe to wait a few minutes
- →Don’t reimage, reboot, or “test if it’s fixed” by reconnecting. A wiped device cannot be examined afterward.
- →Do write down what you saw and when. What looked wrong, what time, who noticed, what’s different from normal. Plain notes are real evidence.
- →Do preserve logs if you can do it safely — controller event logs, workstation logs, firewall logs, alarm history. Copy them off before normal rotation overwrites them.
- →Do isolate, if it doesn’t compromise safe operation. Pull the affected segment rather than everything, if that’s enough and operations allow it.
- →Do coordinate off the suspect network. If an intruder is in your email, planning the response there tells them exactly what you know and how fast you’re moving. Use phones or accounts that don’t live on the affected systems until you know what’s clean.
Free help that already exists
Several of the phases above can be done with one of these rather than by yourself. All verified from live sources on 2026-08-03 — programs change, so confirm availability rather than assuming.
CISA — report an incident
The federal 24/7 operations centre for reporting cyber incidents and anomalous activity, any sector.
cisa.gov/reporting-cyber-incident →EPA Technical Assistance
Free expert help with training, device and account security, vulnerability management and policy. Explicitly covers small and rural systems with no dedicated security staff. Ticket-based intake.
epa.gov/cyberwater →WaterISAC
Threat alerts, twice-weekly security updates, incident reporting and analysis. A complimentary 12-month membership pilot has existed for NRWA members serving under 10,000 people — apply directly rather than assuming it is still open.
waterisac.org/nrwa →AWWA
No-cost water-sector cybersecurity risk-management guidance, a self-assessment tool, and material sized specifically for small and rural systems.
awwa.org →And ask about funding before assuming this comes out of the operating budget. Clean Water and Drinking Water State Revolving Funds are named in CISA’s own materials as routes for exactly this kind of hardening work.
On the State and Local Cybersecurity Grant Program specifically — as checked on 2026-08-03, its authorisation had lapsed on 2026-01-30 with reauthorisation pending, while several states were still administering rounds after that date. The accurate description is lapsed-and-pending, not closed — ask your state’s administering agency whether a round is actually open rather than assuming it either way. The State Revolving Funds are a separate authority and are unaffected.
What this plan is honest about
It is a plan, not a deployment. A real rollout needs your specific device inventory and each device’s capability. Those determine whether Phase 3 is even reachable, and this page cannot know them.
The cryptographic fix is hardware-gated. This is deliberately built to deliver value without it, so a utility on legacy equipment is not left with “nothing until you buy new controllers.”
Staffing reality is assumed, not wished away. Every step is meant to be achievable by a small operator plus an integrator. Where a step genuinely needs a specialist — certificate authority setup, IDS tuning — it is called out as an integrator task rather than quietly assumed in-house.
And where a control has a cost, the cost is named. Phase 3 blinds Phase 2’s monitoring. Phase 3 introduces an expiry failure mode that did not exist before. Those are real trade-offs, and a plan that hides them is not being kind to you — it is setting you up to be surprised at 3 a.m.
Run a small utility, or help one?
All of it is free and open — the runnable proof of concept, the standards mapping, the phased rollout, the worksheets, and the resource list. If you want help applying it to a specific environment, or a walk-through of where you actually stand, reach out. No fear-selling — a straight technical conversation about living safely with a flaw that will not be patched.