Vulnerability Remediation SLAs: A Practical Workbook
A remediation SLA is an operating clock from confirmed finding to verified outcome. It must define when triage starts, when mitigation is required, who delivers a durable fix, and how retest closes the work. It is not a public promise or one universal patch deadline. NIST describes enterprise patch management as an organised process; CISA KEV is a prioritisation input where known exploitation changes urgency.
Set clocks from evidence
Start the clock when a finding is confirmed against an owned asset, not when an unverified scanner result appears. Record discovery time, confirmation time, priority rationale, asset owner, remediation owner, and due times. Use categories matched to local risk. Example internal targets: actively exploited and exposed—contain immediately, plan same day, remediate at earliest safe window; critical reachable issue—mitigate rapidly and fix on accelerated change; normal issue—fix in planned maintenance. These are examples, not universal obligations.
Separate mitigation from remediation. Mitigation reduces immediate exposure: route restriction, feature disablement, vendor workaround, credential rotation, or monitoring. Remediation removes or permanently resolves root cause: vendor update, configuration correction, replacement, or code change. A mitigation must have an owner, evidence, and review date so it does not become permanent by accident.
Make retest a required stage
A change ticket proves work was attempted, not that exposure is gone. Retest matches the original condition: check installed version or configuration, repeat safe reachability test, review service health, and collect new evidence. Mark pass only when expected fixed behaviour and required service function both hold. Mark fail when control still permits original path, evidence is missing, or change broke required operation.
| Stage | Required record | Completed example |
|---|---|---|
| Triage | asset, owner, priority, due date | exposed asset confirmed; owner accepts ticket |
| Mitigation | control, time, evidence | inbound route restricted; log retained |
| Remediation | change and rollback plan | vendor update applied in approved window |
| Retest | original condition plus health test | old version absent; service transaction passes |
| Closure | reviewer and evidence location | reviewer marks pass; record linked |
Control exceptions
An exception is a bounded decision, never a blank due date. Include affected finding and assets, reason fix cannot occur, risk owner, approving authority, compensating controls, mitigation verification, review cadence, and expiry. The system should flag expiry before it passes. On expiry, remediate, renew through a new decision, or return item to overdue status. “Vendor dependency” describes a problem; it is not an exception record.
For KEV entries, retain current catalog evidence and vendor action guidance alongside local exposure facts. If patching cannot proceed, record the chosen mitigation and why it changes exposure. Do not claim compliance with CISA directives unless they apply and are independently verified.
Evidence-driven reporting
Show clock performance as records, not marketing metrics: open count by priority, overdue items, mitigated-but-not-fixed items, stale exceptions, and retest status. Keep names and sensitive implementation detail out of broad reports. A useful SLA workbook lets a reviewer answer five questions quickly: what is affected, who owns it, what is due, what reduced risk now, and what evidence proves closure.
Escalation path
Escalate when triage lacks an owner, mitigation misses its clock, a change window is repeatedly deferred, or retest fails. Escalation asks for a decision: add resources, accept a bounded exception, change the mitigation, or remove the asset from exposure. It does not merely send a reminder. Preserve the decision and new due date in the same record. Review recurring misses for systemic causes such as inventory gaps, unsupported platforms, approval delays, or inadequate maintenance windows. Correcting those causes is remediation of the process, while fixing a single asset is remediation of the finding.
Measure decision quality
Review a sample of closed records monthly. Confirm original evidence remains linked, mitigation was reviewed until permanent fix, retest used correct condition, and exception expiry was handled. Compare due date changes with approval history. Repeated date movement without changed evidence is a failed governance signal. Report blockers by category to management: vendor dependency, capacity, change freeze, asset ownership, or failed remediation. This helps owners improve operating process without converting examples in this workbook into universal service promises. A clock is useful only when it leads to accountable decision and recorded result.
Completed SLA record
A closed example must be reconstructable: VULN-2026-014, asset EDGE-01, confirmation evidence URI, priority rationale, asset owner, remediation owner, mitigation owner, mitigation due, remediation due, change ID, exception ID or none, exception expiry, original-condition retest, service-health retest, reviewer, and closure timestamp. If a change is deferred, record new due date with approving risk owner and concrete compensating control; do not overwrite original date. Closure requires both original exposure no longer present and required service behavior verified.