Currently free during beta - premium features coming soon. Subscribe now to lock in early access.

arXiv: ICS Cybersecurity Datasets: A Systematic Meta-Review of Coverage, Evaluation Practice, and Structural Gaps

AI_SAFETY AI Security & Safety · · arxiv_cscr

AI Analysis

A new systematic meta-review of cybersecurity datasets for industrial control systems (ICS) was published on arXiv, offering a critical assessment of existing research resources. The paper, titled "ICS Cybersecurity Datasets: A Systematic Meta-Review of Coverage, Evaluation Practice, and Structural Gaps," examines the breadth, quality, and practical usability of datasets used to train and test AI-based threat detection systems for operational technology. Its main finding is that current datasets suffer from significant structural gaps, including limited coverage of modern attack vectors, inconsistent evaluation methodologies, and a lack of real-world operational fidelity, which undermines the reliability of AI models deployed in critical infrastructure.

This publication directly affects compliance teams in sectors governed by the EU's NIS2 Directive, the Cyber Resilience Act, and sector-specific rules for energy, water, transport, and manufacturing. Organizations that rely on AI-driven intrusion detection or anomaly monitoring for their ICS environments should treat this as a warning that their validation evidence may be based on incomplete or unrealistic data. Regulators are increasingly demanding demonstrable effectiveness of security controls, and this review provides a basis for questioning the robustness of AI models that have not been tested against diverse, current, and operationally relevant datasets.

Compliance teams should immediately conduct an inventory of any AI-based security tools used in OT environments, requesting from vendors the specific datasets used for training and testing. Next, they should map those datasets against the gaps identified in the review, particularly around protocol diversity and attack coverage. Finally, they should update their risk assessment and vendor due diligence processes to require evidence of dataset currency and independent evaluation, and begin planning for supplementary testing using their own operational data or newly developed benchmark sets. This is a proactive step to avoid future enforcement actions or liability under the proposed AI liability framework.

Get notified about AI_SAFETY changes

Subscribe to our free weekly digest covering 24 compliance frameworks.