Infrastructure
Paragraphs

ABSTRACT: Resilience – thriving through adversity – has become a central theme in discussions about energy and other critical infrastructure sectors. Now that energy systems are controlled, automated and networked through information and communications technologies, the definition of energy sector resilience has expanded to include cyber resilience. This focus commonly emphasizes technology and technological solutions aimed at preventing, mitigating or rebounding from crises. For instance, the most prominent United States definition of cyber resilience (from the Department of Commerce) is largely reactive – “withstand, recover from, and adapt” – and technology-centric, “to compromises on systems that use or are enabled by cyber resources.” 

The terms “cyber” and “resilience” are both words that mean different things to different people. For some, cyber is purely technical, while for others cyber refers to cyber-enabled disinformation. Resilience might refer to the reliability of a single system, or “the ability to continue to function, perhaps in a degraded state, when that system is unavailable.” [1] 

This paper advocates for a broader concept of energy sector cyber resilience for NATO allies and partners. A concept of energy sector cyber resilience that encompasses social-technical and operational dimensions is needed to better addresses the complex and interdependent nature of these systems, and the potential context of hostilities.

Download the full report here.

All Publications button
1
Publication Type
Reports
Publication Date
Authors
Karen Guttieri
Paragraphs

Abstract:

AI models are increasingly prevalent in high-stakes environments, necessitating thorough assessment of their capabilities and risks. Benchmarks are popular for measuring these attributes and for comparing model performance, tracking progress, and identifying weaknesses in foundation and non-foundation models. They can inform model selection for downstream tasks and influence policy initiatives. However, not all benchmarks are the same: their quality depends on their design and usability. In this paper, we develop an assessment framework considering 40 best practices across a benchmark's life cycle and evaluate 25 AI benchmarks against it. We find that there exist large quality differences and that commonly used benchmarks suffer from significant issues. We further find that most benchmarks do not report statistical significance of their results nor can results be easily replicated. To support benchmark developers in aligning with best practices, we provide a checklist for minimum quality assurance based on our assessment. We also develop a living repository of benchmark assessments to support benchmark comparability.

Find the full article here

All Publications button
1
Publication Type
Reports
Publication Date
Authors
Max Lamparth
Subscribe to Infrastructure