One-size-fits-all Evaluation of LLMs for Safety Assurance Considered Harmful

Simon Diemert, Torin Viger, Jeffrey Joyce, Marsha Chechik
SAFECOMP 2025 Position Papers · 2025 · Conference/Workshop

Abstract

Large Language Models (LLMs) have been proposed to support the development and maintenance of assurance cases (ACs), but their use comes with risk. This position paper calls on academic and industrial interest holders to establish methodological standards to evaluate the application of LLMs to ACs. Our position is that there is no "one size-fits-all" evaluation method, and that evaluation must be tailored to the specific objectives and risk profiles of different LLM use cases. As a first step, we outline a preliminary taxonomy of characteristics, risk, and evaluation methods for LLM use cases.