One-size-fits-all Evaluation of LLMs for Safety Assurance Considered Harmful
Abstract
Large Language Models (LLMs) have been proposed to support the development and maintenance of assurance cases (ACs), but their use comes with risk. This position paper calls on academic and industrial interest holders to establish methodological standards to evaluate the application of LLMs to ACs. Our position is that there is no "one size-fits-all" evaluation method, and that evaluation must be tailored to the specific objectives and risk profiles of different LLM use cases. As a first step, we outline a preliminary taxonomy of characteristics, risk, and evaluation methods for LLM use cases.