Evaluating AI-supported eliminative argumentation for developing reliable assurance cases

Torin Viger, Logan Murphy, Simon Diemert, Claudio Menghi, Aren A. Babikian, Jeff Joyce, Alessio Di Sandro, Naweed Anwari, Erin Cyffka, Marsha Chechik
Empirical Software Engineering · 2026 · Journal

Abstract

As software systems become increasingly complex and more prominently used in society, it has become essential for systems involving software to operate safely and reliably. Assurance cases (ACs) are structured arguments that are often used to show why cyber-physical systems are acceptably safe for use in their intended environments. Designing ACs is a complex process involving multiple parties, and ACs can be subject to human errors in their development such as confirmation bias and reasoning errors. AI-Supported Eliminative Argumentation (AI-EA) is a recently proposed framework that leverages Large Language Models (LLMs) to support the development of ACs. Specifically, AI-EA uses LLMs to identify defeaters, i.e., reasons to doubt an AC, so that these doubts can be better understood and mitigated. In this paper, we detail our implementation of AI-EA using GPT-4 and evaluate 225 LLM-generated defeaters in collaboration with safety domain experts. We then evaluate the usefulness of AI-EA in practice through a user study in which eleven AC developers used AI-EA to develop portions of an AC. Our findings show that (i) AI-EA can generate reasonable, informative, relevant, and useful defeaters for ACs, (ii) practitioners found AI-EA to be useful in supporting the AC development process, and (iii) using AI-EA in practice helped AC developers identify a broader range of doubts in the ACs they created.