Artificial Intelligence and Data Science
Online ISSN : 2435-9262
Case study on detecting LLM hallucinations based on self-consistency across multiple responses
Takumi KOBAYASHIShiori FUJISAWATatsukuni TAKEDAMichio OHSUMI
Author information
JOURNAL OPEN ACCESS

2026 Volume 7 Issue 2 Pages 179-185

Details
Abstract

As a means of supporting decision-making in periodic inspections of road bridges and emergency investigations during disasters, his study considered the application of large language models (LLMs) and examined methods for detecting hallucinations. In this research, we investigated the applicability of a self-consistency–based approach that does not require pre-prepared ground-truth data, using Sentence-BERT and Natural Language Inference (NLI). The results indicate that, although they depend on the conditions assumed in this study, by assuming the self-consistency of LLM outputs it may be possible to detect variations in the notation of technical terms, numerical discrepancies, logical inconsistencies, and irrelevant content across multiple responses, even without preparing ground-truth data in advance. On the other hand, the algorithm examined in this study has the limitation that it cannot detect errors when the same incorrect information is produced consistently across all responses.

Content from these authors
© 2026 Japan Society of Civil Engineers
Previous article Next article
feedback
Top