Artificial Intelligence and Data Science
Online ISSN : 2435-9262
A statistical approach for validating LLM-as-a-judge
Jun KAWAMURADaisuke SUGETAKenta HAKOISHI
Author information
JOURNAL OPEN ACCESS

2025 Volume 6 Issue 3 Pages 406-419

Details
Abstract

RAG (Retrieval-Augmented Generation) is being introduced into construction and civil engineering workflows, enabling generative AI to answer questions based on specific documents. As its adoption grows, reliable evaluation methods are increasingly needed. Recent studies have explored LLM-as-a-judge, a cost-effective approach where the AI evaluates its own responses. However, comparisons with human judgment—commonly used in prior research—suffer from low reproducibility, making validation difficult. This study proposes a statistical framework that expresses evaluation stability and discriminative power using confidence level and statistical power, allowing LLM-as-a-judge to be validated without relying on human comparison.

Content from these authors
© 2025 Japan Society of Civil Engineers
Previous article Next article
feedback
Top