2025 Volume 6 Issue 3 Pages 406-419
RAG (Retrieval-Augmented Generation) is being introduced into construction and civil engineering workflows, enabling generative AI to answer questions based on specific documents. As its adoption grows, reliable evaluation methods are increasingly needed. Recent studies have explored LLM-as-a-judge, a cost-effective approach where the AI evaluates its own responses. However, comparisons with human judgment—commonly used in prior research—suffer from low reproducibility, making validation difficult. This study proposes a statistical framework that expresses evaluation stability and discriminative power using confidence level and statistical power, allowing LLM-as-a-judge to be validated without relying on human comparison.