2026 Volume 7 Issue 2 Pages 171-178
This study conducted a preliminary investigation into how qualitative expressions in engineering standards for road bridge inspection, such as “significant” or “extensive,” affect damage assessments made by large language models (LLMs). In addition to analyzing the generated texts, the study examined probability distributions in the models’ internal Softmax layers and evaluated the consistency between these internal probability distributions and the externally generated outputs. The results showed that, as model size increased, the entropy of the final judgment distribution for qualitative expressions tended to decrease, while the semantic similarity of responses to the reference criteria and the consistency of judgments tended to improve. Although the findings are limited to the conditions examined in this study, including relatively short input texts and the use of quantized Qwen2.5 models, the results suggest that models with 7B parameters or larger tend to produce more stable damage assessments.