2026 Volume 7 Issue 2 Pages 195-208
Development of benchmarks that qualify the large language models for bridge inspection and diagnosis is necessary to guarantee the quality of output. As a fundamental study toward developing a benchmark for evaluating the understanding of positional relationships among bridge components, this study proposes a quantitative method for assessing the extent to which information corresponding to positional relationships can be linearly decoded from the internal representations of a model. The proposed method was applied to a cross section of a conventional steel plate girder bridge with four main girders, in which vertical and lateral positional relationships among components were labeled. The results showed that high prediction accuracy was achieved when the positional relationships were explicitly described in the prompt. In contrast, the prediction accuracy decreased when such relationships were not explicitly described, and the accuracy for lateral relationships fell below the random prediction baseline. It should be noted that the linear decodability evaluated in this study does not directly assess whether large language models understand positional relationships among bridge components.