Proceedings of the Annual Conference of JSAI
Online ISSN : 2758-7347
39th (2025)
Session ID : 1Win4-18
Conference information

Internal Representations of Familiarity Judgments in Language Models
*Kai SATORyosuke TAKAHASHIBenjamin HEINZERLINGKenshiro TANAKAYufeng ZHAOYoshihiro SAKAINaoya INOUEInui KENTARO
Author information
CONFERENCE PROCEEDINGS FREE ACCESS

Details
Abstract

The knowledge acquisition capabilities of language models (LMs) have been extensively studied; however, the mechanisms by which LMs judge the familiarity of acquired knowledge remain insufficiently understood. In this study, we employ a LM to perform an analysis of their internal states during familiarity judgment. Our findings reveal that (1) the information required to judge familiarity is embedded within the internal representations at the time the knowledge is learned, and (2) it exhibits different activation patterns when predicting knowledge as familiar versus unfamiliar. These findings provide insights into the mechanisms underlying familiarity judgment in language models.

Content from these authors
© 2025 The Japanese Society for Artificial Intelligence
Previous article Next article
feedback
Top