Proceedings of the Fuzzy System Symposium
37th Fuzzy System Symposium
Session ID : TA4-3
Conference information

proceeding
Analysis of Text Extraction from PDF Documents of Local Assembly Materials for Constructing Structured Data
*Hokuto OtotakeYuzu UchidaKeiichi TakamaruYasutomo Kimura
Author information
CONFERENCE PROCEEDINGS FREE ACCESS

Details
Abstract

Local governments publish information about parliamentary activities and financial conditions on their Web sites, many of which are available in PDF document format. PDF documents have a high level of visual expression such as layout. However, it may be difficult to extract text data such as letters and word order from PDF documents because of high flexibility of their internal data representation. Since local governments publish PDF documents with their own format, the methods and difficulty of extracting text data may vary. To be useful as a language resource or Linked Open Data, it is desirable to structure documents on a plain text basis, such as JSON. In this study, we analyze PDF documents of assembly materials published local governments in Fukuoka Prefecture from the viewpoint of the method and difficulty of extracting plain text.

Content from these authors
© 2021 Japan Society for Fuzzy Theory and Intelligent Informatics
Previous article Next article
feedback
Top