IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences
Online ISSN : 1745-1337
Print ISSN : 0916-8508
Integrating Pre-trained Model and Automatic Correction for Biomedical Multi-label Topic Classification
Jieqiong ZHENGNing JIARuixia CAOJunxiang SONG
Author information
JOURNAL FREE ACCESS Advance online publication

Article ID: 2025EAP1237

Details
Abstract

Automatic knowledge extraction is one of the most important goals of Natural Language Processing. Especially after the emergence of COVID-19, the number of related literature is growing by about ten thousand per month, significantly challenging manual annotation and downstream tasks. In this paper, we describe a system for biomedical multi-label topic classification. Firstly, BERT is pre-trained on biomedical corpora which helps to capture deep semantic information. Furthermore, we fine-tune the pre-trained BERT on the COVID-19 literature from the Lit-Covid Database. Finally, automatic correction for Biomedical Multi-label Topic Classification method is introduced to our system to effectively take advantage of the domain expert experience. Our Act-BERT model achieves a micro F-score of 91.75% and macro F-score of 89.28% in the test set, and the macro F-score is 0.53% higher than the best system, which demonstrates the potential and effectiveness of the proposed framework.

Content from these authors
© 2026 The Institute of Electronics, Information and Communication Engineers
Previous article Next article
feedback
Top