Article ID: 2025EAP1237
Automatic knowledge extraction is one of the most important goals of Natural Language Processing. Especially after the emergence of COVID-19, the number of related literature is growing by about ten thousand per month, significantly challenging manual annotation and downstream tasks. In this paper, we describe a system for biomedical multi-label topic classification. Firstly, BERT is pre-trained on biomedical corpora which helps to capture deep semantic information. Furthermore, we fine-tune the pre-trained BERT on the COVID-19 literature from the Lit-Covid Database. Finally, automatic correction for Biomedical Multi-label Topic Classification method is introduced to our system to effectively take advantage of the domain expert experience. Our Act-BERT model achieves a micro F-score of 91.75% and macro F-score of 89.28% in the test set, and the macro F-score is 0.53% higher than the best system, which demonstrates the potential and effectiveness of the proposed framework.