Automated Labeling of Entities in CVE Vulnerability Descriptions with Natural Language Processing

Kensuke SUMOTO; Kenta KANAKOGI; Hironori WASHIZAKI; Naohiko TSUDA; Nobukazu YOSHIOKA; Yoshiaki FUKAZAWA; Hideyuki KANUKA

doi:10.1587/transinf.2023DAP0013

Abstract

Security-related issues have become more significant due to the proliferation of IT. Collating security-related information in a database improves security. For example, Common Vulnerabilities and Exposures (CVE) is a security knowledge repository containing descriptions of vulnerabilities about software or source code. Although the descriptions include various entities, there is not a uniform entity structure, making security analysis difficult using individual entities. Developing a consistent entity structure will enhance the security field. Herein we propose a method to automatically label select entities from CVE descriptions by applying the Named Entity Recognition (NER) technique. We manually labeled 3287 CVE descriptions and conducted experiments using a machine learning model called BERT to compare the proposed method to labeling with regular expressions. Machine learning using the proposed method significantly improves the labeling accuracy. It has an f1 score of about 0.93, precision of about 0.91, and recall of about 0.95, demonstrating that our method has potential to automatically label select entities from CVE descriptions.

Content from these authors

Favorites & Alerts

Corresponding author

Register with J-STAGE for free!