Genome Informatics
Online ISSN : 2185-842X
Print ISSN : 0919-9454
ISSN-L : 0919-9454
Detecting Gene Symbols and Names in Biological Texts
A First Step toward Pertinent Information Extraction
Denys ProuxFrancois RechenmannLaurent JulliardViolaine PilletBernard Jacq
著者情報
ジャーナル フリー

1998 年 9 巻 p. 72-80

詳細
抄録
Gathering data on molecular interactions to be fed into a specialized database has motivated the development of a computer system to help extracting pertinent information from texts, relying on advanced linguistic tools, completed with object-oriented knowledge modeling capabilities. As a first step toward this challenging objective, a program for the identification of gene symbols and names inside sentences has been devised. The main difficulty is that these names and symbols do not appear to follow construction rules. The program is thus made up of a series of sieves of different natures, lexical, morphological and semantic, to distinguish among the words of a sentence those which can only be potential gene symbols or names. Its performance has been evaluated, in terms of coverage and precision ratios, on a corpus of texts concerning D. melanogaster for which the list of names of known genes is available for checking.
著者関連情報
© Japanese Society for Bioinformatics
前の記事 次の記事
feedback
Top