View article

Enhancing a biomedical information extraction system with dictionary mining and context disambiguation

Authors

Sougata Mukherjea, L Venkata Subramaniam, Gaurav Chanda, Sriram Sankararaman, Ravi Kothari, Vishal Batra, Deo Bhardwaj, Biplav Srivastava

Publication date

2004/9

Journal

IBM Journal of Research and Development

Volume

Issue

5.6

Pages

693-701

Publisher

IBM

Description

Journals and conference proceedings represent the dominant mechanisms for reporting new biomedical results. The unstructured nature of such publications makes it difficult to utilize data mining or automated knowledge discovery techniques. Annotation (or markup) of these unstructured documents represents the first step in making these documents machine-analyzable. Often, however, the use of similar (or the same) labels for different entities and the use of different labels for the same entity makes entity extraction difficult in biomedical literature. In this paper we present a system called BioAnnotator for identifying and classifying biological terms in documents. BioAnnotator uses domain-based dictionary lookup for recognizing known terms and a rule engine for discovering new terms. We explain how the system uses a biomedical dictionary to learn extraction patterns for the rule engine and how it …

Total citations

Cited by 46

200420052006200720082009201020112012201320142015201620172018201920202021202220231 5 4 5 7 2 2 2 1 2 4 2 1 2 1 1 2 1 1

Scholar articles

Enhancing a biomedical information extraction system with dictionary mining and context disambiguation

S Mukherjea, LV Subramaniam, G Chanda… - IBM Journal of Research and Development, 2004