Authors
Soumen Chakrabarti, Byron Dom, Rakesh Agrawal, Prabhakar Raghavan
Publication date
1998/8/25
Journal
The VLDB journal
Volume
7
Issue
3
Pages
163-178
Publisher
Springer Berlin/Heidelberg
Description
We explore how to organize large text databases hierarchically by topic to aid better searching, browsing and filtering. Many corpora, such as internet directories, digital libraries, and patent databases are manually organized into topic hierarchies, also called taxonomies. Similar to indices for relational data, taxonomies make search and access more efficient. However, the exponential growth in the volume of on-line textual information makes it nearly impossible to maintain such taxonomic organization for large, fast-changing corpora by hand. We describe an automatic system that starts with a small sample of the corpus in which topics have been assigned by hand, and then updates the database with new documents as the corpus grows, assigning topics to these new documents with high speed and accuracy. To do this, we use techniques from statistical pattern recognition to efficiently separate the …
Total citations
19992000200120022003200420052006200720082009201020112012201320142015201620172018201920202021202220232024914172732192324192019192014131711369631343