-
2
pages
-
English
-
Documents
Description
Managing Information Extraction[Tutorial Outline]1 2 3AnHai Doan , Raghu Ramakrishnan , Shivakumar Vaithyanathan1 2 3University of Illinois, University of Wisconsin, IBM Research at Almadenanhai@cs.uiuc.edu, raghu@cs.wisc.edu, shiv@almaden.ibm.com1. INTRODUCTION unstructured data and the data extracted from it must beaddressed from scratch, in an extremely labor-intensive andMany applications increasingly involve a large amount oferror-prone process. Only recently has a consensus startedunstructured data. Examples of such data include email,to build on the need for re-usable tools, and for a unifledtext, Web pages, newsgroup postings, news articles, call-management of the entire extraction process, including ex-center text records, business reports, research papers, andtraction, storage, indexing, querying, and maintenance ofso on. In its raw form, the data has limited value since weboth the original raw data and the extracted information.can do little with it beyond keyword search. Consequently,In particular, we believe that such unifled management ofover the past two decades, signiflcant efiorts have focusedextraction isalogicalnextstepindatabasesupportfortext,on the problem of extracting structured information (e.g.,goingbeyondintegrationofinvertedindexesandsupportforresearchers, publications, co-author and advising relation-keyword search in RDBMSs. It is a unique opportunity forships, etc.) from such data. The extracted information ...
-
Publié par
-
Langue
English