-
11
pages
-
English
-
Documents
-
2013
Description
AFrameworkforBenchmarkingEntity-Annotation Systems Marco Cornolti Paolo Ferragina Massimiliano Ciaramita Dipartimento di Informatica Dipartimento di Informatica Google Research University of Pisa, Italy University of Pisa, Italy Zürich, Switzerland massi@google.comcornolti@di.unipi.it ferragina@di.unipi.it ABSTRACT ture but low coverage (such as WordNet, CYC, TAP), and a large text collection with wide coverage but unstructuredIn this paper we design and implement a benchmarking and noisy content (like the whole Web). The process of en-framework for fair and exhaustive comparison of entity-an- tity annotation involves three main steps: (1) parsing of thenotation systems. The framework is based upon the de - input text, which is the task to detect candidate entity men-nition of a set of problems related to the entity-annotation tions and link each of them to all possible entities they couldtask, a set of measures to evaluate systems performance, mention; (2) disambiguation of mentions, which is the taskand a systematic comparative evaluation involving all pub- of selecting the most pertinent Wikipedia page (i.e., entity)licly available datasets, containing texts of various types that best describes each mention; (3) pruning of a mention,such as news, tweets and Web pages.
-
Publié par
-
Publié le
19 mai 2013
-
Langue
English