-
111
pages
-
English
-
Documents
-
2006
Description
Unsupervised Duplicate DetectionUsing Sample Non-DuplicatesVom Fachbereich Informatikder Technischen Universit¨at Darmstadtzur Erlangung des akademischen Grades einesDoktor-Ingenieurs (Dr.-Ing.)genehmigteDissertationvonDiplom-IngenieurPatrick Lehtiaus DarmstadtReferent: Prof. Dr.-Ing. Erich NeuholdKorreferent: Prof. Dr. Thomas HofmannTag der Einreichung: 6. Februar 2006Tag der mu¨ndlichen Pru¨fung: 17. Mai 2006Darmstadt 2006D17AbstractThe problem of identifying objects in databases that refer to the same realworld entity, is known, among others, as duplicate detection or record link-age. Objects may be duplicates, even though they are not identical due toerrors and missing data.Traditional scenarios for duplicate detection are data warehouses, whichare populated from several data sources. Duplicate detection here is part ofthe data cleansing process to improve data quality for the data warehouse.More recently in application scenarios like web portals, that offer users uni-fied access to several data sources, or meta search engines, that distributea search to several other resources and finally merge the individual results,the problem of duplicate detection is also present. In such scenarios no longand expensive data cleansing process can be carried out, but good duplicateestimations must be available directly.
-
Publié par
-
Publié le
01 janvier 2006
-
Langue
English
-
Poids de l'ouvrage
1 Mo