-
2
pages
-
English
-
Documents
-
2012
Description
A Comparison of Document Clustering Techniques Michael Steinbach George Karypis Vipin Kumar Department of Computer Science / Army HPC Research Center, University of Minnesota 4-192 EE/CSci Building, 200 Union Street SE Minneapolis, Minnesota 55455 steinbac@cs.umn.edu karypis@cs.umn.edu kumar@cs.umn.edu ABSTRACT Table 1: Summary description of document sets. This paper presents the results of an experimental study Data Set Source Documents Classes Words of some common document clustering techniques: agglomerative re0 Reuters 1504 13 11465 hierarchical clustering and K-means. (We used both a “standard” re1 Reuters1657 25 3758 K-means algorithm and a “bisecting” K-means algorithm.) Our wap WebAce 1560 20 8460 results indicate that the bisecting K-means technique is better than tr31 TREC927 7 10128 the standard K-means approach and (somewhat surprisingly) as tr45 TREC 690 10 8261 good or better than the hierarchical approaches that we tested. fbis TREC2463 17 2000 Keywords la1 TREC 3204 6 31472 K-means, hierarchical clustering, document clustering. la2 TREC3075 2. Evaluation of Cluster Quality 1. INTRODUCTION We use two metrics for evaluating cluster quality: Hierarchical clustering is often portrayed as the better entropy, which provides a measure of “goodness” for un-nested quality clustering approach, but is limited because of its quadratic clusters or for the clusters at one level of a hierarchical clustering, time complexity.
-
Publié par
-
Publié le
31 mai 2012
-
Langue
English