-
12
pages
-
English
-
Documents
-
2011
Description
PLANET: Massively Parallel Learning of Tree Ensembleswith MapReduceBiswanath Panda, Joshua S. Herbach, Sugato Basu, Roberto J. BayardoGoogle, Inc.[bpanda, jsherbach, sugato]@google.com, bayardo@alum.mit.eduABSTRACT plexities such as data partitioning, scheduling tasks acrossmany machines, handling machine failures, and perform-Classificationandregressiontreelearningonmassivedatasetsing inter-machine communication. These properties haveis a common data mining task at Google, yet many statemotivated many technology companies to run MapReduceof the art tree learning algorithms require training data toframeworks on their compute clusters for data analysis andreside in memory on a single machine. While more scal-other data management tasks. MapReduce has become inable implementations of tree learning have been proposed,some sense an industry standard. For example, there arethey typically require specialized parallel computing archi-open source implementations such as Hadoop that can betectures. In contrast, the majority of Google’s computingrun either in-house or on cloud computing services such asinfrastructure is based on commodity hardware.1 2Amazon EC2. Startups like Cloudera offer software andInthispaper,wedescribePLANET:ascalabledistributedservices to simplify Hadoop deployment, and companies in-frameworkforlearningtreemodelsoverlargedatasets. PLA-cluding Google, IBM and Yahoo! have granted several ...
-
Publié par
-
Publié le
24 juin 2011
-
Langue
English