-
12
pages
-
English
-
Documents
-
2013
Description
UnderstandingNetworkFailuresinDataCenters: Measurement,Analysis,andImplications Phillipa Gill Navendu Jain Nachiappan Nagappan University of Toronto Microsoft Research Microsoft Research navendu@microsoft.com nachin@microsoft.comphillipa@cs.toronto.edu ABSTRACT lessons learned from this study to guide the design of future data center networks.We present the first large-scale analysis of failures in a data cen- Motivated by issues encountered by network operators, weter network. Through our analysis, we seek to answer several fun- study network reliability along three dimensions:damental questions: which devices/links are most unreliable, what causes failures, how do failures impact network traffic and how ef- ? Characterizing the most failure prone network elements. To fective is network redundancy? We answer these questions using achieve high availability amidst multiple failure sources such as multiple data sources commonly collected by network operators. hardware, software, and human errors, operators need to focus The key findings of our study are that (1) data center networks on fixing the most unreliable devices and links in the network. show high reliability, (2) commodity switches such as ToRs and Tothisend,wecharacterizefailurestoidentifynetworkelements AggS are highly reliable, (3) load balancers dominate in terms of with high impact on network reliability e.g.
-
Publié par
-
Publié le
21 mars 2013
-
Langue
English
-
Poids de l'ouvrage
1 Mo