Hucheng Zhou

dblp:75/7061 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
1since 2021 · last 2021
0000-0002-1894-3897ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 2 first-authorSystems, architecture and hardware · 6Artificial intelligence and machine learning · 3Computer networks · 3Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
11 papers
Energy-efficient computing · 35% High-performance computing · 20% Parallel and multicore computing · 15%
Databases, data mining, and information retrieval
4 papers
Information retrieval · 58% Data mining · 15% Recommender systems · 15%
Software engineering, system software, and programming languages
7 papers
Program analysis · 34% Compilers and program optimization · 28% Empirical software engineering · 17%
Human-computer interaction and pervasive computing
2 papers
Ubiquitous computing and smart environments · 81% Wearable and physiological sensing · 19%

Topics — the 29 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing › power management
display power management
0.422015
Demo: Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015
Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015
Energy-efficient computing › power management › display power management
dynamic resolution scaling
0.422015
Demo: Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015
Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015
Energy-efficient computing
power management
0.422015
Demo: Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015
Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015
Information retrieval › ranking
learning to rank
0.312018
RapidScorer: Fast Tree Ensemble Evaluation by Maximizing Compactness in Data Level Parallelization · KDD 2018
Information retrieval
ranking
0.312018
RapidScorer: Fast Tree Ensemble Evaluation by Maximizing Compactness in Data Level Parallelization · KDD 2018
Information retrieval
search engines
0.312018
RapidScorer: Fast Tree Ensemble Evaluation by Maximizing Compactness in Data Level Parallelization · KDD 2018
Program analysis
static analysis
0.332015
Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE · IEEE Trans. Parallel Distributed Syst. 2015
Cybertron: pushing the limit on I/O reduction in data-parallel programs · OOPSLA 2014
Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE · OSDI 2012
High-performance computing › numerical linear algebra
matrix multiplication
0.312017
Improving Execution Concurrency of Large-Scale Matrix Multiplication on Distributed Data-Parallel Platforms · IEEE Trans. Parallel Distributed Syst. 2017
Empirical software engineering
mining software repositories
0.322015
An Empirical Study on Quality Issues of Production Big Data Platform · ICSE (2) 2015
A characteristic study on failures of production distributed data-parallel programs · ICSE 2013
Recommender systems › click-through rate prediction
feature interaction
0.212016
Multi-view Machines · WSDM 2016
Data mining
multi-view learning
0.212016
Multi-view Machines · WSDM 2016
Distributed systems
fault tolerance
0.222015
A characteristic study on failures of production distributed data-parallel programs · ICSE 2013
An Empirical Study on Quality Issues of Production Big Data Platform · ICSE (2) 2015
Software maintenance and evolution
performance diagnosis
0.212015
Log2: A Cost-Aware Logging Mechanism for Performance Diagnosis · USENIX ATC 2015
Program analysis
symbolic execution
0.212015
Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE · IEEE Trans. Parallel Distributed Syst. 2015
Cloud and datacenter computing
big data platform
0.212015
An Empirical Study on Quality Issues of Production Big Data Platform · ICSE (2) 2015
Cloud and datacenter computing
quality of service
0.212015
An Empirical Study on Quality Issues of Production Big Data Platform · ICSE (2) 2015
Parallel and multicore computing › data parallelism
data-parallel systems
0.212013
A characteristic study on failures of production distributed data-parallel programs · ICSE 2013
Storage systems › storage reliability
failure characterization
0.212013
A characteristic study on failures of production distributed data-parallel programs · ICSE 2013
Compilers and program optimization
compiler optimization
0.112012
Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE · OSDI 2012
Parallel and multicore computing
data-parallel programming
0.112012
Optimizing Data Shuffling in Data-Parallel Computation by Understanding User-Defined Functions · NSDI 2012
Distributed systems › distributed data processing
data shuffling
0.112012
Optimizing Data Shuffling in Data-Parallel Computation by Understanding User-Defined Functions · NSDI 2012
Parallel and multicore computing
parallel programming models
0.112012
Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE · OSDI 2012
Compilers and program optimization › compiler optimization › redundancy elimination
partial redundancy elimination
0.112011
An SSA-based algorithm for optimal speculative code motion under an execution profile · PLDI 2011
Compilers and program optimization › code motion
speculative code motion
0.112011
An SSA-based algorithm for optimal speculative code motion under an execution profile · PLDI 2011
Parallel and multicore computing › data-parallel programming
data-level parallelization
0.112018
RapidScorer: Fast Tree Ensemble Evaluation by Maximizing Compactness in Data Level Parallelization · KDD 2018
Wearable and physiological sensing › acoustic sensing
ultrasonic sensing
0.112015
Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015
Debugging and program repair
failure analysis
0.012013
A characteristic study on failures of production distributed data-parallel programs · ICSE 2013
Compilers and program optimization › intermediate representation
static single assignment form
0.012011
An SSA-based algorithm for optimal speculative code motion under an execution profile · PLDI 2011
Machine learning and data management
machine learning for systems
0.012009
Machine learning-based prefetch optimization for data center applications · SC 2009

Methods — techniques the papers use, named apart from their topics

ultrasonic distance detection · 0.9empirical study · 0.8run-length encoding · 0.7data parallelization · 0.7bitvector representation · 0.7system-level optimization · 0.6replication-based execution strategy · 0.6symbolic execution · 0.4dynamic resolution scaling · 0.4dead code elimination · 0.4factorization machines · 0.2incident management analysis · 0.2runtime profiling · 0.2constraint-based encoding · 0.2static analysis · 0.1code optimization · 0.1minimum cut · 0.1flow network · 0.1
YearPublicationVenuePosition
2021 Multiple interleaving interests modeling of sequential user behaviors in e-commerce platform
Yuqiang Han, Qian Li 0016, Hucheng Zhou, Zhenglu Yang, Jian Wu 0001
World Wide Web4
2018 RapidScorer: Fast Tree Ensemble Evaluation by Maximizing Compactness in Data Level Parallelization
abstract
Relevance ranking models based on additive ensembles of regression trees have shown quite good effectiveness in web search engines. In the era of big data, tree ensemble models grow large in both tree depth and ensemble size to provide even better search relevance and user experience. However, the computational cost for their scoring process is high, such that it becomes a challenging issue to apply the big tree ensemble models in a search engine which needs to answer thousands of queries per second. Although several works have been proposed to improve the scoring process, the challenge is still great especially when the model size grows large. In this paper, we present RapidScorer , a novel framework for speeding up the scoring process of industry-scale tree ensemble models, without hurting the quality of scoring results. RapidScorer introduces a modified run length encoding called epitome to the bitvector representation of the tree nodes. Epitome can greatly reduce the computation cost to traverse the tree ensemble, and work with several other proposed strategies to maximize the compactness of data units in memory. The achieved compactness makes it possible to fully utilize data parallelization to improve model scalability. Experiments on two web search benchmarks show that, RapidScorer achieves significant speed-up over the state-of-the-art methods: V-QuickScorer , ranging from 1.3x to 3.5x; QuickScorer , ranging from 2.1x to 25.0x; VPred , ranging from 2.3x to 18.3x; and XGBoost , ranging from 2.6x to 42.5x.
Ting Ye, Hucheng Zhou, Will Y. Zou, Bin Gao 0001, Ruofei Zhang
KDD2
2017 FxpNet: Training a deep convolutional neural network in fixed-point representation
abstract
We introduce FxpNet, a framework to train deep convolutional neural networks with low bit-width arithmetics in both forward pass and backward pass. During training FxpNet further reduces the bit-width of stored parameters (also known as primal parameters) by adaptively updating their fixed-point formats. These primal parameters are usually represented in the full resolution of floating-point values in previous binarized and quantized neural networks. In FxpNet, during forward pass fixed-point primal weights and activations are first binarized before computation, while in backward pass all gradients are represented as low resolution fixed-point values and then accumulated to corresponding fixed-point primal parameters. To have highly efficient implementations in FPGAs, ASICs and other dedicated devices, FxpNet introduces Integer Batch Normalization (IBN) and Fixed-point ADAM (FxpADAM) methods to further reduce the required floating-point operations, which will save considerable power and chip area. The evaluation on CIFAR-10 dataset indicates the effectiveness that FxpNet with 12-bit primal parameters and 12-bit gradients achieves comparable prediction accuracy with state-of-the-art binarized and quantized neural networks.
Xi Chen 0107, Hucheng Zhou, Ningyi Xu
IJCNN3
2017 Improving Execution Concurrency of Large-Scale Matrix Multiplication on Distributed Data-Parallel Platforms
abstract
Matrix multiplication is a dominant but very time-consuming operation in many big data analytic applications. Thus its performance optimization is an important and fundamental research issue. The performance of large-scale matrix multiplication on distributed data-parallel platforms is determined by both computation and IO costs. For existing matrix multiplication execution strategies, when the execution concurrency scales up above a threshold, their execution performance deteriorates quickly because the increase of the IO cost outweighs the decrease of the computation cost. This paper presents a novel parallel execution strategy CRMM (Concurrent Replication-based Matrix Multiplication) along with a parallel algorithm, Marlin, for large-scale matrix multiplication on data-parallel platforms. The CRMM strategy exploits higher execution concurrency for sub-block matrix multiplication with the same IO cost. To further improve the performance of Marlin, we also propose a number of novel system-level optimizations, including increasing the concurrency of local data exchange by calling native library in batch, reducing the overhead of block matrix transformation, and reducing disk heavy shuffle operations by exploiting the semantics of matrix computation. We have implemented Marlin as a library along with a set of related matrix operations on Spark and also contributed Marlin to the open-source community. For large-sized matrix multiplication, Marlin outperforms existing systems including Spark MLlib, SystemML and SciDB, with about 1.29×, 3.53× and 2.21× speedup on average, respectively. The evaluation upon a real-world DNN workload also indicates that Marlin outperforms above systems by about 12.8×, 5.1× and 27.2× speedup, respectively.
Rong Gu 0001, Chen Tian 0001, Hucheng Zhou, Guanru Li, Yihua Huang 0001
IEEE Trans. Parallel Distributed Syst.4
2016 The Improvement of the Trustworthiness of Android App Stores in China
abstract
The absence of Google Play has created a booming area for Android app distribution through third-party app stores in China. Since the study showed that the trustworthy level of app stores was fairly low in 2014, much attention should be paid on the changes of the trustworthiness of Android app stores. In this paper, we present a method to analyze the changes of trustworthiness of the top popular Android app stores in China. In this method, we evaluate the target app stores by analyzing the sampled apps hosted in them. Further more, we have used this method to track the changes of trustworthy level of Android app stores in China about two years. The results indicate that the trustworthy level of the top popular Android app stores in China has been improved 24% on average. It can be seen that the positive changes may be related to the development of the China's mobile market, the improvement of Android system, and the introduced policies. Although the trustworthy level of top popular Android app stores in China is still low, it is predicted to be improving in the future.
Yiying Ng, Hucheng Zhou
APSEC3
2016 Multi-view Machines
abstract
With rapidly growing amount of data available on the web, it becomes increasingly likely to obtain data from different perspectives for multi-view learning. Some successive examples of web applications include recommendation and target advertising. Specifically, to predict whether a user will click an ad in a query context, there are available features extracted from user profile, ad information and query description, and each of them can only capture part of the task signals from a particular aspect/view. Different views provide complementary information to learn a practical model for these applications. Therefore, an effective integration of the multi-view information is critical to facilitate the learning performance.
Bokai Cao, Hucheng Zhou, Guoqiang Li 0008, Philip S. Yu
WSDM2
2015 An Empirical Study on Quality Issues of Production Big Data Platform
abstract
Big Data computing platform has evolved to be a multi-tenant service. The service quality matters because system failure or performance slowdown could adversely affect business and user experience. There is few study in literature on service quality issues of production Big Data computing platform. In this paper, we present an empirical study on the service quality issues of Microsoft ProductA, which is a company-wide multi-tenant Big Data computing platform, serving thousands of customers from hundreds of teams. ProductA has a well-defined incident management process, which helps customers report and mitigate service quality issues on 24/7 basis. This paper explores the common symptom, causes and mitigation of service quality issues in Big Data computing. We conduct an empirical study on 210 real service quality issues in ProductA. Our major findings include (1) 21.0% of escalations are caused by hardware faults; (2) 36.2% are caused by system side defects; (3) 37.2% are due to customer side faults. We also studied the general diagnosis process and the commonly adopted mitigation solutions. Our findings can help improve current development and maintenance practice of Big Data computing platform, and motivate tool support.
Hucheng Zhou, Jian-Guang Lou, Hongyu Zhang 0002, Haoxiang Lin, Tingting Qin
ICSE (2)1
2015 Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling
abstract
The extremely-high display density of modern smartphones imposes a significant burden on power consumption, yet does not always provide an improved user experience and may even lead to a compromised user experience. As human visually-perceivable ability highly depends on the user-screen distance, a reduced display resolution may still achieve the same user experience when the user-screen distance is large. This provides new power-saving opportunities. In this paper, we present a flexible dynamic resolution scaling system for smartphones. The system adopts an ultrasonic-based approach to accurately detect the user-screen distance at low-power cost and makes scaling decisions automatically for maximum user experience and power saving. App developers or users can also adjust the resolution manually as their needs. Our system is able to work on existing commercial smartphones and support legacy apps, without requiring re-building the ROM or any changes of apps. An end-to-end dynamic resolution scaling system is implemented on the Galaxy S5 LTE-A and Nexus 6 smartphones, and the correctness and effectiveness are evaluated against 30 games and benchmarks. Experimental results show that all the 30 apps can run successfully with per-frame, real-time dynamic resolution scaling. The energy per frame can be reduced by 30.1% on average and up to 60.5\% at most when the resolution is halved, for 15 apps. A user study with 10 users indicates that our system remains good user experience, as none of the 10 users could perceive the resolution changes in the user study.
Songtao He, Yunxin Liu 0001, Hucheng Zhou
MobiCom3
2015 Demo: Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling
abstract
The extremely-high display density of modern smartphones imposes a significant burden on power consumption, yet does not always provide an improved user experience and may even lead to a compromised user experience. As human visually-perceivable ability highly depends on the user-screen distance, a reduced display resolution may still achieve the same user experience when the user-screen distance is large. This provides new power-saving opportunities. We present a flexible dynamic resolution scaling system for smartphones. The system adopts an ultrasonic-based approach to detect the user-screen distance at low-power cost and makes scaling decisions automatically for maximum user experience and power saving. App developers or users can also adjust the resolution manually and dynamically as their needs. Our system is able to work on the existing commercial smartphones and support the legacy apps, without requiring re-building the ROM or any changes from apps.
Songtao He, Yunxin Liu 0001, Hucheng Zhou
MobiCom3
2015 Log2: A Cost-Aware Logging Mechanism for Performance Diagnosis
Rui Ding 0001, Hucheng Zhou, Jian-Guang Lou, Hongyu Zhang 0002, Qingwei Lin, Qiang Fu 0015, Dongmei Zhang 0001, Tao Xie 0001
USENIX ATC2
2015 Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE
abstract
To minimize the amount of data-shuffling I/O that occurs between the pipeline stages of a distributed data-parallel program, its procedural code must be optimized with full awareness of the pipeline that it executes in. Unfortunately, neither pipeline optimizers nor traditional compilers examine both the pipeline and procedural code of a data-parallel program so programmers must either hand-optimize their program across pipeline stages or live with poor performance. To resolve this tension between performance and programmability, this paper describes PeriSCOPE, which automatically optimizes a data-parallel program's procedural code in the context of data flow that is reconstructed from the program's pipeline topology. Such optimizations eliminate unnecessary code and data, perform early data filtering, and calculate small derived values (e.g., predicates) earlier in the pipeline, so that less data - sometimes much less data - is transferred between pipeline stages. PeriSCOPE further leverages symbolic execution to enlarge the scope of such optimizations by eliminating dead code. We describe how PeriSCOPE is implemented and evaluate its effectiveness on real production jobs.
Xuepeng Fan, Hai Jin 0001, Xiaofei Liao, Hucheng Zhou, Sean McDirmid, Wei Lin 0016, Jingren Zhou 0001, Lidong Zhou
IEEE Trans. Parallel Distributed Syst.6
2014 Automating Distributed Partial Aggregation
abstract
Partial aggregation is of great importance in many distributed data-parallel systems. Most notably, it is commonly applied by MapReduce programs to optimize I/O by successively aggregating partially reduced results into a final result, as opposed to aggregating all input records at once. In spite of its importance, programmers currently enable partial aggregation by tediously encoding their reduce functionality into separate reduce and combine functions. This is error prone and often leads to missed optimization opportunities.
Chang Liu 0021, Hucheng Zhou, Sean McDirmid, Thomas Moscibroda
SoCC3
2014 Which Android App Store Can Be Trusted in China?
abstract
China has the world's largest Android population with 270 million active users. However, Google Play is only accessible by about 30% of them, and third-party app stores are thus used by 70% of them for daily Android apps (applications) discovery. The trustworthiness of Android app stores in China is still an open question. In this paper, we present a comprehensive study on the trustworthy level of top popular Android app stores in China, by discovering the identicalness and content differences between the APK files hosted in the app stores and the corresponding official APK files. First, we have selected 25 top apps that have the highest installations in China and have the corresponding official ones downloaded from their official websites as oracle, and have collected total 506 APK files across 21 top popular app stores (20 top third party stores as well as Google Play). Afterwards, APK identical checking and APK difference analysis are conducted against the corresponding official versions. Next, assessment is applied to rank the severity of APK files. All the apps are classified into 3 severity levels, ranging from safe (identical and higher level), warning (lower version or modifications on resource related files) to critical (modifications on permission file and/or application codes). Finally, the severity levels contribute to the final trustworthy ranking score of the 21 stores. The study indicates that about only 26.09% of level APK files are safe, 37.74% of them are at warning level, and 36.17% of them are surprisingly at critical level. We have also found out that 10 (about 2%) APK files are modified and resigned by unknown third-parties. In addition, the average trustworthy ranking score (47.37 over 100) has also highlighted that the trustworthy level of the Android app stores in China is relatively low. In conclusion, we suggest Android users to download APK files from its corresponding official websites or use the highest ranked third-party app stores, and we appeal app stores to ensure all hosting APK files are trustworthy enough to provide a "safe-to-download" environment.
Yiying Ng, Hucheng Zhou
COMPSAC2
2014 Cybertron: pushing the limit on I/O reduction in data-parallel programs
abstract
I/O reduction has been a major focus in optimizing data-parallel programs for big-data processing. While the current state-of-the-art techniques use static program analysis to reduce I/O, Cybertron proposes a new direction that incorporates runtime mechanisms to push the limit further on I/O reduction. In particular, Cybertron tracks how data is used in the computation accurately at runtime to filter unused data at finer granularity dynamically, beyond what current static-analysis based mechanisms are capable of, and to facilitate a new mechanism called constraint based encoding for more efficient encoding. Cybertron has been implemented and applied to production data-parallel programs; our extensive evaluations on real programs and real data have shown its effectiveness on I/O reduction over the existing mechanisms at reasonable CPU cost, and its improvement on end-to-end performance in various network environments.
Tian Xiao, Hucheng Zhou, Xu Zhao 0004, Chencheng Ye 0001, Xi Wang 0005, Wei Lin 0016, Lidong Zhou
OOPSLA3
2013 A characteristic study on failures of production distributed data-parallel programs
abstract
SCOPE is adopted by thousands of developers from tens of different product teams in Microsoft Bing for daily web-scale data processing, including index building, search ranking, and advertisement display. A SCOPE job is composed of declarative SQL-like queries and imperative C# user-defined functions (UDFs), which are executed in pipeline by thousands of machines. There are tens of thousands of SCOPE jobs executed on Microsoft clusters per day, while some of them fail after a long execution time and thus waste tremendous resources. Reducing SCOPE failures would save significant resources. This paper presents a comprehensive characteristic study on 200 SCOPE failures/fixes and 50 SCOPE failures with debugging statistics from Microsoft Bing, investigating not only major failure types, failure sources, and fixes, but also current debugging practice. Our major findings include (1) most of the failures (84.5%) are caused by defects in data processing rather than defects in code logic; (2) table-level failures (22.5%) are mainly caused by programmers' mistakes and frequent data-schema changes while row-level failures (62%) are mainly caused by exceptional data; (3) 93% fixes do not change data processing logic; (4) there are 8% failures with root cause not at the failure-exposing stage, making current debugging practice insufficient in this case. Our study results provide valuable guidelines for future development of data-parallel programs. We believe that these guidelines are not limited to SCOPE, but can also be generalized to other similar data-parallel platforms.
Hucheng Zhou, Haoxiang Lin, Tian Xiao, Wei Lin 0016, Tao Xie 0001
ICSE2
2012 Optimizing Data Shuffling in Data-Parallel Computation by Understanding User-Defined Functions
Hucheng Zhou, Rishan Chen, Xuepeng Fan, Haoxiang Lin, Jack Li 0001, Wei Lin 0016, Jingren Zhou 0001, Lidong Zhou
NSDI2
2012 Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE
Xuepeng Fan, Rishan Chen, Hucheng Zhou, Sean McDirmid, Chang Liu 0021, Wei Lin 0016, Jingren Zhou 0001, Lidong Zhou
OSDI5
2011 An SSA-based algorithm for optimal speculative code motion under an execution profile
abstract
To derive maximum optimization benefits from partial redundancy elimination (PRE),it is necessary to go beyond its safety constraint. Algorithms for optimal speculative code motion have been developed based on the application of minimum cut to flow networks formed out of the control flow graph. These previous techniques did not take advantage of the SSA form, which is a popular program representation widely used in modern-day compilers. We have developed the MC-SSAPRE algorithm that enables an SSA-based compiler to take full advantage of SSA to perform optimal speculative code motion efficiently when an execution profile is available. Our work shows that it is possible to form flow networks out of SSA graphs, and the min-cut technique can be applied equally well on these flow networks to find the optimal code placement. We provide proofs of the correctness and computational and lifetime optimality of MC-SSAPRE. We analyze its time complexity to show its efficiency advantage. We have implemented MC-SSAPRE in the open-sourced Path64 compiler. Our experimental data based on the full SPEC CPU2006 Benchmark Suite show that MC-SSAPRE can further improve program performance over traditional SSAPRE, and that our sparse approach to the problem does result in smaller problem sizes.
Hucheng Zhou, Fred C. Chow
PLDI1
2009 Prefetch optimizations on large-scale applications via parameter value prediction
abstract
A typical data center application requires the processor cycles of thousands of machines. Even a single-digit performance improvement can significantly reduce the cost and power consumption of a data center. Unfortunately, achieving sustained improvement, even if modest, is difficult. Data centers are dynamic environments where applications are frequently released and servers are continually upgraded. For maintainability and fault tolerance, the physical capabilities and configuration of the servers are abstracted from the application programmer.
Shih-Wei Liao, Tzu-Han Hung, Donald Nguyen, Hucheng Zhou, Chinyen Chou, Chia-Heng Tu
ICS4
2009 Machine learning-based prefetch optimization for data center applications
abstract
Performance tuning for data centers is essential and complicated. It is important since a data center comprises thousands of machines and thus a single-digit performance improvement can significantly reduce cost and power consumption. Unfortunately, it is extremely difficult as data centers are dynamic environments where applications are frequently released and servers are continually upgraded.
Shih-Wei Liao, Tzu-Han Hung, Donald Nguyen, Chinyen Chou, Chia-Heng Tu, Hucheng Zhou
SC6