EDBT 2026 Demo / reviewers in the wild / expert
Hucheng Zhou
dblp:75/7061
· DBLP profile ↗
20ranked-venue papers
2as first author
1since 2021 · last 2021
0000-0002-1894-3897ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 2 first-authorSystems, architecture and hardware · 6Artificial intelligence and machine learning · 3Computer networks · 3Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Energy-efficient computing · 35% High-performance computing · 20% Parallel and multicore computing · 15% | |
| Databases, data mining, and information retrieval
4 papers |
Information retrieval · 58% Data mining · 15% Recommender systems · 15% | |
| Software engineering, system software, and programming languages
7 papers |
Program analysis · 34% Compilers and program optimization · 28% Empirical software engineering · 17% | |
| Human-computer interaction and pervasive computing
2 papers |
Ubiquitous computing and smart environments · 81% Wearable and physiological sensing · 19% |
Topics — the 29 heaviest of 37, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing › power management
display power management |
0.4 | 2 | 2015 | Demo: Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015 Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015 |
Energy-efficient computing › power management › display power management
dynamic resolution scaling |
0.4 | 2 | 2015 | Demo: Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015 Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015 |
Energy-efficient computing
power management |
0.4 | 2 | 2015 | Demo: Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015 Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015 |
Information retrieval › ranking
learning to rank |
0.3 | 1 | 2018 | RapidScorer: Fast Tree Ensemble Evaluation by Maximizing Compactness in Data Level Parallelization · KDD 2018 |
Information retrieval
ranking |
0.3 | 1 | 2018 | RapidScorer: Fast Tree Ensemble Evaluation by Maximizing Compactness in Data Level Parallelization · KDD 2018 |
Information retrieval
search engines |
0.3 | 1 | 2018 | RapidScorer: Fast Tree Ensemble Evaluation by Maximizing Compactness in Data Level Parallelization · KDD 2018 |
Program analysis
static analysis |
0.3 | 3 | 2015 | Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE · IEEE Trans. Parallel Distributed Syst. 2015 Cybertron: pushing the limit on I/O reduction in data-parallel programs · OOPSLA 2014 Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE · OSDI 2012 |
High-performance computing › numerical linear algebra
matrix multiplication |
0.3 | 1 | 2017 | Improving Execution Concurrency of Large-Scale Matrix Multiplication on Distributed Data-Parallel Platforms · IEEE Trans. Parallel Distributed Syst. 2017 |
Empirical software engineering
mining software repositories |
0.3 | 2 | 2015 | An Empirical Study on Quality Issues of Production Big Data Platform · ICSE (2) 2015 A characteristic study on failures of production distributed data-parallel programs · ICSE 2013 |
Recommender systems › click-through rate prediction
feature interaction |
0.2 | 1 | 2016 | Multi-view Machines · WSDM 2016 |
Data mining
multi-view learning |
0.2 | 1 | 2016 | Multi-view Machines · WSDM 2016 |
Distributed systems
fault tolerance |
0.2 | 2 | 2015 | A characteristic study on failures of production distributed data-parallel programs · ICSE 2013 An Empirical Study on Quality Issues of Production Big Data Platform · ICSE (2) 2015 |
Software maintenance and evolution
performance diagnosis |
0.2 | 1 | 2015 | Log2: A Cost-Aware Logging Mechanism for Performance Diagnosis · USENIX ATC 2015 |
Program analysis
symbolic execution |
0.2 | 1 | 2015 | Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE · IEEE Trans. Parallel Distributed Syst. 2015 |
Cloud and datacenter computing
big data platform |
0.2 | 1 | 2015 | An Empirical Study on Quality Issues of Production Big Data Platform · ICSE (2) 2015 |
Cloud and datacenter computing
quality of service |
0.2 | 1 | 2015 | An Empirical Study on Quality Issues of Production Big Data Platform · ICSE (2) 2015 |
Parallel and multicore computing › data parallelism
data-parallel systems |
0.2 | 1 | 2013 | A characteristic study on failures of production distributed data-parallel programs · ICSE 2013 |
Storage systems › storage reliability
failure characterization |
0.2 | 1 | 2013 | A characteristic study on failures of production distributed data-parallel programs · ICSE 2013 |
Compilers and program optimization
compiler optimization |
0.1 | 1 | 2012 | Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE · OSDI 2012 |
Parallel and multicore computing
data-parallel programming |
0.1 | 1 | 2012 | Optimizing Data Shuffling in Data-Parallel Computation by Understanding User-Defined Functions · NSDI 2012 |
Distributed systems › distributed data processing
data shuffling |
0.1 | 1 | 2012 | Optimizing Data Shuffling in Data-Parallel Computation by Understanding User-Defined Functions · NSDI 2012 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2012 | Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE · OSDI 2012 |
Compilers and program optimization › compiler optimization › redundancy elimination
partial redundancy elimination |
0.1 | 1 | 2011 | An SSA-based algorithm for optimal speculative code motion under an execution profile · PLDI 2011 |
Compilers and program optimization › code motion
speculative code motion |
0.1 | 1 | 2011 | An SSA-based algorithm for optimal speculative code motion under an execution profile · PLDI 2011 |
Parallel and multicore computing › data-parallel programming
data-level parallelization |
0.1 | 1 | 2018 | RapidScorer: Fast Tree Ensemble Evaluation by Maximizing Compactness in Data Level Parallelization · KDD 2018 |
Wearable and physiological sensing › acoustic sensing
ultrasonic sensing |
0.1 | 1 | 2015 | Optimizing Smartphone Power Consumption through Dynamic Resolution Scaling · MobiCom 2015 |
Debugging and program repair
failure analysis |
0.0 | 1 | 2013 | A characteristic study on failures of production distributed data-parallel programs · ICSE 2013 |
Compilers and program optimization › intermediate representation
static single assignment form |
0.0 | 1 | 2011 | An SSA-based algorithm for optimal speculative code motion under an execution profile · PLDI 2011 |
Machine learning and data management
machine learning for systems |
0.0 | 1 | 2009 | Machine learning-based prefetch optimization for data center applications · SC 2009 |
Methods — techniques the papers use, named apart from their topics
ultrasonic distance detection · 0.9empirical study · 0.8run-length encoding · 0.7data parallelization · 0.7bitvector representation · 0.7system-level optimization · 0.6replication-based execution strategy · 0.6symbolic execution · 0.4dynamic resolution scaling · 0.4dead code elimination · 0.4factorization machines · 0.2incident management analysis · 0.2runtime profiling · 0.2constraint-based encoding · 0.2static analysis · 0.1code optimization · 0.1minimum cut · 0.1flow network · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Multiple interleaving interests modeling of sequential user behaviors in e-commerce platform
Yuqiang Han, Qian Li 0016, Hucheng Zhou, Zhenglu Yang, Jian Wu 0001 |
World Wide Web | 4 |
| 2018 | RapidScorer: Fast Tree Ensemble Evaluation by Maximizing Compactness in Data Level ParallelizationabstractRelevance ranking models based on additive ensembles of regression trees have shown quite good effectiveness in web search engines. In the era of big data, tree ensemble models grow large in both tree depth and ensemble size to provide even better search relevance and user experience. However, the computational cost for their scoring process is high, such that it becomes a challenging issue to apply the big tree ensemble models in a search engine which needs to answer thousands of queries per second. Although several works have been proposed to improve the scoring process, the challenge is still great especially when the model size grows large. In this paper, we present RapidScorer , a novel framework for speeding up the scoring process of industry-scale tree ensemble models, without hurting the quality of scoring results. RapidScorer introduces a modified run length encoding called epitome to the bitvector representation of the tree nodes. Epitome can greatly reduce the computation cost to traverse the tree ensemble, and work with several other proposed strategies to maximize the compactness of data units in memory. The achieved compactness makes it possible to fully utilize data parallelization to improve model scalability. Experiments on two web search benchmarks show that, RapidScorer achieves significant speed-up over the state-of-the-art methods: V-QuickScorer , ranging from 1.3x to 3.5x; QuickScorer , ranging from 2.1x to 25.0x; VPred , ranging from 2.3x to 18.3x; and XGBoost , ranging from 2.6x to 42.5x. Ting Ye, Hucheng Zhou, Will Y. Zou, Bin Gao 0001, Ruofei Zhang |
KDD | 2 |
| 2017 | FxpNet: Training a deep convolutional neural network in fixed-point representationabstractWe introduce FxpNet, a framework to train deep convolutional neural networks with low bit-width arithmetics in both forward pass and backward pass. During training FxpNet further reduces the bit-width of stored parameters (also known as primal parameters) by adaptively updating their fixed-point formats. These primal parameters are usually represented in the full resolution of floating-point values in previous binarized and quantized neural networks. In FxpNet, during forward pass fixed-point primal weights and activations are first binarized before computation, while in backward pass all gradients are represented as low resolution fixed-point values and then accumulated to corresponding fixed-point primal parameters. To have highly efficient implementations in FPGAs, ASICs and other dedicated devices, FxpNet introduces Integer Batch Normalization (IBN) and Fixed-point ADAM (FxpADAM) methods to further reduce the required floating-point operations, which will save considerable power and chip area. The evaluation on CIFAR-10 dataset indicates the effectiveness that FxpNet with 12-bit primal parameters and 12-bit gradients achieves comparable prediction accuracy with state-of-the-art binarized and quantized neural networks. Xi Chen 0107, Hucheng Zhou, Ningyi Xu |
IJCNN | 3 |
| 2017 | Improving Execution Concurrency of Large-Scale Matrix Multiplication on Distributed Data-Parallel PlatformsabstractMatrix multiplication is a dominant but very time-consuming operation in many big data analytic applications. Thus its performance optimization is an important and fundamental research issue. The performance of large-scale matrix multiplication on distributed data-parallel platforms is determined by both computation and IO costs. For existing matrix multiplication execution strategies, when the execution concurrency scales up above a threshold, their execution performance deteriorates quickly because the increase of the IO cost outweighs the decrease of the computation cost. This paper presents a novel parallel execution strategy CRMM (Concurrent Replication-based Matrix Multiplication) along with a parallel algorithm, Marlin, for large-scale matrix multiplication on data-parallel platforms. The CRMM strategy exploits higher execution concurrency for sub-block matrix multiplication with the same IO cost. To further improve the performance of Marlin, we also propose a number of novel system-level optimizations, including increasing the concurrency of local data exchange by calling native library in batch, reducing the overhead of block matrix transformation, and reducing disk heavy shuffle operations by exploiting the semantics of matrix computation. We have implemented Marlin as a library along with a set of related matrix operations on Spark and also contributed Marlin to the open-source community. For large-sized matrix multiplication, Marlin outperforms existing systems including Spark MLlib, SystemML and SciDB, with about 1.29×, 3.53× and 2.21× speedup on average, respectively. The evaluation upon a real-world DNN workload also indicates that Marlin outperforms above systems by about 12.8×, 5.1× and 27.2× speedup, respectively. Rong Gu 0001, Chen Tian 0001, Hucheng Zhou, Guanru Li, Yihua Huang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | The Improvement of the Trustworthiness of Android App Stores in ChinaabstractThe absence of Google Play has created a booming area for Android app distribution through third-party app stores in China. Since the study showed that the trustworthy level of app stores was fairly low in 2014, much attention should be paid on the changes of the trustworthiness of Android app stores. In this paper, we present a method to analyze the changes of trustworthiness of the top popular Android app stores in China. In this method, we evaluate the target app stores by analyzing the sampled apps hosted in them. Further more, we have used this method to track the changes of trustworthy level of Android app stores in China about two years. The results indicate that the trustworthy level of the top popular Android app stores in China has been improved 24% on average. It can be seen that the positive changes may be related to the development of the China's mobile market, the improvement of Android system, and the introduced policies. Although the trustworthy level of top popular Android app stores in China is still low, it is predicted to be improving in the future. Yiying Ng, Hucheng Zhou |
APSEC | 3 |
| 2016 | Multi-view MachinesabstractWith rapidly growing amount of data available on the web, it becomes increasingly likely to obtain data from different perspectives for multi-view learning. Some successive examples of web applications include recommendation and target advertising. Specifically, to predict whether a user will click an ad in a query context, there are available features extracted from user profile, ad information and query description, and each of them can only capture part of the task signals from a particular aspect/view. Different views provide complementary information to learn a practical model for these applications. Therefore, an effective integration of the multi-view information is critical to facilitate the learning performance. Bokai Cao, Hucheng Zhou, Guoqiang Li 0008, Philip S. Yu |
WSDM | 2 |
| 2015 | An Empirical Study on Quality Issues of Production Big Data PlatformabstractBig Data computing platform has evolved to be a multi-tenant service. The service quality matters because system failure or performance slowdown could adversely affect business and user experience. There is few study in literature on service quality issues of production Big Data computing platform. In this paper, we present an empirical study on the service quality issues of Microsoft ProductA, which is a company-wide multi-tenant Big Data computing platform, serving thousands of customers from hundreds of teams. ProductA has a well-defined incident management process, which helps customers report and mitigate service quality issues on 24/7 basis. This paper explores the common symptom, causes and mitigation of service quality issues in Big Data computing. We conduct an empirical study on 210 real service quality issues in ProductA. Our major findings include (1) 21.0% of escalations are caused by hardware faults; (2) 36.2% are caused by system side defects; (3) 37.2% are due to customer side faults. We also studied the general diagnosis process and the commonly adopted mitigation solutions. Our findings can help improve current development and maintenance practice of Big Data computing platform, and motivate tool support. Hucheng Zhou, Jian-Guang Lou, Hongyu Zhang 0002, Haoxiang Lin, Tingting Qin |
ICSE (2) | 1 |
| 2015 | Optimizing Smartphone Power Consumption through Dynamic Resolution ScalingabstractThe extremely-high display density of modern smartphones imposes a significant burden on power consumption, yet does not always provide an improved user experience and may even lead to a compromised user experience. As human visually-perceivable ability highly depends on the user-screen distance, a reduced display resolution may still achieve the same user experience when the user-screen distance is large. This provides new power-saving opportunities. In this paper, we present a flexible dynamic resolution scaling system for smartphones. The system adopts an ultrasonic-based approach to accurately detect the user-screen distance at low-power cost and makes scaling decisions automatically for maximum user experience and power saving. App developers or users can also adjust the resolution manually as their needs. Our system is able to work on existing commercial smartphones and support legacy apps, without requiring re-building the ROM or any changes of apps. An end-to-end dynamic resolution scaling system is implemented on the Galaxy S5 LTE-A and Nexus 6 smartphones, and the correctness and effectiveness are evaluated against 30 games and benchmarks. Experimental results show that all the 30 apps can run successfully with per-frame, real-time dynamic resolution scaling. The energy per frame can be reduced by 30.1% on average and up to 60.5\% at most when the resolution is halved, for 15 apps. A user study with 10 users indicates that our system remains good user experience, as none of the 10 users could perceive the resolution changes in the user study. Songtao He, Yunxin Liu 0001, Hucheng Zhou |
MobiCom | 3 |
| 2015 | Demo: Optimizing Smartphone Power Consumption through Dynamic Resolution ScalingabstractThe extremely-high display density of modern smartphones imposes a significant burden on power consumption, yet does not always provide an improved user experience and may even lead to a compromised user experience. As human visually-perceivable ability highly depends on the user-screen distance, a reduced display resolution may still achieve the same user experience when the user-screen distance is large. This provides new power-saving opportunities. We present a flexible dynamic resolution scaling system for smartphones. The system adopts an ultrasonic-based approach to detect the user-screen distance at low-power cost and makes scaling decisions automatically for maximum user experience and power saving. App developers or users can also adjust the resolution manually and dynamically as their needs. Our system is able to work on the existing commercial smartphones and support the legacy apps, without requiring re-building the ROM or any changes from apps. Songtao He, Yunxin Liu 0001, Hucheng Zhou |
MobiCom | 3 |
| 2015 | Log2: A Cost-Aware Logging Mechanism for Performance Diagnosis
Rui Ding 0001, Hucheng Zhou, Jian-Guang Lou, Hongyu Zhang 0002, Qingwei Lin, Qiang Fu 0015, Dongmei Zhang 0001, Tao Xie 0001 |
USENIX ATC | 2 |
| 2015 | Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPEabstractTo minimize the amount of data-shuffling I/O that occurs between the pipeline stages of a distributed data-parallel program, its procedural code must be optimized with full awareness of the pipeline that it executes in. Unfortunately, neither pipeline optimizers nor traditional compilers examine both the pipeline and procedural code of a data-parallel program so programmers must either hand-optimize their program across pipeline stages or live with poor performance. To resolve this tension between performance and programmability, this paper describes PeriSCOPE, which automatically optimizes a data-parallel program's procedural code in the context of data flow that is reconstructed from the program's pipeline topology. Such optimizations eliminate unnecessary code and data, perform early data filtering, and calculate small derived values (e.g., predicates) earlier in the pipeline, so that less data - sometimes much less data - is transferred between pipeline stages. PeriSCOPE further leverages symbolic execution to enlarge the scope of such optimizations by eliminating dead code. We describe how PeriSCOPE is implemented and evaluate its effectiveness on real production jobs. Xuepeng Fan, Hai Jin 0001, Xiaofei Liao, Hucheng Zhou, Sean McDirmid, Wei Lin 0016, Jingren Zhou 0001, Lidong Zhou |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2014 | Automating Distributed Partial AggregationabstractPartial aggregation is of great importance in many distributed data-parallel systems. Most notably, it is commonly applied by MapReduce programs to optimize I/O by successively aggregating partially reduced results into a final result, as opposed to aggregating all input records at once. In spite of its importance, programmers currently enable partial aggregation by tediously encoding their reduce functionality into separate reduce and combine functions. This is error prone and often leads to missed optimization opportunities. Chang Liu 0021, Hucheng Zhou, Sean McDirmid, Thomas Moscibroda |
SoCC | 3 |
| 2014 | Which Android App Store Can Be Trusted in China?abstractChina has the world's largest Android population with 270 million active users. However, Google Play is only accessible by about 30% of them, and third-party app stores are thus used by 70% of them for daily Android apps (applications) discovery. The trustworthiness of Android app stores in China is still an open question. In this paper, we present a comprehensive study on the trustworthy level of top popular Android app stores in China, by discovering the identicalness and content differences between the APK files hosted in the app stores and the corresponding official APK files. First, we have selected 25 top apps that have the highest installations in China and have the corresponding official ones downloaded from their official websites as oracle, and have collected total 506 APK files across 21 top popular app stores (20 top third party stores as well as Google Play). Afterwards, APK identical checking and APK difference analysis are conducted against the corresponding official versions. Next, assessment is applied to rank the severity of APK files. All the apps are classified into 3 severity levels, ranging from safe (identical and higher level), warning (lower version or modifications on resource related files) to critical (modifications on permission file and/or application codes). Finally, the severity levels contribute to the final trustworthy ranking score of the 21 stores. The study indicates that about only 26.09% of level APK files are safe, 37.74% of them are at warning level, and 36.17% of them are surprisingly at critical level. We have also found out that 10 (about 2%) APK files are modified and resigned by unknown third-parties. In addition, the average trustworthy ranking score (47.37 over 100) has also highlighted that the trustworthy level of the Android app stores in China is relatively low. In conclusion, we suggest Android users to download APK files from its corresponding official websites or use the highest ranked third-party app stores, and we appeal app stores to ensure all hosting APK files are trustworthy enough to provide a "safe-to-download" environment. Yiying Ng, Hucheng Zhou |
COMPSAC | 2 |
| 2014 | Cybertron: pushing the limit on I/O reduction in data-parallel programsabstractI/O reduction has been a major focus in optimizing data-parallel programs for big-data processing. While the current state-of-the-art techniques use static program analysis to reduce I/O, Cybertron proposes a new direction that incorporates runtime mechanisms to push the limit further on I/O reduction. In particular, Cybertron tracks how data is used in the computation accurately at runtime to filter unused data at finer granularity dynamically, beyond what current static-analysis based mechanisms are capable of, and to facilitate a new mechanism called constraint based encoding for more efficient encoding. Cybertron has been implemented and applied to production data-parallel programs; our extensive evaluations on real programs and real data have shown its effectiveness on I/O reduction over the existing mechanisms at reasonable CPU cost, and its improvement on end-to-end performance in various network environments. Tian Xiao, Hucheng Zhou, Xu Zhao 0004, Chencheng Ye 0001, Xi Wang 0005, Wei Lin 0016, Lidong Zhou |
OOPSLA | 3 |
| 2013 | A characteristic study on failures of production distributed data-parallel programsabstractSCOPE is adopted by thousands of developers from tens of different product teams in Microsoft Bing for daily web-scale data processing, including index building, search ranking, and advertisement display. A SCOPE job is composed of declarative SQL-like queries and imperative C# user-defined functions (UDFs), which are executed in pipeline by thousands of machines. There are tens of thousands of SCOPE jobs executed on Microsoft clusters per day, while some of them fail after a long execution time and thus waste tremendous resources. Reducing SCOPE failures would save significant resources. This paper presents a comprehensive characteristic study on 200 SCOPE failures/fixes and 50 SCOPE failures with debugging statistics from Microsoft Bing, investigating not only major failure types, failure sources, and fixes, but also current debugging practice. Our major findings include (1) most of the failures (84.5%) are caused by defects in data processing rather than defects in code logic; (2) table-level failures (22.5%) are mainly caused by programmers' mistakes and frequent data-schema changes while row-level failures (62%) are mainly caused by exceptional data; (3) 93% fixes do not change data processing logic; (4) there are 8% failures with root cause not at the failure-exposing stage, making current debugging practice insufficient in this case. Our study results provide valuable guidelines for future development of data-parallel programs. We believe that these guidelines are not limited to SCOPE, but can also be generalized to other similar data-parallel platforms. Hucheng Zhou, Haoxiang Lin, Tian Xiao, Wei Lin 0016, Tao Xie 0001 |
ICSE | 2 |
| 2012 | Optimizing Data Shuffling in Data-Parallel Computation by Understanding User-Defined Functions
Hucheng Zhou, Rishan Chen, Xuepeng Fan, Haoxiang Lin, Jack Li 0001, Wei Lin 0016, Jingren Zhou 0001, Lidong Zhou |
NSDI | 2 |
| 2012 | Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE
Xuepeng Fan, Rishan Chen, Hucheng Zhou, Sean McDirmid, Chang Liu 0021, Wei Lin 0016, Jingren Zhou 0001, Lidong Zhou |
OSDI | 5 |
| 2011 | An SSA-based algorithm for optimal speculative code motion under an execution profileabstractTo derive maximum optimization benefits from partial redundancy elimination (PRE),it is necessary to go beyond its safety constraint. Algorithms for optimal speculative code motion have been developed based on the application of minimum cut to flow networks formed out of the control flow graph. These previous techniques did not take advantage of the SSA form, which is a popular program representation widely used in modern-day compilers. We have developed the MC-SSAPRE algorithm that enables an SSA-based compiler to take full advantage of SSA to perform optimal speculative code motion efficiently when an execution profile is available. Our work shows that it is possible to form flow networks out of SSA graphs, and the min-cut technique can be applied equally well on these flow networks to find the optimal code placement. We provide proofs of the correctness and computational and lifetime optimality of MC-SSAPRE. We analyze its time complexity to show its efficiency advantage. We have implemented MC-SSAPRE in the open-sourced Path64 compiler. Our experimental data based on the full SPEC CPU2006 Benchmark Suite show that MC-SSAPRE can further improve program performance over traditional SSAPRE, and that our sparse approach to the problem does result in smaller problem sizes. Hucheng Zhou, Fred C. Chow |
PLDI | 1 |
| 2009 | Prefetch optimizations on large-scale applications via parameter value predictionabstractA typical data center application requires the processor cycles of thousands of machines. Even a single-digit performance improvement can significantly reduce the cost and power consumption of a data center. Unfortunately, achieving sustained improvement, even if modest, is difficult. Data centers are dynamic environments where applications are frequently released and servers are continually upgraded. For maintainability and fault tolerance, the physical capabilities and configuration of the servers are abstracted from the application programmer. Shih-Wei Liao, Tzu-Han Hung, Donald Nguyen, Hucheng Zhou, Chinyen Chou, Chia-Heng Tu |
ICS | 4 |
| 2009 | Machine learning-based prefetch optimization for data center applicationsabstractPerformance tuning for data centers is essential and complicated. It is important since a data center comprises thousands of machines and thus a single-digit performance improvement can significantly reduce cost and power consumption. Unfortunately, it is extremely difficult as data centers are dynamic environments where applications are frequently released and servers are continually upgraded. Shih-Wei Liao, Tzu-Han Hung, Donald Nguyen, Chinyen Chou, Chia-Heng Tu, Hucheng Zhou |
SC | 6 |