Shuangde Fang

dblp:85/11029 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Cloud and datacenter computing · 28% GPUs and heterogeneous computing · 20% Hardware accelerators and domain-specific architectures · 20%
Software engineering, system software, and programming languages
4 papers
Compilers and program optimization · 93% Empirical software engineering · 7%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
datacenter workloads
0.322015
Practical Iterative Optimization for the Data Center · ACM Trans. Archit. Code Optim. 2015
Iterative optimization for the data center · ASPLOS 2012
Hardware accelerators and domain-specific architectures › accelerator orchestration
accelerator selection
0.212014
Performance Portability Across Heterogeneous SoCs Using a Generalized Library-Based Approach · ACM Trans. Archit. Code Optim. 2014
Storage systems
hardware parameter tuning
0.212014
Performance Portability Across Heterogeneous SoCs Using a Generalized Library-Based Approach · ACM Trans. Archit. Code Optim. 2014
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous soc
0.212014
Performance Portability Across Heterogeneous SoCs Using a Generalized Library-Based Approach · ACM Trans. Archit. Code Optim. 2014
Parallel and multicore computing › data-parallel programming
mapreduce
0.112015
Practical Iterative Optimization for the Data Center · ACM Trans. Archit. Code Optim. 2015
Performance modeling and evaluation
benchmarking
0.012012
Deconstructing iterative optimization · ACM Trans. Archit. Code Optim. 2012

Methods — techniques the papers use, named apart from their topics

iterative optimization · 0.7dynamic aggressiveness adjustment · 0.4simulated annealing · 0.4machine learning · 0.4dataset suite construction · 0.3case study · 0.3
YearPublicationVenuePosition
2015 Practical Iterative Optimization for the Data Center
abstract
Iterative optimization is a simple but powerful approach that searches the best possible combination of compiler optimizations for a given workload. However, iterative optimization is plagued by several practical issues that prevent it from being widely used in practice: a large number of runs are required to find the best combination, the optimum combination is dataset dependent, and the exploration process incurs significant overhead that needs to be compensated for by performance benefits. Therefore, although iterative optimization has been shown to have a significant performance potential, it seldom is used in production compilers. In this article, we propose iterative optimization for the data center (IODC): we show that the data center offers a context in which all of the preceding hurdles can be overcome. The basic idea is to spawn different combinations across workers and recollect performance statistics at the master, which then evolves to the optimum combination of compiler optimizations. IODC carefully manages costs and benefits, and it is transparent to the end user. To bring IODC to practice, we evaluate it in the presence of co-runners to better reflect real-life data center operation with multiple applications co-running per server. We enhance IODC with the capability to find compatible co-runners along with a mechanism to dynamically adjust the level of aggressiveness to improve its robustness in the presence of co-running applications. We evaluate IODC using both MapReduce and compute-intensive throughput server applications. To reflect the large number of users interacting with the system, we gather a very large collection of datasets (up to hundreds of millions of unique datasets per program), for a total storage of 16.4TB and 850 days of CPU time. We report an average performance improvement of 1.48 × and up to 2.08 × for five MapReduce applications, and 1.12 × and up to 1.39 × for nine server applications. Furthermore, our experiments demonstrate that IODC is effective in the presence of co-runners, improving performance by greater than 13% compared to the worst possible co-runner schedule.
Shuangde Fang, Lieven Eeckhout, Olivier Temam, Yunji Chen, Chengyong Wu, Xiaobing Feng 0002
ACM Trans. Archit. Code Optim.1
2014 Performance Portability Across Heterogeneous SoCs Using a Generalized Library-Based Approach
abstract
Because of tight power and energy constraints, industry is progressively shifting toward heterogeneous system-on-chip (SoC) architectures composed of a mix of general-purpose cores along with a number of accelerators. However, such SoC architectures can be very challenging to efficiently program for the vast majority of programmers, due to numerous programming approaches and languages. Libraries, on the other hand, provide a simple way to let programmers take advantage of complex architectures, which does not require programmers to acquire new accelerator-specific or domain-specific languages. Increasingly, library-based, also called algorithm-centric, programming approaches propose to generalize the usage of libraries and to compose programs around these libraries, instead of using libraries as mere complements. In this article, we present a software framework for achieving performance portability by leveraging a generalized library-based approach. Inspired by the notion of a component, as employed in software engineering and HW/SW codesign, we advocate nonexpert programmers to write simple wrapper code around existing libraries to provide simple but necessary semantic information to the runtime. To achieve performance portability, the runtime employs machine learning (simulated annealing) to select the most appropriate accelerator and its parameters for a given algorithm. This selection factors in the possibly complex composition of algorithms used in the application, the communication among the various accelerators, and the tradeoff between different objectives (i.e., accuracy, performance, and energy). Using a set of benchmarks run on a real heterogeneous SoC composed of a multicore processor and a GPU, we show that the runtime overhead is fairly small at 5.1% for the GPU and 6.4% for the multi-core. We then apply our accelerator selection approach to a simulated SoC platform containing multiple inexact accelerators. We show that accelerator selection together with hardware parameter tuning achieves an average 46.2% energy reduction and a speedup of 2.1× while meeting the desired application error target.
Shuangde Fang, Zidong Du, Yuntan Fang, Yuanjie Huang, Lieven Eeckhout, Olivier Temam, Huawei Li 0001, Yunji Chen, Chengyong Wu
ACM Trans. Archit. Code Optim.1
2012 Iterative optimization for the data center
abstract
Iterative optimization is a simple but powerful approach that searches for the best possible combination of compiler optimizations for a given workload. However, each program, if not each data set, potentially favors a different combination. As a result, iterative optimization is plagued by several practical issues that prevent it from being widely used in practice: a large number of runs are required for finding the best combination; the process can be data set dependent; and the exploration process incurs significant overhead that needs to be compensated for by performance benefits.Therefore, while iterative optimization has been shown to have significant performance potential, it is seldomly used in production compilers.
Shuangde Fang, Lieven Eeckhout, Olivier Temam, Chengyong Wu
ASPLOS2
2012 Deconstructing iterative optimization
abstract
Iterative optimization is a popular compiler optimization approach that has been studied extensively over the past decade. In this article, we deconstruct iterative optimization by evaluating whether it works across datasets and by analyzing why it works. Up to now, most iterative optimization studies are based on a premise which was never truly evaluated: that it is possible to learn the best compiler optimizations across datasets. In this article, we evaluate this question for the first time with a very large number of datasets. We therefore compose KDataSets, a dataset suite with 1000 datasets for 32 programs, which we release to the public. We characterize the diversity of KDataSets, and subsequently use it to evaluate iterative optimization. For all 32 programs, we find that there exists at least one combination of compiler optimizations that achieves at least 83% or more of the best possible speedup across all datasets on two widely used compilers (Intel's ICC and GNU's GCC). This optimal combination is program-specific and yields speedups up to 3.75× (averaged across datasets of a program) over the highest optimization level of the compilers (-O3 for GCC and -fast for ICC). This finding suggests that optimizing programs across datasets might be much easier than previously anticipated. In addition, we evaluate the idea of introducing compiler choice as part of iterative optimization. We find that it can further improve the performance of iterative optimization because different programs favor different compilers. We also investigate why iterative optimization works by analyzing the optimal combinations. We find that only a handful optimizations yield most of the speedup. Finally, we show that optimizations interact in a complex and sometimes counterintuitive way through two case studies, which confirms that iterative optimization is an irreplaceable and important compiler strategy.
Shuangde Fang, Yuanjie Huang, Lieven Eeckhout, Grigori Fursin, Olivier Temam, Chengyong Wu
ACM Trans. Archit. Code Optim.2