VLDB 2026 Research / reviewers in the wild / expert
Chengyong Wu
dblp:47/382
· DBLP profile ↗
19ranked-venue papers
0as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17Software engineering, systems software and programming languages · 5Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
Hardware accelerators and domain-specific architectures · 34% Memory systems · 21% Emerging computing paradigms · 18% | |
| Software engineering, system software, and programming languages
8 papers |
Compilers and program optimization · 56% Operating systems · 31% Empirical software engineering · 13% |
Topics — the 25 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.8 | 4 | 2015 | A Small-Footprint Accelerator for Large-Scale Neural Networks · ACM Trans. Comput. Syst. 2015 Leveraging the Error Resilience of Neural Networks for Designing Highly Energy Efficient Accelerators · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 Neuromorphic accelerators: a comparison between neuroscience and machine-learning approaches · MICRO 2015 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 2 | 2015 | A Small-Footprint Accelerator for Large-Scale Neural Networks · ACM Trans. Comput. Syst. 2015 DianNao: a small-footprint high-throughput accelerator for ubiquitous machine-learning · ASPLOS 2014 |
Operating systems › resource management
memory management |
0.4 | 3 | 2016 | Rethinking Memory Management in Modern Operating System: Horizontal, Vertical or Random? · IEEE Trans. Computers 2016 BPM/BPM+: Software-based dynamic memory partitioning mechanisms for mitigating DRAM bank-/channel-level interferences in multicore systems · ACM Trans. Archit. Code Optim. 2014 Going vertical in memory management: Handling multiplicity by multi-policy · ISCA 2014 |
Cloud and datacenter computing
datacenter workloads |
0.3 | 2 | 2015 | Practical Iterative Optimization for the Data Center · ACM Trans. Archit. Code Optim. 2015 Iterative optimization for the data center · ASPLOS 2012 |
Memory systems
memory hierarchy |
0.2 | 1 | 2016 | Rethinking Memory Management in Modern Operating System: Horizontal, Vertical or Random? · IEEE Trans. Computers 2016 |
Emerging computing paradigms
approximate computing |
0.2 | 1 | 2015 | Leveraging the Error Resilience of Neural Networks for Designing Highly Energy Efficient Accelerators · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Hardware accelerators and domain-specific architectures › approximate computing accelerator
approximate neural network accelerator |
0.2 | 1 | 2015 | Leveraging the Error Resilience of Neural Networks for Designing Highly Energy Efficient Accelerators · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Emerging computing paradigms › approximate computing
inexact computing |
0.2 | 1 | 2015 | Leveraging the Error Resilience of Neural Networks for Designing Highly Energy Efficient Accelerators · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Emerging computing paradigms
neuromorphic computing |
0.2 | 1 | 2015 | Neuromorphic accelerators: a comparison between neuroscience and machine-learning approaches · MICRO 2015 |
Emerging computing paradigms
neuromorphic hardware |
0.2 | 1 | 2015 | Neuromorphic accelerators: a comparison between neuroscience and machine-learning approaches · MICRO 2015 |
Hardware accelerators and domain-specific architectures › accelerator orchestration
accelerator selection |
0.2 | 1 | 2014 | Performance Portability Across Heterogeneous SoCs Using a Generalized Library-Based Approach · ACM Trans. Archit. Code Optim. 2014 |
Storage systems
hardware parameter tuning |
0.2 | 1 | 2014 | Performance Portability Across Heterogeneous SoCs Using a Generalized Library-Based Approach · ACM Trans. Archit. Code Optim. 2014 |
GPUs and heterogeneous computing › heterogeneous architecture
heterogeneous soc |
0.2 | 1 | 2014 | Performance Portability Across Heterogeneous SoCs Using a Generalized Library-Based Approach · ACM Trans. Archit. Code Optim. 2014 |
Memory systems
memory management |
0.2 | 1 | 2014 | Going vertical in memory management: Handling multiplicity by multi-policy · ISCA 2014 |
Memory systems › memory controller
memory scheduling |
0.2 | 1 | 2014 | BPM/BPM+: Software-based dynamic memory partitioning mechanisms for mitigating DRAM bank-/channel-level interferences in multicore systems · ACM Trans. Archit. Code Optim. 2014 |
Memory systems › virtual memory management
page coloring |
0.2 | 1 | 2014 | BPM/BPM+: Software-based dynamic memory partitioning mechanisms for mitigating DRAM bank-/channel-level interferences in multicore systems · ACM Trans. Archit. Code Optim. 2014 |
Storage systems › data placement
vertical partitioning |
0.2 | 1 | 2014 | Going vertical in memory management: Handling multiplicity by multi-policy · ISCA 2014 |
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture |
0.2 | 1 | 2013 | Elastic CGRAs · FPGA 2013 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 2 | 2015 | A Small-Footprint Accelerator for Large-Scale Neural Networks · ACM Trans. Comput. Syst. 2015 DianNao: a small-footprint high-throughput accelerator for ubiquitous machine-learning · ASPLOS 2014 |
Empirical software engineering › software engineering research methodology
empirical study |
0.1 | 1 | 2010 | Evaluating iterative optimization across 1000 datasets · PLDI 2010 |
Performance modeling and evaluation
benchmarking |
0.1 | 2 | 2012 | Deconstructing iterative optimization · ACM Trans. Archit. Code Optim. 2012 Evaluating iterative optimization across 1000 datasets · PLDI 2010 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2016 | Rethinking Memory Management in Modern Operating System: Horizontal, Vertical or Random? · IEEE Trans. Computers 2016 |
Hardware reliability and fault tolerance
error resilience |
0.1 | 1 | 2015 | Leveraging the Error Resilience of Neural Networks for Designing Highly Energy Efficient Accelerators · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Parallel and multicore computing › data-parallel programming
mapreduce |
0.1 | 1 | 2015 | Practical Iterative Optimization for the Data Center · ACM Trans. Archit. Code Optim. 2015 |
Processor architecture and microarchitecture
superscalar processor |
0.0 | 1 | 2013 | Elastic CGRAs · FPGA 2013 |
Methods — techniques the papers use, named apart from their topics
iterative optimization · 0.7workload correlation analysis · 0.5kernel implementation · 0.5layout at 65nm · 0.4dynamic aggressiveness adjustment · 0.4ASIC design · 0.4neural network training · 0.2layout comparison · 0.2inexact computing · 0.2energy and area analysis · 0.2workload characterization · 0.2simulated annealing · 0.2performance-monitoring unit monitoring · 0.2machine learning · 0.2kernel module · 0.2data mining · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Pragma Directed Shared Memory Centric Optimizations on GPUs
Lei Liu 0030, Xiang-Hua Liu, Xiaobing Feng 0002, Chengyong Wu |
J. Comput. Sci. Technol. | 7 |
| 2016 | Rethinking Memory Management in Modern Operating System: Horizontal, Vertical or Random?abstractOn modern multicore machines, the memory management typically combines address interleaving in hardware and random allocation in the operating system (OS) to improve performance of both memory and cache. The conventional solutions, however, are increasingly strained as a wide variety of workloads run on complicated memory hierarchy and cause contention at multiple levels. We describe a new framework (named HVR) in OS memory management to support a flexible policy space for tackling diverse application needs, integrating vertical partitioning across layers, horizontal partitioning and random-interleaved allocation at a single layer. We exhaustively study the performance of these policies for over 2,000 workloads and correlate performance with application characteristics. Based on this correlation we derive several practical rules of memory allocation that we integrate into the unified HVR framework to guide resource partitioning and sharing for dynamic and diverse workloads. We implement our approach in Linux kernel 2.6.32 as a restructured page indexing system plus a series of kernel modules. Experimental results show that our framework consistently outperforms the unmodified Linux kernel, with up to 21 percent performance gains, and outperforms prior solutions at individual levels of the memory hierarchy. Lei Liu 0030, Chen Ding 0001, Chengyong Wu |
IEEE Trans. Computers | 5 |
| 2015 | Retraining-based timing error mitigation for hardware neural networks
Jiachao Deng, Yuntan Fang, Zidong Du, Ying Wang 0001, Huawei Li 0001, Olivier Temam, Paolo Ienne, David Novo, Xiaowei Li 0001, Yunji Chen, Chengyong Wu |
DATE | 11 |
| 2015 | Neuromorphic accelerators: a comparison between neuroscience and machine-learning approachesabstractA vast array of devices, ranging from industrial robots to self-driven cars or smartphones, require increasingly sophisticated processing of real-world input data (image, voice, radio, ...). Interestingly, hardware neural network accelerators are emerging again as attractive candidate architectures for such tasks. The neural network algorithms considered come from two, largely separate, domains: machine-learning and neuroscience. These neural networks have very different characteristics, so it is unclear which approach should be favored for hardware implementation. Yet, few studies compare them from a hardware perspective. We implement both types of networks down to the layout, and we compare the relative merit of each approach in terms of energy, speed, area cost, accuracy and functionality. Zidong Du, Daniel Ben Dayan Rubin, Yunji Chen, Liqiang He, Tianshi Chen 0002, Lei Zhang 0008, Chengyong Wu, Olivier Temam |
MICRO | 7 |
| 2015 | Practical Iterative Optimization for the Data CenterabstractIterative optimization is a simple but powerful approach that searches the best possible combination of compiler optimizations for a given workload. However, iterative optimization is plagued by several practical issues that prevent it from being widely used in practice: a large number of runs are required to find the best combination, the optimum combination is dataset dependent, and the exploration process incurs significant overhead that needs to be compensated for by performance benefits. Therefore, although iterative optimization has been shown to have a significant performance potential, it seldom is used in production compilers. In this article, we propose iterative optimization for the data center (IODC): we show that the data center offers a context in which all of the preceding hurdles can be overcome. The basic idea is to spawn different combinations across workers and recollect performance statistics at the master, which then evolves to the optimum combination of compiler optimizations. IODC carefully manages costs and benefits, and it is transparent to the end user. To bring IODC to practice, we evaluate it in the presence of co-runners to better reflect real-life data center operation with multiple applications co-running per server. We enhance IODC with the capability to find compatible co-runners along with a mechanism to dynamically adjust the level of aggressiveness to improve its robustness in the presence of co-running applications. We evaluate IODC using both MapReduce and compute-intensive throughput server applications. To reflect the large number of users interacting with the system, we gather a very large collection of datasets (up to hundreds of millions of unique datasets per program), for a total storage of 16.4TB and 850 days of CPU time. We report an average performance improvement of 1.48 × and up to 2.08 × for five MapReduce applications, and 1.12 × and up to 1.39 × for nine server applications. Furthermore, our experiments demonstrate that IODC is effective in the presence of co-runners, improving performance by greater than 13% compared to the worst possible co-runner schedule. Shuangde Fang, Lieven Eeckhout, Olivier Temam, Yunji Chen, Chengyong Wu, Xiaobing Feng 0002 |
ACM Trans. Archit. Code Optim. | 7 |
| 2015 | Leveraging the Error Resilience of Neural Networks for Designing Highly Energy Efficient AcceleratorsabstractIn recent years, inexact computing has been increasingly regarded as one of the most promising approaches for slashing energy consumption in many applications that can tolerate a certain degree of inaccuracy. Driven by the principle of trading tolerable amounts of application accuracy in return for significant resource savings-the energy consumed, the (critical path) delay, and the (silicon) area-this approach has been limited to application-specified integrated circuits (ASICs) so far. These ASIC realizations have a narrow application scope and are often rigid in their tolerance to inaccuracy, as currently designed; the latter often determining the extent of resource savings we would achieve. In this paper, we propose to improve the application scope, error resilience and the energy savings of inexact computing by combining it with hardware neural networks. These neural networks are fast emerging as popular candidate accelerators for future heterogeneous multicore platforms and have flexible error resilience limits owing to their ability to be trained. Our results in 65-nm technology demonstrate that the proposed inexact neural network accelerator could achieve 1.78-2.67× savings in energy consumption (with corresponding delay and area savings being 1.23 and 1.46×, respectively) when compared to the existing baseline neural network implementation, at the cost of a small accuracy loss (mean squared error increases from 0.14 to 0.20 on average). Zidong Du, Lingamneni Avinash, Yunji Chen, Krishna V. Palem, Olivier Temam, Chengyong Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2015 | A Small-Footprint Accelerator for Large-Scale Neural NetworksabstractMachine-learning tasks are becoming pervasive in a broad range of domains, and in a broad range of systems (from embedded systems to data centers). At the same time, a small set of machine-learning algorithms (especially Convolutional and Deep Neural Networks, i.e., CNNs and DNNs) are proving to be state-of-the-art across many applications. As architectures evolve toward heterogeneous multicores composed of a mix of cores and accelerators, a machine-learning accelerator can achieve the rare combination of efficiency (due to the small number of target algorithms) and broad application scope. Until now, most machine-learning accelerator designs have been focusing on efficiently implementing the computational part of the algorithms. However, recent state-of-the-art CNNs and DNNs are characterized by their large size. In this study, we design an accelerator for large-scale CNNs and DNNs, with a special emphasis on the impact of memory on accelerator design, performance, and energy. We show that it is possible to design an accelerator with a high throughput, capable of performing 452 GOP/s (key NN operations such as synaptic weight multiplications and neurons outputs additions) in a small footprint of 3.02mm2 and 485mW; compared to a 128-bit 2GHz SIMD processor, the accelerator is 117.87 × faster, and it can reduce the total energy by 21.08 ×. The accelerator characteristics are obtained after layout at 65nm. Such a high throughput in a small footprint can open up the usage of state-of-the-art machine-learning algorithms in a broad set of systems and for a broad set of applications. Tianshi Chen 0002, Shijin Zhang, Shaoli Liu, Zidong Du, Dongsheng Wang 0002, Chengyong Wu, Ninghui Sun, Yunji Chen, Olivier Temam |
ACM Trans. Comput. Syst. | 9 |
| 2014 | Leveraging the error resilience of machine-learning applications for designing highly energy efficient acceleratorsabstractIn recent years, inexact computing has been increasingly regarded as one of the most promising approaches for reducing energy consumption in many applications that can tolerate a degree of inaccuracy. Driven by the principle of trading tolerable amounts of application accuracy in return for significant resource savings - the energy consumed, the (critical path) delay and the (silicon) area being the resources - this approach has been limited to certain application domains. In this paper, we propose to expand the application scope, error tolerance as well as the energy savings of inexact computing systems through neural network architectures. Such neural networks are fast emerging as popular candidate accelerators for future heterogeneous multi-core platforms, and have flexible error tolerance limits owing to their ability to be trained. Our results based on simulated 65nm technology designs demonstrate that the proposed inexact neural network accelerator could achieve 43.91%-62.49% savings in energy consumption (with corresponding delay and area savings being 18.79% and 31.44% respectively) when compared to existing baseline neural network implementation, at the cost of an accuracy loss (quantified as the Mean Square Error (MSE) which increases from 0.14 to 0.20 on average). Zidong Du, Krishna V. Palem, Lingamneni Avinash, Olivier Temam, Yunji Chen, Chengyong Wu |
ASP-DAC | 6 |
| 2014 | DianNao: a small-footprint high-throughput accelerator for ubiquitous machine-learningabstractMachine-Learning tasks are becoming pervasive in a broad range of domains, and in a broad range of systems (from embedded systems to data centers). At the same time, a small set of machine-learning algorithms (especially Convolutional and Deep Neural Networks, i.e., CNNs and DNNs) are proving to be state-of-the-art across many applications. As architectures evolve towards heterogeneous multi-cores composed of a mix of cores and accelerators, a machine-learning accelerator can achieve the rare combination of efficiency (due to the small number of target algorithms) and broad application scope. Tianshi Chen 0002, Zidong Du, Ninghui Sun, Chengyong Wu, Yunji Chen, Olivier Temam |
ASPLOS | 5 |
| 2014 | A low-cost memory interface for high-throughput acceleratorsabstractHeterogeneous multi-cores, a mix of cores and accelerators, are becoming prevalent. These accelerators are designed for both speed and energy improvements, and thus, they increasingly come with a large number of load/store ports for achieving a high degree of parallelism. However, beyond GPG-PUs, accelerators such as ASICs and CGRAs are increasingly capable of accelerating computations with irregular control flow and memory accesses; as a result, such accelerators need to be plugged to caches instead of scratchpads, and few studies focus on accelerator-to-cache interfaces. The main existing alternative are Load/Store Queues (LSQs) traditionally used to connect superscalar processors to caches and memory, but in the context of accelerators, they are overkill and could significantly reduce the area and power benefits of accelerators. Moreover, we show that they are just not fit for accelerators plugged to multi-banked caches. Yuanjie Huang, Olivier Temam, Paolo Ienne, Yunji Chen, Chengyong Wu |
CASES | 6 |
| 2014 | Going vertical in memory management: Handling multiplicity by multi-policyabstractMany emerging applications from various domains often exhibit heterogeneous memory characteristics. When running in combination on parallel platforms, these applications present a daunting variety of workload behaviors that challenge the effectiveness of any memory allocation strategy. Prior partitioning-based or random memory allocation schemes typically manage only one level of the memory hierarchy and often target specific workloads. To handle diverse and dynamically changing memory and cache allocation needs, we augment existing “horizontal” cache/DRAM bank partitioning with vertical partitioning and explore the resulting multi-policy space. We study the performance of these policies for over 2000 workloads and correlate the results with application characteristics via a data mining approach. Based on this correlation we derive several practical memory allocation rules that we integrate into a unified multi-policy framework to guide resources partitioning and coalescing for dynamic and diverse multi-programmed/threaded workloads. We implement our approach in Linux kernel 2.6.32 as a restructured page indexing system plus a series of kernel modules. Extensive experiments show that, in practice, our framework can select proper memory allocation policy and consistently outperforms the unmodified Linux kernel, achieving up to 11% performance gains compared to prior techniques. Lei Liu 0030, Zehan Cui, Yungang Bao, Mingyu Chen 0001, Chengyong Wu |
ISCA | 6 |
| 2014 | Performance Portability Across Heterogeneous SoCs Using a Generalized Library-Based ApproachabstractBecause of tight power and energy constraints, industry is progressively shifting toward heterogeneous system-on-chip (SoC) architectures composed of a mix of general-purpose cores along with a number of accelerators. However, such SoC architectures can be very challenging to efficiently program for the vast majority of programmers, due to numerous programming approaches and languages. Libraries, on the other hand, provide a simple way to let programmers take advantage of complex architectures, which does not require programmers to acquire new accelerator-specific or domain-specific languages. Increasingly, library-based, also called algorithm-centric, programming approaches propose to generalize the usage of libraries and to compose programs around these libraries, instead of using libraries as mere complements. In this article, we present a software framework for achieving performance portability by leveraging a generalized library-based approach. Inspired by the notion of a component, as employed in software engineering and HW/SW codesign, we advocate nonexpert programmers to write simple wrapper code around existing libraries to provide simple but necessary semantic information to the runtime. To achieve performance portability, the runtime employs machine learning (simulated annealing) to select the most appropriate accelerator and its parameters for a given algorithm. This selection factors in the possibly complex composition of algorithms used in the application, the communication among the various accelerators, and the tradeoff between different objectives (i.e., accuracy, performance, and energy). Using a set of benchmarks run on a real heterogeneous SoC composed of a multicore processor and a GPU, we show that the runtime overhead is fairly small at 5.1% for the GPU and 6.4% for the multi-core. We then apply our accelerator selection approach to a simulated SoC platform containing multiple inexact accelerators. We show that accelerator selection together with hardware parameter tuning achieves an average 46.2% energy reduction and a speedup of 2.1× while meeting the desired application error target. Shuangde Fang, Zidong Du, Yuntan Fang, Yuanjie Huang, Lieven Eeckhout, Olivier Temam, Huawei Li 0001, Yunji Chen, Chengyong Wu |
ACM Trans. Archit. Code Optim. | 10 |
| 2014 | BPM/BPM+: Software-based dynamic memory partitioning mechanisms for mitigating DRAM bank-/channel-level interferences in multicore systemsabstractThe main memory system is a shared resource in modern multicore machines that can result in serious interference leading to reduced throughput and unfairness. Many new memory scheduling mechanisms have been proposed to address the interference problem. However, these mechanisms usually employ relative complex scheduling logic and need modifications to Memory Controllers (MCs), which incur expensive hardware design and manufacturing overheads. This article presents a practical software approach to effectively eliminate the interference without any hardware modifications. The key idea is to modify the OS memory management system and adopt a page-coloring-based Bank-level Partitioning Mechanism (BPM) that allocates dedicated DRAM banks to each core (or thread). By using BPM, memory requests from distinct programs are segregated across multiple memory banks to promote locality/fairness and reduce interference. We further extend BPM to BPM+ by incorporating channel-level partitioning, on which we demonstrate additional gain over BPM in many cases. To achieve benefits in the presence of diverse application memory needs and avoid performance degradation due to resource underutilization, we propose a dynamic mechanism upon BPM/BPM+ that assigns appropriate bank/channel resources based on application memory/bandwidth demands monitored through PMU (performance-monitoring unit) and a low-overhead OS page table scanning process. We implement BPM/BPM+ in Linux 2.6.32.15 kernel and evaluate the technique on four-core and eight-core real machines by running a large amount of randomly generated multiprogrammed and multithreaded workloads. Experimental results show that BPM/BPM+ can improve the overall system throughput by 4.7%/5.9%, on average, (up to 8.6%/9.5%) and reduce the unfairness by an average of 4.2%/6.1% (up to 15.8%/13.9%). Lei Liu 0030, Zehan Cui, Yungang Bao, Mingyu Chen 0001, Chengyong Wu |
ACM Trans. Archit. Code Optim. | 6 |
| 2013 | Elastic CGRAsabstractVital technology trends such as voltage scaling and homogeneous multicore scaling have reached their limits and architects turn to alternate computing paradigms, such as heterogeneous and domain-specialized solutions. Coarse-Grain Reconfigurable Arrays (CGRAs) promise the performance of massively spatial computing while offering interesting trade-offs of flexibility versus energy efficiency. Yet, configuring and scheduling execution for CGRAs generally runs into the classic difficulties that have hampered Very-Long Instruction Word (VLIW) architectures: efficient schedules are difficult to generate, especially for applications with complex control flow and data structures, and they are inherently static - thus, in adapted to variable-latency components (such as the read ports of caches). Over the years, VLIWs have been relegated to important but specific application domains where such issues are more under the control of the designers; similarly, statically-scheduled CGRAs may prove inadequate for future general-purpose computing systems. In this paper, we introduce Elastic CGRAs, the superscalar processors of computing fabrics: no complex schedule needs to be computed at configuration time, and the operations execute dynamically in the CGRA when data are ready, thus exploiting the data parallelism that an application offers. We designed, down to a manufacturable layout, a simple CGRA where we demonstrated and optimized our elastic control circuitry. We also built a complete compilation toolchain that transforms arbitrary C code in a configuration for the array. The area overhead (26.2%), critical path overhead (8.2%) and energy overhead (53.6%) of Elastic CGRAs over non-elastic CGRAs are significantly lower than the overhead of superscalar processors over VLIWs, while providing the same benefits. At such moderate costs, elasticity may prove to be one of the key enablers to make the adoption of CGRAs widespread. Yuanjie Huang, Paolo Ienne, Olivier Temam, Yunji Chen, Chengyong Wu |
FPGA | 5 |
| 2012 | A software memory partition approach for eliminating bank-level interference in multicore systemsabstractMain memory system is a shared resource in modern multicore machines, resulting in serious interference, which causes performance degradation in terms of throughput slowdown and unfairness. Numerous new memory scheduling algorithms have been proposed to address the interference problem. However, these algorithms usually employ complex scheduling logic and need hardware modification to memory controllers, as a result, industrial venders seem to have some hesitation in adopting them. Lei Liu 0030, Zehan Cui, Mingjie Xing, Yungang Bao, Mingyu Chen 0001, Chengyong Wu |
PACT | 6 |
| 2012 | Iterative optimization for the data centerabstractIterative optimization is a simple but powerful approach that searches for the best possible combination of compiler optimizations for a given workload. However, each program, if not each data set, potentially favors a different combination. As a result, iterative optimization is plagued by several practical issues that prevent it from being widely used in practice: a large number of runs are required for finding the best combination; the process can be data set dependent; and the exploration process incurs significant overhead that needs to be compensated for by performance benefits.Therefore, while iterative optimization has been shown to have significant performance potential, it is seldomly used in production compilers. Shuangde Fang, Lieven Eeckhout, Olivier Temam, Chengyong Wu |
ASPLOS | 5 |
| 2012 | Deconstructing iterative optimizationabstractIterative optimization is a popular compiler optimization approach that has been studied extensively over the past decade. In this article, we deconstruct iterative optimization by evaluating whether it works across datasets and by analyzing why it works. Up to now, most iterative optimization studies are based on a premise which was never truly evaluated: that it is possible to learn the best compiler optimizations across datasets. In this article, we evaluate this question for the first time with a very large number of datasets. We therefore compose KDataSets, a dataset suite with 1000 datasets for 32 programs, which we release to the public. We characterize the diversity of KDataSets, and subsequently use it to evaluate iterative optimization. For all 32 programs, we find that there exists at least one combination of compiler optimizations that achieves at least 83% or more of the best possible speedup across all datasets on two widely used compilers (Intel's ICC and GNU's GCC). This optimal combination is program-specific and yields speedups up to 3.75× (averaged across datasets of a program) over the highest optimization level of the compilers (-O3 for GCC and -fast for ICC). This finding suggests that optimizing programs across datasets might be much easier than previously anticipated. In addition, we evaluate the idea of introducing compiler choice as part of iterative optimization. We find that it can further improve the performance of iterative optimization because different programs favor different compilers. We also investigate why iterative optimization works by analyzing the optimal combinations. We find that only a handful optimizations yield most of the speedup. Finally, we show that optimizations interact in a complex and sometimes counterintuitive way through two case studies, which confirms that iterative optimization is an irreplaceable and important compiler strategy. Shuangde Fang, Yuanjie Huang, Lieven Eeckhout, Grigori Fursin, Olivier Temam, Chengyong Wu |
ACM Trans. Archit. Code Optim. | 7 |
| 2010 | Evaluating iterative optimization across 1000 datasetsabstractWhile iterative optimization has become a popular compiler optimization approach, it is based on a premise which has never been truly evaluated: that it is possible to learn the best compiler optimizations across data sets. Up to now, most iterative optimization studies find the best optimizations through repeated runs on the same data set. Only a handful of studies have attempted to exercise iterative optimization on a few tens of data sets. Yuanjie Huang, Lieven Eeckhout, Grigori Fursin, Olivier Temam, Chengyong Wu |
PLDI | 7 |
| 2008 | Global Tiling for Communication Minimal Parallelization on Distributed Memory Systems
Lei Liu 0030, Chengyong Wu, Xiaobing Feng 0002 |
Euro-Par | 3 |