VLDB 2026 Research / reviewers in the wild / expert
Wei Ding 0008
dblp:59/622-8
· DBLP profile ↗
23ranked-venue papers
11as first author
0since 2021 · last 2017
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 10 first-authorSoftware engineering, systems software and programming languages · 6 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
Memory systems · 68% Storage systems · 14% Parallel and multicore computing · 7% | |
| Software engineering, system software, and programming languages
5 papers |
Compilers and program optimization · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 100% |
Topics — the 30 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache |
0.4 | 3 | 2014 | CApRI: CAche-conscious data reordering for irregular codes · SIGMETRICS 2014 Data layout optimization for GPGPU architectures · PPoPP 2013 Compiler Support for Optimizing Memory Bank-Level Parallelism · MICRO 2014 |
Memory systems
data layout optimization |
0.3 | 2 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Data integration and cleaning
query mapping |
0.3 | 1 | 2017 | Cache Hierarchy-Aware Query Mapping on Emerging Multicore Architectures · IEEE Trans. Computers 2017 |
Compilers and program optimization
compiler optimization |
0.3 | 2 | 2015 | Network footprint reduction through data access and computation placement in NoC-based manycores · DAC 2015 A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Storage systems
data placement |
0.2 | 1 | 2015 | Optimizing off-chip accesses in multicores · PLDI 2015 |
Compilers and program optimization
loop optimization |
0.2 | 1 | 2014 | Compiler Support for Optimizing Memory Bank-Level Parallelism · MICRO 2014 |
Memory systems › DRAM › DRAM microarchitecture
bank-level parallelism |
0.2 | 1 | 2014 | Compiler Support for Optimizing Memory Bank-Level Parallelism · MICRO 2014 |
Memory systems › data layout optimization
cache-conscious data structure layout |
0.2 | 1 | 2014 | CApRI: CAche-conscious data reordering for irregular codes · SIGMETRICS 2014 |
Memory systems › cache
cache miss reduction |
0.2 | 1 | 2014 | CApRI: CAche-conscious data reordering for irregular codes · SIGMETRICS 2014 |
Memory systems › data layout optimization
data reordering |
0.2 | 1 | 2014 | CApRI: CAche-conscious data reordering for irregular codes · SIGMETRICS 2014 |
Memory systems › memory access optimization
memory-level parallelism |
0.2 | 1 | 2014 | Compiler Support for Optimizing Memory Bank-Level Parallelism · MICRO 2014 |
GPUs and heterogeneous computing
GPU programming |
0.2 | 1 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 |
Memory systems › memory hierarchy
memory hierarchy optimization |
0.2 | 1 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 |
Compilers and program optimization › vectorization
superword level parallelism |
0.1 | 1 | 2012 | A compiler framework for extracting superword level parallelism · PLDI 2012 |
Memory systems
cache coherence |
0.1 | 1 | 2012 | A hybrid NoC design for cache coherence optimization for chip multiprocessors · DAC 2012 |
Storage systems › file systems › file organization
file layout optimization |
0.1 | 1 | 2012 | Compiler-directed file layout optimization for hierarchical storage systems · SC 2012 |
Storage systems
i/o optimization |
0.1 | 1 | 2012 | Compiler-directed file layout optimization for hierarchical storage systems · SC 2012 |
Interconnection networks and networks-on-chip
network-on-chip design |
0.1 | 1 | 2012 | A hybrid NoC design for cache coherence optimization for chip multiprocessors · DAC 2012 |
Memory systems › cache coherence › cache coherence protocol
snoopy coherence |
0.1 | 1 | 2012 | A hybrid NoC design for cache coherence optimization for chip multiprocessors · DAC 2012 |
Memory systems
cache design |
0.1 | 1 | 2011 | A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Parallel and multicore computing › task allocation
computation-to-core mapping |
0.1 | 1 | 2011 | A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Memory systems › memory hierarchy › cache hierarchy
non-uniform cache access |
0.1 | 1 | 2011 | A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Memory systems › memory hierarchy
cache hierarchy |
0.1 | 1 | 2017 | Cache Hierarchy-Aware Query Mapping on Emerging Multicore Architectures · IEEE Trans. Computers 2017 |
Parallel and multicore computing › data parallelism
data-parallel applications |
0.1 | 1 | 2015 | Optimizing off-chip accesses in multicores · PLDI 2015 |
Parallel and multicore computing › thread-level parallelism
multithreaded applications |
0.1 | 1 | 2015 | Optimizing off-chip accesses in multicores · PLDI 2015 |
Memory systems
cache miss prediction |
0.1 | 1 | 2014 | Compiler Support for Optimizing Memory Bank-Level Parallelism · MICRO 2014 |
Compilers and program optimization › loop optimization
loop nest optimization |
0.0 | 1 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 |
Compilers and program optimization › memory optimization
data layout optimization |
0.0 | 1 | 2012 | A compiler framework for extracting superword level parallelism · PLDI 2012 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2012 | A hybrid NoC design for cache coherence optimization for chip multiprocessors · DAC 2012 |
Processor architecture and microarchitecture › SIMD
SIMD instructions |
0.0 | 1 | 2012 | A compiler framework for extracting superword level parallelism · PLDI 2012 |
Methods — techniques the papers use, named apart from their topics
integer linear programming · 0.6graph partitioning · 0.6computation decomposition · 0.4tile scheduling · 0.4cache miss equations · 0.4affine loop nest analysis · 0.3experimental evaluation · 0.2compiler-based data localization · 0.2locality model · 0.2data layout reorganization · 0.2data localization · 0.2statement scheduling · 0.1statement grouping · 0.1data layout optimization · 0.1full-system simulation · 0.1array tiling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | DEMM: A Dynamic Energy-Saving Mechanism for Multicore MemoriesabstractSince main memory system contributes to a large and increasing fraction of server/datacenter energy consumption, there have been several efforts to reduce its power and energy consumption. DVFS schemes have been used to reduce the memory power, but they come with a performance penalty. In this work, we propose DEMM, an OS-based, high performance DVFS mechanism that reduces memory power by dynamically scaling individual memory channel frequencies/voltages. Our strategy also involves clustering the running applications based on their sensitivities to memory latency, and assigning memory channels to the application clusters. We introduce a new metric called Discrete Misses per Kilo Cycle (DMPKC) to capture the performance sensitivities of the applications to memory frequency modulation. DEMM allows us to save power in the memory system with negligible impact on performance. We demonstrate around 25% savings in the memory system energy and 10% savings in the total system energy, with only a 4% loss in workload performance. Akbar Sharifi, Wei Ding 0008, Diana R. Guttman, Hui Zhao 0013, Xulong Tang, Mahmut T. Kandemir, Chita R. Das |
MASCOTS | 2 |
| 2017 | Cache Hierarchy-Aware Query Mapping on Emerging Multicore ArchitecturesabstractOne of the important characteristics of emerging multicores/manycores is the existence of “shared on-chip caches,” through which different threads/processes can share data (help each other) or displace each other's data (hurt each other). Most of current commercial multicore systems on the market have on-chip cache hierarchies with multiple layers (typically, in the form of L1, L2 and L3, the last two being either fully or partially shared). In the context of database workloads, exploiting full potential of these caches can be critical. Motivated by this observation, our main contribution in this work is to present and experimentally evaluate a cache hierarchy-aware query mapping scheme targeting workloads that consist of batch queries to be executed on emerging multicores. Our proposed scheme distributes a given batch of queries across the cores of a target multicore architecture based on the affinity relations among the queries. The primary goal behind this scheme is to maximize the utilization of the underlying on-chip cache hierarchy while keeping the load nearly balanced across domain affinities. Each domain affinity in this context corresponds to a cache structure bounded by a particular level of the cache hierarchy. A graph partitioning-based method is employed to distribute queries across cores, and an integer linear programming (ILP) formulation is used to address locality and load balancing concerns. We evaluate our scheme using the TPC-H benchmarks on an Intel Xeon based multicore. Our solution achieves up to 25 percent improvement in individual query execution times and 15-19 percent improvement in throughput over the default Linux-based process scheduler. Ozcan Ozturk 0001, Umut Orhan, Wei Ding 0008, Praveen Yedlapalli, Mahmut T. Kandemir |
IEEE Trans. Computers | 3 |
| 2015 | Reactive tilingabstractTo fully exploit the power of emerging multicore architectures, managing shared resources (i.e., caches) across applications and over time is critical. However, to our knowledge, most prior efforts view this problem from the OS/hardware side, and do not consider whether applications themselves can also participate in this process of managing shared resources. In this paper, we show how an application can react to OS/hardware-based resource management decisions by adapting itself (called reactive application), with the goal of maximizing the utilization of the shared resources allocated to it. Specifically, we present a framework that can generate code for adaptive (reactive) tiling, and propose an execution model in which a reactive application can react to the modulations in its cache space allocations to prevent its performance from degrading. One can expect two potential benefits from this approach. First, matching tile size to available cache capacity dynamically (during execution) improves performance of the target application. Second and equally important, better utilization of shared cache space reduces pressure on other applications (co-runners) that execute concurrently with the target application. Our experimental results show that the proposed scheme improves the performance of applications (over the best static tiles) by 8.4%, on average, when using synthetic cache allocations. Further with dynamic cache allocations determined by the utility-based cache partitioning (a state-of-the-art cache partitioning scheme), it improves performance of a set of eleven HPC applications by 11.3%. Jithendra Srinivas, Wei Ding 0008, Mahmut T. Kandemir |
CGO | 2 |
| 2015 | Network footprint reduction through data access and computation placement in NoC-based manycoresabstractTargeting network-on-chip based manycores, we propose a novel compiler framework to optimize the network latencies experienced by off-chip data accesses in reaching the target memory controllers. Our framework consists of two main components: data access placement and computation placement. In the data access placement, we separate the data access nodes from the computation nodes, with the goal of minimizing the number of links that need to be visited by the request messages. In the computation placement, we introduce computation decomposition and select appropriate computation nodes, to reduce the amount of data sent in the response messages and also to minimize the number of communication links visited. We performed an experimental evaluation of our proposed approach, and the results show an average execution time improvement of 21.1%, while reducing the network latency by 67.3%. Jun Liu 0008, Jagadish Kotra, Wei Ding 0008, Mahmut T. Kandemir |
DAC | 3 |
| 2015 | Optimizing off-chip accesses in multicoresabstractIn a network-on-chip (NoC) based manycore architecture, an off-chip data access (main memory access) needs to travel through the on-chip network, spending considerable amount of time within the chip (in addition to the memory access latency). In addition, it contends with on-chip (cache) accesses as both use the same NoC resources. In this paper, focusing on data-parallel, multithreaded applications, we propose a compiler-based off-chip data access localization strategy, which places data elements in the memory space such that an off-chip access traverses a minimum number of links (hops) to reach the memory controller that handles this access. This brings three main benefits. First, the network latency of off-chip accesses gets reduced; second, the network latency of on-chip accesses gets reduced; and finally, the memory latency of off-chip accesses improves, due to reduced queue latencies. We present an experimental evaluation of our optimization strategy using a set of 13 multithreaded application programs under both private and shared last-level caches. The results collected emphasize the importance of optimizing the off-chip data accesses. Wei Ding 0008, Xulong Tang, Mahmut T. Kandemir, Emre Kultursay |
PLDI | 1 |
| 2014 | Trading cache hit rate for memory performanceabstractMost of the prior compiler based data locality optimization works target exclusively cache locality optimization, and row-buffer locality in DRAM banks received much less attention. In particular, to the best of our knowledge, there is no single compiler based approach that can improve row-buffer locality in executing irregular applications. This presents a critical problem considering the fact that executing irregular applications in a power and performance efficient manner will be a key requirement to extract maximum benefits from emerging multicore machines and exascale systems. Motivated by these observations, this paper makes the following contributions. First, it presents a compiler-runtime cooperative data layout optimization approach that takes as input an irregular program that has already been optimized for cache locality and generates an output code with the same cache performance but better row-buffer locality (lower number of row-buffer misses). Second, it discusses a more aggressive strategy that sacrifices some cache performance in order to further improve row-buffer performance (i.e., it trades cache performance for memory system performance). The ultimate goal of this strategy is to find the right tradeoff point between cache performance and row-buffer performance so that the overall application performance is improved. Third, the paper performs a detailed evaluation of these two approaches using both an AMD Opteron based multicore system and a multicore simulator. The experimental results, collected using five real-world irregular applications, show that (i) conventional cache optimizations do not improve row-buffer locality significantly; (ii) our first approach achieves about 9.8% execution time improvement by keeping the number of cache misses the same as a cache-optimized code but reducing the number of row-buffer misses; and (iii) our second approach achieves even higher execution time improvements (13.8% on average) by sacrificing cache performance for additional memory performance. Wei Ding 0008, Mahmut T. Kandemir, Diana R. Guttman, Adwait Jog, Chita R. Das, Praveen Yedlapalli |
PACT | 1 |
| 2014 | Quantifying and Optimizing the Impact of Victim Cache Line Selection in Manycore SystemsabstractIn both architecture and software, the main goal of data locality-oriented optimizations has always been "minimizing the number of cache misses" (especially, costly last-level cache misses). However, this paper shows that other metrics such as the distance between the last-level cache and memory controller as well as the memory queuing latency can play an equally important role, as far as application performance is concerned. Focusing on a large set of multithreaded applications, we first show that the last-level cache "write backs" (memory writes due to displacement of a victim block from the last-level cache) can exhibit significant latencies as well as variances, and then make a case for "relaxing" the strict LRU policy to save (write back) cycles in both the on-chip network and memory queues. Specifically, we explore novel architecture-level schemes that optimize on-chip network latency, memory queuing latency or both, of the write back messages, by carefully selecting the victim block to write back at the time of cache replacement. Our extensive experimental evaluations using 15 multithreaded applications and a cycle-accurate simulation infrastructure clearly demonstrate that this tradeoffs (between cache hit rate and on-chip network/memory queuing latency) pays off in most of the cases, leading to about 12.2% execution time improvement and 14.9% energy savings, in our default 64-core system with 6 memory controllers. Mahmut T. Kandemir, Wei Ding 0008, Diana R. Guttman |
MASCOTS | 2 |
| 2014 | Compiler Support for Optimizing Memory Bank-Level ParallelismabstractMany prior compiler-based optimization schemes focused exclusively on cache data locality. However, cache locality is only one part of the overall performance of applications running on emerging multicores or many cores. For example, memory stalls could constitute a very large fraction of execution time even in cache-optimized codes, and one of the main reasons for this is lack of memory-level parallelism. Motivated by this, we propose a compiler-based Bank-Level Parallelism (BLP) optimization scheme that uses loop tile scheduling. More specifically, we first use Cache Miss Equations to predict where the last-level cache miss will happen in each tile, and then identify the set of memory banks that will be accessed in each tile. Using this information, two tile scheduling algorithms are proposed to maximize BLP, each targeting a different scenario. We further discuss how our compiler-based scheme can be enhanced to consider memory controller-level parallelism and row-buffer locality. Our experimental evaluation using 11 multithreaded applications shows that the proposed BLP optimization can improve average BLP by 17.1% on average, resulting in a 9.2% reduction in average memory access latency. Furthermore, considering memory controller-level parallelism and row-buffer locality (in addition to BLP) takes our average improvement in memory access latency to 22.2%. Wei Ding 0008, Diana R. Guttman, Mahmut T. Kandemir |
MICRO | 1 |
| 2014 | CApRI: CAche-conscious data reordering for irregular codesabstractCaches play a critical role in today's computer systems and optimizing their performance has been a critical objective in the last couple of decades. Unfortunately, compared to a plethora of work in software and hardware directed code/data optimizations, much less effort has been spent in understanding the fundamental characteristics of data access patterns exhibited by application programs and their interaction with the underlying cache hardware. Therefore, in general it is hard to reason about cache behavior of a program running on a target system. Motivated by this observation, we first set up a "locality model" that can help us determine the theoretical bounds of the cache misses caused by irregular data accesses. We then explain how this locality model can be used for different data locality optimization purposes. After that, based on our model, we propose a data reordering (data layout reorganization) scheme that can be applied after any existing data reordering schemes for irregular applications to improve cache performance by further reducing the cache misses. We evaluate the effectiveness of our scheme using a set of 8 programs with irregular data accesses, and show that it brings significant improvements over the state-of-the-art on two commercial multicore machines. Wei Ding 0008, Mahmut T. Kandemir |
SIGMETRICS | 1 |
| 2013 | Reshaping cache misses to improve row-buffer locality in multicore systemsabstractOptimizing cache locality has always been important since the emergence of caches, and numerous cache locality optimization schemes have been published in compiler literature. However, in modern architectures, cache locality is not the only factor that determines memory system performance. Many emerging multicores employ banked memory systems and each bank is attached a row-buffer that holds the most-recently accessed memory row (page). A last-level cache miss that also misses in the row-buffer can experience much higher latency than a cache miss that hits in the row-buffer. Consequently, optimizing for row-buffer locality can be as important as optimizing for cache locality. Targeting emerging multicores and multithreaded applications, this paper presents a compiler-directed row-buffer locality optimization strategy. This strategy modifies the memory layout of data to increase the number of row-buffer hits without increasing the number of misses in the on-chip cache hierarchy. We implemented our proposed optimization strategy in an open-source compiler and tested its effectiveness in improving the row-buffer performance using a set of multithreaded applications. Our results indicate that the proposed approach improves the average data access latency by about 29%, and this translates, on average, to about 15% improvement in execution time. Wei Ding 0008, Jun Liu 0008, Mahmut T. Kandemir, Mary Jane Irwin |
PACT | 1 |
| 2013 | Locality-aware mapping and scheduling for multicoresabstractThis paper presents a cache hierarchy-aware code mapping and scheduling strategy for multicore architectures. Our mapping strategy determines a loop iteration-to-core mapping by taking into account application data access patterns and on-chip cache hierarchy. It employs a novel concept called “core vectors” to obtain a mapping matrix which exploits data reuses at different layers of the cache hierarchy based on their reuse distances, with the goal of maximizing data locality at each level, while minimizing data dependences across the cores. Our scheduling strategy on the other hand determines a schedule for the iterations assigned to each core, with the goal of reducing data reuse distances across the cores for dependence-free loop nests. Our experimental evaluation shows that the proposed mapping scheme reduces miss rates at all levels of caches and application execution time significantly, and when supported by scheduling, the reduction in cache miss rates and execution time become much larger. Wei Ding 0008, Mahmut T. Kandemir, Jithendra Srinivas, Praveen Yedlapalli |
CGO | 1 |
| 2013 | Data layout optimization for GPGPU architecturesabstractGPUs are being widely used in accelerating general-purpose applications, leading to the emergence of GPGPU architectures. New programming models, e.g., Compute Unified Device Architecture (CUDA), have been proposed to facilitate programming general-purpose computations in GPGPUs. However, writing high-performance CUDA codes manually is still tedious and difficult. In particular, the organization of the data in the memory space can greatly affect the performance due to the unique features of a custom GPGPU memory hierarchy. In this work, we propose an automatic data layout transformation framework to solve the key issues associated with a GPGPU memory hierarchy (i.e., channel skewing, data coalescing, and bank conflicts). Our approach employs a widely applicable strategy based on a novel concept called data localization. Specifically, we try to optimize the layout of the arrays accessed in affine loop nests, for both the device memory and shared memory, at both coarse grain and fine grain parallelization levels. We performed an experimental evaluation of our data layout optimization strategy using 15 benchmarks on an NVIDIA CUDA GPU device. The results show that the proposed data transformation approach brings around 4.3X speedup on average. Jun Liu 0008, Wei Ding 0008, Ohyoung Jang, Mahmut T. Kandemir |
PPoPP | 2 |
| 2012 | Off-chip access localization for NoC-based multicoresabstractIn a network-on-chip based multicore, an off-chip data access needs to travel through the on-chip network, spending considerable amount of time within the chip (in addition to the memory access itself). Further, it also causes additional delays for on-chip accesses by creating contention on network resources. In this paper, we propose a compiler-guided off-chip data access localization strategy to ensure that, an off-chip access traverses a small number of links (hops) to reach the memory controller which governs the memory bank that holds the requested data. We present an extensive evaluation of this strategy using a set of 12 multithreaded application programs. The results collected clearly emphasize the importance of localizing off-chip accesses. Wei Ding 0008, Mahmut T. Kandemir, Emre Kultursay |
PACT | 1 |
| 2012 | A hybrid NoC design for cache coherence optimization for chip multiprocessorsabstractOn chip many-core systems, evolving from prior multi-processor systems, are considered as a promising solution to the performance scalability and power consumption problems. The long communication distance between the traditional multi-processors makes directory-based cache coherence protocols better solutions compared to bus-based snooping protocols even with the overheads from indirections. However, much smaller distances between the CMP cores enhance the reachability of buses, revitalizing the applicability of snooping protocols for cache-to-cache transfers. In this work, we propose a hybrid NoC design to provide optimized support for cache coherency. In our design, on-chip links can be dynamically configured as either point-to-point links between NoC nodes or short buses to facilitate localized snooping. By taking advantage of the best of both worlds, bus-based snooping coherency and NoC-based directory coherency, our approach brings both power and performance benefits. Hui Zhao 0013, Ohyoung Jang, Wei Ding 0008, Mahmut T. Kandemir, Mary Jane Irwin |
DAC | 3 |
| 2012 | Improving last level cache locality by integrating loop and data transformationsabstractMotivated by the observation that most existing data locality optimizations do not specifically target shared last-level caches of emerging multicores and that even multicore-specific locality-oriented techniques employ either loop or data layout optimizations but not both, in this paper we present an integrated loop and data layout optimization strategy, with the goal of improving the last-level cache performance of multicores that execute multithreaded applications. We present a detailed mathematical formulation of our locality optimization strategy and present experimental data from our current implementation. Our results, collected using 14 application programs, clearly show that the proposed integrated approach is very successful in practice, and outperforms both pure loop optimization and pure data layout optimization based alternatives. Our results also indicate that the savings achieved increase with increased core count and larger data set sizes. Wei Ding 0008, Mahmut T. Kandemir |
ICCAD | 1 |
| 2012 | A compiler framework for extracting superword level parallelismabstractSIMD (single-instruction multiple-data) instruction set extensions are quite common today in both high performance and embedded microprocessors, and enable the exploitation of a specific type of data parallelism called SLP (Superword Level Parallelism). While prior research shows that significant performance savings are possible when SLP is exploited, placing SIMD instructions in an application code manually can be very difficult and error prone. In this paper, we propose a novel automated compiler framework for improving superword level parallelism exploitation. The key part of our framework consists of two stages: superword statement generation and data layout optimization. The first stage is our main contribution and has two phases, statement grouping and statement scheduling, of which the primary goals are to increase SIMD parallelism and, more importantly, capture more superword reuses among the superword statements through global data access and reuse pattern analysis. Further, as a complementary optimization, our data layout optimization organizes data in memory space such that the price of memory operations for SLP is minimized. The results from our compiler implementation and tests on two systems indicate performance improvements as high as 15.2% over a state-of-the-art SLP optimization algorithm. Jun Liu 0008, Ohyoung Jang, Wei Ding 0008, Mahmut T. Kandemir |
PLDI | 4 |
| 2012 | Compiler-directed file layout optimization for hierarchical storage systemsabstractFile layout of array data is a critical factor that effects the behavior of storage caches, and has so far taken not much attention in the context of hierarchical storage systems. The main contribution of this paper is a compiler-driven file layout optimization scheme for hierarchical storage caches. This approach, fully automated within an optimizing compiler, analyzes a multi-threaded application code and determines a file layout for each disk-resident array referenced by the code, such that the performance of the target storage cache hierarchy is maximized. We tested our approach using 16 I/O intensive application programs and compared its performance against two previously proposed approaches under different cache space management schemes. Our experimental results show that the proposed approach improves the execution time of these parallel applications by 23.7% on average. Wei Ding 0008, Mahmut T. Kandemir, Seung Woo Son 0001 |
SC | 1 |
| 2011 | Compiler Directed Data Locality Optimization for Multicore ArchitecturesabstractEmerging multicore architectures differ from prior multi- processor based systems in that they typically employ on- chip cache hierarchies where subsets of caches are shared by subsets of cores. This cache sharing can result in opportunities as well as problems in a multithreaded execution, depending on how compatible the data access and data sharing patterns exhibited by threads are with the physical cache sharing imposed by the underlying architecture. Investigating this relationship and exploiting it to improve application performance through program transformation are the main objectives of this paper. Wei Ding 0008, Jithendra Srinivas, Mahmut T. Kandemir, Mustafa Karaköy |
PACT | 1 |
| 2011 | Optimizing Data Layouts for Parallel Computation on MulticoresabstractThe emergence of multicore platforms offers several opportunities for boosting application performance. These opportunities, which include parallelism and data locality benefits, require strong support from compilers as well as operating systems. Current compiler research targeting multicores mostly focuses on code restructuring and mapping. In this work, we explore automatic data layout transformation targeting multithreaded applications running on multicores. Our transformation considers both data access patterns exhibited by different threads of a multithreaded application and the on-chip cache topology of the target multicore architecture. It automatically determines a customized memory layout for each target array to minimize potential cache conflicts across threads. Our experiments show that, our optimization brings significant benefits over state-of-the-art data locality optimization strategies when tested using 30 benchmark programs on an Intel multicore machine. The results also indicate that this strategy is able to scale to larger core counts and it performs better with increased data set sizes. Wei Ding 0008, Jun Liu 0008, Mahmut T. Kandemir |
PACT | 2 |
| 2011 | On-chip cache hierarchy-aware tile scheduling for multicore machinesabstractIteration space tiling and scheduling is an important technique for optimizing loops that constitute a large fraction of execution times in computation kernels of both scientific codes and embedded applications. While tiling has been studied extensively in the context of both uniprocessor and multiprocessor platforms, prior research has paid less attention to tile scheduling, especially when targeting multicore machines with deep on-chip cache hierarchies. In this paper, we propose a cache hierarchy-aware tile scheduling algorithm for multicore machines, with the purpose of maximizing both horizontal and vertical data reuses in on-chip caches, and balancing the workloads across different cores. This scheduling algorithm is one of the key components in a source-to-source translation tool that we developed for automatic loop parallelization and multithreaded code generation from sequential codes. To the best of our knowledge, this is the first effort that develops a fully-automated tile scheduling strategy customized for on-chip cache topologies of multicore machines. The experimental results collected by executing twelve application programs on three commercial Intel machines (Nehalem, Dunnington, and Harpertown) reveal that our cache-aware tile scheduling brings about 27.9% reduction in cache misses, and on average, 13.5% improvement in execution times over an alternate method tested. Jun Liu 0008, Wei Ding 0008, Mahmut T. Kandemir |
CGO | 3 |
| 2011 | Optimizing data locality using array tilingabstractData transformation is one of the key optimizations in maximizing cache locality. Traditional data transformation strategies employ linear data layouts, e.g., row-major or column-major, for multidimensional arrays. Although a linear layout matches the linear memory space well in most cases, it can only optimize for self-spatial locality for individual references. In this work, we propose a novel data layout transformation framework that is able to determine a tiled layout for each array in an application program. Tiled layout can exploit the group-spatial locality among different references and improve cache line utilization. In our strategy, the data elements accessed by different references in one loop iteration are placed into a tile and fetched into the same cache line at runtime. This helps minimizing conflict misses in caches. We evaluated our data layout transformation framework using 30 benchmarks on a commercial multicore machine. The experimental results show that our approach outperforms state-of-the-art data transformation strategies and works well with large core counts. Wei Ding 0008, Jun Liu 0008, Mahmut T. Kandemir |
ICCAD | 1 |
| 2011 | Exploring heterogeneous NoC design spaceabstractThe Network-on-Chip (NoC) plays a crucial role in designing low cost chip multiprocessors (CMPs) as the number of cores on a chip keeps increasing. However, buffers in NoC routers increase the cost of CMPs in terms of both area and power. Recently, bufferless routers have been proposed to reduce such costs by removing buffers from the routers. However, bufferless routers can provide competitive performance only when network utilization is moderate. In this paper, we propose a novel heterogeneous design that employs both buffered and bufferless routers in the same NoC to achieve high performance at low cost. We evaluate a variety of plans to place buffered and bufferless routers in an NoC based CMP according to performance requirements and power allowances. In order to take full advantage of these heterogeneous NoCs, we also propose novel strategies for buffered-router-aware application thread mapping and a routing algorithm (once the router placement is fixed). Our evaluations show that, by utilizing the techniques we proposed, a heterogeneous NoC does not only achieve performance comparable to that of the NoCs with buffered routers but also reduces buffer costs and energy consumption. Hui Zhao 0013, Mahmut T. Kandemir, Wei Ding 0008, Mary Jane Irwin |
ICCAD | 3 |
| 2011 | A data layout optimization framework for NUCA-based multicoresabstractFuture multicore architectures are likely to include a large number of cores connected using an on-chip network with Non-uniform Cache Access (NUCA). In such architectures, whether a data request is satisfied from a local cache or a remote cache can make an important difference. To exploit this NUCA property, prior research explored both architectural enhancements as well as compiler-based code optimization strategies. In this work, we take an alternate view, and explore data layout optimizations to improve locality of data accesses in a NUCA-based system. Our proposed approach includes three steps: array tiling, computation-to-core mapping, and layout customization. The first of these tries to identify the affinity between data and computation taking into account parallelization information, with the goal of minimizing remote accesses. The second step maps computations (and their associated data) to cores with the goal of minimizing average distance-to-data, and the last step further customizes the memory layout taking into account the data placement policy adopted by the underlying architecture. We evaluated the success of this three-step approach in enhancing on-chip cache behavior using all application programs from the SPECOMP suite on a full-system simulator. Our results show that the proposed approach improves on average data access latency and execution time by 24.7% and 18.4%, respectively, in the case of static NUCA, and 18.1% and 12.7%, respectively, in the case of dynamic NUCA. Wei Ding 0008, Mahmut T. Kandemir, Jun Liu 0008, Ohyoung Jang |
MICRO | 2 |