VLDB 2026 Research / reviewers in the wild / expert
Jun Liu 0008
dblp:95/3736-8
· DBLP profile ↗
13ranked-venue papers
5as first author
0since 2021 · last 2015
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-authorSoftware engineering, systems software and programming languages · 4 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Memory systems · 64% GPUs and heterogeneous computing · 12% Storage systems · 9% | |
| Software engineering, system software, and programming languages
4 papers |
Compilers and program optimization · 100% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
data layout optimization |
0.3 | 2 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Compilers and program optimization
compiler optimization |
0.3 | 2 | 2015 | Network footprint reduction through data access and computation placement in NoC-based manycores · DAC 2015 A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Memory systems
cache |
0.2 | 1 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 |
GPUs and heterogeneous computing
GPU programming |
0.2 | 1 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 |
Memory systems › memory hierarchy
memory hierarchy optimization |
0.2 | 1 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 |
Compilers and program optimization › vectorization
superword level parallelism |
0.1 | 1 | 2012 | A compiler framework for extracting superword level parallelism · PLDI 2012 |
Memory systems
cache design |
0.1 | 1 | 2011 | A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Parallel and multicore computing › task allocation
computation-to-core mapping |
0.1 | 1 | 2011 | A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Memory systems › memory hierarchy › cache hierarchy
non-uniform cache access |
0.1 | 1 | 2011 | A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Compilers and program optimization › loop optimization
loop nest optimization |
0.0 | 1 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 |
Compilers and program optimization › memory optimization
data layout optimization |
0.0 | 1 | 2012 | A compiler framework for extracting superword level parallelism · PLDI 2012 |
Processor architecture and microarchitecture › SIMD
SIMD instructions |
0.0 | 1 | 2012 | A compiler framework for extracting superword level parallelism · PLDI 2012 |
Compilers and program optimization › memory optimization
data locality optimization |
0.0 | 1 | 2011 | A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Energy-efficient computing
power management |
0.0 | 1 | 2011 | Software-directed data access scheduling for reducing disk energy consumption · HPDC 2011 |
Methods — techniques the papers use, named apart from their topics
computation decomposition · 0.4data localization · 0.3affine loop nest analysis · 0.3statement scheduling · 0.3statement grouping · 0.3data layout optimization · 0.3layout customization · 0.2full-system simulation · 0.2array tiling · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Network footprint reduction through data access and computation placement in NoC-based manycoresabstractTargeting network-on-chip based manycores, we propose a novel compiler framework to optimize the network latencies experienced by off-chip data accesses in reaching the target memory controllers. Our framework consists of two main components: data access placement and computation placement. In the data access placement, we separate the data access nodes from the computation nodes, with the goal of minimizing the number of links that need to be visited by the request messages. In the computation placement, we introduce computation decomposition and select appropriate computation nodes, to reduce the amount of data sent in the response messages and also to minimize the number of communication links visited. We performed an experimental evaluation of our proposed approach, and the results show an average execution time improvement of 21.1%, while reducing the network latency by 67.3%. Jun Liu 0008, Jagadish Kotra, Wei Ding 0008, Mahmut T. Kandemir |
DAC | 1 |
| 2013 | Reshaping cache misses to improve row-buffer locality in multicore systemsabstractOptimizing cache locality has always been important since the emergence of caches, and numerous cache locality optimization schemes have been published in compiler literature. However, in modern architectures, cache locality is not the only factor that determines memory system performance. Many emerging multicores employ banked memory systems and each bank is attached a row-buffer that holds the most-recently accessed memory row (page). A last-level cache miss that also misses in the row-buffer can experience much higher latency than a cache miss that hits in the row-buffer. Consequently, optimizing for row-buffer locality can be as important as optimizing for cache locality. Targeting emerging multicores and multithreaded applications, this paper presents a compiler-directed row-buffer locality optimization strategy. This strategy modifies the memory layout of data to increase the number of row-buffer hits without increasing the number of misses in the on-chip cache hierarchy. We implemented our proposed optimization strategy in an open-source compiler and tested its effectiveness in improving the row-buffer performance using a set of multithreaded applications. Our results indicate that the proposed approach improves the average data access latency by about 29%, and this translates, on average, to about 15% improvement in execution time. Wei Ding 0008, Jun Liu 0008, Mahmut T. Kandemir, Mary Jane Irwin |
PACT | 2 |
| 2013 | Data layout optimization for GPGPU architecturesabstractGPUs are being widely used in accelerating general-purpose applications, leading to the emergence of GPGPU architectures. New programming models, e.g., Compute Unified Device Architecture (CUDA), have been proposed to facilitate programming general-purpose computations in GPGPUs. However, writing high-performance CUDA codes manually is still tedious and difficult. In particular, the organization of the data in the memory space can greatly affect the performance due to the unique features of a custom GPGPU memory hierarchy. In this work, we propose an automatic data layout transformation framework to solve the key issues associated with a GPGPU memory hierarchy (i.e., channel skewing, data coalescing, and bank conflicts). Our approach employs a widely applicable strategy based on a novel concept called data localization. Specifically, we try to optimize the layout of the arrays accessed in affine loop nests, for both the device memory and shared memory, at both coarse grain and fine grain parallelization levels. We performed an experimental evaluation of our data layout optimization strategy using 15 benchmarks on an NVIDIA CUDA GPU device. The results show that the proposed data transformation approach brings around 4.3X speedup on average. Jun Liu 0008, Wei Ding 0008, Ohyoung Jang, Mahmut T. Kandemir |
PPoPP | 1 |
| 2012 | Panacea: towards holistic optimization of MapReduce applicationsabstractMapReduce has emerged as one of the most popular programming models for data parallel enterprise applications. Despite advances in runtime, the opportunities for optimizing MapReduce applications remain largely unexplored. In this paper, we present a framework for performing holistic compiler optimizations on legacy MapReduce applications. We have identified and implemented two optimizations and evaluated them with a set of Hadoop applications on a cluster of Xeon servers. Our experiments show that performance gains of more than 3X can be achieved without user involvement. Jun Liu 0008, Nishkam Ravi, Srimat T. Chakradhar, Mahmut T. Kandemir |
CGO | 1 |
| 2012 | Software-Directed Data Access Scheduling for Reducing Disk Energy ConsumptionabstractMost existing research in disk power management has focused on exploiting idle periods of disks. Both hardware power-saving mechanisms (such as spin-down disks and multi-speed disks) and complementary software strategies (such as code and data layout transformations to increase the length of idle periods) have been explored. However, while hardware power-saving mechanisms cannot handle short idle periods of high-performance parallel applications, prior code/data reorganization strategies typically require extensive code modifications. In this paper, we propose and evaluate a compiler-directed data access (I/O call) scheduling framework for saving disk energy, which groups as many data requests as possible in a shorter period, thus creating longer disk idle periods for improving the effectiveness of hardware power-saving mechanisms. As compared to prior software based efforts, it requires no code or data restructuring. We evaluate our approach using six application programs in a cluster-based simulation environment. The experimental results show that it improves the effectiveness of both spin-down disks and multi-speed disks with doubled power savings on average. Jun Liu 0008, Mahmut T. Kandemir |
ICDCS | 2 |
| 2012 | A compiler framework for extracting superword level parallelismabstractSIMD (single-instruction multiple-data) instruction set extensions are quite common today in both high performance and embedded microprocessors, and enable the exploitation of a specific type of data parallelism called SLP (Superword Level Parallelism). While prior research shows that significant performance savings are possible when SLP is exploited, placing SIMD instructions in an application code manually can be very difficult and error prone. In this paper, we propose a novel automated compiler framework for improving superword level parallelism exploitation. The key part of our framework consists of two stages: superword statement generation and data layout optimization. The first stage is our main contribution and has two phases, statement grouping and statement scheduling, of which the primary goals are to increase SIMD parallelism and, more importantly, capture more superword reuses among the superword statements through global data access and reuse pattern analysis. Further, as a complementary optimization, our data layout optimization organizes data in memory space such that the price of memory operations for SLP is minimized. The results from our compiler implementation and tests on two systems indicate performance improvements as high as 15.2% over a state-of-the-art SLP optimization algorithm. Jun Liu 0008, Ohyoung Jang, Wei Ding 0008, Mahmut T. Kandemir |
PLDI | 1 |
| 2011 | Optimizing Data Layouts for Parallel Computation on MulticoresabstractThe emergence of multicore platforms offers several opportunities for boosting application performance. These opportunities, which include parallelism and data locality benefits, require strong support from compilers as well as operating systems. Current compiler research targeting multicores mostly focuses on code restructuring and mapping. In this work, we explore automatic data layout transformation targeting multithreaded applications running on multicores. Our transformation considers both data access patterns exhibited by different threads of a multithreaded application and the on-chip cache topology of the target multicore architecture. It automatically determines a customized memory layout for each target array to minimize potential cache conflicts across threads. Our experiments show that, our optimization brings significant benefits over state-of-the-art data locality optimization strategies when tested using 30 benchmark programs on an Intel multicore machine. The results also indicate that this strategy is able to scale to larger core counts and it performs better with increased data set sizes. Wei Ding 0008, Jun Liu 0008, Mahmut T. Kandemir |
PACT | 3 |
| 2011 | Neighborhood-aware data locality optimization for NoC-based multicoresabstractData locality optimization is a critical issue for NoC (network-on-chip) based multicore systems. In this paper, focusing on a two-dimensional NoC-based multicore and data-intensive multithreaded applications, we first discuss a data locality aware scheduling algorithm for any given computation-to-core mapping, and then propose an integrated mapping+scheduling algorithm that performs both tasks together. Both our algorithms consider temporal (time-wise) and spatial (neighborhood-aware) data reuse, and try to minimize distance-to-data in on-chip cache accesses. We test the effectiveness of our compiler algorithms using a set of twelve application programs. Our experiments indicate that the proposed algorithms achieve significant improvements in data access latencies (42.7% on average) and overall execution times (24.1% on average). We also conduct a sensitivity analysis where we change the number of cores, on-chip cache capacities, and data movement (migration) strategies. These experiments show that our proposed algorithms generate consistently good results. Mahmut T. Kandemir, Jun Liu 0008, Taylan Yemliha |
CGO | 3 |
| 2011 | On-chip cache hierarchy-aware tile scheduling for multicore machinesabstractIteration space tiling and scheduling is an important technique for optimizing loops that constitute a large fraction of execution times in computation kernels of both scientific codes and embedded applications. While tiling has been studied extensively in the context of both uniprocessor and multiprocessor platforms, prior research has paid less attention to tile scheduling, especially when targeting multicore machines with deep on-chip cache hierarchies. In this paper, we propose a cache hierarchy-aware tile scheduling algorithm for multicore machines, with the purpose of maximizing both horizontal and vertical data reuses in on-chip caches, and balancing the workloads across different cores. This scheduling algorithm is one of the key components in a source-to-source translation tool that we developed for automatic loop parallelization and multithreaded code generation from sequential codes. To the best of our knowledge, this is the first effort that develops a fully-automated tile scheduling strategy customized for on-chip cache topologies of multicore machines. The experimental results collected by executing twelve application programs on three commercial Intel machines (Nehalem, Dunnington, and Harpertown) reveal that our cache-aware tile scheduling brings about 27.9% reduction in cache misses, and on average, 13.5% improvement in execution times over an alternate method tested. Jun Liu 0008, Wei Ding 0008, Mahmut T. Kandemir |
CGO | 1 |
| 2011 | Software-directed data access scheduling for reducing disk energy consumption
Jun Liu 0008, Ellis Herbert Wilson, Mahmut T. Kandemir |
HPDC | 2 |
| 2011 | Optimizing data locality using array tilingabstractData transformation is one of the key optimizations in maximizing cache locality. Traditional data transformation strategies employ linear data layouts, e.g., row-major or column-major, for multidimensional arrays. Although a linear layout matches the linear memory space well in most cases, it can only optimize for self-spatial locality for individual references. In this work, we propose a novel data layout transformation framework that is able to determine a tiled layout for each array in an application program. Tiled layout can exploit the group-spatial locality among different references and improve cache line utilization. In our strategy, the data elements accessed by different references in one loop iteration are placed into a tile and fetched into the same cache line at runtime. This helps minimizing conflict misses in caches. We evaluated our data layout transformation framework using 30 benchmarks on a commercial multicore machine. The experimental results show that our approach outperforms state-of-the-art data transformation strategies and works well with large core counts. Wei Ding 0008, Jun Liu 0008, Mahmut T. Kandemir |
ICCAD | 3 |
| 2011 | A data layout optimization framework for NUCA-based multicoresabstractFuture multicore architectures are likely to include a large number of cores connected using an on-chip network with Non-uniform Cache Access (NUCA). In such architectures, whether a data request is satisfied from a local cache or a remote cache can make an important difference. To exploit this NUCA property, prior research explored both architectural enhancements as well as compiler-based code optimization strategies. In this work, we take an alternate view, and explore data layout optimizations to improve locality of data accesses in a NUCA-based system. Our proposed approach includes three steps: array tiling, computation-to-core mapping, and layout customization. The first of these tries to identify the affinity between data and computation taking into account parallelization information, with the goal of minimizing remote accesses. The second step maps computations (and their associated data) to cores with the goal of minimizing average distance-to-data, and the last step further customizes the memory layout taking into account the data placement policy adopted by the underlying architecture. We evaluated the success of this three-step approach in enhancing on-chip cache behavior using all application programs from the SPECOMP suite on a full-system simulator. Our results show that the proposed approach improves on average data access latency and execution time by 24.7% and 18.4%, respectively, in the case of static NUCA, and 18.1% and 12.7%, respectively, in the case of dynamic NUCA. Wei Ding 0008, Mahmut T. Kandemir, Jun Liu 0008, Ohyoung Jang |
MICRO | 4 |
| 2010 | Scalable Parallelization Strategies to Accelerate NuFFT Data Translation on Multicores
Jun Liu 0008, Emre Kultursay, Mahmut T. Kandemir, Nikos Pitsianis, Xiaobai Sun |
Euro-Par (2) | 2 |