Yi Yang 0018

dblp:33/4854-18 · DBLP profile ↗
← Back
28ranked-venue papers
11as first author
3since 2021 · last 2022
0000-0003-1462-5100ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 9 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
GPUs and heterogeneous computing · 52% Parallel and multicore computing · 17% Memory systems · 13%
Software engineering, system software, and programming languages
7 papers
Compilers and program optimization · 100%

Topics — the 22 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU programming
0.432014
CUDA-NP: realizing nested thread-level parallelism in GPGPU applications · PPoPP 2014
A unified optimizing compiler framework for different GPGPU architectures · ACM Trans. Archit. Code Optim. 2012
A GPGPU compiler for memory optimization and parallelism management · PLDI 2010
Compilers and program optimization › accelerator compilation
GPU compiler
0.322012
A unified optimizing compiler framework for different GPGPU architectures · ACM Trans. Archit. Code Optim. 2012
An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010
GPUs and heterogeneous computing
GPU architecture
0.212014
Warp-level divergence in GPUs: Characterization, impact, and mitigation · HPCA 2014
Cloud and datacenter computing
resource management
0.212014
Warp-level divergence in GPUs: Characterization, impact, and mitigation · HPCA 2014
Parallel and multicore computing
thread-level parallelism
0.212014
Warp-level divergence in GPUs: Characterization, impact, and mitigation · HPCA 2014
GPUs and heterogeneous computing › GPU scheduling
warp scheduling
0.212014
Warp-level divergence in GPUs: Characterization, impact, and mitigation · HPCA 2014
Hardware accelerators and domain-specific architectures
many-core accelerator
0.212013
Semi-automatic restructuring of offloadable tasks for many-core accelerators · SC 2013
Parallel and multicore computing
work distribution
0.212013
Semi-automatic restructuring of offloadable tasks for many-core accelerators · SC 2013
Compilers and program optimization › memory optimization
memory hierarchy optimization
0.112012
A unified optimizing compiler framework for different GPGPU architectures · ACM Trans. Archit. Code Optim. 2012
Memory systems
cache
0.112012
CPU-assisted GPGPU on fused CPU-GPU architectures · HPCA 2012
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
fused CPU-GPU architecture
0.112012
CPU-assisted GPGPU on fused CPU-GPU architectures · HPCA 2012
Compilers and program optimization › accelerator compilation
GPU compiler optimization
0.112010
A GPGPU compiler for memory optimization and parallelism management · PLDI 2010
Compilers and program optimization › memory optimization
memory access optimization
0.112010
An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010
Memory systems › data locality
data reuse
0.112010
An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010
GPUs and heterogeneous computing
GPU computing
0.112010
An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010
Memory systems › memory hierarchy
memory hierarchy optimization
0.112010
A GPGPU compiler for memory optimization and parallelism management · PLDI 2010
GPUs and heterogeneous computing › GPU communication
host-device data transfer
0.112014
COMP: Compiler Optimizations for Manycore Processors · MICRO 2014
Parallel and multicore computing › parallel scheduling
parallel loop scheduling
0.112014
CUDA-NP: realizing nested thread-level parallelism in GPGPU applications · PPoPP 2014
High-performance computing › performance optimization
auto-tuning
0.012013
Semi-automatic restructuring of offloadable tasks for many-core accelerators · SC 2013
Parallel and multicore computing › parallel scheduling
runtime scheduling
0.012013
Semi-automatic restructuring of offloadable tasks for many-core accelerators · SC 2013
Compilers and program optimization › compiler construction
compiler algorithms
0.012012
CPU-assisted GPGPU on fused CPU-GPU architectures · HPCA 2012
GPUs and heterogeneous computing › GPU memory access
memory coalescing
0.012010
An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010

Methods — techniques the papers use, named apart from their topics

source-to-source compilation · 0.4dynamic parallelism · 0.4OpenMP-like pragmas · 0.4compiler directives · 0.3l2 prefetcher · 0.3instruction-level parallelism · 0.3memory access optimization · 0.2data layout transformation · 0.2warp-level resource allocation · 0.2dynamic resource release · 0.2pre-execution · 0.1kernel generation · 0.1architecture-specific optimization · 0.1auto-tuning · 0.1
YearPublicationVenuePosition
2022 DyCo: Dynamic, Contextualized AI Models
abstract
Devices with limited computing resources use smaller AI models to achieve low-latency inferencing. However, model accuracy is typically much lower than the accuracy of a bigger model that is trained and deployed in places where the computing resources are relatively abundant. We describe DyCo, a novel system that ensures privacy of stream data and dynamically improves the accuracy of small models used in devices. Unlike knowledge distillation or federated learning, DyCo treats AI models as black boxes. DyCo uses a semi-supervised approach to leverage existing training frameworks and network model architectures to periodically train contextualized, smaller models for resource-constrained devices. DyCo uses a bigger, highly accurate model in the edge-cloud to auto-label data received from each sensor stream. Training in the edge-cloud (as opposed to the public cloud) ensures data privacy, and bespoke models for thousands of live data streams can be designed in parallel by using multiple edge-clouds. DyCo uses the auto-labeled data to periodically re-train, stream-specific, bespoke small models. To reduce the periodic training costs, DyCo uses different policies that are based on stride, accuracy, and confidence information. We evaluate our system, and the contextualized models, by using two object detection models for vehicles and people, and two datasets (a public benchmark and another real-world proprietary dataset). Our results show that DyCo increases the mAP accuracy measure of small models by an average of 16.3% (and up to 20%) for the public benchmark and an average of 19.0% (and up to 64.9%) for the real-world dataset. DyCo also decreases the training costs for contextualized models by more than an order of magnitude.
Yi Yang 0018, Murugan Sankaradass, Srimat T. Chakradhar
ACM Trans. Embed. Comput. Syst.1
2021 Magic-Pipe: self-optimizing video analytics pipelines
abstract
Microservices-based video analytics pipelines routinely use multiple deep convolutional neural networks. We observe that the best allocation of resources to deep learning engines (or microservices) in a pipeline, and the best configuration of parameters for each engine vary over time, often at a timescale of minutes or even seconds based on the dynamic content in the video. We leverage these observations to develop Magic-Pipe, a self-optimizing video analytic pipeline that leverages AI techniques to periodically self-optimize. First, we propose a new, adaptive resource allocation technique to dynamically balance the resource usage of different microservices, based on dynamic video content. Then, we propose an adaptive microservice parameter tuning technique to balance the accuracy and performance of a microservice, also based on video content. Finally, we propose two different approaches to reduce unnecessary computations due to unavoidable mismatch of independently designed, re-usable deep-learning engines: a deep learning approach to improve the feature extractor performance by filtering inputs for which no features can be extracted, and a low-overhead graph-theoretic approach to minimize redundant computations across frames. Our evaluation of Magic-Pipe shows that pipelines augmented with self-optimizing capability exhibit application response times that are an order of magnitude better than the original pipelines, while using the same hardware resources, and achieving similar high accuracy.
Giuseppe Coviello, Yi Yang 0018, Kunal Rao, Srimat T. Chakradhar
Middleware2
2021 F3S: Free Flow Fever Screening
abstract
Identification of people with elevated body temperature can reduce or dramatically slow down the spread of infectious diseases like COVID-19. We present a novel fever-screening system, F3S, that uses edge machine learning techniques to accurately measure core body temperatures of multiple individuals in a free-flow setting. F3S performs real-time sensor fusion of visual camera with thermal camera data streams to detect elevated body temperature, and it has several unique features: (a) visual and thermal streams represent very different modalities, and we dynamically associate semantically-equivalent regions across visual and thermal frames by using a new, dynamic alignment technique that analyzes content and context in real-time, (b) we track people through occlusions, identify the eye (inner canthus), forehead, face and head regions where possible, and provide an accurate temperature reading by using a prioritized refinement algorithm, and (c) we robustly detect elevated body temperature even in the presence of personal protective equipment like masks, or sunglasses or hats, all of which can be affected by hot weather and lead to spurious temperature readings. F3S has been deployed at over a dozen large commercial establishments, providing contact-less, free-flow, real-time fever screening for thousands of employees and customers in indoors and outdoor settings.
Kunal Rao, Giuseppe Coviello, Min Feng 0001, Biplob Debnath, Wang-Pin Hsiung, Murugan Sankaradass, Yi Yang 0018, Oliver Po, Utsav Drolia, Srimat T. Chakradhar
SMARTCOMP7
2017 Accelerating deep neural network training with inconsistent stochastic gradient descent
Linnan Wang, Yi Yang 0018, Martin Renqiang Min, Srimat T. Chakradhar
Neural Networks2
2016 HppCnn: A High-Performance, Portable Deep-Learning Library for GPGPUs
abstract
The massively parallel computation capability has made GPGPUs a promising platform for convolutional neural networks (CNNs). In this paper, we present HppCnn, a CNN library achieves both the high performance and portability on GPGPUs. In HppCnn, we propose a novel three-step approach to implement convolutional kernels using Nvidia cuBLAS efficiently. To overcome limitations of our three-step approach, we improve cuBLAS by enabling nested parallelism, and implement a low-cost auto-tuning module to leveraging existing libraries in the runtime. The experiments show HppCnn achieves significant speedups over both other cuBLAS-based and hand-optimized solutions. The results also show our solution delivers near-optimal performance on GPUs with the portability.
Yi Yang 0018, Min Feng 0001, Srimat T. Chakradhar
ICPP1
2016 BLASX: A High Performance Level-3 BLAS Library for Heterogeneous Multi-GPU Computing
abstract
Basic Linear Algebra Subprograms (BLAS) are a set of low level linear algebra kernels widely adopted by applications involved with the deep learning and scientific computing. The massive and economic computing power brought forth by the emerging GPU architectures drives interest in implementation of compute-intensive level 3 BLAS on multi-GPU systems. In this paper, we investigate existing multi-GPU level 3 BLAS and present that 1) issues, such as the improper load balancing, inefficient communication, insufficient GPU stream level concurrency and data caching, impede current implementations from fully harnessing heterogeneous computing resources; 2) and the inter-GPU Peer-to-Peer(P2P) communication remains unexplored. We then present BLASX: a highly optimized multi-GPU level-3 BLAS. We adopt the concepts of algorithms-by-tiles treating a matrix tile as the basic data unit and operations on tiles as the basic task. Tasks are guided with a dynamic asynchronous runtime, which is cache and locality aware. The communication cost under BLASX becomes trivial as it perfectly overlaps communication and computation across multiple streams during asynchronous task progression. It also takes the current tile cache scheme one step further by proposing an innovative 2-level hierarchical tile cache, taking advantage of inter-GPU P2P communication. As a result, linear speedup is observable with BLASX under multi-GPU configurations; and the extensive benchmarks demonstrate that BLASX consistently outperforms the related leading industrial and academic implementations such as cuBLAS-XT, SuperMatrix, MAGMA.
Linnan Wang, Wei Wu 0016, Zenglin Xu, Jianxiong Xiao, Yi Yang 0018
ICS5
2016 Optimizing memory efficiency for deep convolutional neural networks on GPUs
abstract
Leveraging large data sets, deep Convolutional Neural Networks (CNNs) achieve state-of-the-art recognition accuracy. Due to the substantial compute and memory operations, however, they require significant execution time. The massive parallel computing capability of GPUs make them as one of the ideal platforms to accelerate CNNs and a number of GPU-based CNN libraries have been developed. While existing works mainly focus on the computational efficiency of CNNs, the memory efficiency of CNNs have been largely overlooked. Yet CNNs have intricate data structures and their memory behavior can have significant impact on the performance. In this work, we study the memory efficiency of various CNN layers and reveal the performance implication from both data layouts and memory access patterns. Experiments show the universal effect of our proposed optimizations on both single layers and various networks, with up to 27.9× for a single layer and up to 5.6× on the whole networks.
Chao Li 0004, Yi Yang 0018, Min Feng 0001, Srimat T. Chakradhar, Huiyang Zhou
SC2
2015 Revisiting ILP Designs for Throughput-Oriented GPGPU Architecture
abstract
Many-core architectures such as graphics processing units (GPUs) rely on thread-level parallelism (TLP)to overcome pipeline hazards. Consequently, each core in a many-core processor employs a relatively simple in-order pipeline with limited capability to exploit instruction-level parallelism (ILP). In this paper, we study the ILP impact on the throughput-oriented many-core architecture, including data bypassing, score boarding and branch prediction. We show that these ILP techniques significantly reduce the performance dependency on TLP. This is especially useful for applications, whose resource usage limits the hardware to run a high number of threads concurrently. Furthermore, ILP techniques reduce the demand on on-chip resource to support high TLP. Given the workload-dependent impact from ILP, we propose heterogeneous GPGPU architecture, consisting of both the cores designed for high TLP and those customized with ILPtechniques. Our results show that our heterogeneous GPUarchitecture achieves high throughput as well as high energy and area-efficiency compared to homogenous designs.
Ping Xiang, Yi Yang 0018, Mike Mantor, Norman Rubin, Huiyang Zhou
CCGRID2
2015 Automatic data placement into GPU on-chip memory resources
abstract
Although graphics processing units (GPUs) rely on thread-level parallelism to hide long off-chip memory access latency, judicious utilization of on-chip memory resources, including register files, shared memory, and data caches, is critical to application performance. However, explicitly managing GPU on-chip memory resources is a non-trivial task for application developers. More importantly, as on-chip memory resources vary among different GPU generations, performance portability has become a daunting challenge. In this paper, we tackle this problem with compiler-driven automatic data placement. We focus on programs that have already been reasonably optimized either manually by programmers or automatically by compiler tools. Our proposed compiler algorithms refine these programs by revising data placement across different types of GPU on-chip resources to achieve both performance enhancement and performance portability. Among 12 benchmarks in our study, our proposed compiler algorithm improves the performance by 1.76× on average on Nvidia GTX480, and by 1.61× on average on GTX680.
Chao Li 0004, Yi Yang 0018, Huiyang Zhou
CGO2
2015 CUDA-NP: Realizing Nested Thread-Level Parallelism in GPGPU Applications
Yi Yang 0018, Chao Li 0004, Huiyang Zhou
J. Comput. Sci. Technol.1
2014 Warp-level divergence in GPUs: Characterization, impact, and mitigation
abstract
High throughput architectures rely on high thread-level parallelism (TLP) to hide execution latencies. In state-of-art graphics processing units (GPUs), threads are organized in a grid of thread blocks (TBs) and each TB contains tens to hundreds of threads. With a TB-level resource management scheme, all the resource required by a TB is allocated/released when it is dispatched to / finished in a streaming multiprocessor (SM). In this paper, we highlight that such TB-level resource management can severely affect the TLP that may be achieved in the hardware. First, different warps in a TB may finish at different times, which we refer to as `warp-level divergence'. Due to TB-level resource management, the resources allocated to early finished warps are essentially wasted as they need to wait for the longest running warp in the same TB to finish. Second, TB-level management can lead to resource fragmentation. For example, the maximum number of threads to run on an SM in an NVIDIA GTX 480 GPU is 1536. For an application with a TB containing 1024 threads, only 1 TB can run on the SM even though it has sufficient resource for a few hundreds more threads. To overcome these inefficiencies, we propose to allocate and release resources at the warp level. Warps are dispatched to an SM as long as it has sufficient resource for a warp rather than a TB. Furthermore, whenever a warp is completed, its resource is released and can accommodate a new warp. This way, we effectively increase the number of active warps without actually increasing the size of critical resources. We present our lightweight architectural support for our proposed warp-level resource management. The experimental results show that our approach achieves up to 76.0% and an average of 16.0% performance gains and up to 21.7% and an average of 6.7% energy savings at minor hardware overhead.
Ping Xiang, Yi Yang 0018, Huiyang Zhou
HPCA2
2014 Automating and optimizing data transfers for many-core coprocessors
abstract
Orchestrating data transfers between CPUs and a coprocessor manually is cumbersome, particularly for multi-dimensional arrays and other data structures with multi-level pointers, which are common in scientific computations. This work describes a system that includes both compile-time and runtime solutions for this problem, with the overarching goal of improving programmer productivity while maintaining performance.
Bin Ren 0002, Nishkam Ravi, Yi Yang 0018, Min Feng 0001, Gagan Agrawal, Srimat T. Chakradhar
ICS3
2014 A Case for a Flexible Scalar Unit in SIMT Architecture
abstract
The wide availability and the Single-Instruction Multiple-Thread (SIMT)-style programming model have made graphics processing units (GPUs) a promising choice for high performance computing. However, because of the SIMT style processing, an instruction will be executed in every thread even if the operands are identical for all the threads. To overcome this inefficiency, the AMD's latest Graphics Core Next (GCN) architecture integrates a scalar unit into a SIMT unit. In GCN, both the SIMT unit and the scalar unit share a single SIMT style instruction stream. Depending on its type, an instruction is issued to either a scalar or a SIMT unit. In this paper, we propose to extend the scalar unit so that it can either share the instruction stream with the SIMT unit or execute a separate instruction stream. The program to be executed by the scalar unit is referred to as a scalar program and its purpose is to assist SIMT-unit execution. The scalar programs are either generated from SIMT programs automatically by the compiler or manually developed by expert developers. We make a case for our proposed flexible scalar unit through three collaborative execution paradigms: data prefetching, control divergence elimination, and scalar-workload extraction. Our experimental results show that significant performance gains can be achieved using our proposed approaches compared to the state-of-art SIMT style processing.
Yi Yang 0018, Ping Xiang, Mike Mantor, Norman Rubin, Lisa R. Hsu, Qunfeng Dong, Huiyang Zhou
IPDPS1
2014 Understanding the tradeoffs between software-managed vs. hardware-managed caches in GPUs
abstract
On-chip caches are commonly used in computer systems to hide long off-chip memory access latencies. To manage on-chip caches, either software-managed or hardware-managed schemes can be employed. State-of-art accelerators, such as the NVIDIA Fermi or Kepler GPUs and Intel's forthcoming MIC “Knights Landing” (KNL), support both software-managed caches, aka. shared memory (GPUs) or near memory (KNL), and hardware-managed L1 data caches (D-caches). Furthermore, shared memory and the L1 D-cache on a GPU utilize the same physical storage and their capacity can be configured at runtime (same for KNL). In this paper, we present an in-depth study to reveal interesting and sometimes unexpected tradeoffs between shared memory and the hardware-managed L1 D- caches in GPU architecture. In our study, the kernels utilizing the L1 D-caches are generated from those leveraging shared memory to ensure that the same optimizations such as tiling are applied equally in both versions. Our detailed analyses reveal that rather than cache hit rates, the following tradeoffs often have more profound performance impacts. On one hand, the kernels utilizing the L1 caches may support higher degrees of thread-level parallelism, offer more opportunities for data to be allocated in registers, and sometimes result in lower dynamic instruction counts. On the other hand, the applications utilizing shared memory enable more coalesced accesses and tend to achieve higher degrees of memory-level parallelism. Overall, our results show that most benchmarks perform significantly better with shared memory than the L1 D-caches due to the high impact of memory-level parallelism and memory coalescing.
Chao Li 0004, Yi Yang 0018, Hongwen Dai, Shengen Yan, Frank Mueller 0001, Huiyang Zhou
ISPASS2
2014 COMP: Compiler Optimizations for Manycore Processors
abstract
Applications executing on multicore processors can now easily offload computations to many core processors, such as Intel Xeon Phi coprocessors. However, it requires high levels of expertise and effort to tune such offloaded applications to realize high-performance execution. Previous efforts have focused on optimizing the execution of offloaded computations on many core processors. However, we observe that the data transfer overhead between multicore and many core processors, and the limited device memories of many core processors often constrain the performance gains that are possible by offloading computations. In this paper, we present three source-to-source compiler optimizations that can significantly improve the performance of applications that offload computations to many core processors. The first optimization automatically transforms offloaded codes to enable data streaming, which overlaps data transfer between multicore and many core processors with computations on these processors to hide data transfer overhead. This optimization is also designed to minimize the memory usage on many core processors, while achieving the optimal performance. The second compiler optimization re-orders computations to regularize irregular memory accesses. It enables data streaming and factorization on many core processors, even when the memory access patterns in the original source codes are irregular. Finally, our new shared memory mechanism provides efficient support for transferring large pointer-based data structures between hosts and many core processors. Our evaluation shows that the proposed compiler optimizations benefit 9 out of 12 benchmarks. Compared with simply offloading the original parallel implementations of these benchmarks, we can achieve 1.16x-52.21x speedups.
Linhai Song, Min Feng 0001, Nishkam Ravi, Yi Yang 0018, Srimat T. Chakradhar
MICRO4
2014 CUDA-NP: realizing nested thread-level parallelism in GPGPU applications
abstract
Parallel programs consist of series of code sections with different thread-level parallelism (TLP). As a result, it is rather common that a thread in a parallel program, such as a GPU kernel in CUDA programs, still contains both se-quential code and parallel loops. In order to leverage such parallel loops, the latest Nvidia Kepler architecture intro-duces dynamic parallelism, which allows a GPU thread to start another GPU kernel, thereby reducing the overhead of launching kernels from a CPU. However, with dynamic parallelism, a parent thread can only communicate with its child threads through global memory and the overhead of launching GPU kernels is non-trivial even within GPUs. In this paper, we first study a set of GPGPU benchmarks that contain parallel loops, and highlight that these bench-marks do not have a very high loop count or high degrees of TLP. Consequently, the benefits of leveraging such par-allel loops using dynamic parallelism are too limited to offset its overhead. We then present our proposed solution to exploit nested parallelism in CUDA, referred to as CUDA-NP. With CUDA-NP, we initially enable a high number of threads when a GPU program starts, and use control flow to activate different numbers of threads for different code sections. We implemented our proposed CUDA-NP framework using a directive-based compiler approach. For a GPU kernel, an application developer only needs to add OpenMP-like pragmas for parallelizable code sections. Then, our CUDA-NP compiler automatically gen-erates the optimized GPU kernels. It supports both the reduction and the scan primitives, explores different ways to distribute parallel loop iterations into threads, and effi-ciently manages on-chip resource. Our experiments show that for a set of GPGPU benchmarks, which have already been optimized and contain nested parallelism, our pro-posed CUDA-NP framework further improves the perfor-mance by up to 6.69 times and 2.18 times on average.
Yi Yang 0018, Huiyang Zhou
PPoPP1
2013 Exploiting uniform vector instructions for GPGPU performance, energy efficiency, and opportunistic reliability enhancement
abstract
State-of-art graphics processing units (GPUs) employ the single-instruction multiple-data (SIMD) style execution to achieve both high computational throughput and energy efficiency. As previous works have shown, there exists significant computational redundancy in SIMD execution, where different execution lanes operate on the same operand values. Such value locality is referred to as uniform vectors. In this paper, we first show that besides redundancy within a uniform vector, different vectors can also have the identical values. Then, we propose detailed architecture designs to exploit both types of redundancy. For redundancy within a uniform vector, we propose to either extend the vector register file with token bits or add a separate small scalar register file to eliminate redundant computations as well as redundant data storage. For redundancy across different uniform vectors, we adopt instruction reuse, proposed originally for CPU architectures, to detect and eliminate redundancy. The elimination of redundant computations and data storage leads to both significant energy savings and performance improvement. Furthermore, we propose to leverage such redundancy to protect arithmetic-logic units (ALUs) and register files against hardware errors. Our detailed evaluation shows that our proposed design has low hardware overhead and achieves performance gains, up to 23.9% and 12.0% on average, along with energy savings, up to 24.8% and 12.6% on average, as well as a 21.1% and 14.1% protection coverage for ALUs and register files, respectively.
Ping Xiang, Yi Yang 0018, Mike Mantor, Norman Rubin, Lisa R. Hsu, Huiyang Zhou
ICS2
2013 Semi-automatic restructuring of offloadable tasks for many-core accelerators
abstract
Work division between the processor and accelerator is a common theme in modern heterogenous computing. Recent efforts (such as LEO and OpenAcc) provide directives that allow the developer to mark code regions in the original application from which offloadable tasks can be generated by the compiler. Auto-tuners and runtime schedulers work with the options (i.e., offloadable tasks) generated at compile time, which is limited by the directives specified by the developer. There is no provision for offload restructuring.
Nishkam Ravi, Yi Yang 0018, Srimat T. Chakradhar
SC2
2013 Locality principle revisited: A probability-based quantitative approach
Saurabh Gupta 0002, Ping Xiang, Yi Yang 0018, Huiyang Zhou
J. Parallel Distributed Comput.3
2012 Many-thread aware instruction-level parallelism: architecting shader cores for GPU computing
abstract
No abstract available.
Ping Xiang, Yi Yang 0018, Mike Mantor, Norman Rubin, Huiyang Zhou
PACT2
2012 Shared memory multiplexing: a novel way to improve GPGPU throughput
abstract
On-chip shared memory (a.k.a. local data share) is a critical resource to many GPGPU applications. In current GPUs, the shared memory is allocated when a thread block (also called a workgroup) is dispatched to a streaming multiprocessor (SM) and is released when the thread block is completed. As a result, the limited capacity of shared memory becomes a bottleneck for a GPU to host a high number of thread blocks, limiting the otherwise available thread-level parallelism (TLP). In this paper, we propose software and/or hardware approaches to multiplex the shared memory among multiple thread blocks.
Yi Yang 0018, Ping Xiang, Mike Mantor, Norman Rubin, Huiyang Zhou
PACT1
2012 CPU-assisted GPGPU on fused CPU-GPU architectures
abstract
This paper presents a novel approach to utilize the CPU resource to facilitate the execution of GPGPU programs on fused CPU-GPU architectures. In our model of fused architectures, the GPU and the CPU are integrated on the same die and share the on-chip L3 cache and off-chip memory, similar to the latest Intel Sandy Bridge and AMD accelerated processing unit (APU) platforms. In our proposed CPU-assisted GPGPU, after the CPU launches a GPU program, it executes a pre-execution program, which is generated automatically from the GPU kernel using our proposed compiler algorithms and contains memory access instructions of the GPU kernel for multiple thread-blocks. The CPU pre-execution program runs ahead of GPU threads because (1) the CPU pre-execution thread only contains memory fetch instructions from GPU kernels and not floating-point computations, and (2) the CPU runs at higher frequencies and exploits higher degrees of instruction-level parallelism than GPU scalar cores. We also leverage the prefetcher at the L2-cache on the CPU side to increase the memory traffic from CPU. As a result, the memory accesses of GPU threads hit in the L3 cache and their latency can be drastically reduced. Since our pre-execution is directly controlled by user-level applications, it enjoys both high accuracy and flexibility. Our experiments on a set of benchmarks show that our proposed pre-execution improves the performance by up to 113% and 21.4% on average.
Yi Yang 0018, Ping Xiang, Mike Mantor, Huiyang Zhou
HPCA1
2012 Fixing Performance Bugs: An Empirical Study of Open-Source GPGPU Programs
abstract
Given the extraordinary computational power of modern graphics processing units (GPUs), general purpose computation on GPUs (GPGPU) has become an increasingly important platform for high performance computing. To better understand how well the GPU resource has been utilized by application developers and then to facilitate them to develop high performance GPGPU code, we conduct an empirical study on GPGPU programs from ten open-source projects. These projects span a wide range of disciplines and many are designed as high performance libraries. Among these projects, we found various performance 'bugs', i.e., code segments leading to inefficient use of GPU hardware. We characterize these performance bugs, and propose the bug fixes. Our experiments confirm both significant performance gains and energy savings from our fixes and reveal interesting insights on different GPUs.
Yi Yang 0018, Ping Xiang, Mike Mantor, Huiyang Zhou
ICPP1
2012 Apricot: an optimizing compiler and productivity tool for x86-compatible many-core coprocessors
abstract
Intel MIC (Many Integrated Core) is the first x86-based coprocessor architecture aimed at accelerating multi-core HPC applications. In the most common usage model, parallel code sections are offloaded to the MIC coprocessor using LEO (Language Extensions for Offload). The developer is responsible for identifying and specifying offloadable code regions, managing data transfers between the CPU and MIC and optimizing the application for performance, which requires some amount of effort and experimentation. In this paper, we present Apricot, an optimizing compiler and productivity tool for x86-compatible many-core coprocessors (such as Intel MIC) that minimizes developer effort by (i) automatically inserting LEO clauses for parallelizable code regions, (ii) selectively offloading some of the code regions to the coprocessor at runtime based on a cost model that we have developed, (iii) applying a set ofoptimizations for minimizing the data communication overhead and improving overall performance. Apricot is intended to assist programmers in porting existing multi-core applications and writing new ones to take advantage of the many-core coprocessor, while maximizing overall performance. Experiments with SpecOMP and NAS Parallel benchmarks show that Apricot can successfully transform OpenMP applications to run on the MIC coprocessor with good performance gains.
Nishkam Ravi, Yi Yang 0018, Srimat T. Chakradhar
ICS2
2012 Locality Principle Revisited: A Probability-Based Quantitative Approach
abstract
This paper revisits the fundamental concept of the locality of references and proposes to quantify it as a conditional probability: in an address stream, given the condition that an address is accessed, how likely the same address (temporal locality) or an address within its neighborhood (spatial locality) will be accessed in the near future. Based on this definition, spatial locality is a function of two parameters, the neighborhood size and the scope of near future, and can be visualized with a 3D mesh. Temporal locality becomes a special case of spatial locality with the neighborhood size being zero byte. Previous works on locality analysis use stack/reuse distances to compute distance histograms as a measure of temporal locality. For spatial locality, some ad-hoc metrics have been proposed as a quantitative measure. In contrast, our conditional probability-based locality measure has a clear mathematical meaning, offers justification for distance histograms, and provides a theoretically sound and unified way to quantify both temporal and spatial locality. The proposed locality measure clearly exhibits the inherent application characteristics, from which we can easily derive information such as the sizes of the working data sets and how locality can be exploited. We showcase that our quantified locality visualized in 3D-meshes can be used to evaluate compiler optimizations, to analyze the locality at different levels of memory hierarchy, to optimize the cache architecture to effectively leverage the locality, and to examine the effect of data prefetching mechanisms. A GPU-based parallel algorithm is also presented to accelerate the locality computation for large address traces.
Saurabh Gupta 0002, Ping Xiang, Yi Yang 0018, Huiyang Zhou
IPDPS3
2012 A unified optimizing compiler framework for different GPGPU architectures
abstract
This article presents a novel optimizing compiler for general purpose computation on graphics processing units (GPGPU). It addresses two major challenges of developing high performance GPGPU programs: effective utilization of GPU memory hierarchy and judicious management of parallelism. The input to our compiler is a naïve GPU kernel function, which is functionally correct but without any consideration for performance optimization. The compiler generates two kernels, one optimized for global memories and the other for texture memories. The proposed compilation process is effective for both AMD/ATI and NVIDIA GPUs. The experiments show that our optimized code achieves very high performance, either superior or very close to highly fine-tuned libraries.
Yi Yang 0018, Ping Xiang, Jingfei Kong, Mike Mantor, Huiyang Zhou
ACM Trans. Archit. Code Optim.1
2010 A GPGPU compiler for memory optimization and parallelism management
abstract
This paper presents a novel optimizing compiler for general purpose computation on graphics processing units (GPGPU). It addresses two major challenges of developing high performance GPGPU programs: effective utilization of GPU memory hierarchy and judicious management of parallelism.
Yi Yang 0018, Ping Xiang, Jingfei Kong, Huiyang Zhou
PLDI1
2010 An optimizing compiler for GPGPU programs with input-data sharing
abstract
Developing high performance GPGPU programs is challenging for application developers since the performance is dependent upon how well the code leverages the hardware features of specific graphics processors. To solve this problem and relieve application developers of low-level hardware-specific optimizations, we introduce a novel compiler to optimize GPGPU programs. Our compiler takes a naive GPU kernel function, which is functionally correct but without any consideration for performance optimization. The compiler then analyzes the code, identifies memory access patterns, and generates optimized code. The proposed compiler optimizations target at one category of scientific and media processing algorithms, which has the characteristics of input-data sharing when computing neighboring output pixels/elements. Many commonly used algorithms, such as matrix multiplication, convolution, etc., share such characteristics. For these algorithms, novel approaches are proposed to enforce memory coalescing and achieve effective data reuse. Data prefetching and hardware-specific tuning are also performed automatically with our compiler framework. The experimental results based on a set of applications show that our compiler achieves very high performance, either superior or very close to the highly fine-tuned library, NVIDIA CUBLAS 2.1.
Yi Yang 0018, Ping Xiang, Jingfei Kong, Huiyang Zhou
PPoPP1