EDBT 2026 Demo / reviewers in the wild / expert
Janghaeng Lee
dblp:51/5477
· DBLP profile ↗
11ranked-venue papers
5as first author
0since 2021 · last 2019
0009-0005-1915-6077ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
GPUs and heterogeneous computing · 47% Parallel and multicore computing · 30% Emerging computing paradigms · 21% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Emerging computing paradigms
approximate computing |
0.4 | 3 | 2014 | Scaling Performance via Self-Tuning Approximation for Graphics Engines · ACM Trans. Comput. Syst. 2014 Paraprox: pattern-based approximation for data parallel applications · ASPLOS 2014 SAGE: self-tuning approximation for graphics engines · MICRO 2013 |
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU co-processing |
0.2 | 1 | 2015 | SKMD: Single Kernel on Multiple Devices for Transparent CPU-GPU Collaboration · ACM Trans. Comput. Syst. 2015 |
Parallel and multicore computing
data-parallel programming |
0.2 | 1 | 2014 | Paraprox: pattern-based approximation for data parallel applications · ASPLOS 2014 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 1 | 2014 | Paraprox: pattern-based approximation for data parallel applications · ASPLOS 2014 |
Parallel and multicore computing
speculative parallelization |
0.2 | 1 | 2014 | Leveraging GPUs using cooperative loop speculation · ACM Trans. Archit. Code Optim. 2014 |
GPUs and heterogeneous computing › GPU computing
approximate GPU arithmetic |
0.2 | 1 | 2013 | SAGE: self-tuning approximation for graphics engines · MICRO 2013 |
Compilers and program optimization › dynamic optimization
adaptive compilation |
0.1 | 1 | 2012 | Adaptive input-aware compilation for graphics engines · PLDI 2012 |
GPUs and heterogeneous computing
GPU compilation |
0.1 | 1 | 2012 | Adaptive input-aware compilation for graphics engines · PLDI 2012 |
Parallel and multicore computing › parallel programming runtimes
thread management |
0.1 | 1 | 2010 | Thread tailor: dynamically weaving threads together for efficient, adaptive parallel applications · ISCA 2010 |
Parallel and multicore computing
parallel programming models |
0.1 | 2 | 2015 | SKMD: Single Kernel on Multiple Devices for Transparent CPU-GPU Collaboration · ACM Trans. Comput. Syst. 2015 Thread tailor: dynamically weaving threads together for efficient, adaptive parallel applications · ISCA 2010 |
GPUs and heterogeneous computing
GPU performance optimization |
0.0 | 1 | 2012 | Adaptive input-aware compilation for graphics engines · PLDI 2012 |
Memory systems › memory hierarchy
memory hierarchy optimization |
0.0 | 1 | 2012 | Adaptive input-aware compilation for graphics engines · PLDI 2012 |
Parallel and multicore computing
thread-level parallelism |
0.0 | 1 | 2010 | Thread tailor: dynamically weaving threads together for efficient, adaptive parallel applications · ISCA 2010 |
Methods — techniques the papers use, named apart from their topics
runtime kernel selection · 0.5thread fusion · 0.4static compiler · 0.4distributed conflict detection · 0.4data packing · 0.4workload partitioning · 0.2output merging · 0.2kernel generation · 0.2runtime tuning · 0.2pattern-based approximation · 0.2streaming language · 0.1runtime specialization · 0.1adaptive compilation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Multi-objective Exploration for Practical Optimization Decisions in Binary TranslationabstractIn the design of mobile systems, hardware/software (HW/SW) co-design has important advantages by creating specialized hardware for the performance or power optimizations. Dynamic binary translation (DBT) is a key component in co-design. During the translation, a dynamic optimizer in the DBT system applies various software optimizations to improve the quality of the translated code. With dynamic optimization, optimization time is an exposed run-time overhead and useful analyses are often restricted due to their high costs. Thus, a dynamic optimizer needs to make smart decisions with limited analysis information, which complicates the design of optimization decision models and often causes failures in human-made heuristics. In mobile systems, this problem is even more challenging because of strict constraints on computing capabilities and memory size. To overcome the challenge, we investigate an opportunity to build practical optimization decision models for DBT by using machine learning techniques. As the first step, loop unrolling is chosen as the representative optimization. We base our approach on the industrial strength DBT infrastructure and conduct evaluation with 17,116 unrollable loops collected from 200 benchmarks and real-life programs across various domains. By utilizing all available features that are potentially important for loop unrolling decision, we identify the best classification algorithm for our infrastructure with consideration for both prediction accuracy and cost. The greedy feature selection algorithm is then applied to the classification algorithm to distinguish its significant features and cut down the feature space. By maintaining significant features only, the best affordable classifier, which satisfies the budgets allocated to the decision process, shows 74.5% of prediction accuracy for the optimal unroll factor and realizes an average 20.9% reduction in dynamic instruction count during the steady-state translated code execution. For comparison, the best baseline heuristic achieves 46.0% prediction accuracy with an average 13.6% instruction count reduction. Given that the infrastructure is already highly optimized and the ideal upper bound for instruction reduction is observed at 23.8%, we believe this result is noteworthy. Sunghyun Park 0004, Youfeng Wu, Janghaeng Lee, Amir Ayupov, Scott A. Mahlke |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | Orchestrating Multiple Data-Parallel Kernels on Multiple DevicesabstractTraditionally, programmers and software tools have focused on mapping a single data-parallel kernel onto a heterogeneous computing system consisting of multiple general-purpose processors (CPUS) and graphics processing units (GPUs). These methodologies break down as application complexity grows to contain multiple communicating data-parallel kernels. This paper introduces MKMD, an automatic system for mapping multiple kernels across multiple computing devices in a seamless manner. MKMD is a two phased approach that combines coarse grain scheduling of indivisible kernels followed by opportunistic fine-grained workgroup-level partitioning to exploit idle resources. During this process, MKMD considers kernel dependencies and the underlying systems along with the execution time model built with a few sets of profile data. With the scheduling decision, MKMD transparently manages the order of executions and data transfers for each device. On a real machine with one CPU and two different GPUs, MKMD achieves a mean speedup of 1.89x compared to the in-order execution on the fastest device for a set of applications with multiple kernels. 53% of this speedup comes from the coarse-grained scheduling and the other 47% is the result of the fine-grained partitioning. Janghaeng Lee, Mehrzad Samadi, Scott A. Mahlke |
PACT | 1 |
| 2015 | SKMD: Single Kernel on Multiple Devices for Transparent CPU-GPU CollaborationabstractHeterogeneous computing on CPUs and GPUs has traditionally used fixed roles for each device: the GPU handles data parallel work by taking advantage of its massive number of cores while the CPU handles non data-parallel work, such as the sequential code or data transfer management. This work distribution can be a poor solution as it underutilizes the CPU, has difficulty generalizing beyond the single CPU-GPU combination, and may waste a large fraction of time transferring data. Further, CPUs are performance competitive with GPUs on many workloads, thus simply partitioning work based on the fixed roles may be a poor choice. In this article, we present the single-kernel multiple devices (SKMD) system, a framework that transparently orchestrates collaborative execution of a single data-parallel kernel across multiple asymmetric CPUs and GPUs. The programmer is responsible for developing a single data-parallel kernel in OpenCL, while the system automatically partitions the workload across an arbitrary set of devices, generates kernels to execute the partial workloads, and efficiently merges the partial outputs together. The goal is performance improvement by maximally utilizing all available resources to execute the kernel. SKMD handles the difficult challenges of exposed data transfer costs and the performance variations GPUs have with respect to input size. On real hardware, SKMD achieves an average speedup of 28% on a system with one multicore CPU and two asymmetric GPUs compared to a fastest device execution strategy for a set of popular OpenCL kernels. Janghaeng Lee, Mehrzad Samadi, Yongjun Park 0001, Scott A. Mahlke |
ACM Trans. Comput. Syst. | 1 |
| 2014 | VAST: the illusion of a large memory space for GPUsabstractHeterogeneous systems equipped with traditional processors (CPUs) and graphics processing units (GPUs) have enabled processing large data sets. With new programming models, such as OpenCL and CUDA, programmers are encouraged to offload data parallel workloads to GPUs as much as possible in order to fully utilize the available resources. Unfortunately, offloading work is strictly limited by the size of the physical memory on a specific GPU. In this paper, we present Virtual Address Space for Throughput processors (VAST), an automatic GPU memory management system that provides an OpenCL program with the illusion of a virtual memory space. Based on the available physical memory on the target GPU, VAST does the following: automatically partitions the data parallel workload into chunks; efficiently extracts the precise working set required for the divided workload; rearranges the working set in contiguous memory space; and, transforms the kernel to operate on the reorganized working set. With VAST, the programmer is responsible for developing a data parallel kernel in OpenCL without concern for physical memory space limitations of individual GPUs. VAST transparently handles code generation dealing with the constraints of the actual physical memory and improves the retargetability of the OpenCL with moderate overhead. Experiments demonstrate that a real GPU, NVIDIA GTX 760 with 2 GB of memory, can compute any size of data without program changes achieving 2.6x speedup over CPU exeuction, which is a realistic alternative for large data computation. Janghaeng Lee, Mehrzad Samadi, Scott A. Mahlke |
PACT | 1 |
| 2014 | Paraprox: pattern-based approximation for data parallel applicationsabstractApproximate computing is an approach where reduced accuracy of results is traded off for increased speed, throughput, or both. Loss of accuracy is not permissible in all computing domains, but there are a growing number of data-intensive domains where the output of programs need not be perfectly correct to provide useful results or even noticeable differences to the end user. These soft domains include multimedia processing, machine learning, and data mining/analysis. An important challenge with approximate computing is transparency to insulate both software and hardware developers from the time, cost, and difficulty of using approximation. This paper proposes a software-only system, Paraprox, for realizing transparent approximation of data-parallel programs that operates on commodity hardware systems. Paraprox starts with a data-parallel kernel implemented using OpenCL or CUDA and creates a parameterized approximate kernel that is tuned at runtime to maximize performance subject to a target output quality (TOQ) that is supplied by the user. Approximate kernels are created by recognizing common computation idioms found in data-parallel programs (e.g., Map, Scatter/Gather, Reduction, Scan, Stencil, and Partition) and substituting approximate implementations in their place. Across a set of 13 soft data-parallel applications with at most 10% quality degradation, Paraprox yields an average performance gain of 2.7x on a NVIDIA GTX 560 GPU and 2.5x on an Intel Core i7 quad-core processor compared to accurate execution on each platform. Mehrzad Samadi, Davoud Anoushe Jamshidi, Janghaeng Lee, Scott A. Mahlke |
ASPLOS | 3 |
| 2014 | Leveraging GPUs using cooperative loop speculationabstractGraphics processing units, or GPUs, provide TFLOPs of additional performance potential in commodity computer systems that frequently go unused by most applications. Even with the emergence of languages such as CUDA and OpenCL, programming GPUs remains a difficult challenge for a variety of reasons, including the inherent algorithmic characteristics and data structure choices used by applications as well as the tedious performance optimization cycle that is necessary to achieve high performance. The goal of this work is to increase the applicability of GPUs beyond CUDA/OpenCL to implicitly data-parallel applications written in C/C++ using speculative parallelization. To achieve this goal, we propose Paragon : a static/dynamic compiler platform to speculatively run possibly data-parallel portions of sequential applications on the GPU while cooperating with the system CPU. For such loops, Paragon utilizes the GPU in an opportunistic way while orchestrating a cooperative relation between the CPU and GPU to reduce the overhead of miss-speculations. Paragon monitors the dependencies for the loops running speculatively on the GPU and nonspeculatively on the CPU using a lightweight distributed conflict detection designed specifically for GPUs, and transfers the execution to the CPU in case a conflict is detected. Paragon resumes the execution on the GPU after the CPU resolves the dependency. Our experiments show that Paragon achieves 4x on average and up to 30x speedup compared to unsafe CPU execution with four threads and 7x on average and up to 64x speedup versus sequential execution across a set of sequential but implicitly data-parallel applications. Mehrzad Samadi, Amir Hormati, Janghaeng Lee, Scott A. Mahlke |
ACM Trans. Archit. Code Optim. | 3 |
| 2014 | Scaling Performance via Self-Tuning Approximation for Graphics EnginesabstractApproximate computing, where computation accuracy is traded off for better performance or higher data throughput, is one solution that can help data processing keep pace with the current and growing abundance of information. For particular domains, such as multimedia and learning algorithms, approximation is commonly used today. We consider automation to be essential to provide transparent approximation, and we show that larger benefits can be achieved by constructing the approximation techniques to fit the underlying hardware. Our target platform is the GPU because of its high performance capabilities and difficult programming challenges that can be alleviated with proper automation. Our approach—SAGE—combines a static compiler that automatically generates a set of CUDA kernels with varying levels of approximation with a runtime system that iteratively selects among the available kernels to achieve speedup while adhering to a target output quality set by the user. The SAGE compiler employs three optimization techniques to generate approximate kernels that exploit the GPU microarchitecture: selective discarding of atomic operations, data packing, and thread fusion. Across a set of machine learning and image processing kernels, SAGE's approximation yields an average of 2.5× speedup with less than 10% quality loss compared to the accurate execution on a NVIDIA GTX 560 GPU. Mehrzad Samadi, Janghaeng Lee, Davoud Anoushe Jamshidi, Scott A. Mahlke, Amir Hormati |
ACM Trans. Comput. Syst. | 2 |
| 2013 | Transparent CPU-GPU collaboration for data-parallel kernels on heterogeneous systemsabstractHeterogeneous computing on CPUs and GPUs has traditionally used fixed roles for each device: the GPU handles data parallel work by taking advantage of its massive number of cores while the CPU handles non data-parallel work, such as the sequential code or data transfer management. Unfortunately, this work distribution can be a poor solution as it under utilizes the CPU, has difficulty generalizing beyond the single CPU-GPU combination, and may waste a large fraction of time transferring data. Further, CPUs are performance competitive with GPUs on many workloads, thus simply partitioning work based on the fixed roles may be a poor choice. In this paper, we present the single kernel multiple devices (SKMD) system, a framework that transparently orchestrates collaborative execution of a single data-parallel kernel across multiple asymmetric CPUs and GPUs. The programmer is responsible for developing a single data-parallel kernel in OpenCL, while the system automatically partitions the workload across an arbitrary set of devices, generates kernels to execute the partial workloads, and efficiently merges the partial outputs together. The goal is performance improvement by maximally utilizing all available resources to execute the kernel. SKMD handles the difficult challenges of exposed data transfer costs and the performance variations GPUs have with respect to input size. On real hardware, SKMD achieves an average speedup of 29% on a system with one multicore CPU and two asymmetric GPUs compared to a fastest device execution strategy for a set of popular OpenCL kernels. Janghaeng Lee, Mehrzad Samadi, Yongjun Park 0001, Scott A. Mahlke |
PACT | 1 |
| 2013 | SAGE: self-tuning approximation for graphics enginesabstractApproximate computing, where computation accuracy is traded off for better performance or higher data throughput, is one solution that can help data processing keep pace with the current and growing overabundance of information. For particular domains such as multimedia and learning algorithms, approximation is commonly used today. We consider automation to be essential to provide transparent approximation and we show that larger benefits can be achieved by constructing the approximation techniques to fit the underlying hardware. Our target platform is the GPU because of its high performance capabilities and difficult programming challenges that can be alleviated with proper automation. Our approach, SAGE, combines a static compiler that automatically generates a set of CUDA kernels with varying levels of approximation with a run-time system that iteratively selects among the available kernels to achieve speedup while adhering to a target output quality set by the user. The SAGE compiler employs three optimization techniques to generate approximate kernels that exploit the GPU microarchitecture: selective discarding of atomic operations, data packing, and thread fusion. Across a set of machine learning and image processing kernels, SAGE's approximation yields an average of 2.5x speedup with less than 10% quality loss compared to the accurate execution on a NVIDIA GTX 560 GPU. Mehrzad Samadi, Janghaeng Lee, Davoud Anoushe Jamshidi, Amir Hormati, Scott A. Mahlke |
MICRO | 2 |
| 2012 | Adaptive input-aware compilation for graphics enginesabstractWhile graphics processing units (GPUs) provide low-cost and efficient platforms for accelerating high performance computations, the tedious process of performance tuning required to optimize applications is an obstacle to wider adoption of GPUs. In addition to the programmability challenges posed by GPU's complex memory hierarchy and parallelism model, a well-known application design problem is target portability across different GPUs. However, even for a single GPU target, changing a program's input characteristics can make an already-optimized implementation of a program perform poorly. In this work, we propose Adaptic, an adaptive input-aware compilation system to tackle this important, yet overlooked, input portability problem. Using this system, programmers develop their applications in a high-level streaming language and let Adaptic undertake the difficult task of input portable optimizations and code generation. Several input-aware optimizations are introduced to make efficient use of the memory hierarchy and customize thread composition. At runtime, a properly optimized version of the application is executed based on the actual program input. We perform a head-to-head comparison between the Adaptic generated and hand-optimized CUDA programs. The results show that Adaptic is capable of generating codes that can perform on par with their hand-optimized counterparts over certain input ranges and outperform them when the input falls out of the hand-optimized programs' "comfort zone". Furthermore, we show that input-aware results are sustainable across different GPU targets making it possible to write and optimize applications once and run them anywhere. Mehrzad Samadi, Amir Hormati, Mojtaba Mehrara, Janghaeng Lee, Scott A. Mahlke |
PLDI | 4 |
| 2010 | Thread tailor: dynamically weaving threads together for efficient, adaptive parallel applicationsabstractExtracting performance from modern parallel architectures requires that applications be divided into many different threads of execution. Unfortunately selecting the appropriate number of threads for an application is a daunting task. Having too many threads can quickly saturate shared resources, such as cache capacity or memory bandwidth, thus degrading performance. On the other hand, having too few threads makes inefficient use of the resources available. Beyond static resource assignment, the program inputs and dynamic system state (e.g., what other applications are executing in the system) can have a significant impact on the right number of threads to use for a particular application. Janghaeng Lee, Haicheng Wu, Madhumitha Ravichandran, Nathan Clark |
ISCA | 1 |