VLDB 2026 Research / reviewers in the wild / expert
Andrew Kerr
dblp:59/6625
· DBLP profile ↗
8ranked-venue papers
2as first author
1since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 92% Programming languages and type systems · 8% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
GPUs and heterogeneous computing · 68% Processor architecture and microarchitecture · 24% Parallel and multicore computing · 7% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
deep learning compiler |
0.8 | 1 | 2024 | EVT: Accelerating Deep Learning Training with Epilogue Visitor Tree · ASPLOS (3) 2024 |
Compilers and program optimization › deep learning compiler
operator fusion |
0.8 | 1 | 2024 | EVT: Accelerating Deep Learning Training with Epilogue Visitor Tree · ASPLOS (3) 2024 |
GPUs and heterogeneous computing › GPU compilation
GPU kernel generation |
0.2 | 1 | 2024 | EVT: Accelerating Deep Learning Training with Epilogue Visitor Tree · ASPLOS (3) 2024 |
Programming languages and type systems
control flow |
0.1 | 1 | 2011 | SIMD re-convergence at thread frontiers · MICRO 2011 |
GPUs and heterogeneous computing › control flow divergence
branch divergence |
0.1 | 1 | 2011 | SIMD re-convergence at thread frontiers · MICRO 2011 |
Processor architecture and microarchitecture
SIMD |
0.1 | 1 | 2011 | SIMD re-convergence at thread frontiers · MICRO 2011 |
Parallel and multicore computing
data-parallel programming |
0.0 | 1 | 2011 | SIMD re-convergence at thread frontiers · MICRO 2011 |
Methods — techniques the papers use, named apart from their topics
kernel fusion · 1.5epilogue visitor tree · 1.5thread frontier · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | EVT: Accelerating Deep Learning Training with Epilogue Visitor TreeabstractAs deep learning models become increasingly complex, the deep learning compilers are critical for enhancing the system efficiency and unlocking hidden optimization opportunities. Although excellent speedups have been achieved in inference workloads, existing compilers face significant limitations in training. Firstly, the training computation graph involves intricate operations challenging to fuse, such as normalization, loss functions, and reductions, which limit optimization opportunities like kernel fusion. Secondly, the training graph's additional edges connecting forward and backward operators pose challenges in finding optimal and feasible partitions for kernel fusion. More importantly, existing compilers cannot either generate kernels with state-of-the-art performance on modern GPUs or accommodate diverse fusion patterns. Zhaodong Chen 0001, Andrew Kerr, Richard Cai, Jack Kosaian, Haicheng Wu, Yufei Ding 0001, Yuan Xie 0001 |
ASPLOS (3) | 2 |
| 2014 | Upper limb movement analysis via marker tracking with a single-camera systemabstractOptical motion capture systems have been widely adopted for human motion analysis in stroke rehabilitation because of real-time processing and high-accuracy features. However, these systems require a large laboratory space and multiple cameras and thus can be expensive and not transportable. In this paper, we propose a portable, cheap, single-camera motion analysis system to implement upper limb movement analysis. The proposed system consists of video acquisition, camera calibration, marker tracking, autonomous joint angle calculation, visualization, validation and classification. The validation with a state-of-the-art optical motion analysis system using Bland-Altman plot, a typical clinical measure, indicates that the proposed system can accurately capture elbow movement, trunk-tilt, and shoulder movement for diagnosis. Furthermore, the volunteers are explicitly classified into healthy and stroke groups via a support vector machine trained on statistics of the trunk-tilt and shoulder movement. Experimental results show that the proposed system can accurately capture the upper limb movement patterns, automatically classify stroke survivors using ordinal scale classification of upper limb impairment, and offer a convenient and inexpensive solution for upper limb movement analysis. Cheng Yang 0003, Andrew Kerr, Vladimir Stankovic 0001, Lina Stankovic, Philip J. Rowe |
ICIP | 2 |
| 2012 | Dynamic compilation of data-parallel kernels for vector processorsabstractModern processors enjoy augmented throughput and power efficiency through specialized functional units leveraged via instruction set extensions. These functional units accelerate performance for specific types of operations but must be programmed explicitly. Moreover, applications targeting these specialized units will not take advantage of future ISA extensions and tend not to be portable across multiple ISAs. As architecture designers increasingly rely on heterogeneity for performance improvements, the challenges of leveraging specialized functional units will only become more critical. In particular, exploiting software parallelism without sacrificing portability across the spectrum of commodity and multi-core SIMD processors remains elusive. Andrew Kerr, Gregory Frederick Diamos, Sudhakar Yalamanchili |
CGO | 1 |
| 2012 | Eiger: A framework for the automated synthesis of statistical performance modelsabstractAs processor architectures continue to evolve to increasingly heterogeneous and asymmetric designs, the construction of accurate performance models of execution time and energy consumption has become increasingly more challenging. Models that are constructed, are quickly invalidated by new features in the next generation of processors while many interactions between application and architecture parameters are often simply not obvious or even apparent. Consequently, we foresee a need for an automated methodology for the systematic construction of performance models of heterogeneous processors. The methodology should be founded on rigorous mathematical techniques yet leave room for the exploration and adaptation of a space of analytic models. Our current effort toward creating such an extensible, targeted methodology is Eiger. This paper describes the methodology implemented in Eiger, the specifics of Eiger's extensible implementation and the results of one scenario in which Eiger has been applied - the synthesis of performance models for use in the simulation-based design space exploration of Exascale architectures. Andrew Kerr, Eric Anger, Gilbert Hendry, Sudhakar Yalamanchili |
HiPC | 1 |
| 2012 | Lynx: A dynamic instrumentation system for data-parallel applications on GPGPU architecturesabstractAs parallel execution platforms continue to proliferate, there is a growing need for real-time introspection tools to provide insight into platform behavior for performance debugging, correctness checks, and to drive effective resource management schemes. To address this need, we present the Lynx dynamic instrumentation system. Lynx provides the capability to write instrumentation routines that are (1) selective, instrumenting only what is needed, (2) transparent, without changes to the applications' source code, (3) customizable, and (4) efficient. Lynx is embedded into the broader GPU Ocelot system, which provides run-time code generation of CUDA programs for heterogeneous architectures. This paper describes (1) the Lynx framework and implementation, (2) its language constructs geared to the Single Instruction Multiple Data (SIMD) model of data-parallel programming used in current general-purpose GPU (GPGPU) based systems, and (3) useful performance metrics described via Lynx's instrumentation language that provide insights into the design of effective instrumentation routines for GPGPU systems. The paper concludes with a comparative analysis of Lynx with existing GPU profiling tools and a quantitative assessment of Lynx's instrumentation performance, providing insights into optimization opportunities for running instrumented GPU kernels. Naila Farooqui, Andrew Kerr, Greg Eisenhauer, Karsten Schwan, Sudhakar Yalamanchili |
ISPASS | 2 |
| 2011 | SIMD re-convergence at thread frontiersabstractHardware and compiler techniques for mapping data-parallel programs with divergent control flow to SIMD architectures have recently enabled the emergence of new GPGPU programming models such as CUDA, OpenCL, and DirectX Compute. The impact of branch divergence can be quite different depending upon whether the program's control flow is structured or unstructured. In this paper, we show that unstructured control flow occurs frequently in applications and can lead to significant code expansion when executed using existing approaches for handling branch divergence. Gregory Frederick Diamos, Benjamin Ashbaugh, Subramaniam Maiyuran, Andrew Kerr, Haicheng Wu, Sudhakar Yalamanchili |
MICRO | 4 |
| 2010 | Ocelot: a dynamic optimization framework for bulk-synchronous applications in heterogeneous systemsabstractOcelot is a dynamic compilation framework designed to map the explicitly data parallel execution model used by NVIDIA CUDA applications onto diverse multithreaded platforms. Ocelot includes a dynamic binary translator from Parallel Thread eXecution ISA (PTX) to many-core processors that leverages the Low Level Virtual Machine (LLVM) code generator to target x86 and other ISAs. The dynamic compiler is able to execute existing CUDA binaries without recompilation from source and supports switching between execution on an NVIDIA GPU and a many-core CPU at runtime. It has been validated against over 130 applications taken from the CUDA SDK, the UIUC Parboil benchmarks [1], the Virginia Rodinia benchmarks [2], the GPU-VSIPL signal and image processing library [3], the Thrust library [4], and several domain specific applications. Gregory Frederick Diamos, Andrew Kerr, Sudhakar Yalamanchili, Nathan Clark |
PACT | 2 |
| 2003 | The Potential for Image Guided Radiation Therapy with Cobalt-60 Tomotherapy
L. John Schreiner, Andrew Kerr, Gregory Salomons, Christine Dyck, George Hajdok |
MICCAI (2) | 2 |