VLDB 2026 Research / reviewers in the wild / expert
Dana Schaa
dblp:85/4409
· DBLP profile ↗
8ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Memory systems · 34% Processor architecture and microarchitecture · 14% GPUs and heterogeneous computing · 14% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
memory hierarchy |
0.2 | 1 | 2016 | UMH: A Hardware-Based Unified Memory Hierarchy for Systems with Multiple Discrete GPUs · ACM Trans. Archit. Code Optim. 2016 |
Memory systems › virtual memory management
unified memory |
0.2 | 1 | 2016 | UMH: A Hardware-Based Unified Memory Hierarchy for Systems with Multiple Discrete GPUs · ACM Trans. Archit. Code Optim. 2016 |
Electronic design automation
design space exploration |
0.2 | 1 | 2014 | Exploring the Heterogeneous Design Space for both Performance and Reliability · DAC 2014 |
Embedded and real-time systems
heterogeneous multi-core systems |
0.2 | 1 | 2014 | Exploring the Heterogeneous Design Space for both Performance and Reliability · DAC 2014 |
Processor architecture and microarchitecture
multicore design |
0.2 | 1 | 2014 | Exploring the Heterogeneous Design Space for both Performance and Reliability · DAC 2014 |
Hardware reliability and fault tolerance
reliability modeling |
0.2 | 1 | 2014 | Exploring the Heterogeneous Design Space for both Performance and Reliability · DAC 2014 |
Processor architecture and microarchitecture
data-parallel architecture |
0.1 | 1 | 2011 | Exploiting Memory Access Patterns to Improve Memory Performance in Data-Parallel Architectures · IEEE Trans. Parallel Distributed Syst. 2011 |
GPUs and heterogeneous computing › GPU memory management
GPU memory access optimization |
0.1 | 1 | 2011 | Exploiting Memory Access Patterns to Improve Memory Performance in Data-Parallel Architectures · IEEE Trans. Parallel Distributed Syst. 2011 |
Compilers and program optimization › vectorization
loop vectorization |
0.1 | 1 | 2010 | Data transformations enabling loop vectorization on multithreaded data parallel architectures · PPoPP 2010 |
Parallel and multicore computing
data parallelism |
0.1 | 1 | 2010 | Data transformations enabling loop vectorization on multithreaded data parallel architectures · PPoPP 2010 |
Parallel and multicore computing › data parallelism
SIMD vectorization |
0.1 | 1 | 2010 | Data transformations enabling loop vectorization on multithreaded data parallel architectures · PPoPP 2010 |
Memory systems
cache coherence |
0.1 | 1 | 2016 | UMH: A Hardware-Based Unified Memory Hierarchy for Systems with Multiple Discrete GPUs · ACM Trans. Archit. Code Optim. 2016 |
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU memory management |
0.1 | 1 | 2016 | UMH: A Hardware-Based Unified Memory Hierarchy for Systems with Multiple Discrete GPUs · ACM Trans. Archit. Code Optim. 2016 |
Memory systems › cache coherence
directory-based coherence |
0.1 | 1 | 2016 | UMH: A Hardware-Based Unified Memory Hierarchy for Systems with Multiple Discrete GPUs · ACM Trans. Archit. Code Optim. 2016 |
GPUs and heterogeneous computing
multi-GPU computing |
0.1 | 1 | 2016 | UMH: A Hardware-Based Unified Memory Hierarchy for Systems with Multiple Discrete GPUs · ACM Trans. Archit. Code Optim. 2016 |
Performance modeling and evaluation
performability analysis |
0.1 | 1 | 2014 | Exploring the Heterogeneous Design Space for both Performance and Reliability · DAC 2014 |
GPUs and heterogeneous computing
GPU computing |
0.0 | 1 | 2011 | Exploiting Memory Access Patterns to Improve Memory Performance in Data-Parallel Architectures · IEEE Trans. Parallel Distributed Syst. 2011 |
Methods — techniques the papers use, named apart from their topics
NMOESI coherence protocol · 0.2mathematical modeling of memory access patterns · 0.2performance modeling · 0.2design space exploration · 0.2memory access pattern analysis · 0.1loop body characterization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | UMH: A Hardware-Based Unified Memory Hierarchy for Systems with Multiple Discrete GPUsabstractIn this article, we describe how to ease memory management between a Central Processing Unit (CPU) and one or multiple discrete Graphic Processing Units (GPUs) by architecting a novel hardware-based Unified Memory Hierarchy (UMH). Adopting UMH, a GPU accesses the CPU memory only if it does not find its required data in the directories associated with its high-bandwidth memory, or the NMOESI coherency protocol limits the access to that data. Using UMH with NMOESI improves performance of a CPU-multiGPU system by at least 1.92 × in comparison to alternative software-based approaches. It also allows the CPU to access GPUs modified data by at least 13 × faster. Amir Kavyan Ziabari, Yifan Sun 0002, Yenai Ma, Dana Schaa, José L. Abellán, Rafael Ubal, John Kim 0001, Ajay Joshi, David R. Kaeli |
ACM Trans. Archit. Code Optim. | 4 |
| 2014 | Exploring the Heterogeneous Design Space for both Performance and ReliabilityabstractAs we move into a new era of heterogeneous multi-core systems, our ability to tune the performance and understand the reliability of both hardware and software becomes more challenging. Given the multiplicity of different design trade-offs in hardware and software, and the rate of introduction of new architectures and hardware/software features, it becomes difficult to properly model emerging heterogeneous platforms. Rafael Ubal, Dana Schaa, Perhaad Mistry, Yash Ukidave, Zhongliang Chen, Gunar Schirner, David R. Kaeli |
DAC | 2 |
| 2014 | Runtime Support for Adaptive Spatial Partitioning and Inter-Kernel Communication on GPUsabstractGPUs have gained tremendous popularity in a broad range of application domains. These applications possess varying grains of parallelism and place high demands on compute resources -- many times imposing real-time constraints, requiring flexible work schedules, and relying on concurrent execution of multiple kernels on the device. These requirements present a number of challenges when targeting current GPUs. To support this class of applications, and to take full advantage of the large number of compute cores present on the GPU, we need a new mechanism to support concurrent execution and provide flexible mapping of compute kernels to the GPU. In this paper, we describe a new scheduling mechanism for dynamic spatial partitioning of the GPU, which adapts to the current execution state of compute workloads on the device. To enable this functionality, we extend the OpenCL runtime environment to map multiple command queues to a single device, and effectively partitioning the device. The result is that kernels that can benefit from concurrent execution on a partitioned device can effectively utilize the full compute resources on the GPU. To accelerate next-generation workloads, we also support an inter-kernel communication mechanism that enables concurrent kernels to interact in a producer-consumer relationship. The proposed partitioning mechanism is evaluated using real world applications taken from signal and image processing, linear algebra, and data mining domains. For these performance-hungry applications we achieve a 3.1X performance speedup using a combination of the proposed scheduling scheme and inter-kernel communication, versus relying on the conventional GPU runtime. Yash Ukidave, Charu Kalra, David R. Kaeli, Perhaad Mistry, Dana Schaa |
SBAC-PAD | 5 |
| 2012 | Multi2Sim: a simulation framework for CPU-GPU computingabstractAccurate simulation is essential for the proper design and evaluation of any computing platform. Upon the current move toward the CPU-GPU heterogeneous computing era, researchers need a simulation framework that can model both kinds of computing devices and their interaction. In this paper, we present Multi2Sim, an open-source, modular, and fully configurable toolset that enables ISA-level simulation of an x86 CPU and an AMD Evergreen GPU. Focusing on a model of the AMD Radeon 5870 GPU, we address program emulation correctness, as well as architectural simulation accuracy, using AMD's OpenCL benchmark suite. Simulation capabilities are demonstrated with a preliminary architectural exploration study, and workload characterization examples. The project source code, benchmark packages, and a detailed user's guide are publicly available at www.multi2sim.org. Rafael Ubal, Byunghyun Jang, Perhaad Mistry, Dana Schaa, David R. Kaeli |
PACT | 4 |
| 2011 | Exploiting Memory Access Patterns to Improve Memory Performance in Data-Parallel ArchitecturesabstractThe introduction of General-Purpose computation on GPUs (GPGPUs) has changed the landscape for the future of parallel computing. At the core of this phenomenon are massively multithreaded, data-parallel architectures possessing impressive acceleration ratings, offering low-cost supercomputing together with attractive power budgets. Even given the numerous benefits provided by GPGPUs, there remain a number of barriers that delay wider adoption of these architectures. One major issue is the heterogeneous and distributed nature of the memory subsystem commonly found on data-parallel architectures. Application acceleration is highly dependent on being able to utilize the memory subsystem effectively so that all execution units remain busy. In this paper, we present techniques for enhancing the memory efficiency of applications on data-parallel architectures, based on the analysis and characterization of memory access patterns in loop bodies; we target vectorization via data transformation to benefit vector-based architectures (e.g., AMD GPUs) and algorithmic memory selection for scalar-based architectures (e.g., NVIDIA GPUs). We demonstrate the effectiveness of our proposed methods with kernels from a wide range of benchmark suites. For the benchmark kernels studied, we achieve consistent and significant performance improvements (up to 11.4× and 13.5× over baseline GPU implementations on each platform, respectively) by applying our proposed methodology. Byunghyun Jang, Dana Schaa, Perhaad Mistry, David R. Kaeli |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2010 | Data transformations enabling loop vectorization on multithreaded data parallel architecturesabstractLoop vectorization, a key feature exploited to obtain high performance on Single Instruction Multiple Data (SIMD) vector architectures, is significantly hindered by irregular memory access patterns in the data stream. This paper describes data transformations that allow us to vectorize loops targeting massively multithreaded data parallel architectures. We present a mathematical model that captures loop-based memory access patterns and computes the most appropriate data transformations in order to enable vectorization. Our experimental results show that the proposed data transformations can significantly increase the number of loops that can be vectorized and enhance the data-level parallelism of applications. Our results also show that the overhead associated with our data transformations can be easily amortized as the size of the input data set increases. For the set of high performance benchmark kernels studied, we achieve consistent and significant performance improvements (up to 11.4X) by applying vectorization using our data transformation approach. Byunghyun Jang, Perhaad Mistry, Dana Schaa, Rodrigo Dominguez, David R. Kaeli |
PPoPP | 3 |
| 2009 | Exploring the multiple-GPU design spaceabstractGraphics processing units (GPUs) have been growing in popularity due to their impressive processing capabilities, and with general purpose programming languages such as NVIDIA's CUDA interface, are becoming the platform of choice in the scientific computing community. Previous studies that used GPUs focused on obtaining significant performance gains from execution on a single GPU. These studies employed low-level, architecture-specific tuning in order to achieve sizeable benefits over multicore CPU execution. In this paper, we consider the benefits of running on multiple (parallel) GPUs to provide further orders of performance speedup. Our methodology allows developers to accurately predict execution time for GPU applications while varying the number and configuration of the GPUs, and the size of the input data set. This is a natural next step in GPU computing because it allows researchers to determine the most appropriate GPU configuration for an application without having to purchase hardware, or write the code for a multiple-GPU implementation. When used to predict performance on six scientific applications, our framework produces accurate performance estimates (11% difference on average and 40% maximum difference in a single case) for a range of short and long running scientific programs. Dana Schaa, David R. Kaeli |
IPDPS | 1 |
| 2007 | Exploring Novel Parallelization Technologies for 3-D Imaging ApplicationsabstractMulti-dimensional imaging techniques involve the processing of high resolution images commonly used in medical, civil and remote-sensing applications. A barrier commonly encountered in this class of applications is the time required to carry out repetitive operations on large matrices. Partitioning these large datasets can help improve performance, and lends the data to more efficient parallel execution. In this paper we describe our experience exploring two novel parallelization technologies: 1) a graphical processor unit (GPU)-based approach which utilizes 128 cores on a single GPU accelerator card, and 2) a middleware approach for semi-automatic parallelization on a cluster of multiple multi-core processors. We investigate these two platforms and describe their strengths and limitations. In addition, we provide some guidance to the programmer on which platform to use when porting multi-dimensional imaging applications. Using a 3-D application taken from a clinical image reconstruction algorithm, we demonstrate the degree of speedup we can obtain from these two approaches. Dana Schaa, Micha Moffie, David R. Kaeli |
SBAC-PAD | 2 |