VLDB 2026 Research / reviewers in the wild / expert
Sara S. Baghsorkhi
dblp:48/4607
· DBLP profile ↗
11ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Hardware accelerators and domain-specific architectures · 32% Processor architecture and microarchitecture · 30% GPUs and heterogeneous computing · 15% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% |
Topics — the 26 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.8 | 2 | 2020 | SAVE: Sparsity-Aware Vector Engine for Accelerating DNN Training and Inference on CPUs · MICRO 2020 C3-Flow: Compute Compression Co-Design Flow for Deep Neural Networks · DAC 2019 |
Hardware accelerators and domain-specific architectures
sparse matrix multiplication accelerator |
0.4 | 1 | 2020 | SAVE: Sparsity-Aware Vector Engine for Accelerating DNN Training and Inference on CPUs · MICRO 2020 |
Processor architecture and microarchitecture › vector processor
vector processing unit |
0.4 | 1 | 2020 | SAVE: Sparsity-Aware Vector Engine for Accelerating DNN Training and Inference on CPUs · MICRO 2020 |
Machine learning › Efficient and distributed learning › model compression
low-rank approximation |
0.4 | 1 | 2019 | C3-Flow: Compute Compression Co-Design Flow for Deep Neural Networks · DAC 2019 |
Machine learning › Efficient and distributed learning
model compression |
0.4 | 1 | 2019 | C3-Flow: Compute Compression Co-Design Flow for Deep Neural Networks · DAC 2019 |
Compilers and program optimization › vectorization
irregular loop vectorization |
0.2 | 1 | 2016 | FlexVec: auto-vectorization for irregular loops · PLDI 2016 |
Compilers and program optimization
vectorization |
0.2 | 1 | 2016 | FlexVec: auto-vectorization for irregular loops · PLDI 2016 |
Processor architecture and microarchitecture
instruction set architecture |
0.2 | 1 | 2016 | FlexVec: auto-vectorization for irregular loops · PLDI 2016 |
Processor architecture and microarchitecture › SIMD
SIMD extensions |
0.2 | 1 | 2016 | FlexVec: auto-vectorization for irregular loops · PLDI 2016 |
Memory systems › memory hierarchy
cache hierarchy |
0.1 | 1 | 2012 | Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors · PPoPP 2012 |
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy |
0.1 | 1 | 2012 | Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors · PPoPP 2012 |
Performance modeling and evaluation
performance monitoring |
0.1 | 1 | 2012 | Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors · PPoPP 2012 |
Processor architecture and microarchitecture
SIMD |
0.1 | 1 | 2020 | SAVE: Sparsity-Aware Vector Engine for Accelerating DNN Training and Inference on CPUs · MICRO 2020 |
High-performance computing › performance optimization
auto-tuning |
0.1 | 1 | 2011 | Auto-tuning of fast fourier transform on graphics processors · PPoPP 2011 |
GPUs and heterogeneous computing
GPU kernel optimization |
0.1 | 1 | 2011 | Auto-tuning of fast fourier transform on graphics processors · PPoPP 2011 |
Performance modeling and evaluation
analytical modeling |
0.1 | 1 | 2010 | An adaptive performance modeling tool for GPU architectures · PPoPP 2010 |
Performance modeling and evaluation › performance prediction
GPU performance prediction |
0.1 | 1 | 2010 | An adaptive performance modeling tool for GPU architectures · PPoPP 2010 |
GPUs and heterogeneous computing › GPU kernel optimization
CUDA optimization |
0.1 | 1 | 2008 | Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008 |
GPUs and heterogeneous computing
GPU performance optimization |
0.1 | 1 | 2008 | Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008 |
GPUs and heterogeneous computing
GPU programming |
0.1 | 1 | 2008 | Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008 |
Memory systems
memory access optimization |
0.1 | 1 | 2008 | Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008 |
GPUs and heterogeneous computing
GPU architecture |
0.1 | 2 | 2012 | Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors · PPoPP 2012 An adaptive performance modeling tool for GPU architectures · PPoPP 2010 |
Parallel and multicore computing › data parallelism
SIMD vectorization |
0.1 | 1 | 2016 | FlexVec: auto-vectorization for irregular loops · PLDI 2016 |
Processor architecture and microarchitecture
many-core architecture |
0.1 | 1 | 2007 | Implicitly Parallel Programming Models for Thousand-Core Microprocessors · DAC 2007 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2007 | Implicitly Parallel Programming Models for Thousand-Core Microprocessors · DAC 2007 |
Compilers and program optimization › parallelization
automatic parallelization |
0.0 | 1 | 2007 | Implicitly Parallel Programming Models for Thousand-Core Microprocessors · DAC 2007 |
Methods — techniques the papers use, named apart from their topics
non-uniform low-rank approximation · 0.8co-design · 0.8dependence analysis · 0.5code generation · 0.5unstructured sparsity skipping · 0.4mixed-precision kernels · 0.4statistical bounds · 0.1monte carlo simulation · 0.1memory trace collection · 0.1auto-tuning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | SAVE: Sparsity-Aware Vector Engine for Accelerating DNN Training and Inference on CPUsabstractGeneral Matrix Multiplication (GEMM) is the key operation in Deep Neural Networks (DNNs). While dense GEMM uses SIMD CPUs efficiently, sparse GEMM is much less efficient, especially at the modest levels of unstructured sparsity common in DNN inference/training. Thus, most DNNs use dense GEMM.In this paper, we propose SAVE, a novel vector engine for CPUs that efficiently skips ineffectual computation due to sparsity in dense DNN implementations. SAVE's hardware extensions to the vector pipeline are transparent to software. SAVE accelerates FP32 and mixed-precision kernels with unstructured sparsity from both weights and activations. Further, SAVE is not DNN-specific and can potentially speed-up any vector workload with sparsity. To evaluate SAVE, we use simulations of a 28-core machine and run VGG16, ResNet-50, and GNMT, with and without pruning. With realistic sparsity, SAVE accelerates inference by 1.37x-1.68x and end-to-end training by 1.28x-1.64x. Zhangxiaowen Gong, Houxiang Ji, Christopher W. Fletcher, Christopher J. Hughes, Sara S. Baghsorkhi, Josep Torrellas |
MICRO | 5 |
| 2019 | C3-Flow: Compute Compression Co-Design Flow for Deep Neural NetworksabstractExisting approaches to neural network compression have failed to holistically address algorithmic (training accuracy) and computational (inference performance) demands of real-world systems, particularly on resource-constrained devices. We present C3-Flow, a new approach adding non-uniformity to low-rank approximations and designed specifically to enable highly-efficient computation on common hardware architectures while retaining more accuracy than competing methods. Evaluation on two state-of-the-art acoustic models (versus existing work, empirical limit study approaches, and hand-tuned models) demonstrates up to 60% lower error. Finally, we show that our co-design approach achieves up to 14X inference speedup across three Haswell- and Broadwell-based platforms. Matthew Sotoudeh, Sara S. Baghsorkhi |
DAC | 2 |
| 2018 | Automating efficient variable-grained resiliency for low-power IoT systemsabstractNew trends in edge computing encourage pushing more of the compute and analytics to the outer edge and processing most of the data locally. We explore how to transparently provide resiliency for heavy duty edge applications running on low-power devices that must deal with frequent and unpredictable power disruptions. Complicating this process further are (a) memory usage restrictions in tiny low-power devices, that affect not only performance but efficacy of the resiliency techniques, and (b) differing resiliency requirements across deployment environments. Nevertheless, an application developer wants the ability to write an application once, and have it be reusable across all low-power platforms and across all different deployment settings. In response to these challenges, we have devised a transparent roll-back recovery mechanism that performs incremental checkpoints with minimal execution time overhead and at variable granularities. Our solution includes the co-design of firmware, runtime and compiler transformations for providing seamless fault-tolerance, along with an auto-tuning layer that automatically generates multiple resilient variants of an application. Each variant spreads application’s execution over atomic transactional regions of a certain granularity. Variants with smaller regions provide better resiliency, but incur higher overhead; thus, there is no single best option, but rather a Pareto optimal set of configurations. We apply these strategies across a variety of edge device applications and measure the execution time overhead of the framework on a TI MSP430FR6989. When we restrict unin- terrupted atomic intervals to 100ms, our framework keeps geomean overhead below 2.48x. Sara S. Baghsorkhi, Christos Margiolas |
CGO | 1 |
| 2016 | FlexVec: auto-vectorization for irregular loopsabstractTraditional vectorization techniques build a dependence graph with distance and direction information to determine whether a loop is vectorizable. Since vectorization reorders the execution of instructions across iterations, in general instructions involved in a strongly connected component (SCC) are deemed not vectorizable unless the SCC can be eliminated using techniques such as scalar expansion or privatization. Therefore, traditional vectorization techniques are limited in their ability to efficiently handle loops with dynamic cross-iteration dependencies or complex control flow interweaved within the dependence cycles. When potential dependencies do not occur very often, the end-result is under utilization of the SIMD hardware. In this paper, we propose FlexVec architecture that combines new vector instructions with novel code generation techniques to dynamically adjusts vector length for loop statements affected by cross-iteration dependencies that happen at runtime. We have designed and implemented FlexVec's new ISA as extensions to the recently released AVX-512 ISA. We have evaluated the performance improvements enabled by FlexVec vectorization for 11 C/C++ SPEC 2006 benchmarks and 7 real applications with AVX-512 vectorization as baseline. We show that FlexVec vectorization technique produces a Geomean speedup of 9% for SPEC 2006 and a Geomean speedup of 11% for 7 real applications. Sara S. Baghsorkhi, Nalini Vasudevan, Youfeng Wu |
PLDI | 1 |
| 2012 | Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processorsabstractWith the emergence of highly multithreaded architectures, performance monitoring techniques face new challenges in efficiently locating sources of performance discrepancies in the program source code. For example, the state-of-the-art performance counters in highly multithreaded graphics processing units (GPUs) report only the overall occurrences of microarchitecture events at the end of program execution. Furthermore, even if supported, any fine-grained sampling of performance counters will distort the actual program behavior and will make the sampled values inaccurate. On the other hand, it is difficult to achieve high resolution performance information at low sampling rates in the presence of thousands of concurrently running threads. In this paper, we present a novel software-based approach for monitoring the memory hierarchy performance in highly multithreaded general-purpose graphics processors. The proposed analysis is based on memory traces collected for snapshots of an application execution. A trace-based memory hierarchy model with a Monte Carlo experimental methodology generates statistical bounds of performance measures without being concerned about the exact inter-thread ordering of individual events but rather studying the behavior of the overall system. The statistical approach overcomes the classical problem of disturbed execution timing due to fine-grained instrumentation. The approach scales well as we deploy an efficient parallel trace collection technique to reduce the trace generation overhead and a simple memory hierarchy model to reduce the simulation time. The proposed scheme also keeps track of individual memory operations in the source code and can quantify their efficiency with respect to the memory system. A cross-validation of our results shows close agreement with the values read from the hardware performance counters on an NVIDIA Tesla C2050 GPU. Based on the high resolution profile data produced by our model we optimized memory accesses in the sparse matrix vector multiply kernel and achieved speedups ranging from 2.4 to 14.8 depending on the characteristics of the input matrices. Sara S. Baghsorkhi, Isaac Gelado, Matthieu Delahaye, Wen-Mei W. Hwu |
PPoPP | 1 |
| 2011 | Auto-tuning of fast fourier transform on graphics processorsabstractWe present an auto-tuning framework for FFTs on graphics processors (GPUs). Due to complex design of the memory and compute subsystems on GPUs, the performance of FFT kernels over the range of possible input parameters can vary widely. We generate several variants for each component of the FFT kernel that, for different cases, are likely to perform well. Our auto-tuner composes variants to generate kernels and selects the best ones. We present heuristics to prune the search space and profile only a small fraction of all possible kernels. We compose optimized kernels to improve the performance of larger FFT computations. We implement the system using the NVIDIA CUDA API and compare its performance to the state-of-the-art FFT libraries. On a range of NVIDIA GPUs and input sizes, our auto-tuned FFTs outperform the NVIDIA CUFFT 3.0 library by up to 38x and deliver up to 3x higher performance compared to a manually-tuned FFT. Yuri Dotsenko, Sara S. Baghsorkhi, Brandon Lloyd, Naga K. Govindaraju |
PPoPP | 2 |
| 2010 | An adaptive performance modeling tool for GPU architecturesabstractThis paper presents an analytical model to predict the performance of Sara S. Baghsorkhi, Matthieu Delahaye, Sanjay J. Patel, William Gropp, Wen-Mei W. Hwu |
PPoPP | 1 |
| 2008 | Program optimization space pruning for a multithreaded gpuabstractProgram optimization for highly-parallel systems has historically been considered an art, with experts doing much of the performance tuning by hand. With the introduction of inexpensive, single-chip, massively parallel platforms, more developers will be creating highly-parallel applications for these platforms, who lack the substantial experience and knowledge needed to maximize their performance. This creates a need for more structured optimization methods with means to estimate their performance effects. Furthermore these methods need to be understandable by most programmers. This paper shows the complexity involved in optimizing applications for one such system and one relatively simple methodology for reducing the workload involved in the optimization process. Shane Ryoo, Christopher I. Rodrigues, Sam S. Stone, Sara S. Baghsorkhi, Sain-Zee Ueng, John A. Stratton, Wen-Mei W. Hwu |
CGO | 4 |
| 2008 | Optimization principles and application performance evaluation of a multithreaded GPU using CUDAabstractGPUs have recently attracted the attention of many application developers as commodity data-parallel coprocessors. The newest generations of GPU architecture provide easier programmability and increased generality while maintaining the tremendous memory bandwidth and computational power of traditional GPUs. This opportunity should redirect efforts in GPGPU research from ad hoc porting of applications to establishing principles and strategies that allow efficient mapping of computation to graphics hardware. In this work we discuss the GeForce 8800 GTX processor's organization, features, and generalized optimization strategies. Key to performance on this platform is using massive multithreading to utilize the large number of cores and hide global memory latency. To achieve this, developers face the challenge of striking the right balance between each thread's resource usage and the number of simultaneously active threads. The resources to manage include the number of registers and the amount of on-chip memory used per thread, number of threads per multiprocessor, and global memory bandwidth. We also obtain increased performance by reordering accesses to off-chip memory to combine requests to the same or contiguous memory locations and apply classical optimizations to reduce the number of executed operations. We apply these strategies across a variety of applications and domains and achieve between a 10.5X to 457X speedup in kernel codes and between 1.16X to 431X total application speedup. Shane Ryoo, Christopher I. Rodrigues, Sara S. Baghsorkhi, Sam S. Stone, David Blair Kirk, Wen-Mei W. Hwu |
PPoPP | 3 |
| 2008 | Program optimization carving for GPU computing
Shane Ryoo, Christopher I. Rodrigues, Sam S. Stone, John A. Stratton, Sain-Zee Ueng, Sara S. Baghsorkhi, Wen-Mei W. Hwu |
J. Parallel Distributed Comput. | 6 |
| 2007 | Implicitly Parallel Programming Models for Thousand-Core MicroprocessorsabstractThis paper argues for an implicitly parallel programming model for many-core microprocessors, and provides initial technical approaches towards this goal. In an implicitly parallel programming model, programmers maximize algorithm-level parallelism, express their parallel algorithms by asserting high-level properties on top of a traditional sequential programming language, and rely on parallelizing compilers and hardware support to perform parallel execution under the hood. In such a model, compilers and related tools require much more advanced program analysis capabilities and programmer assertions than what are currently available so that a comprehensive understanding of the input program's concurrency can be derived. Such an understanding is then used to drive automatic or interactive parallel code generation tools for a diverse set of parallel hardware organizations. The chip-level architecture and hardware should maintain parallel execution state in such a way that a strictly sequential execution state can always be derived for the purpose of verifying and debugging the program. We argue that implicitly parallel programming models are critical for addressing the software development crises and software scalability challenges for many-core microprocessors. Wen-Mei W. Hwu, Shane Ryoo, Sain-Zee Ueng, John H. Kelm, Isaac Gelado, Sam S. Stone, Robert E. Kidd, Sara S. Baghsorkhi, Aqeel Mahesri, Stephanie C. Tsao, Nacho Navarro, Steven S. Lumetta, Matthew I. Frank, Sanjay J. Patel |
DAC | 8 |