Sara S. Baghsorkhi

dblp:48/4607 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Hardware accelerators and domain-specific architectures · 32% Processor architecture and microarchitecture · 30% GPUs and heterogeneous computing · 15%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%

Topics — the 26 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.822020
SAVE: Sparsity-Aware Vector Engine for Accelerating DNN Training and Inference on CPUs · MICRO 2020
C3-Flow: Compute Compression Co-Design Flow for Deep Neural Networks · DAC 2019
Hardware accelerators and domain-specific architectures
sparse matrix multiplication accelerator
0.412020
SAVE: Sparsity-Aware Vector Engine for Accelerating DNN Training and Inference on CPUs · MICRO 2020
Processor architecture and microarchitecture › vector processor
vector processing unit
0.412020
SAVE: Sparsity-Aware Vector Engine for Accelerating DNN Training and Inference on CPUs · MICRO 2020
Machine learning › Efficient and distributed learning › model compression
low-rank approximation
0.412019
C3-Flow: Compute Compression Co-Design Flow for Deep Neural Networks · DAC 2019
Machine learning › Efficient and distributed learning
model compression
0.412019
C3-Flow: Compute Compression Co-Design Flow for Deep Neural Networks · DAC 2019
Compilers and program optimization › vectorization
irregular loop vectorization
0.212016
FlexVec: auto-vectorization for irregular loops · PLDI 2016
Compilers and program optimization
vectorization
0.212016
FlexVec: auto-vectorization for irregular loops · PLDI 2016
Processor architecture and microarchitecture
instruction set architecture
0.212016
FlexVec: auto-vectorization for irregular loops · PLDI 2016
Processor architecture and microarchitecture › SIMD
SIMD extensions
0.212016
FlexVec: auto-vectorization for irregular loops · PLDI 2016
Memory systems › memory hierarchy
cache hierarchy
0.112012
Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors · PPoPP 2012
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
0.112012
Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors · PPoPP 2012
Performance modeling and evaluation
performance monitoring
0.112012
Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors · PPoPP 2012
Processor architecture and microarchitecture
SIMD
0.112020
SAVE: Sparsity-Aware Vector Engine for Accelerating DNN Training and Inference on CPUs · MICRO 2020
High-performance computing › performance optimization
auto-tuning
0.112011
Auto-tuning of fast fourier transform on graphics processors · PPoPP 2011
GPUs and heterogeneous computing
GPU kernel optimization
0.112011
Auto-tuning of fast fourier transform on graphics processors · PPoPP 2011
Performance modeling and evaluation
analytical modeling
0.112010
An adaptive performance modeling tool for GPU architectures · PPoPP 2010
Performance modeling and evaluation › performance prediction
GPU performance prediction
0.112010
An adaptive performance modeling tool for GPU architectures · PPoPP 2010
GPUs and heterogeneous computing › GPU kernel optimization
CUDA optimization
0.112008
Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008
GPUs and heterogeneous computing
GPU performance optimization
0.112008
Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008
GPUs and heterogeneous computing
GPU programming
0.112008
Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008
Memory systems
memory access optimization
0.112008
Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008
GPUs and heterogeneous computing
GPU architecture
0.122012
Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors · PPoPP 2012
An adaptive performance modeling tool for GPU architectures · PPoPP 2010
Parallel and multicore computing › data parallelism
SIMD vectorization
0.112016
FlexVec: auto-vectorization for irregular loops · PLDI 2016
Processor architecture and microarchitecture
many-core architecture
0.112007
Implicitly Parallel Programming Models for Thousand-Core Microprocessors · DAC 2007
Parallel and multicore computing
parallel programming models
0.112007
Implicitly Parallel Programming Models for Thousand-Core Microprocessors · DAC 2007
Compilers and program optimization › parallelization
automatic parallelization
0.012007
Implicitly Parallel Programming Models for Thousand-Core Microprocessors · DAC 2007

Methods — techniques the papers use, named apart from their topics

non-uniform low-rank approximation · 0.8co-design · 0.8dependence analysis · 0.5code generation · 0.5unstructured sparsity skipping · 0.4mixed-precision kernels · 0.4statistical bounds · 0.1monte carlo simulation · 0.1memory trace collection · 0.1auto-tuning · 0.1
YearPublicationVenuePosition
2020 SAVE: Sparsity-Aware Vector Engine for Accelerating DNN Training and Inference on CPUs
abstract
General Matrix Multiplication (GEMM) is the key operation in Deep Neural Networks (DNNs). While dense GEMM uses SIMD CPUs efficiently, sparse GEMM is much less efficient, especially at the modest levels of unstructured sparsity common in DNN inference/training. Thus, most DNNs use dense GEMM.In this paper, we propose SAVE, a novel vector engine for CPUs that efficiently skips ineffectual computation due to sparsity in dense DNN implementations. SAVE's hardware extensions to the vector pipeline are transparent to software. SAVE accelerates FP32 and mixed-precision kernels with unstructured sparsity from both weights and activations. Further, SAVE is not DNN-specific and can potentially speed-up any vector workload with sparsity. To evaluate SAVE, we use simulations of a 28-core machine and run VGG16, ResNet-50, and GNMT, with and without pruning. With realistic sparsity, SAVE accelerates inference by 1.37x-1.68x and end-to-end training by 1.28x-1.64x.
Zhangxiaowen Gong, Houxiang Ji, Christopher W. Fletcher, Christopher J. Hughes, Sara S. Baghsorkhi, Josep Torrellas
MICRO5
2019 C3-Flow: Compute Compression Co-Design Flow for Deep Neural Networks
abstract
Existing approaches to neural network compression have failed to holistically address algorithmic (training accuracy) and computational (inference performance) demands of real-world systems, particularly on resource-constrained devices. We present C3-Flow, a new approach adding non-uniformity to low-rank approximations and designed specifically to enable highly-efficient computation on common hardware architectures while retaining more accuracy than competing methods. Evaluation on two state-of-the-art acoustic models (versus existing work, empirical limit study approaches, and hand-tuned models) demonstrates up to 60% lower error. Finally, we show that our co-design approach achieves up to 14X inference speedup across three Haswell- and Broadwell-based platforms.
Matthew Sotoudeh, Sara S. Baghsorkhi
DAC2
2018 Automating efficient variable-grained resiliency for low-power IoT systems
abstract
New trends in edge computing encourage pushing more of the compute and analytics to the outer edge and processing most of the data locally. We explore how to transparently provide resiliency for heavy duty edge applications running on low-power devices that must deal with frequent and unpredictable power disruptions. Complicating this process further are (a) memory usage restrictions in tiny low-power devices, that affect not only performance but efficacy of the resiliency techniques, and (b) differing resiliency requirements across deployment environments. Nevertheless, an application developer wants the ability to write an application once, and have it be reusable across all low-power platforms and across all different deployment settings. In response to these challenges, we have devised a transparent roll-back recovery mechanism that performs incremental checkpoints with minimal execution time overhead and at variable granularities. Our solution includes the co-design of firmware, runtime and compiler transformations for providing seamless fault-tolerance, along with an auto-tuning layer that automatically generates multiple resilient variants of an application. Each variant spreads application’s execution over atomic transactional regions of a certain granularity. Variants with smaller regions provide better resiliency, but incur higher overhead; thus, there is no single best option, but rather a Pareto optimal set of configurations. We apply these strategies across a variety of edge device applications and measure the execution time overhead of the framework on a TI MSP430FR6989. When we restrict unin- terrupted atomic intervals to 100ms, our framework keeps geomean overhead below 2.48x.
Sara S. Baghsorkhi, Christos Margiolas
CGO1
2016 FlexVec: auto-vectorization for irregular loops
abstract
Traditional vectorization techniques build a dependence graph with distance and direction information to determine whether a loop is vectorizable. Since vectorization reorders the execution of instructions across iterations, in general instructions involved in a strongly connected component (SCC) are deemed not vectorizable unless the SCC can be eliminated using techniques such as scalar expansion or privatization. Therefore, traditional vectorization techniques are limited in their ability to efficiently handle loops with dynamic cross-iteration dependencies or complex control flow interweaved within the dependence cycles. When potential dependencies do not occur very often, the end-result is under utilization of the SIMD hardware. In this paper, we propose FlexVec architecture that combines new vector instructions with novel code generation techniques to dynamically adjusts vector length for loop statements affected by cross-iteration dependencies that happen at runtime. We have designed and implemented FlexVec's new ISA as extensions to the recently released AVX-512 ISA. We have evaluated the performance improvements enabled by FlexVec vectorization for 11 C/C++ SPEC 2006 benchmarks and 7 real applications with AVX-512 vectorization as baseline. We show that FlexVec vectorization technique produces a Geomean speedup of 9% for SPEC 2006 and a Geomean speedup of 11% for 7 real applications.
Sara S. Baghsorkhi, Nalini Vasudevan, Youfeng Wu
PLDI1
2012 Efficient performance evaluation of memory hierarchy for highly multithreaded graphics processors
abstract
With the emergence of highly multithreaded architectures, performance monitoring techniques face new challenges in efficiently locating sources of performance discrepancies in the program source code. For example, the state-of-the-art performance counters in highly multithreaded graphics processing units (GPUs) report only the overall occurrences of microarchitecture events at the end of program execution. Furthermore, even if supported, any fine-grained sampling of performance counters will distort the actual program behavior and will make the sampled values inaccurate. On the other hand, it is difficult to achieve high resolution performance information at low sampling rates in the presence of thousands of concurrently running threads. In this paper, we present a novel software-based approach for monitoring the memory hierarchy performance in highly multithreaded general-purpose graphics processors. The proposed analysis is based on memory traces collected for snapshots of an application execution. A trace-based memory hierarchy model with a Monte Carlo experimental methodology generates statistical bounds of performance measures without being concerned about the exact inter-thread ordering of individual events but rather studying the behavior of the overall system. The statistical approach overcomes the classical problem of disturbed execution timing due to fine-grained instrumentation. The approach scales well as we deploy an efficient parallel trace collection technique to reduce the trace generation overhead and a simple memory hierarchy model to reduce the simulation time. The proposed scheme also keeps track of individual memory operations in the source code and can quantify their efficiency with respect to the memory system. A cross-validation of our results shows close agreement with the values read from the hardware performance counters on an NVIDIA Tesla C2050 GPU. Based on the high resolution profile data produced by our model we optimized memory accesses in the sparse matrix vector multiply kernel and achieved speedups ranging from 2.4 to 14.8 depending on the characteristics of the input matrices.
Sara S. Baghsorkhi, Isaac Gelado, Matthieu Delahaye, Wen-Mei W. Hwu
PPoPP1
2011 Auto-tuning of fast fourier transform on graphics processors
abstract
We present an auto-tuning framework for FFTs on graphics processors (GPUs). Due to complex design of the memory and compute subsystems on GPUs, the performance of FFT kernels over the range of possible input parameters can vary widely. We generate several variants for each component of the FFT kernel that, for different cases, are likely to perform well. Our auto-tuner composes variants to generate kernels and selects the best ones. We present heuristics to prune the search space and profile only a small fraction of all possible kernels. We compose optimized kernels to improve the performance of larger FFT computations. We implement the system using the NVIDIA CUDA API and compare its performance to the state-of-the-art FFT libraries. On a range of NVIDIA GPUs and input sizes, our auto-tuned FFTs outperform the NVIDIA CUFFT 3.0 library by up to 38x and deliver up to 3x higher performance compared to a manually-tuned FFT.
Yuri Dotsenko, Sara S. Baghsorkhi, Brandon Lloyd, Naga K. Govindaraju
PPoPP2
2010 An adaptive performance modeling tool for GPU architectures
abstract
This paper presents an analytical model to predict the performance of
Sara S. Baghsorkhi, Matthieu Delahaye, Sanjay J. Patel, William Gropp, Wen-Mei W. Hwu
PPoPP1
2008 Program optimization space pruning for a multithreaded gpu
abstract
Program optimization for highly-parallel systems has historically been considered an art, with experts doing much of the performance tuning by hand. With the introduction of inexpensive, single-chip, massively parallel platforms, more developers will be creating highly-parallel applications for these platforms, who lack the substantial experience and knowledge needed to maximize their performance. This creates a need for more structured optimization methods with means to estimate their performance effects. Furthermore these methods need to be understandable by most programmers. This paper shows the complexity involved in optimizing applications for one such system and one relatively simple methodology for reducing the workload involved in the optimization process.
Shane Ryoo, Christopher I. Rodrigues, Sam S. Stone, Sara S. Baghsorkhi, Sain-Zee Ueng, John A. Stratton, Wen-Mei W. Hwu
CGO4
2008 Optimization principles and application performance evaluation of a multithreaded GPU using CUDA
abstract
GPUs have recently attracted the attention of many application developers as commodity data-parallel coprocessors. The newest generations of GPU architecture provide easier programmability and increased generality while maintaining the tremendous memory bandwidth and computational power of traditional GPUs. This opportunity should redirect efforts in GPGPU research from ad hoc porting of applications to establishing principles and strategies that allow efficient mapping of computation to graphics hardware. In this work we discuss the GeForce 8800 GTX processor's organization, features, and generalized optimization strategies. Key to performance on this platform is using massive multithreading to utilize the large number of cores and hide global memory latency. To achieve this, developers face the challenge of striking the right balance between each thread's resource usage and the number of simultaneously active threads. The resources to manage include the number of registers and the amount of on-chip memory used per thread, number of threads per multiprocessor, and global memory bandwidth. We also obtain increased performance by reordering accesses to off-chip memory to combine requests to the same or contiguous memory locations and apply classical optimizations to reduce the number of executed operations. We apply these strategies across a variety of applications and domains and achieve between a 10.5X to 457X speedup in kernel codes and between 1.16X to 431X total application speedup.
Shane Ryoo, Christopher I. Rodrigues, Sara S. Baghsorkhi, Sam S. Stone, David Blair Kirk, Wen-Mei W. Hwu
PPoPP3
2008 Program optimization carving for GPU computing
Shane Ryoo, Christopher I. Rodrigues, Sam S. Stone, John A. Stratton, Sain-Zee Ueng, Sara S. Baghsorkhi, Wen-Mei W. Hwu
J. Parallel Distributed Comput.6
2007 Implicitly Parallel Programming Models for Thousand-Core Microprocessors
abstract
This paper argues for an implicitly parallel programming model for many-core microprocessors, and provides initial technical approaches towards this goal. In an implicitly parallel programming model, programmers maximize algorithm-level parallelism, express their parallel algorithms by asserting high-level properties on top of a traditional sequential programming language, and rely on parallelizing compilers and hardware support to perform parallel execution under the hood. In such a model, compilers and related tools require much more advanced program analysis capabilities and programmer assertions than what are currently available so that a comprehensive understanding of the input program's concurrency can be derived. Such an understanding is then used to drive automatic or interactive parallel code generation tools for a diverse set of parallel hardware organizations. The chip-level architecture and hardware should maintain parallel execution state in such a way that a strictly sequential execution state can always be derived for the purpose of verifying and debugging the program. We argue that implicitly parallel programming models are critical for addressing the software development crises and software scalability challenges for many-core microprocessors.
Wen-Mei W. Hwu, Shane Ryoo, Sain-Zee Ueng, John H. Kelm, Isaac Gelado, Sam S. Stone, Robert E. Kidd, Sara S. Baghsorkhi, Aqeel Mahesri, Stephanie C. Tsao, Nacho Navarro, Steven S. Lumetta, Matthew I. Frank, Sanjay J. Patel
DAC8