Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Davoud Anoushe Jamshidi

dblp:138/4183 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
GPUs and heterogeneous computing · 36% Emerging computing paradigms · 18% Parallel and multicore computing · 17%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms
approximate computing
0.432014
Scaling Performance via Self-Tuning Approximation for Graphics Engines · ACM Trans. Comput. Syst. 2014
Paraprox: pattern-based approximation for data parallel applications · ASPLOS 2014
SAGE: self-tuning approximation for graphics engines · MICRO 2013
GPUs and heterogeneous computing
GPU architecture
0.312017
Regless: just-in-time operand staging for GPUs · MICRO 2017
Processor architecture and microarchitecture
register file
0.312017
Regless: just-in-time operand staging for GPUs · MICRO 2017
Parallel and multicore computing › task scheduling
memory-aware scheduling
0.212015
Mascar: Speeding up GPU warps by reducing memory pitstops · HPCA 2015
GPUs and heterogeneous computing › GPU scheduling
warp scheduling
0.212015
Mascar: Speeding up GPU warps by reducing memory pitstops · HPCA 2015
Parallel and multicore computing
data-parallel programming
0.212014
Paraprox: pattern-based approximation for data parallel applications · ASPLOS 2014
GPUs and heterogeneous computing
GPU computing
0.212014
Paraprox: pattern-based approximation for data parallel applications · ASPLOS 2014
GPUs and heterogeneous computing › GPU computing
approximate GPU arithmetic
0.212013
SAGE: self-tuning approximation for graphics engines · MICRO 2013
Cloud and datacenter computing
computation offloading
0.112012
COMET: Code Offload by Migrating Execution Transparently · OSDI 2012
Energy-efficient computing › energy-efficient architecture
GPU energy efficiency
0.112017
Regless: just-in-time operand staging for GPUs · MICRO 2017
Energy-efficient computing
power management
0.112017
Regless: just-in-time operand staging for GPUs · MICRO 2017
Cloud and datacenter computing
mobile cloud computing
0.012012
COMET: Code Offload by Migrating Execution Transparently · OSDI 2012

Methods — techniques the papers use, named apart from their topics

runtime kernel selection · 0.5thread fusion · 0.4static compiler · 0.4data packing · 0.4warp scheduling · 0.2cache re-execution · 0.2runtime tuning · 0.2pattern-based approximation · 0.2CUDA · 0.2
YearPublicationVenuePosition
2017 Regless: just-in-time operand staging for GPUs
abstract
The register file is one of the largest and most power-hungry structures in a Graphics Processing Unit (GPU), because massive multithreading requires all the register state for every active thread to be available. Previous approaches to making register accesses more efficient have optimized how registers are stored, but they must keep all values for active threads in a large, high-bandwidth structure. If operand storage is to be reduced further, there will not be enough capacity for every live value to be stored at the same time. Our insight is that computation graphs can be sliced into regions and operand storage can be allocated to these regions as they are encountered at run time, allowing a small operand staging unit to replace the register file. Most operand values have a short lifetime that is contained in one region, so their value does not need to persist in the staging unit past the end of that region. The small number of longer-lived operands can be stored in lower-bandwidth global memory, but the hardware must anticipate their use to fetch them early enough to avoid stalls. In RegLess, hardware uses compiler annotations to anticipate warps' operand usage at run time, allowing the register file to be replaced with an operand staging unit 25% of the size, saving 75% of register file energy and 11% of total GPU energy with no average performance loss.
John Kloosterman, Jonathan Beaumont, Davoud Anoushe Jamshidi, Jonathan Bailey, Trevor N. Mudge, Scott A. Mahlke
MICRO3
2015 Mascar: Speeding up GPU warps by reducing memory pitstops
abstract
With the prevalence of GPUs as throughput engines for data parallel workloads, the landscape of GPU computing is changing significantly. Non-graphics workloads with high memory intensity and irregular access patterns are frequently targeted for acceleration on GPUs. While GPUs provide large numbers of compute resources, the resources needed for memory intensive workloads are more scarce. Therefore, managing access to these limited memory resources is a challenge for GPUs. We propose a novel Memory Aware Scheduling and Cache Access Re-execution (Mascar) system on GPUs tailored for better performance for memory intensive workloads. This scheme detects memory saturation and prioritizes memory requests among warps to enable better overlapping of compute and memory accesses. Furthermore, it enables limited re-execution of memory instructions to eliminate structural hazards in the memory subsystem and take advantage of cache locality in cases where requests cannot be sent to the memory due to memory saturation. Our results show that Mascar provides a 34% speedup over the baseline round-robin scheduler and 10% speedup over the state of the art warp schedulers for memory intensive workloads. Mascar also achieves an average of 12% savings in energy for such workloads.
Ankit Sethia, Davoud Anoushe Jamshidi, Scott A. Mahlke
HPCA2
2014 D2MA: accelerating coarse-grained data transfer for GPUs
abstract
To achieve high performance on many-core architectures like GPUs, it is crucial to efficiently utilize the available memory bandwidth. Currently, it is common to use fast, on-chip scratchpad memories, like the shared memory available on GPUs' shader cores, to buffer data for computation. This buffering, however, has some sources of inefficiency that hinder it from most efficiently utilizing the available memory resources. These issues stem from shader resources being used for repeated, regular address calculations, a need to shuffle data multiple times between a physically unified on-chip memory, and forcing all threads to synchronize to ensure RAW consistency based on the speed of the slowest threads. To address these inefficiencies, we propose Data-Parallel DMA, or D2MA. D2MA is a reimagination of traditional DMA that addresses the challenges of extending DMA to thousands of concurrently executing threads. D2MA de-couples address generation from the shader's computational resources, provides a more direct and efficient path for data in global memory to travel into the shared memory, and introduces a novel dynamic synchronization scheme that is transparent to the programmer. These advancements allow D2MA to achieve speedups as high as 2.29x, and reduces the average time to buffer data by 81% on average.
Davoud Anoushe Jamshidi, Mehrzad Samadi, Scott A. Mahlke
PACT1
2014 Paraprox: pattern-based approximation for data parallel applications
abstract
Approximate computing is an approach where reduced accuracy of results is traded off for increased speed, throughput, or both. Loss of accuracy is not permissible in all computing domains, but there are a growing number of data-intensive domains where the output of programs need not be perfectly correct to provide useful results or even noticeable differences to the end user. These soft domains include multimedia processing, machine learning, and data mining/analysis. An important challenge with approximate computing is transparency to insulate both software and hardware developers from the time, cost, and difficulty of using approximation. This paper proposes a software-only system, Paraprox, for realizing transparent approximation of data-parallel programs that operates on commodity hardware systems. Paraprox starts with a data-parallel kernel implemented using OpenCL or CUDA and creates a parameterized approximate kernel that is tuned at runtime to maximize performance subject to a target output quality (TOQ) that is supplied by the user. Approximate kernels are created by recognizing common computation idioms found in data-parallel programs (e.g., Map, Scatter/Gather, Reduction, Scan, Stencil, and Partition) and substituting approximate implementations in their place. Across a set of 13 soft data-parallel applications with at most 10% quality degradation, Paraprox yields an average performance gain of 2.7x on a NVIDIA GTX 560 GPU and 2.5x on an Intel Core i7 quad-core processor compared to accurate execution on each platform.
Mehrzad Samadi, Davoud Anoushe Jamshidi, Janghaeng Lee, Scott A. Mahlke
ASPLOS2
2014 Scaling Performance via Self-Tuning Approximation for Graphics Engines
abstract
Approximate computing, where computation accuracy is traded off for better performance or higher data throughput, is one solution that can help data processing keep pace with the current and growing abundance of information. For particular domains, such as multimedia and learning algorithms, approximation is commonly used today. We consider automation to be essential to provide transparent approximation, and we show that larger benefits can be achieved by constructing the approximation techniques to fit the underlying hardware. Our target platform is the GPU because of its high performance capabilities and difficult programming challenges that can be alleviated with proper automation. Our approach—SAGE—combines a static compiler that automatically generates a set of CUDA kernels with varying levels of approximation with a runtime system that iteratively selects among the available kernels to achieve speedup while adhering to a target output quality set by the user. The SAGE compiler employs three optimization techniques to generate approximate kernels that exploit the GPU microarchitecture: selective discarding of atomic operations, data packing, and thread fusion. Across a set of machine learning and image processing kernels, SAGE's approximation yields an average of 2.5× speedup with less than 10% quality loss compared to the accurate execution on a NVIDIA GTX 560 GPU.
Mehrzad Samadi, Janghaeng Lee, Davoud Anoushe Jamshidi, Scott A. Mahlke, Amir Hormati
ACM Trans. Comput. Syst.3
2013 SAGE: self-tuning approximation for graphics engines
abstract
Approximate computing, where computation accuracy is traded off for better performance or higher data throughput, is one solution that can help data processing keep pace with the current and growing overabundance of information. For particular domains such as multimedia and learning algorithms, approximation is commonly used today. We consider automation to be essential to provide transparent approximation and we show that larger benefits can be achieved by constructing the approximation techniques to fit the underlying hardware. Our target platform is the GPU because of its high performance capabilities and difficult programming challenges that can be alleviated with proper automation. Our approach, SAGE, combines a static compiler that automatically generates a set of CUDA kernels with varying levels of approximation with a run-time system that iteratively selects among the available kernels to achieve speedup while adhering to a target output quality set by the user. The SAGE compiler employs three optimization techniques to generate approximate kernels that exploit the GPU microarchitecture: selective discarding of atomic operations, data packing, and thread fusion. Across a set of machine learning and image processing kernels, SAGE's approximation yields an average of 2.5x speedup with less than 10% quality loss compared to the accurate execution on a NVIDIA GTX 560 GPU.
Mehrzad Samadi, Janghaeng Lee, Davoud Anoushe Jamshidi, Amir Hormati, Scott A. Mahlke
MICRO3
2012 COMET: Code Offload by Migrating Execution Transparently
Mark S. Gordon, Davoud Anoushe Jamshidi, Scott A. Mahlke, Z. Morley Mao, Xu Chen 0028
OSDI2