EDBT 2026 Demo / reviewers in the wild / expert
Davoud Anoushe Jamshidi
dblp:138/4183
· DBLP profile ↗
7ranked-venue papers
1as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
GPUs and heterogeneous computing · 36% Emerging computing paradigms · 18% Parallel and multicore computing · 17% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Emerging computing paradigms
approximate computing |
0.4 | 3 | 2014 | Scaling Performance via Self-Tuning Approximation for Graphics Engines · ACM Trans. Comput. Syst. 2014 Paraprox: pattern-based approximation for data parallel applications · ASPLOS 2014 SAGE: self-tuning approximation for graphics engines · MICRO 2013 |
GPUs and heterogeneous computing
GPU architecture |
0.3 | 1 | 2017 | Regless: just-in-time operand staging for GPUs · MICRO 2017 |
Processor architecture and microarchitecture
register file |
0.3 | 1 | 2017 | Regless: just-in-time operand staging for GPUs · MICRO 2017 |
Parallel and multicore computing › task scheduling
memory-aware scheduling |
0.2 | 1 | 2015 | Mascar: Speeding up GPU warps by reducing memory pitstops · HPCA 2015 |
GPUs and heterogeneous computing › GPU scheduling
warp scheduling |
0.2 | 1 | 2015 | Mascar: Speeding up GPU warps by reducing memory pitstops · HPCA 2015 |
Parallel and multicore computing
data-parallel programming |
0.2 | 1 | 2014 | Paraprox: pattern-based approximation for data parallel applications · ASPLOS 2014 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 1 | 2014 | Paraprox: pattern-based approximation for data parallel applications · ASPLOS 2014 |
GPUs and heterogeneous computing › GPU computing
approximate GPU arithmetic |
0.2 | 1 | 2013 | SAGE: self-tuning approximation for graphics engines · MICRO 2013 |
Cloud and datacenter computing
computation offloading |
0.1 | 1 | 2012 | COMET: Code Offload by Migrating Execution Transparently · OSDI 2012 |
Energy-efficient computing › energy-efficient architecture
GPU energy efficiency |
0.1 | 1 | 2017 | Regless: just-in-time operand staging for GPUs · MICRO 2017 |
Energy-efficient computing
power management |
0.1 | 1 | 2017 | Regless: just-in-time operand staging for GPUs · MICRO 2017 |
Cloud and datacenter computing
mobile cloud computing |
0.0 | 1 | 2012 | COMET: Code Offload by Migrating Execution Transparently · OSDI 2012 |
Methods — techniques the papers use, named apart from their topics
runtime kernel selection · 0.5thread fusion · 0.4static compiler · 0.4data packing · 0.4warp scheduling · 0.2cache re-execution · 0.2runtime tuning · 0.2pattern-based approximation · 0.2CUDA · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Regless: just-in-time operand staging for GPUsabstractThe register file is one of the largest and most power-hungry structures in a Graphics Processing Unit (GPU), because massive multithreading requires all the register state for every active thread to be available. Previous approaches to making register accesses more efficient have optimized how registers are stored, but they must keep all values for active threads in a large, high-bandwidth structure. If operand storage is to be reduced further, there will not be enough capacity for every live value to be stored at the same time. Our insight is that computation graphs can be sliced into regions and operand storage can be allocated to these regions as they are encountered at run time, allowing a small operand staging unit to replace the register file. Most operand values have a short lifetime that is contained in one region, so their value does not need to persist in the staging unit past the end of that region. The small number of longer-lived operands can be stored in lower-bandwidth global memory, but the hardware must anticipate their use to fetch them early enough to avoid stalls. In RegLess, hardware uses compiler annotations to anticipate warps' operand usage at run time, allowing the register file to be replaced with an operand staging unit 25% of the size, saving 75% of register file energy and 11% of total GPU energy with no average performance loss. John Kloosterman, Jonathan Beaumont, Davoud Anoushe Jamshidi, Jonathan Bailey, Trevor N. Mudge, Scott A. Mahlke |
MICRO | 3 |
| 2015 | Mascar: Speeding up GPU warps by reducing memory pitstopsabstractWith the prevalence of GPUs as throughput engines for data parallel workloads, the landscape of GPU computing is changing significantly. Non-graphics workloads with high memory intensity and irregular access patterns are frequently targeted for acceleration on GPUs. While GPUs provide large numbers of compute resources, the resources needed for memory intensive workloads are more scarce. Therefore, managing access to these limited memory resources is a challenge for GPUs. We propose a novel Memory Aware Scheduling and Cache Access Re-execution (Mascar) system on GPUs tailored for better performance for memory intensive workloads. This scheme detects memory saturation and prioritizes memory requests among warps to enable better overlapping of compute and memory accesses. Furthermore, it enables limited re-execution of memory instructions to eliminate structural hazards in the memory subsystem and take advantage of cache locality in cases where requests cannot be sent to the memory due to memory saturation. Our results show that Mascar provides a 34% speedup over the baseline round-robin scheduler and 10% speedup over the state of the art warp schedulers for memory intensive workloads. Mascar also achieves an average of 12% savings in energy for such workloads. Ankit Sethia, Davoud Anoushe Jamshidi, Scott A. Mahlke |
HPCA | 2 |
| 2014 | D2MA: accelerating coarse-grained data transfer for GPUsabstractTo achieve high performance on many-core architectures like GPUs, it is crucial to efficiently utilize the available memory bandwidth. Currently, it is common to use fast, on-chip scratchpad memories, like the shared memory available on GPUs' shader cores, to buffer data for computation. This buffering, however, has some sources of inefficiency that hinder it from most efficiently utilizing the available memory resources. These issues stem from shader resources being used for repeated, regular address calculations, a need to shuffle data multiple times between a physically unified on-chip memory, and forcing all threads to synchronize to ensure RAW consistency based on the speed of the slowest threads. To address these inefficiencies, we propose Data-Parallel DMA, or D2MA. D2MA is a reimagination of traditional DMA that addresses the challenges of extending DMA to thousands of concurrently executing threads. D2MA de-couples address generation from the shader's computational resources, provides a more direct and efficient path for data in global memory to travel into the shared memory, and introduces a novel dynamic synchronization scheme that is transparent to the programmer. These advancements allow D2MA to achieve speedups as high as 2.29x, and reduces the average time to buffer data by 81% on average. Davoud Anoushe Jamshidi, Mehrzad Samadi, Scott A. Mahlke |
PACT | 1 |
| 2014 | Paraprox: pattern-based approximation for data parallel applicationsabstractApproximate computing is an approach where reduced accuracy of results is traded off for increased speed, throughput, or both. Loss of accuracy is not permissible in all computing domains, but there are a growing number of data-intensive domains where the output of programs need not be perfectly correct to provide useful results or even noticeable differences to the end user. These soft domains include multimedia processing, machine learning, and data mining/analysis. An important challenge with approximate computing is transparency to insulate both software and hardware developers from the time, cost, and difficulty of using approximation. This paper proposes a software-only system, Paraprox, for realizing transparent approximation of data-parallel programs that operates on commodity hardware systems. Paraprox starts with a data-parallel kernel implemented using OpenCL or CUDA and creates a parameterized approximate kernel that is tuned at runtime to maximize performance subject to a target output quality (TOQ) that is supplied by the user. Approximate kernels are created by recognizing common computation idioms found in data-parallel programs (e.g., Map, Scatter/Gather, Reduction, Scan, Stencil, and Partition) and substituting approximate implementations in their place. Across a set of 13 soft data-parallel applications with at most 10% quality degradation, Paraprox yields an average performance gain of 2.7x on a NVIDIA GTX 560 GPU and 2.5x on an Intel Core i7 quad-core processor compared to accurate execution on each platform. Mehrzad Samadi, Davoud Anoushe Jamshidi, Janghaeng Lee, Scott A. Mahlke |
ASPLOS | 2 |
| 2014 | Scaling Performance via Self-Tuning Approximation for Graphics EnginesabstractApproximate computing, where computation accuracy is traded off for better performance or higher data throughput, is one solution that can help data processing keep pace with the current and growing abundance of information. For particular domains, such as multimedia and learning algorithms, approximation is commonly used today. We consider automation to be essential to provide transparent approximation, and we show that larger benefits can be achieved by constructing the approximation techniques to fit the underlying hardware. Our target platform is the GPU because of its high performance capabilities and difficult programming challenges that can be alleviated with proper automation. Our approach—SAGE—combines a static compiler that automatically generates a set of CUDA kernels with varying levels of approximation with a runtime system that iteratively selects among the available kernels to achieve speedup while adhering to a target output quality set by the user. The SAGE compiler employs three optimization techniques to generate approximate kernels that exploit the GPU microarchitecture: selective discarding of atomic operations, data packing, and thread fusion. Across a set of machine learning and image processing kernels, SAGE's approximation yields an average of 2.5× speedup with less than 10% quality loss compared to the accurate execution on a NVIDIA GTX 560 GPU. Mehrzad Samadi, Janghaeng Lee, Davoud Anoushe Jamshidi, Scott A. Mahlke, Amir Hormati |
ACM Trans. Comput. Syst. | 3 |
| 2013 | SAGE: self-tuning approximation for graphics enginesabstractApproximate computing, where computation accuracy is traded off for better performance or higher data throughput, is one solution that can help data processing keep pace with the current and growing overabundance of information. For particular domains such as multimedia and learning algorithms, approximation is commonly used today. We consider automation to be essential to provide transparent approximation and we show that larger benefits can be achieved by constructing the approximation techniques to fit the underlying hardware. Our target platform is the GPU because of its high performance capabilities and difficult programming challenges that can be alleviated with proper automation. Our approach, SAGE, combines a static compiler that automatically generates a set of CUDA kernels with varying levels of approximation with a run-time system that iteratively selects among the available kernels to achieve speedup while adhering to a target output quality set by the user. The SAGE compiler employs three optimization techniques to generate approximate kernels that exploit the GPU microarchitecture: selective discarding of atomic operations, data packing, and thread fusion. Across a set of machine learning and image processing kernels, SAGE's approximation yields an average of 2.5x speedup with less than 10% quality loss compared to the accurate execution on a NVIDIA GTX 560 GPU. Mehrzad Samadi, Janghaeng Lee, Davoud Anoushe Jamshidi, Amir Hormati, Scott A. Mahlke |
MICRO | 3 |
| 2012 | COMET: Code Offload by Migrating Execution Transparently
Mark S. Gordon, Davoud Anoushe Jamshidi, Scott A. Mahlke, Z. Morley Mao, Xu Chen 0028 |
OSDI | 2 |