EDBT 2026 Demo / reviewers in the wild / expert
Naser Sedaghati
dblp:53/1705 · also Naser Sedaghati-Mokhtari
· DBLP profile ↗
11ranked-venue papers
3as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Energy-efficient computing · 37% Performance modeling and evaluation · 18% GPUs and heterogeneous computing · 13% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 11 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.2 | 1 | 2016 | Core tunneling: Variation-aware voltage noise mitigation in GPUs · HPCA 2016 |
Energy-efficient computing
power management |
0.2 | 1 | 2016 | Core tunneling: Variation-aware voltage noise mitigation in GPUs · HPCA 2016 |
Energy-efficient computing › power delivery
voltage noise mitigation |
0.2 | 1 | 2016 | Core tunneling: Variation-aware voltage noise mitigation in GPUs · HPCA 2016 |
Performance modeling and evaluation
bottleneck analysis |
0.2 | 1 | 2014 | On Using the Roofline Model with Lower Bounds on Data Movement · ACM Trans. Archit. Code Optim. 2014 |
GPUs and heterogeneous computing › GPU computing › GPU sparse computation
GPU sparse linear algebra |
0.2 | 1 | 2014 | Fast Sparse Matrix-Vector Multiplication on GPUs for Graph Applications · SC 2014 |
Performance modeling and evaluation › analytical modeling
roofline model |
0.2 | 1 | 2014 | On Using the Roofline Model with Lower Bounds on Data Movement · ACM Trans. Archit. Code Optim. 2014 |
High-performance computing › sparse linear algebra › sparse matrix computation
sparse matrix-vector multiplication |
0.2 | 1 | 2014 | Fast Sparse Matrix-Vector Multiplication on GPUs for Graph Applications · SC 2014 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2012 | Booster: Reactive core acceleration for mitigating the effects of process variation and application imbalance in low-voltage chips · HPCA 2012 |
Electronic design automation › yield analysis
process variation modeling |
0.1 | 1 | 2012 | Booster: Reactive core acceleration for mitigating the effects of process variation and application imbalance in low-voltage chips · HPCA 2012 |
GPUs and heterogeneous computing
GPU reliability |
0.1 | 1 | 2016 | Core tunneling: Variation-aware voltage noise mitigation in GPUs · HPCA 2016 |
Hardware reliability and fault tolerance
process variation |
0.1 | 1 | 2016 | Core tunneling: Variation-aware voltage noise mitigation in GPUs · HPCA 2016 |
Methods — techniques the papers use, named apart from their topics
critical path monitoring · 0.2clock gating · 0.2lower bounds on data movement · 0.2cache capacity analysis · 0.2gating circuit · 0.1dynamic voltage scaling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Core tunneling: Variation-aware voltage noise mitigation in GPUsabstractVoltage noise and manufacturing process variation represent significant reliability challenges for modern microprocessors. Voltage noise is caused by rapid changes in processor activity that can lead to timing violations and errors. Process variation is caused by manufacturing challenges in low-nanometer technologies and can lead to significant heterogeneity in performance and reliability across the chip. To ensure correct execution under worst-case conditions, chip designers generally add operating margins that are often unnecessarily conservative for most use cases, which results in wasted energy. This paper investigates the combined effects of process variation and voltage noise on modern GPU architectures. A distributed power delivery and process variation model at functional unit granularity was developed and used to simulate supply voltage behavior in a multicore GPU system. We observed that, just like in CPUs, large changes in power demand can lead to significant voltage droops. We also note that process variation makes some cores much more vulnerable to noise than others in the same GPU. Therefore, protecting the chip against large voltage droops by using fixed and uniform voltage guardbands is costly and inefficient. This paper presents core tunneling, a variation-aware solution for dynamically reducing voltage margins. The system relies on hardware critical path monitors to detect voltage noise conditions and quickly reacts by clock-gating vulnerable cores to prevent timing violations. This allows a substantial reduction in voltage margins. Since clock gating is enabled infrequently and only on the most vulnerable cores, the performance impact of core tunneling is very low. On average, core tunneling reduces energy consumption by 15%. Renji Thomas, Kristin Barber, Naser Sedaghati, Li Zhou 0012, Radu Teodorescu |
HPCA | 3 |
| 2016 | EmerGPU: Understanding and mitigating resonance-induced voltage noise in GPU architecturesabstractThis paper characterizes voltage noise in GPU architectures running general purpose workloads. In particular, it focuses on resonance-induced voltage noise, which is caused by workload-induced fluctuations in power demand that occur at the resonance frequency of the chip's power delivery network. A distributed power delivery model at functional unit granularity was developed and used to simulate supply voltage behavior in a GPU system. We observe that resonance noise can lead to very large voltage droops and protecting against these droops by using voltage guardbands is costly and inefficient. We propose EmerGPU, a solution that detects and mitigates resonance noise in GPUs. EmerGPU monitors workload activity levels and detects oscillations in power demand that approach resonance frequencies. When such conditions are detected, EmerGPU deploys a mitigation mechanism implemented in the warp scheduler that disrupts the resonance activity pattern. EmerGPU has no impact on performance and a small power cost. Reducing voltage noise improves system reliability and allows for smaller voltage margins to be used, reducing overall energy consumption by an average of 21%. Renji Thomas, Naser Sedaghati, Radu Teodorescu |
ISPASS | 2 |
| 2015 | Automatic Selection of Sparse Matrix Representation on GPUsabstractSparse matrix-vector multiplication (SpMV) is a core kernel in numerous applications, ranging from physics simulation and large-scale solvers to data analytics. Many GPU implementations of SpMV have been proposed, targeting several sparse representations and aiming at maximizing overall performance. No single sparse matrix representation is uniformly superior, and the best performing representation varies for sparse matrices with different sparsity patterns. Naser Sedaghati, Te Mu, Louis-Noël Pouchet, Srinivasan Parthasarathy 0001, P. Sadayappan |
ICS | 1 |
| 2015 | A model-driven blocking strategy for load balanced sparse matrix-vector multiplication on GPUs
Arash Ashari, Naser Sedaghati, John Eisenlohr, P. Sadayappan |
J. Parallel Distributed Comput. | 2 |
| 2014 | An efficient two-dimensional blocking strategy for sparse matrix-vector multiplication on GPUsabstractSparse matrix-vector multiplication (SpMV) is one of the key operations in linear algebra. Overcoming thread divergence, load imbalance and non-coalesced and indirect memory access due to sparsity and irregularity are challenges to optimizing SpMV on GPUs. Arash Ashari, Naser Sedaghati, John Eisenlohr, P. Sadayappan |
ICS | 2 |
| 2014 | Fast Sparse Matrix-Vector Multiplication on GPUs for Graph ApplicationsabstractSparse matrix-vector multiplication (SpMV) is a widely used computational kernel. The most commonly used format for a sparse matrix is CSR (Compressed Sparse Row), but a number of other representations have recently been developed that achieve higher SpMV performance. However, the alternative representations typically impose a significant preprocessing overhead. While a high preprocessing overhead can be amortized for applications requiring many iterative invocations of SpMV that use the same matrix, it is not always feasible -- for instance when analyzing large dynamically evolving graphs. This paper presents ACSR, an adaptive SpMV algorithm that uses the standard CSR format but reduces thread divergence by combining rows into groups (bins) which have a similar number of non-zero elements. Further, for rows in bins that span a wide range of non zero counts, dynamic parallelism is leveraged. A significant benefit of ACSR over other proposed SpMV approaches is that it works directly with the standard CSR format, and thus avoids significant preprocessing overheads. A CUDA implementation of ACSR is shown to outperform SpMV implementations in the NVIDIA CUSP and cuSPARSE libraries on a set of sparse matrices representing power-law graphs. We also demonstrate the use of ACSR for the analysis of dynamic graphs, where the improvement over extant approaches is even higher. Arash Ashari, Naser Sedaghati, John Eisenlohr, Srinivasan Parthasarathy 0001, P. Sadayappan |
SC | 2 |
| 2014 | On Using the Roofline Model with Lower Bounds on Data MovementabstractThe roofline model is a popular approach for “bound and bottleneck” performance analysis. It focuses on the limits to the performance of processors because of limited bandwidth to off-chip memory. It models upper bounds on performance as a function of operational intensity, the ratio of computational operations per byte of data moved from/to memory. While operational intensity can be directly measured for a specific implementation of an algorithm on a particular target platform, it is of interest to obtain broader insights on bottlenecks, where various semantically equivalent implementations of an algorithm are considered, along with analysis for variations in architectural parameters. This is currently very cumbersome and requires performance modeling and analysis of many variants. In this article, we address this problem by using the roofline model in conjunction with upper bounds on the operational intensity of computations as a function of cache capacity, derived from lower bounds on data movement. This enables bottleneck analysis that holds across all dependence-preserving semantically equivalent implementations of an algorithm. We demonstrate the utility of the approach in assessing fundamental limits to performance and energy efficiency for several benchmark algorithms across a design space of architectural variations. Venmugil Elango, Naser Sedaghati, Fabrice Rastello, Louis-Noël Pouchet, J. Ramanujam, Radu Teodorescu, P. Sadayappan |
ACM Trans. Archit. Code Optim. | 2 |
| 2012 | Booster: Reactive core acceleration for mitigating the effects of process variation and application imbalance in low-voltage chipsabstractLowering supply voltage is one of the most effective techniques for reducing microprocessor power consumption. Unfortunately, at low voltages, chips are very sensitive to process variation, which can lead to large differences in the maximum frequency achieved by individual cores. This paper presents Booster, a simple, low-overhead framework for dynamically rebalancing performance heterogeneity caused by process variation and application imbalance. The Booster CMP includes two power supply rails set at two very low but different voltages. Each core can be dynamically assigned to either of the two rails using a gating circuit. This allows cores to quickly switch between two different frequencies. An on-chip governor controls the timing of the switching and the time spent on each rail. The governor manages a “boost budget” that dictates how many cores can be sped up (depending on the power constraints) at any given time. We present two implementations of Booster: Booster VAR, which virtually eliminates the effects of core-to-core frequency variation in near-threshold CMPs, and Booster SYNC, which additionally reduces the effects of imbalance in multithreaded applications. Evaluation using PARSEC and SPLASH2 benchmarks running on a simulated 32-core system shows an average performance improvement of 11% for Booster VAR and 23% for Booster SYNC. Timothy N. Miller, Renji Thomas, Naser Sedaghati, Radu Teodorescu |
HPCA | 4 |
| 2011 | StVEC: A Vector Instruction Extension for High Performance Stencil ComputationabstractStencil computations comprise the compute-intensive core of many scientific applications. The data access pattern of stencil computations often requires several adjacent data elements of arrays to be accessed in innermost parallel loops. Although such loops are vectorized by current compilers like GCC and ICC that target short-vector SIMD instruction sets, a number of redundant loads or additional intra-register data shuffle operations are required, reducing the achievable performance. Thus, even when all arrays are cache resident, the peak performance achieved with stencil computations is considerably lower than machine peak. In this paper, we present a hardware-based solution for this problem. We propose an extension to the standard addressing mode of vector floating-point instructions in ISAs such as SSE, AVX, VMX etc. We propose an extended mode of paired-register addressing and its hardware implementation, to overcome the performance limitation of current short-vector SIMD ISA's for stencil computations. Further, we present a code generation approach that can be used by a vectorizing compiler for processors with such an instructions set. Using an optimistic as well as a pessimistic emulation of the proposed instruction extension, we demonstrate the effectiveness of the proposed approach on top of SSE and AVX capable processors. We also synthesize parts of the proposed design using a 45nm CMOS library and show minimal impact on processor cycle time. Naser Sedaghati, Renji Thomas, Louis-Noël Pouchet, Radu Teodorescu, P. Sadayappan |
PACT | 1 |
| 2007 | MDST: Multiprocessor DSP Simulation Toolkit for Voice Processing ApplicationsabstractIn this paper, we propose a multiprocessor DSP simulation toolkit suitable for performance evaluation of data-parallel applications like voice processing. The proposed toolkit uses the benefits of multi-level parallelism and clustering. Different DSP clusters are considered for the multiprocessor DSP simulation engine in which the DSP processors are grouped to cooperate. Satisfying the communication requirements, two global and local communication engines (GCE and LCE) implement the real behavior of intra-and inter-cluster communications. Using efficient abstraction levels for interconnections reduces the simulation time significantly. Abstract communication modeling, cycle-accurate behavior, and multi-level controlling are the most important features of the proposed simulation platform. Performance of the simulator is verified by standard single- and multi-channel voice processing applications such as ITU-T G.729a speech codec. Naser Sedaghati, Mahdi Nazm Bojnordi, Sied Mehdi Fakhraie |
MASCOTS | 1 |
| 2006 | Power efficient sequential multiplication using pre-computationabstractA pre-computation based technique to lower the power consumption of sequential multipliers is presented. This technique also speeds up the multiplication by reducing the number of clock ticks required to complete a multiplication. The proposed technique may be applied to different sequential multiplication schemes. The benchmark data is extracted from typical DSP applications to show the efficiency of the proposed technique in the domain of DSP computations in which the low power computing is of rapidly increasing importance. The results show an average of 25% reduction in the switching activity and 30% reduction in the clock tick count, compared to sequential multipliers without this technique Nima Honarmand, M. Reza Javaheri, Naser Sedaghati, Ali Afzali-Kusha |
ISCAS | 3 |