Naser Sedaghati

dblp:53/1705 · also Naser Sedaghati-Mokhtari · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Energy-efficient computing · 37% Performance modeling and evaluation · 18% GPUs and heterogeneous computing · 13%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing › power management
dynamic voltage and frequency scaling
0.212016
Core tunneling: Variation-aware voltage noise mitigation in GPUs · HPCA 2016
Energy-efficient computing
power management
0.212016
Core tunneling: Variation-aware voltage noise mitigation in GPUs · HPCA 2016
Energy-efficient computing › power delivery
voltage noise mitigation
0.212016
Core tunneling: Variation-aware voltage noise mitigation in GPUs · HPCA 2016
Performance modeling and evaluation
bottleneck analysis
0.212014
On Using the Roofline Model with Lower Bounds on Data Movement · ACM Trans. Archit. Code Optim. 2014
GPUs and heterogeneous computing › GPU computing › GPU sparse computation
GPU sparse linear algebra
0.212014
Fast Sparse Matrix-Vector Multiplication on GPUs for Graph Applications · SC 2014
Performance modeling and evaluation › analytical modeling
roofline model
0.212014
On Using the Roofline Model with Lower Bounds on Data Movement · ACM Trans. Archit. Code Optim. 2014
High-performance computing › sparse linear algebra › sparse matrix computation
sparse matrix-vector multiplication
0.212014
Fast Sparse Matrix-Vector Multiplication on GPUs for Graph Applications · SC 2014
Processor architecture and microarchitecture
chip multiprocessor
0.112012
Booster: Reactive core acceleration for mitigating the effects of process variation and application imbalance in low-voltage chips · HPCA 2012
Electronic design automation › yield analysis
process variation modeling
0.112012
Booster: Reactive core acceleration for mitigating the effects of process variation and application imbalance in low-voltage chips · HPCA 2012
GPUs and heterogeneous computing
GPU reliability
0.112016
Core tunneling: Variation-aware voltage noise mitigation in GPUs · HPCA 2016
Hardware reliability and fault tolerance
process variation
0.112016
Core tunneling: Variation-aware voltage noise mitigation in GPUs · HPCA 2016

Methods — techniques the papers use, named apart from their topics

critical path monitoring · 0.2clock gating · 0.2lower bounds on data movement · 0.2cache capacity analysis · 0.2gating circuit · 0.1dynamic voltage scaling · 0.1
YearPublicationVenuePosition
2016 Core tunneling: Variation-aware voltage noise mitigation in GPUs
abstract
Voltage noise and manufacturing process variation represent significant reliability challenges for modern microprocessors. Voltage noise is caused by rapid changes in processor activity that can lead to timing violations and errors. Process variation is caused by manufacturing challenges in low-nanometer technologies and can lead to significant heterogeneity in performance and reliability across the chip. To ensure correct execution under worst-case conditions, chip designers generally add operating margins that are often unnecessarily conservative for most use cases, which results in wasted energy. This paper investigates the combined effects of process variation and voltage noise on modern GPU architectures. A distributed power delivery and process variation model at functional unit granularity was developed and used to simulate supply voltage behavior in a multicore GPU system. We observed that, just like in CPUs, large changes in power demand can lead to significant voltage droops. We also note that process variation makes some cores much more vulnerable to noise than others in the same GPU. Therefore, protecting the chip against large voltage droops by using fixed and uniform voltage guardbands is costly and inefficient. This paper presents core tunneling, a variation-aware solution for dynamically reducing voltage margins. The system relies on hardware critical path monitors to detect voltage noise conditions and quickly reacts by clock-gating vulnerable cores to prevent timing violations. This allows a substantial reduction in voltage margins. Since clock gating is enabled infrequently and only on the most vulnerable cores, the performance impact of core tunneling is very low. On average, core tunneling reduces energy consumption by 15%.
Renji Thomas, Kristin Barber, Naser Sedaghati, Li Zhou 0012, Radu Teodorescu
HPCA3
2016 EmerGPU: Understanding and mitigating resonance-induced voltage noise in GPU architectures
abstract
This paper characterizes voltage noise in GPU architectures running general purpose workloads. In particular, it focuses on resonance-induced voltage noise, which is caused by workload-induced fluctuations in power demand that occur at the resonance frequency of the chip's power delivery network. A distributed power delivery model at functional unit granularity was developed and used to simulate supply voltage behavior in a GPU system. We observe that resonance noise can lead to very large voltage droops and protecting against these droops by using voltage guardbands is costly and inefficient. We propose EmerGPU, a solution that detects and mitigates resonance noise in GPUs. EmerGPU monitors workload activity levels and detects oscillations in power demand that approach resonance frequencies. When such conditions are detected, EmerGPU deploys a mitigation mechanism implemented in the warp scheduler that disrupts the resonance activity pattern. EmerGPU has no impact on performance and a small power cost. Reducing voltage noise improves system reliability and allows for smaller voltage margins to be used, reducing overall energy consumption by an average of 21%.
Renji Thomas, Naser Sedaghati, Radu Teodorescu
ISPASS2
2015 Automatic Selection of Sparse Matrix Representation on GPUs
abstract
Sparse matrix-vector multiplication (SpMV) is a core kernel in numerous applications, ranging from physics simulation and large-scale solvers to data analytics. Many GPU implementations of SpMV have been proposed, targeting several sparse representations and aiming at maximizing overall performance. No single sparse matrix representation is uniformly superior, and the best performing representation varies for sparse matrices with different sparsity patterns.
Naser Sedaghati, Te Mu, Louis-Noël Pouchet, Srinivasan Parthasarathy 0001, P. Sadayappan
ICS1
2015 A model-driven blocking strategy for load balanced sparse matrix-vector multiplication on GPUs
Arash Ashari, Naser Sedaghati, John Eisenlohr, P. Sadayappan
J. Parallel Distributed Comput.2
2014 An efficient two-dimensional blocking strategy for sparse matrix-vector multiplication on GPUs
abstract
Sparse matrix-vector multiplication (SpMV) is one of the key operations in linear algebra. Overcoming thread divergence, load imbalance and non-coalesced and indirect memory access due to sparsity and irregularity are challenges to optimizing SpMV on GPUs.
Arash Ashari, Naser Sedaghati, John Eisenlohr, P. Sadayappan
ICS2
2014 Fast Sparse Matrix-Vector Multiplication on GPUs for Graph Applications
abstract
Sparse matrix-vector multiplication (SpMV) is a widely used computational kernel. The most commonly used format for a sparse matrix is CSR (Compressed Sparse Row), but a number of other representations have recently been developed that achieve higher SpMV performance. However, the alternative representations typically impose a significant preprocessing overhead. While a high preprocessing overhead can be amortized for applications requiring many iterative invocations of SpMV that use the same matrix, it is not always feasible -- for instance when analyzing large dynamically evolving graphs. This paper presents ACSR, an adaptive SpMV algorithm that uses the standard CSR format but reduces thread divergence by combining rows into groups (bins) which have a similar number of non-zero elements. Further, for rows in bins that span a wide range of non zero counts, dynamic parallelism is leveraged. A significant benefit of ACSR over other proposed SpMV approaches is that it works directly with the standard CSR format, and thus avoids significant preprocessing overheads. A CUDA implementation of ACSR is shown to outperform SpMV implementations in the NVIDIA CUSP and cuSPARSE libraries on a set of sparse matrices representing power-law graphs. We also demonstrate the use of ACSR for the analysis of dynamic graphs, where the improvement over extant approaches is even higher.
Arash Ashari, Naser Sedaghati, John Eisenlohr, Srinivasan Parthasarathy 0001, P. Sadayappan
SC2
2014 On Using the Roofline Model with Lower Bounds on Data Movement
abstract
The roofline model is a popular approach for “bound and bottleneck” performance analysis. It focuses on the limits to the performance of processors because of limited bandwidth to off-chip memory. It models upper bounds on performance as a function of operational intensity, the ratio of computational operations per byte of data moved from/to memory. While operational intensity can be directly measured for a specific implementation of an algorithm on a particular target platform, it is of interest to obtain broader insights on bottlenecks, where various semantically equivalent implementations of an algorithm are considered, along with analysis for variations in architectural parameters. This is currently very cumbersome and requires performance modeling and analysis of many variants. In this article, we address this problem by using the roofline model in conjunction with upper bounds on the operational intensity of computations as a function of cache capacity, derived from lower bounds on data movement. This enables bottleneck analysis that holds across all dependence-preserving semantically equivalent implementations of an algorithm. We demonstrate the utility of the approach in assessing fundamental limits to performance and energy efficiency for several benchmark algorithms across a design space of architectural variations.
Venmugil Elango, Naser Sedaghati, Fabrice Rastello, Louis-Noël Pouchet, J. Ramanujam, Radu Teodorescu, P. Sadayappan
ACM Trans. Archit. Code Optim.2
2012 Booster: Reactive core acceleration for mitigating the effects of process variation and application imbalance in low-voltage chips
abstract
Lowering supply voltage is one of the most effective techniques for reducing microprocessor power consumption. Unfortunately, at low voltages, chips are very sensitive to process variation, which can lead to large differences in the maximum frequency achieved by individual cores. This paper presents Booster, a simple, low-overhead framework for dynamically rebalancing performance heterogeneity caused by process variation and application imbalance. The Booster CMP includes two power supply rails set at two very low but different voltages. Each core can be dynamically assigned to either of the two rails using a gating circuit. This allows cores to quickly switch between two different frequencies. An on-chip governor controls the timing of the switching and the time spent on each rail. The governor manages a “boost budget” that dictates how many cores can be sped up (depending on the power constraints) at any given time. We present two implementations of Booster: Booster VAR, which virtually eliminates the effects of core-to-core frequency variation in near-threshold CMPs, and Booster SYNC, which additionally reduces the effects of imbalance in multithreaded applications. Evaluation using PARSEC and SPLASH2 benchmarks running on a simulated 32-core system shows an average performance improvement of 11% for Booster VAR and 23% for Booster SYNC.
Timothy N. Miller, Renji Thomas, Naser Sedaghati, Radu Teodorescu
HPCA4
2011 StVEC: A Vector Instruction Extension for High Performance Stencil Computation
abstract
Stencil computations comprise the compute-intensive core of many scientific applications. The data access pattern of stencil computations often requires several adjacent data elements of arrays to be accessed in innermost parallel loops. Although such loops are vectorized by current compilers like GCC and ICC that target short-vector SIMD instruction sets, a number of redundant loads or additional intra-register data shuffle operations are required, reducing the achievable performance. Thus, even when all arrays are cache resident, the peak performance achieved with stencil computations is considerably lower than machine peak. In this paper, we present a hardware-based solution for this problem. We propose an extension to the standard addressing mode of vector floating-point instructions in ISAs such as SSE, AVX, VMX etc. We propose an extended mode of paired-register addressing and its hardware implementation, to overcome the performance limitation of current short-vector SIMD ISA's for stencil computations. Further, we present a code generation approach that can be used by a vectorizing compiler for processors with such an instructions set. Using an optimistic as well as a pessimistic emulation of the proposed instruction extension, we demonstrate the effectiveness of the proposed approach on top of SSE and AVX capable processors. We also synthesize parts of the proposed design using a 45nm CMOS library and show minimal impact on processor cycle time.
Naser Sedaghati, Renji Thomas, Louis-Noël Pouchet, Radu Teodorescu, P. Sadayappan
PACT1
2007 MDST: Multiprocessor DSP Simulation Toolkit for Voice Processing Applications
abstract
In this paper, we propose a multiprocessor DSP simulation toolkit suitable for performance evaluation of data-parallel applications like voice processing. The proposed toolkit uses the benefits of multi-level parallelism and clustering. Different DSP clusters are considered for the multiprocessor DSP simulation engine in which the DSP processors are grouped to cooperate. Satisfying the communication requirements, two global and local communication engines (GCE and LCE) implement the real behavior of intra-and inter-cluster communications. Using efficient abstraction levels for interconnections reduces the simulation time significantly. Abstract communication modeling, cycle-accurate behavior, and multi-level controlling are the most important features of the proposed simulation platform. Performance of the simulator is verified by standard single- and multi-channel voice processing applications such as ITU-T G.729a speech codec.
Naser Sedaghati, Mahdi Nazm Bojnordi, Sied Mehdi Fakhraie
MASCOTS1
2006 Power efficient sequential multiplication using pre-computation
abstract
A pre-computation based technique to lower the power consumption of sequential multipliers is presented. This technique also speeds up the multiplication by reducing the number of clock ticks required to complete a multiplication. The proposed technique may be applied to different sequential multiplication schemes. The benchmark data is extracted from typical DSP applications to show the efficiency of the proposed technique in the domain of DSP computations in which the low power computing is of rapidly increasing importance. The results show an average of 25% reduction in the switching activity and 30% reduction in the clock tick count, compared to sequential multipliers without this technique
Nima Honarmand, M. Reza Javaheri, Naser Sedaghati, Ali Afzali-Kusha
ISCAS3