Nagesh B. Lakshminarayana

dblp:00/1316 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Processor architecture and microarchitecture · 42% Memory systems · 26% GPUs and heterogeneous computing · 13%
Software engineering, system software, and programming languages
1 paper
Program analysis · 100%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
memory latency tolerance
0.322014
Spare register aware prefetching for graph algorithms on GPUs · HPCA 2014
Many-Thread Aware Prefetching Mechanisms for GPGPU Applications · MICRO 2010
Memory systems › cache
prefetching
0.322014
Spare register aware prefetching for graph algorithms on GPUs · HPCA 2014
Many-Thread Aware Prefetching Mechanisms for GPGPU Applications · MICRO 2010
Energy-efficient computing › low-power design
low-power processor design
0.212015
Block-Precise Processors: Low-Power Processors with Reduced Operand Store Accesses and Result Broadcasts · IEEE Trans. Computers 2015
Processor architecture and microarchitecture › out-of-order execution
out-of-order processor
0.212015
Block-Precise Processors: Low-Power Processors with Reduced Operand Store Accesses and Result Broadcasts · IEEE Trans. Computers 2015
GPUs and heterogeneous computing
GPU memory access
0.212014
Spare register aware prefetching for graph algorithms on GPUs · HPCA 2014
Memory systems › memory access patterns
irregular memory access
0.212014
Spare register aware prefetching for graph algorithms on GPUs · HPCA 2014
Program analysis
dynamic analysis
0.212013
SD3: An Efficient Dynamic Data-Dependence Profiling Mechanism · IEEE Trans. Computers 2013
Memory systems
cache
0.112010
Many-Thread Aware Prefetching Mechanisms for GPGPU Applications · MICRO 2010
GPUs and heterogeneous computing
GPU computing
0.112010
Many-Thread Aware Prefetching Mechanisms for GPGPU Applications · MICRO 2010
Parallel and multicore computing › parallel scheduling
heterogeneous multiprocessor scheduling
0.112009
Age based scheduling for asymmetric multiprocessors · SC 2009
Electronic design automation › high-level synthesis
scheduling
0.112009
Age based scheduling for asymmetric multiprocessors · SC 2009
Processor architecture and microarchitecture › pipelining
bypass network
0.112015
Block-Precise Processors: Low-Power Processors with Reduced Operand Store Accesses and Result Broadcasts · IEEE Trans. Computers 2015
Processor architecture and microarchitecture › pipelining
instruction pipeline
0.112014
Spare register aware prefetching for graph algorithms on GPUs · HPCA 2014
Processor architecture and microarchitecture
multithreading
0.012010
Many-Thread Aware Prefetching Mechanisms for GPGPU Applications · MICRO 2010
Parallel and multicore computing
thread-level parallelism
0.012010
Many-Thread Aware Prefetching Mechanisms for GPGPU Applications · MICRO 2010
Processor architecture and microarchitecture › multiprocessor architecture
asymmetric multiprocessor
0.012009
Age based scheduling for asymmetric multiprocessors · SC 2009
Processor architecture and microarchitecture
multicore design
0.012009
Age based scheduling for asymmetric multiprocessors · SC 2009

Methods — techniques the papers use, named apart from their topics

stride pattern detection · 0.3memory access compression · 0.3register file caching · 0.2hardware load detection · 0.2compiler-assisted load detection · 0.2software prefetching · 0.1hardware prefetching · 0.1adaptive prefetch throttling · 0.1scheduling algorithm · 0.1
YearPublicationVenuePosition
2015 Block-Precise Processors: Low-Power Processors with Reduced Operand Store Accesses and Result Broadcasts
abstract
Power is a first order design constraint for most processors today. Benefits of low power designs include lower manufacturing and operating costs and a longer battery life. In this work we propose an out of order processor architecture called Block-precise processor (B-Processor) that is designed for low power consumption. The B-Processor consumes lower power than typical processor designs by eliding the write of results of many instructions to the reorder buffer and to the register file, which are power hungry structures. The B-Processor reduces power consumption even further by omitting the broadcast of certain results over multiple levels of the bypass network. Experimental results show that on average the B-Processor spends 15.1 percent less power on register file and reorder buffer accesses and 14.5 percent less power on broadcasting results. In combination with register file caching, on average the B-Processor saves 28.7 percent power for accessing the register file and the reorder buffer.
Nagesh B. Lakshminarayana, Hyesoon Kim
IEEE Trans. Computers1
2014 Spare register aware prefetching for graph algorithms on GPUs
abstract
More and more graph algorithms are being GPU enabled. Graph algorithm implementations on GPUs have irregular control flow and are memory-intensive with many irregular/data-dependent memory accesses. Due to these factors graph algorithms on GPUs have low execution efficiency. In this work we propose a mechanism to improve the execution efficiency of graph algorithms by improving their memory access latency tolerance. We propose a mechanism for prefetching data for load pairs that have one load dependent on the other - such pairs are common in graph algorithms. Our mechanism detects the target loads in hardware and injects instructions into the pipeline to prefetch data into spare registers that are not being used by any active threads. By prefetching data into registers, early eviction of prefetched data can be eliminated. We also propose a mechanism that uses the compiler to identify the target loads. Our mechanism improves performance over no prefetching by 10% on average and upto 51% for nine memory intensive graph algorithm kernels.
Nagesh B. Lakshminarayana, Hyesoon Kim
HPCA1
2014 Power Modeling for GPU Architectures Using McPAT
abstract
Graphics Processing Units (GPUs) are very popular for both graphics and general-purpose applications. Since GPUs operate many processing units and manage multiple levels of memory hierarchy, they consume a significant amount of power. Although several power models for CPUs are available, the power consumption of GPUs has not been studied much yet. In this article we develop a new power model for GPUs by utilizing McPAT, a CPU power tool. We generate initial power model data from McPAT with a detailed GPU configuration, and then adjust the models by comparing them with empirical data. We use the NVIDIA's Fermi architecture for building the power model, and our model estimates the GPU power consumption with an average error of 7.7% and 12.8% for the microbenchmarks and Merge benchmarks, respectively.
Jieun Lim 0001, Nagesh B. Lakshminarayana, Hyesoon Kim, William J. Song, Sudhakar Yalamanchili, Wonyong Sung
ACM Trans. Design Autom. Electr. Syst.2
2013 SD3: An Efficient Dynamic Data-Dependence Profiling Mechanism
abstract
As multicore processors are deployed in mainstream computing, the need for software tools to help parallelize programs is increasing dramatically. Data-dependence profiling is an important program analysis technique to exploit parallelism in serial programs. More specifically, manual, semiautomatic, or automatic parallelization can use the outcomes of data-dependence profiling to guide where and how to parallelize in a program. However, state-of-the-art data-dependence profiling techniques consume extremely huge resources as they suffer from two major issues when profiling large and long-running applications: 1) runtime overhead and 2) memory overhead. Existing data-dependence profilers are either unable to profile large-scale applications with a typical resource budget or only report very limited information. In this paper, we propose an efficient approach to data-dependence profiling that can address both runtime and memory overhead in a single framework. Our technique, called SD$({}^3)$, reduces the runtime overhead by parallelizing the dependence profiling step itself. To reduce the memory overhead, we compress memory accesses that exhibit stride patterns and compute data dependences directly in a compressed format. We demonstrate that SD$({}^3)$ reduces the runtime overhead when profiling SPEC 2006 by a factor of 4.1× and 9.7× on eight cores and 32 cores, respectively. For the memory overhead, we successfully profile 22 SPEC 2006 benchmarks with the reference input, while the previous approaches fail even with the train input. In some cases, we observe more than a 20× improvement in memory consumption and a 16× speedup in profiling time when 32 cores are used. We also demonstrate the usefulness of SD$({}^3)$ by showing manual parallelization followed by data dependence profiling results.
Minjang Kim, Nagesh B. Lakshminarayana, Hyesoon Kim, Chi-Keung Luk
IEEE Trans. Computers2
2010 Many-Thread Aware Prefetching Mechanisms for GPGPU Applications
abstract
We consider the problem of how to improve memory latency tolerance in massively multithreaded GPGPUs when the thread-level parallelism of an application is not sufficient to hide memory latency. One solution used in conventional CPU systems is prefetching, both in hardware and software. However, we show that straightforwardly applying such mechanisms to GPGPU systems does not deliver the expected performance benefits and can in fact hurt performance when not used judiciously. This paper proposes new hardware and software prefetching mechanisms tailored to GPGPU systems, which we refer to as many-thread aware prefetching (MT-prefetching) mechanisms. Our software MT-prefetching mechanism, called inter-thread prefetching, exploits the existence of common memory access behavior among fine-grained threads. For hardware MT-prefetching, we describe a scalable prefetcher training algorithm along with a hardware-based inter-thread prefetching mechanism. In some cases, blindly applying prefetching degrades performance. To reduce such negative effects, we propose an adaptive prefetch throttling scheme, which permits automatic GPGPU application- and hardware-specific adjustment. We show that adaptation reduces the negative effects of prefetching and can even improve performance. Overall, compared to the state-of-the-art software and hardware prefetching, our MT-prefetching improves performance on average by 16%(software pref.)/15% (hardware pref.) on our benchmarks.
Jaekyu Lee, Nagesh B. Lakshminarayana, Hyesoon Kim, Richard W. Vuduc
MICRO2
2009 Age based scheduling for asymmetric multiprocessors
abstract
Asymmetric (or Heterogeneous) Multiprocessors are becoming popular in the current era of multi-cores due to their power efficiency and potential performance and energy efficiency. However, scheduling of multithreaded applications in Asymmetric Multiprocessors is still a challenging problem. Scheduling algorithms for Asymmetric Multiprocessors must not only be aware of asymmetry in processor performance, but have to consider the characteristics of application threads also.
Nagesh B. Lakshminarayana, Jaekyu Lee, Hyesoon Kim
SC1
2008 Understanding performance, power and energy behavior in asymmetric multiprocessors
abstract
Multiprocessor architectures are becoming popular in both desktop and mobile processors. Among multiprocessor architectures, asymmetric architectures show promise in saving energy and power. However, the performance and energy consumption behavior of asymmetric multiprocessors with desktop-oriented multithreaded applications has not been studied widely. In this study, we measure performance and power consumption in asymmetric and symmetric multiprocessors using real 8 and 16 processor systems to understand the relationships between thread interactions and performance/power behavior. We find that when the workload is asymmetric, using an asymmetric multiprocessor can save energy, but for most of the symmetric workloads, using a symmetric multiprocessor (with the highest clock frequency) consumes less energy.
Nagesh B. Lakshminarayana, Hyesoon Kim
ICCD1