VLDB 2026 Research / reviewers in the wild / expert
Nagesh B. Lakshminarayana
dblp:00/1316
· DBLP profile ↗
7ranked-venue papers
4as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Processor architecture and microarchitecture · 42% Memory systems · 26% GPUs and heterogeneous computing · 13% | |
| Software engineering, system software, and programming languages
1 paper |
Program analysis · 100% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
memory latency tolerance |
0.3 | 2 | 2014 | Spare register aware prefetching for graph algorithms on GPUs · HPCA 2014 Many-Thread Aware Prefetching Mechanisms for GPGPU Applications · MICRO 2010 |
Memory systems › cache
prefetching |
0.3 | 2 | 2014 | Spare register aware prefetching for graph algorithms on GPUs · HPCA 2014 Many-Thread Aware Prefetching Mechanisms for GPGPU Applications · MICRO 2010 |
Energy-efficient computing › low-power design
low-power processor design |
0.2 | 1 | 2015 | Block-Precise Processors: Low-Power Processors with Reduced Operand Store Accesses and Result Broadcasts · IEEE Trans. Computers 2015 |
Processor architecture and microarchitecture › out-of-order execution
out-of-order processor |
0.2 | 1 | 2015 | Block-Precise Processors: Low-Power Processors with Reduced Operand Store Accesses and Result Broadcasts · IEEE Trans. Computers 2015 |
GPUs and heterogeneous computing
GPU memory access |
0.2 | 1 | 2014 | Spare register aware prefetching for graph algorithms on GPUs · HPCA 2014 |
Memory systems › memory access patterns
irregular memory access |
0.2 | 1 | 2014 | Spare register aware prefetching for graph algorithms on GPUs · HPCA 2014 |
Program analysis
dynamic analysis |
0.2 | 1 | 2013 | SD3: An Efficient Dynamic Data-Dependence Profiling Mechanism · IEEE Trans. Computers 2013 |
Memory systems
cache |
0.1 | 1 | 2010 | Many-Thread Aware Prefetching Mechanisms for GPGPU Applications · MICRO 2010 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2010 | Many-Thread Aware Prefetching Mechanisms for GPGPU Applications · MICRO 2010 |
Parallel and multicore computing › parallel scheduling
heterogeneous multiprocessor scheduling |
0.1 | 1 | 2009 | Age based scheduling for asymmetric multiprocessors · SC 2009 |
Electronic design automation › high-level synthesis
scheduling |
0.1 | 1 | 2009 | Age based scheduling for asymmetric multiprocessors · SC 2009 |
Processor architecture and microarchitecture › pipelining
bypass network |
0.1 | 1 | 2015 | Block-Precise Processors: Low-Power Processors with Reduced Operand Store Accesses and Result Broadcasts · IEEE Trans. Computers 2015 |
Processor architecture and microarchitecture › pipelining
instruction pipeline |
0.1 | 1 | 2014 | Spare register aware prefetching for graph algorithms on GPUs · HPCA 2014 |
Processor architecture and microarchitecture
multithreading |
0.0 | 1 | 2010 | Many-Thread Aware Prefetching Mechanisms for GPGPU Applications · MICRO 2010 |
Parallel and multicore computing
thread-level parallelism |
0.0 | 1 | 2010 | Many-Thread Aware Prefetching Mechanisms for GPGPU Applications · MICRO 2010 |
Processor architecture and microarchitecture › multiprocessor architecture
asymmetric multiprocessor |
0.0 | 1 | 2009 | Age based scheduling for asymmetric multiprocessors · SC 2009 |
Processor architecture and microarchitecture
multicore design |
0.0 | 1 | 2009 | Age based scheduling for asymmetric multiprocessors · SC 2009 |
Methods — techniques the papers use, named apart from their topics
stride pattern detection · 0.3memory access compression · 0.3register file caching · 0.2hardware load detection · 0.2compiler-assisted load detection · 0.2software prefetching · 0.1hardware prefetching · 0.1adaptive prefetch throttling · 0.1scheduling algorithm · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Block-Precise Processors: Low-Power Processors with Reduced Operand Store Accesses and Result BroadcastsabstractPower is a first order design constraint for most processors today. Benefits of low power designs include lower manufacturing and operating costs and a longer battery life. In this work we propose an out of order processor architecture called Block-precise processor (B-Processor) that is designed for low power consumption. The B-Processor consumes lower power than typical processor designs by eliding the write of results of many instructions to the reorder buffer and to the register file, which are power hungry structures. The B-Processor reduces power consumption even further by omitting the broadcast of certain results over multiple levels of the bypass network. Experimental results show that on average the B-Processor spends 15.1 percent less power on register file and reorder buffer accesses and 14.5 percent less power on broadcasting results. In combination with register file caching, on average the B-Processor saves 28.7 percent power for accessing the register file and the reorder buffer. Nagesh B. Lakshminarayana, Hyesoon Kim |
IEEE Trans. Computers | 1 |
| 2014 | Spare register aware prefetching for graph algorithms on GPUsabstractMore and more graph algorithms are being GPU enabled. Graph algorithm implementations on GPUs have irregular control flow and are memory-intensive with many irregular/data-dependent memory accesses. Due to these factors graph algorithms on GPUs have low execution efficiency. In this work we propose a mechanism to improve the execution efficiency of graph algorithms by improving their memory access latency tolerance. We propose a mechanism for prefetching data for load pairs that have one load dependent on the other - such pairs are common in graph algorithms. Our mechanism detects the target loads in hardware and injects instructions into the pipeline to prefetch data into spare registers that are not being used by any active threads. By prefetching data into registers, early eviction of prefetched data can be eliminated. We also propose a mechanism that uses the compiler to identify the target loads. Our mechanism improves performance over no prefetching by 10% on average and upto 51% for nine memory intensive graph algorithm kernels. Nagesh B. Lakshminarayana, Hyesoon Kim |
HPCA | 1 |
| 2014 | Power Modeling for GPU Architectures Using McPATabstractGraphics Processing Units (GPUs) are very popular for both graphics and general-purpose applications. Since GPUs operate many processing units and manage multiple levels of memory hierarchy, they consume a significant amount of power. Although several power models for CPUs are available, the power consumption of GPUs has not been studied much yet. In this article we develop a new power model for GPUs by utilizing McPAT, a CPU power tool. We generate initial power model data from McPAT with a detailed GPU configuration, and then adjust the models by comparing them with empirical data. We use the NVIDIA's Fermi architecture for building the power model, and our model estimates the GPU power consumption with an average error of 7.7% and 12.8% for the microbenchmarks and Merge benchmarks, respectively. Jieun Lim 0001, Nagesh B. Lakshminarayana, Hyesoon Kim, William J. Song, Sudhakar Yalamanchili, Wonyong Sung |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2013 | SD3: An Efficient Dynamic Data-Dependence Profiling MechanismabstractAs multicore processors are deployed in mainstream computing, the need for software tools to help parallelize programs is increasing dramatically. Data-dependence profiling is an important program analysis technique to exploit parallelism in serial programs. More specifically, manual, semiautomatic, or automatic parallelization can use the outcomes of data-dependence profiling to guide where and how to parallelize in a program. However, state-of-the-art data-dependence profiling techniques consume extremely huge resources as they suffer from two major issues when profiling large and long-running applications: 1) runtime overhead and 2) memory overhead. Existing data-dependence profilers are either unable to profile large-scale applications with a typical resource budget or only report very limited information. In this paper, we propose an efficient approach to data-dependence profiling that can address both runtime and memory overhead in a single framework. Our technique, called SD$({}^3)$, reduces the runtime overhead by parallelizing the dependence profiling step itself. To reduce the memory overhead, we compress memory accesses that exhibit stride patterns and compute data dependences directly in a compressed format. We demonstrate that SD$({}^3)$ reduces the runtime overhead when profiling SPEC 2006 by a factor of 4.1× and 9.7× on eight cores and 32 cores, respectively. For the memory overhead, we successfully profile 22 SPEC 2006 benchmarks with the reference input, while the previous approaches fail even with the train input. In some cases, we observe more than a 20× improvement in memory consumption and a 16× speedup in profiling time when 32 cores are used. We also demonstrate the usefulness of SD$({}^3)$ by showing manual parallelization followed by data dependence profiling results. Minjang Kim, Nagesh B. Lakshminarayana, Hyesoon Kim, Chi-Keung Luk |
IEEE Trans. Computers | 2 |
| 2010 | Many-Thread Aware Prefetching Mechanisms for GPGPU ApplicationsabstractWe consider the problem of how to improve memory latency tolerance in massively multithreaded GPGPUs when the thread-level parallelism of an application is not sufficient to hide memory latency. One solution used in conventional CPU systems is prefetching, both in hardware and software. However, we show that straightforwardly applying such mechanisms to GPGPU systems does not deliver the expected performance benefits and can in fact hurt performance when not used judiciously. This paper proposes new hardware and software prefetching mechanisms tailored to GPGPU systems, which we refer to as many-thread aware prefetching (MT-prefetching) mechanisms. Our software MT-prefetching mechanism, called inter-thread prefetching, exploits the existence of common memory access behavior among fine-grained threads. For hardware MT-prefetching, we describe a scalable prefetcher training algorithm along with a hardware-based inter-thread prefetching mechanism. In some cases, blindly applying prefetching degrades performance. To reduce such negative effects, we propose an adaptive prefetch throttling scheme, which permits automatic GPGPU application- and hardware-specific adjustment. We show that adaptation reduces the negative effects of prefetching and can even improve performance. Overall, compared to the state-of-the-art software and hardware prefetching, our MT-prefetching improves performance on average by 16%(software pref.)/15% (hardware pref.) on our benchmarks. Jaekyu Lee, Nagesh B. Lakshminarayana, Hyesoon Kim, Richard W. Vuduc |
MICRO | 2 |
| 2009 | Age based scheduling for asymmetric multiprocessorsabstractAsymmetric (or Heterogeneous) Multiprocessors are becoming popular in the current era of multi-cores due to their power efficiency and potential performance and energy efficiency. However, scheduling of multithreaded applications in Asymmetric Multiprocessors is still a challenging problem. Scheduling algorithms for Asymmetric Multiprocessors must not only be aware of asymmetry in processor performance, but have to consider the characteristics of application threads also. Nagesh B. Lakshminarayana, Jaekyu Lee, Hyesoon Kim |
SC | 1 |
| 2008 | Understanding performance, power and energy behavior in asymmetric multiprocessorsabstractMultiprocessor architectures are becoming popular in both desktop and mobile processors. Among multiprocessor architectures, asymmetric architectures show promise in saving energy and power. However, the performance and energy consumption behavior of asymmetric multiprocessors with desktop-oriented multithreaded applications has not been studied widely. In this study, we measure performance and power consumption in asymmetric and symmetric multiprocessors using real 8 and 16 processor systems to understand the relationships between thread interactions and performance/power behavior. We find that when the workload is asymmetric, using an asymmetric multiprocessor can save energy, but for most of the symmetric workloads, using a symmetric multiprocessor (with the highest clock frequency) consumes less energy. Nagesh B. Lakshminarayana, Hyesoon Kim |
ICCD | 1 |