Bhanu Shankar

dblp:53/1679 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 80% Performance modeling and evaluation · 20%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
0.011993
An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993
Memory systems › cache
cache miss
0.011993
An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993
Memory systems
memory referencing behavior
0.011993
An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993
Memory systems › cache
prefetching
0.011993
An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993
Performance modeling and evaluation
workload characterization
0.011993
An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993

Methods — techniques the papers use, named apart from their topics

profiling · 0.0cache simulation · 0.0
YearPublicationVenuePosition
2025 Performance Characterization of CXL Memory and Its Use Cases
abstract
Compute eXpress Link (CXL) is emerging as a promising memory interface technology. However, its performance characteristics remain largely unclear due to the limited availability of production hardware. Key questions include: What are the use cases for the CXL memory? What are the impacts of the CXL memory on application performance? How to use the CXL memory in combination with existing memory components? In this work, we study the performance of three genuine CXL memory-expansion cards from different vendors. We characterize the basic performance of the CXL memory, study how HPC applications and large language models (LLM) can benefit from the CXL memory, and study the interplay between memory tiering and page interleaving. We also propose a novel data object-level interleaving policy to match the interleaving policy with memory access patterns. Our findings reveal the challenges and opportunities of using the CXL memory.
Xi Wang 0027, Jie Liu 0096, Shuangyan Yang, Jie Ren 0015, Bhanu Shankar, Dong Li 0001
IPDPS6
2001 Resource Management in Dataflow-Based Multithreaded Execution
Lucas Roh, Bhanu Shankar, A. P. Wim Böhm, Walid A. Najjar
J. Parallel Distributed Comput.2
1996 Generation, Optimization, and Evaluation of Multithreaded Code
Lucas Roh, Walid A. Najjar, Bhanu Shankar, A. P. Wim Böhm
J. Parallel Distributed Comput.3
1995 Control of loop parallelism in multithreaded code
Bhanu Shankar, Lucas Roh, A. P. Wim Böhm, Walid A. Najjar
PACT1
1993 An evaluation of bottom-up and top-down thread generation techniques
abstract
Due to increasing cache-miss latencies, cache control instructions are being implemented for future systems. The authors study the memory referencing behavior of individual machine-level instructions using simulations of fully-associative caches under MIN replacement. Their objective is to obtain a deeper understanding of useful program behavior that can be eventually employed at optimizing programs and to motivate architectural features aimed at improving the efficacy of memory hierarchies. The simulation results show that a very small number of load/store instructions account for a majority of data cache misses. Specifically, fewer than 10 instructions account for half the misses for six out of nine SPEC89 benchmarks. Selectively prefetching data referenced by a small number of instructions identified through profiling can reduce overall miss ratio significantly while only incurring a small number of unnecessary prefetches.>
A. P. Wim Böhm, Walid A. Najjar, Bhanu Shankar, Lucas Roh
MICRO3