VLDB 2026 Research / reviewers in the wild / expert
Sudhanshu Shukla
dblp:173/1935
· DBLP profile ↗
5ranked-venue papers
4as first author
1since 2021 · last 2022
0000-0002-5069-9708ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 54% Processor architecture and microarchitecture · 46% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › out-of-order execution
out-of-order processor |
0.6 | 1 | 2022 | Register file prefetching · ISCA 2022 |
Memory systems
cache coherence |
0.3 | 1 | 2017 | Tiny Directory: Efficient Shared Memory in Many-Core Systems with Ultra-Low-Overhead Coherence Tracking · HPCA 2017 |
Processor architecture and microarchitecture › many-core architecture
many-core chip multiprocessor |
0.3 | 1 | 2017 | Tiny Directory: Efficient Shared Memory in Many-Core Systems with Ultra-Low-Overhead Coherence Tracking · HPCA 2017 |
Memory systems › cache coherence › directory-based coherence
sparse directory |
0.3 | 1 | 2017 | Tiny Directory: Efficient Shared Memory in Many-Core Systems with Ultra-Low-Overhead Coherence Tracking · HPCA 2017 |
Bioinformatics and computational biology › transcriptomics
RNA-seq analysis |
0.2 | 1 | 2016 | Two-pass alignment improves novel splice junction quantification · Bioinform. 2016 |
Bioinformatics and computational biology › transcriptomics › RNA-seq analysis
splice junction quantification |
0.2 | 1 | 2016 | Two-pass alignment improves novel splice junction quantification · Bioinform. 2016 |
Bioinformatics and computational biology
transcriptomics |
0.2 | 1 | 2016 | Two-pass alignment improves novel splice junction quantification · Bioinform. 2016 |
Memory systems › memory access latency
cache access latency |
0.2 | 1 | 2022 | Register file prefetching · ISCA 2022 |
Memory systems
memory wall |
0.2 | 1 | 2022 | Register file prefetching · ISCA 2022 |
Memory systems › memory hierarchy › cache hierarchy
last-level cache |
0.1 | 1 | 2017 | Tiny Directory: Efficient Shared Memory in Many-Core Systems with Ultra-Low-Overhead Coherence Tracking · HPCA 2017 |
Methods — techniques the papers use, named apart from their topics
prefetching · 0.6simulation-based study · 0.3selective spilling · 0.3two-pass alignment · 0.2STAR aligner · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Register file prefetchingabstractThe memory wall continues to limit the performance of modern out-of-order (OOO) processors, despite the expensive provisioning of large multi-level caches and advancements in memory prefetching. In this paper, we put forth an important observation that the memory wall is not monolithic, but is constituted of many latency walls arising due to the latency of each tier of cache/memory. Our results show that even though level-1 (L1) data cache latency is nearly 40X lower than main memory latency, mitigating this latency offers a very similar performance opportunity as the more widely studied, main memory latency. Sudhanshu Shukla, Sumeet Bandishte, Jayesh Gaur, Sreenivas Subramoney |
ISCA | 1 |
| 2017 | Tiny Directory: Efficient Shared Memory in Many-Core Systems with Ultra-Low-Overhead Coherence TrackingabstractThe sparse directory has emerged as a critical component for supporting the shared memory abstraction in multiand many-core chip-multiprocessors. Recent research efforts have explored ways to reduce the number of entries in the sparse directory. These include tracking coherence of private regions at a coarse grain, not tracking blocks that belong to pages identified as private by the operating system (OS), and not tracking a subset of blocks that are speculated to be private by the hardware. These techniques require support for multi-grain coherence, assistance of OS, or broadcast-based recovery on sharing an untracked block that is wrongly speculated as private. In this paper, we design a robust minimally-sized sparse directory that can offer adequate performance while enjoying the simplicity, scalability, and OS-independence of traditional broadcast-free block-grain coherence. We begin our exploration with a naïve design that does not have a sparse directory and the location/sharers of a block are tracked by borrowing a portion of the block's lastlevel cache (LLC) data way. Such a design, however, lengthens the critical path from two transactions to three transactions (two hops to three hops) for the blocks that experience frequent shared read accesses. We address this problem by architecting a tiny sparse directory that dynamically identifies and tracks a selected subset of the blocks that experience a large volume of shared accesses. We augment the tiny directory proposal with an option of selectively spilling into the LLC space for tracking the coherence of the critical shared blocks that the tiny directory fails to accommodate. Detailed simulation-based study on a 128-core system with a large set of multi-threaded applications spanning scientific, general-purpose, and commercial computing shows that our coherence tracking proposal operating with1/32 × to 1/256 × sparse directories offers performance within a percentage of a traditional 2× sparse directory. Sudhanshu Shukla, Mainak Chaudhuri |
HPCA | 1 |
| 2017 | Sharing-Aware Efficient Private Caching in Many-Core Server ProcessorsabstractThe general-purpose cache-coherent many-core server processors are usually designed with a per-core private cache hierarchy and a large shared multi-banked last-level cache (LLC). The round-trip latency and the volume of traffic through the on-die interconnect between the per-core private cache hierarchy and the shared LLC banks can be significantly large. As a result, optimized private caching is important in such architectures. Traditionally, the private cache hierarchy in these processors treats the private and the shared blocks equally. We, however, observe that elimination of all non-compulsory non-coherence core cache misses to a small subset of shared code and data blocks can save a large fraction of the core requests to the LLC indicating large potential for reducing the interconnect traffic in such architectures. We architect a specialized exclusive per-core private L2 cache which serves as a victim cache for the per-core private L1 cache. The proposed victim cache selectively captures a subset of the L1 cache victims. Our best selective victim caching proposal is driven by an online partitioning of the L1 cache victims based on two distinct features, namely, an estimate of sharing degree and an indirect simple estimate of reuse distance. Our proposal learns the collective reuse probability of the blocks in each partition on-the-fly and decides the victim caching candidates based on these probability estimates. Detailed simulation results on a 128-core system running a selected set of multi-threaded commercial and scientific computing applications show that our best victim cache design proposal at 64 KB capacity, on average, saves 44.1% core cache miss requests sent to the LLC and 10.6% execution cycles compared to a baseline system that has no private L2 cache. In contrast, a traditional 128 KB non-inclusive LRU L2 cache saves 42.2% core cache misses sent to the LLC compared to the same baseline while performing slightly worse than the proposed 64 KB victim cache. In summary, our proposal outperforms the traditional design and enjoys lower interconnect traffic while halving the space investment for the per-core private L2 cache. Further, the savings in core cache misses achieved due to introduction of the proposed victim cache are observed to be only 8% less than an optimal victim cache design at 32 KB and 64 KB capacity points. Sudhanshu Shukla, Mainak Chaudhuri |
ICCD | 1 |
| 2016 | Two-pass alignment improves novel splice junction quantificationabstractMOTIVATION: Discovery of novel splicing from RNA sequence data remains a critical and exciting focus of transcriptomics, but reduced alignment power impedes expression quantification of novel splice junctions. RESULTS: Here, we profile performance characteristics of two-pass alignment, which separates splice junction discovery from quantification. Per sample, across a variety of transcriptome sequencing datasets, two-pass alignment improved quantification of at least 94% of simulated novel splice junctions, and provided as much as 1.7-fold deeper median read depth over those splice junctions. We further demonstrate that two-pass alignment works by increasing alignment of reads to splice junctions by short lengths, and that potential alignment errors are readily identifiable by simple classification. Taken together, two-pass alignment promises to advance quantification and discovery of novel splicing events. CONTACT: [email protected], [email protected] AVAILABILITY AND IMPLEMENTATION: Two-pass alignment was implemented here as sequential alignment, genome indexing, and re-alignment steps with STAR. Full parameters are provided in Supplementary Table 2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Brendan A. Veeneman, Sudhanshu Shukla, Saravana M. Dhanasekaran, Arul M. Chinnaiyan, Alexey I. Nesvizhskii |
Bioinform. | 2 |
| 2015 | Pool directory: Efficient coherence tracking with dynamic directory allocation in many-core systemsabstractThe coherence directory in a chip-multiprocessor keeps track of each memory block inside the cache hierarchy and plays a significant role in offering a scalable shared memory abstraction in many-core systems. Multi-threaded applications typically require two types of directory entries, namely, limited pointer entries tracking a few sharers of a block and bitvector entries tracking larger number of sharers for widely shared blocks. Recent proposals aiming to optimize the average number of bits per directory entry have organized the directory as either a static mix of these two types of entries or a collection of relatively short bitvector entries that can encode either a limited number of sharer pointers or a larger number of sharers hierarchically. In this paper, we present a directory organization that facilitates allocation of two different types of directory entries dynamically. Our design maintains a pool of limited pointer entries, where each entry can also double as a segment directory entry encoding the sharers in a cluster of cores. Each tag in the primary sparse directory array has a pointer that can either represent a sharer or point to an entry in the pool. When multiple segment directory entries are needed to encode all the sharers of a block, our pool management protocol guarantees that all these entries are allocated contiguously so that maintaining a pointer to the head entry is enough. Such a design offers significant flexibility in sharer encoding and allows us to independently size the sparse directory array and the pool. Detailed simulation results show that our proposal incorporated in a 128-core system running multi-threaded applications drawn from scientific, general-purpose, and commercial computing domains can offer, on average, 5% improvement in performance and 20% savings in interconnect traffic compared to the state-of-the-art scalable coherence directory (SCD) proposal when using a 1/16 × sparse directory. Sudhanshu Shukla, Mainak Chaudhuri |
ICCD | 1 |