EDBT 2026 Demo / reviewers in the wild / expert
Shubhendra Pal Singhal
dblp:243/3014
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0002-0610-7672ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Distributed systems · 43% Memory systems · 28% High-performance computing · 14% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › memory interference
cache contention |
1.0 | 1 | 2026 | Performance Analysis of Conveyors: Memory Dominates? · HPDC 2026 |
Memory systems › memory interference
memory contention |
1.0 | 1 | 2026 | Performance Analysis of Conveyors: Memory Dominates? · HPDC 2026 |
High-performance computing
message aggregation |
1.0 | 1 | 2026 | Performance Analysis of Conveyors: Memory Dominates? · HPDC 2026 |
Distributed systems › fault tolerance
byzantine fault tolerance |
0.8 | 1 | 2024 | Rashnu: Data-Dependent Order-Fairness · Proc. VLDB Endow. 2024 |
Distributed systems
consensus |
0.8 | 1 | 2024 | Rashnu: Data-Dependent Order-Fairness · Proc. VLDB Endow. 2024 |
Distributed systems › consensus › byzantine agreement
order-fairness |
0.8 | 1 | 2024 | Rashnu: Data-Dependent Order-Fairness · Proc. VLDB Endow. 2024 |
Distributed systems › replication
state machine replication |
0.8 | 1 | 2024 | Rashnu: Data-Dependent Order-Fairness · Proc. VLDB Endow. 2024 |
Algorithmic game theory and mechanism design
influence maximization |
0.8 | 1 | 2024 | Asynchronous Distributed-Memory Parallel Algorithms for Influence Maximization · SC 2024 |
Interconnection networks and networks-on-chip
network bandwidth |
0.3 | 1 | 2026 | Performance Analysis of Conveyors: Memory Dominates? · HPDC 2026 |
Interconnection networks and networks-on-chip › high-speed networks
supercomputer interconnect |
0.3 | 1 | 2026 | Performance Analysis of Conveyors: Memory Dominates? · HPDC 2026 |
Transaction processing and concurrency control › transaction scheduling
transaction ordering |
0.2 | 1 | 2024 | Rashnu: Data-Dependent Order-Fairness · Proc. VLDB Endow. 2024 |
Parallel and multicore computing › parallelization strategies
asynchronous parallelism |
0.2 | 1 | 2024 | Asynchronous Distributed-Memory Parallel Algorithms for Influence Maximization · SC 2024 |
Parallel and multicore computing › parallel algorithms
distributed-memory parallel algorithms |
0.2 | 1 | 2024 | Asynchronous Distributed-Memory Parallel Algorithms for Influence Maximization · SC 2024 |
Methods — techniques the papers use, named apart from their topics
parallel execution · 1.5martingale · 1.5dependency graph · 1.5MPI · 1.5profiling · 1.0performance measurement · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Performance Analysis of Conveyors: Memory Dominates?abstractSmall-message aggregation is critical for scaling irregular, communication intensive applications in high-performance computing. In this paper, contrary to conventional wisdom, we present the first systematic study showing that memory contention, not network bandwidth, is the dominant bottleneck in message aggregation runtimes. Using the state-of-the-art conveyors library as our reference implementation, we conducted extensive experiments on HPC systems featuring Slingshot 11 and InfiniBand interconnects, scaling to 16k cores (256 nodes) and processing 10s–100s GB of data. Our measurements reveal that interference between user data and aggregation buffers drives LLC miss rates to 77%, inflating memory costs by 2–3× over the algorithmic baseline. Consequently, we advocate for dedicated near-memory subsystems to improve the scalability and performance of message aggregation runtimes. This paper also demonstrates up to an order of magnitude higher latency for conveyor termination compared to a traditional HPC barrier, and it examines the impact of communication context isolation and the critical challenge of programmability. Shubhendra Pal Singhal, Aaron Welch, Oscar R. Hernandez, Stephen W. Poole, Akihiro Hayashi, Vivek Sarkar |
HPDC | 1 |
| 2024 | Bottleneck Scenarios in Use of the Conveyors Message Aggregation LibraryabstractOn massively parallel computing systems, applications involving many small message transfers suffer from degradation in performance and scalability. Typically, applications built on OpenSHMEM, Chapel, and Unified Parallel C (UPC) suffer due to poor network message rates, causing under-utilization of network bandwidth. In past work, a memory-efficient aggregation library called Conveyors was developed to address this issue in the PGAS (Partitioned Global Address Space) based SPMD (Single Program Multiple Data) model offering flexible APIs and contracts. The Conveyors library has been used by multiple HPC frameworks, including Chapel, HABU, and HClib. We scope this paper as an investigative study of Conveyors at a time where none of the profilers, such as score-p, Intel Vtune, and CrayPat, measure asynchronous OpenSHMEM calls. Specifically, we conduct a series of experiments and identify cases where certain one-to-one/all-to-all communication patterns can lead to bottlenecks due to specific design choices made by Conveyors. Shubhendra Pal Singhal, Akihiro Hayashi, Vivek Sarkar |
ISPASS | 1 |
| 2024 | Asynchronous Distributed-Memory Parallel Algorithms for Influence MaximizationabstractInfluence maximization (IM) is the problem of finding the k most influential nodes in a graph. We propose distributed-memory parallel algorithms for the two main kernels of a state-of-the-art implementation of one IM algorithm, influence maximization via martingales (IMM). The baseline relies on a bulk-synchronous parallel approach and uses replication to reduce communication and achieve approximate load balance, at the cost of synchronization and high memory requirements. By contrast, our method fully distributes the data, thereby improving memory scalability, and uses fine-grained asynchronous parallelism to improve network utilization and the cost of doing more communication. We show our design and implementation can achieve up to $29.6 \times$ speedup over the MPI-based state-of-the-art on synthetic and real-world network graphs. Moreover, ours is the first implementation that can run IMM to find influencers in the ‘twitter’ graph (41M nodes and 1.4B edges) in 200 seconds using 8 K CPU cores of NERSC Perlmutter supercomputer. Shubhendra Pal Singhal, Souvadra Hati, Jeffrey Young 0001, Vivek Sarkar, Akihiro Hayashi, Richard W. Vuduc |
SC | 1 |
| 2024 | Rashnu: Data-Dependent Order-FairnessabstractDistributed data management systems use state Machine Replication (SMR) to provide fault tolerance. The SMR algorithm enables Byzantine Fault-Tolerant (BFT) protocols to guarantee safety and liveness despite the malicious failure of nodes. However, SMR does not prevent the adversarial manipulation of the order of transactions, where the order assigned by a malicious leader differs from the order in that transactions are received from clients. While order-fairness has been recently studied in a few protocols, such protocols rely on synchronized clocks, suffer from liveness issues, or incur significant performance overhead. This paper presents Rashnu , a high-performance fair ordering protocol. Rashnu is motivated by the fact that fair ordering among two transactions is needed only when both transactions access a shared resource. Based on this observation, we define the notion of data-dependent order fairness where replicas capture only the order of data-dependent transactions and the leader uses these orders to propose a dependency graph that represents fair ordering among transactions. Replicas then execute transactions using the dependency graph, resulting in the parallel execution of independent transactions. We implemented a prototype of Rashnu where our experimental evaluation reveals the low overhead of providing order-fairness in Rashnu. Heena Nagda, Shubhendra Pal Singhal, Mohammad Javad Amiri, Boon Thau Loo |
Proc. VLDB Endow. | 2 |