Souvadra Hati

dblp:392/8148 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0002-5195-3392ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
1 paper
Algorithmic game theory and mechanism design · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Algorithmic game theory and mechanism design
influence maximization
0.812024
Asynchronous Distributed-Memory Parallel Algorithms for Influence Maximization · SC 2024
Parallel and multicore computing › parallelization strategies
asynchronous parallelism
0.212024
Asynchronous Distributed-Memory Parallel Algorithms for Influence Maximization · SC 2024
Parallel and multicore computing › parallel algorithms
distributed-memory parallel algorithms
0.212024
Asynchronous Distributed-Memory Parallel Algorithms for Influence Maximization · SC 2024

Methods — techniques the papers use, named apart from their topics

martingale · 1.5MPI · 1.5
YearPublicationVenuePosition
2025 An Asynchronous Distributed-Memory Parallel Algorithm for $k$-Mer Counting
abstract
This paper describes a new asynchronous algorithm and implementation for the problem of$\boldsymbol{k}$-mer counting$(\text{KC})$, which concerns quantifying the frequency of length$k$substrings in a DNA sequence. This operation is common to many computational biology workloads and can take up to 77 % of the total runtime of de novo genome assembly. The performance and scalability of the current state-of-the-art distributed-memory KC algorithm are hampered by multiple rounds of Many-To-Many collectives. Therefore, we develop an asynchronous algorithm (DAKC) that uses fine-grained, asynchronous messages to obviate most of this global communication while utilizing network bandwidth efficiently via custom message aggregation protocols. DAKC can perform strong scaling up to 256 nodes (512 sockets$/ 6 ~\mathrm{K}$cores) and can count$k$-mers up to$9 \times$faster than the state-of-the-art distributed-memory algorithm, and up to$100 \times$faster than the shared-memory alternative. We also provide an analytical model to understand the hardware resource utilization of our asynchronous KC algorithm and provide insights on the performance.
Souvadra Hati, Akihiro Hayashi, Richard W. Vuduc
IPDPS1
2024 Asynchronous Distributed-Memory Parallel Algorithms for Influence Maximization
abstract
Influence maximization (IM) is the problem of finding the k most influential nodes in a graph. We propose distributed-memory parallel algorithms for the two main kernels of a state-of-the-art implementation of one IM algorithm, influence maximization via martingales (IMM). The baseline relies on a bulk-synchronous parallel approach and uses replication to reduce communication and achieve approximate load balance, at the cost of synchronization and high memory requirements. By contrast, our method fully distributes the data, thereby improving memory scalability, and uses fine-grained asynchronous parallelism to improve network utilization and the cost of doing more communication. We show our design and implementation can achieve up to $29.6 \times$ speedup over the MPI-based state-of-the-art on synthetic and real-world network graphs. Moreover, ours is the first implementation that can run IMM to find influencers in the ‘twitter’ graph (41M nodes and 1.4B edges) in 200 seconds using 8 K CPU cores of NERSC Perlmutter supercomputer.
Shubhendra Pal Singhal, Souvadra Hati, Jeffrey Young 0001, Vivek Sarkar, Akihiro Hayashi, Richard W. Vuduc
SC2