EDBT 2026 Demo / reviewers in the wild / expert
Bapi Chatterjee
dblp:117/6988
· DBLP profile ↗
18ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-2742-4028ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Security and privacy · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Contemporary data-driven innovations in peptide-based therapeutic designabstractIn recent years, peptides have grabbed significant attention across many fields, including pharmaceutical, biomedical, and biotechnological industries, owing to their notable biological activity, low toxicity, and high specificity. Naturally occurring peptides play crucial roles in handling different biological processes (cellular signaling, immune responses, and enzymatic functions, etc.), while laboratory-made synthetic peptides can be adopted for many applied industrial applications. Following the practical problems in large-scale peptide synthesis in the laboratory which demand high cost and time, recent developments have benefited from the best use of machine learning (ML), deep learning, active learning, reinforcement learning (RL), generative artificial intelligence (AI), and large language models (LLMs) to reduce the number of experiments. ML algorithms enable the prediction of the peptide structure-activity relationship, bioavailability, and other drug-like properties with high accuracy. The integration of AI with peptide-based therapeutics design marks a paradigm shift in drug discovery which was otherwise dominated by small organic molecules. This review comprehensively examines AI-driven methodologies, including classical ML approaches, deep generative models, RL, and LLMs, that overcome historical limitations in peptide design, such as structural flexibility, enzymatic degradation, and membrane impermeability. Recent advances in structure-aware algorithms and sequence-based frameworks have accelerated peptide-based therapeutic development across oncology, metabolic disorders, and infectious diseases. Despite challenges in data scarcity and validation gaps, the convergence of computational prediction with experimental automation promises clinical translation of AI-designed peptides in the near future. This review highlights the transformative potential of AI in ushering a new era of precision peptide therapeutics. Lipsa Priyadarsinee, Vyacheslav Kungurtsev, Vibhor Kumar, Bapi Chatterjee, G. Narahari Sastry, Natarajan Arul Murugan |
Briefings Bioinform. | 4 |
| 2025 | Improved Coresets for Vertical Federated Learning: Regularized Linear and Logistic RegressionsabstractCoreset, as a summary of training data, offers an efficient approach for reducing data processing and storage complexity during training. In the emerging vertical federated learning (VFL) setting, where scattered clients store different data features, it directly reduces communication complexity. In this work, we introduce coresets construction for regularized logistic regression both in centralized and VFL settings. Additionally, we improve the coreset size for regularized linear regression in the VFL setting. We also eliminate the dependency of the coreset size on a property of the data due to the VFL setting. The improvement in the coreset sizes is due to our novel coreset construction algorithms that capture the reduced model complexity due to the added regularization and its subsequent analysis. In experiments, we provide extensive empirical evaluation that backs our theoretical claims. We also report the performance of our coresets by comparing the models trained on the complete data and on the coreset. Supratim Shit, Gurmehak Kaur Chadha, Bapi Chatterjee |
ICML | 4 |
| 2024 | Federated SGD with Local AsynchronyabstractParallel SGD in a shared-memory setting is oft-represented by the popular Hogwild! algorithm, in which lock-free updates are asynchronously performed by multiple computing processes. Unfortunately, scaling Hogwild! to distributed workers is largely unexplored. Specifically, it is unknown if any adaptation of Hogwild! to the popular decentralized multi-GPU setting offers any competitive speedup, either empirically or theoretically. In this work, we investigate the potential of decentralizing Hogwild! by incorporating simultaneously (a) asynchronous local gradient updates on the shared memory of GPUs, and (b) non-blocking asynchronous decentralized federated averaging. A naive direct implementation shows degradation in performance, arising from scheduling overheads and concurrent write conflicts on GPUs. To mitigate these drawbacks, we investigate and propose a new method, based on careful block selection rules, which update only portions of the parameter vectors. Our experiments show that the resulting decentralized training method exhibits improved throughput and competitive accuracy for standard image classification benchmarks on the CIFAR-10, CIFAR-100, and Imagenet datasets. On the theoretical side, we prove that our method guarantees sublinear ergodic convergence rates for non-convex objectives. Bapi Chatterjee, Vyacheslav Kungurtsev, Dan Alistarh |
ICDCS | 1 |
| 2024 | Kanva: A Lock-free Learned Search Data StructureabstractLock-free concurrent data structures offer throughput with scalability and guarantees for the completion of operations on multicore computers. Recently, queries using machine learning models trained to predict the data distribution have gained remarkable attention. Yet, to our knowledge, no existing lock-free data structure employs them. This paper introduces a lock-free search structure that supports concurrent updates, membership, and range queries accelerated by a shallow hierarchy of lightweight machine learning models. The proposed approach significantly outperforms the current state-of-the-art lock-free data structures in many workload and data distribution settings. Ours is the first provably linearizable learned lock-free concurrent range search index. Gaurav Bhardwaj, Bapi Chatterjee, Sathya Peri, Siddharth Nayak |
ICPP | 2 |
| 2024 | Brief Announcement: Lock-free Learned Search Data StructureabstractThis paper introduces a lock-free linearizable search structure that supports concurrent updates, membership, and range queries accelerated by a shallow hierarchy of lightweight machine learning (ML) models. The proposed approach significantly outperforms the current state-of-the-art lock-free data structures in many workload and data distribution settings. Gaurav Bhardwaj, Bapi Chatterjee, Sathya Peri, Siddharth Nayak |
SPAA | 2 |
| 2023 | Wait-Free Updates and Range Search Using Uruv
Gaurav Bhardwaj, Bapi Chatterjee, Abhay Jain, Sathya Peri |
SSS | 2 |
| 2021 | Asynchronous Optimization Methods for Efficient Training of Deep Neural Networks with GuaranteesabstractAsynchronous distributed algorithms are a popular way to reduce synchronization costs in large-scale optimization, and in particular for neural network training. However, for nonsmooth and nonconvex objectives, few convergence guarantees exist beyond cases where closed-form proximal operator solutions are available. As training most popular deep neural networks corresponds to optimizing nonsmooth and nonconvex objectives, there is a pressing need for such convergence guarantees. In this paper, we analyze for the first time the convergence of stochastic asynchronous optimization for this general class of objectives. In particular, we focus on stochastic subgradient methods allowing for block variable partitioning, where the shared model is asynchronously updated by concurrent processes. To this end, we use a probabilistic model which captures key features of real asynchronous scheduling between concurrent processes. Under this model, we establish convergence with probability one to an invariant set for stochastic subgradient methods with momentum. From a practical perspective, one issue with the family of algorithms that we consider is that they are not efficiently supported by machine learning frameworks, which mostly focus on distributed data-parallel strategies. To address this, we propose a new implementation strategy for shared-memory based training of deep neural networks for a partitioned but shared model in single- and multi-GPU settings. Based on this implementation, we achieve on average1.2x speed-up in comparison to state-of-the-art training methods for popular image classification tasks, without compromising accuracy. Vyacheslav Kungurtsev, Malcolm Egan, Bapi Chatterjee, Dan Alistarh |
AAAI | 3 |
| 2021 | Elastic Consistency: A Practical Consistency Model for Distributed Stochastic Gradient DescentabstractOne key element behind the recent progress of machine learning has been the ability to train machine learning models in large-scale distributed shared-memory and message-passing environments. Most of these models are trained employing variants of stochastic gradient descent (SGD) based optimization, but most methods involve some type of consistency relaxation relative to sequential SGD, to mitigate its large communication or synchronization costs at scale. In this paper, we introduce a general consistency condition covering communication-reduced and asynchronous distributed SGD implementations. Our framework, called elastic consistency, decouples the system-specific aspects of the implementation from the SGD convergence requirements, giving a general way to obtain convergence bounds for a wide variety of distributed SGD methods used in practice. Elastic consistency can be used to re-derive or improve several previous convergence bounds in message-passing and shared-memory settings, but also to analyze new models and distribution schemes. As a direct application, we propose and analyze a new synchronization-avoiding scheduling scheme for distributed SGD, and show that it can be used to efficiently train deep convolutional models for image classification. Giorgi Nadiradze, Ilia Markov, Bapi Chatterjee, Vyacheslav Kungurtsev, Dan Alistarh |
AAAI | 3 |
| 2021 | Non-Blocking Dynamic Unbounded Graphs with Worst-Case Amortized Bounds
Bapi Chatterjee, Sathya Peri, Muktikanta Sa, Komma Manogna |
OPODIS | 1 |
| 2021 | Brief Announcement: Non-Blocking Dynamic Unbounded Graphs with Worst-Case Amortized BoundsabstractThis paper reports a new concurrent graph data structure that supports updates of both edges and vertices and queries: Breadth-first search, Single-source shortest-path, and Betweenness centrality. The operations are provably linearizable and non-blocking. Bapi Chatterjee, Sathya Peri, Muktikanta Sa |
DISC | 1 |
| 2021 | Concurrent linearizable nearest neighbour search in LockFree-kD-tree
Bapi Chatterjee, Ivan Walulya, Philippas Tsigas |
Theor. Comput. Sci. | 1 |
| 2018 | Efficiently Processing Temporal Queries on Hyperledger FabricabstractIn this paper, we discuss the problem of efficiently handling temporal queries on Hyperledger Fabric, a popular implementation of Blockchain technology. The temporal nature of the data inserted by the Hyperledger Fabric transactions can be leveraged to support various use-cases. This requires that the temporal queries be processed efficiently on this data. Currently this presents significant challenges as this data is organized on file-system, is exposed to users via a limited API and does not support any temporal indexes. We present two models for overcoming these limitations and improving the performance of temporal queries on Fabric. The first model creates a copy of each event inserted by a Fabric transaction and stores temporally close events together on Fabric. The second model keeps the event count intact but tags some metadata to each event being inserted on Fabric s.t. temporally close events share the same metadata. We discuss these two models in detail and show that these two models significantly outperform the naive ways of handling temporal queries on Fabric. We also discuss the performance trade-offs for these two models across various dimensions - data storage, query performance, data ingestion time etc. Sandeep Hans, Kushagra Aggarwal, Sameep Mehta, Bapi Chatterjee, Praveen Jayachandran |
ICDE | 5 |
| 2018 | Concurrent Lock-Free Unbounded Priority Queue with Mutable Priorities
Ivan Walulya, Bapi Chatterjee, Ajoy K. Datta, Rashmi Niyolia, Philippas Tsigas |
SSS | 2 |
| 2017 | Provenance in Context of Hadoop as a Service (HaaS) - State of the Art and Research DirectionsabstractHadoop as a service (HaaS), also known as Hadoop in the cloud, is a big data analytics framework that stores and analyzes data in the cloud using Hadoop/Spark. In this paper, we discuss the importance of providing provenance capabilities in context of Hadoop as a service (HaaS) framework. We first review the state of the art in provenance tracking in context of databases and work-flow processing, in context of cloud and in context of big data analytics frameworks like Hadoop and Spark. We next identify a number of provenance capabilities which have been developed in context of databases and workflow processing but the corresponding solutions have not been developed in context of Hadoop or Spark. We argue that developing these solutions is important so that a comprehensive provenance aware Hadoop as a Service (HaaS) can be provided on cloud. The paper ends by identifying some research challenges in developing these provenance capabilities. Sameep Mehta, Sandeep Hans, Bapi Chatterjee, Pranay Lohia, Rajmohan C |
CLOUD | 4 |
| 2016 | Help-Optimal and Language-Portable Lock-Free Concurrent Data StructuresabstractHelping is a widely used technique to guarantee lock-freedom in many concurrent data structures. An optimized helping strategy improves the overall performance of a lock-free algorithm. In this paper, we propose help-optimality, which essentially implies that no operation step is accounted for exclusive helping in the lock-free synchronization of concurrent operations. To describe the concept, we revisit the designs of a lock-free linked-list and a lock-free binary search tree and present improved algorithms. Our algorithms employ atomic single-word compare-and-swap (CAS) primitives and are linearizable. We design the algorithms without using any language/platformspecific mechanism. Specifically, we use neither bit-stealing froma pointer nor runtime type introspection of objects. Thus, our algorithms are language-portable. Further, to optimize the amortized number of steps per operation, if a CAS execution tomodify a shared pointer fails, we obtain a fresh set of thread-local variables without restarting an operation from scratch. We use several micro-benchmarks in both C/C++ and Java to validate the efficiency of our algorithms against existing state-of-the-art. The experiments show that the algorithms are scalable. Our implementations perform on a par with highly optimizedones and in many cases yield 10%-50% higher throughput. Bapi Chatterjee, Ivan Walulya, Philippas Tsigas |
ICPP | 1 |
| 2014 | Efficient lock-free binary search treesabstractIn this paper we present a novel algorithm for concurrent lock-free internal binary search trees (BST) and implement a Set abstract data type (ADT) based on that. We show that in the presented lock-free BST algorithm the amortized step complexity of each set operation - Add, Remove and Contains - is O(H(n) + c), where H(n) is the height of the BST with n number of nodes and c is the contention during the execution. Our algorithm adapts to contention measures according to read-write load. If the situation is read-heavy, the operations avoid helping the concurrent Remove operations during traversal, and adapt to interval contention. However, for the write-heavy situations we let an operation help a concurrent Remove, even though it is not obstructed. In that case, an operation adapts to point contention. It uses single-word compare-and-swap (CAS) operations. We show that our algorithm has improved disjoint-access-parallelism compared to similar existing algorithms. We prove that the presented algorithm is linearizable. To the best of our knowledge, this is the first algorithm for any concurrent tree data-structure in which the modify operations are performed with an additive term of contention measure. Bapi Chatterjee, Nhan Nguyen Dang, Philippas Tsigas |
PODC | 1 |
| 2013 | A Study of the Behavior of Synchronization Methods in Commonly Used Languages and SystemsabstractSynchronization is a central issue in concurrency and plays an important role in the behavior and performance of modern programmes. Programming languages and hardware designers are trying to provide synchronization constructs and primitives that can handle concurrency and synchronization issues efficiently. Programmers have to find a way to select the most appropriate constructs and primitives in order to gain the desired behavior and performance under concurrency. Several parameters and factors affect the choice, through complex interactions among (i) the language and the language constructs that it supports, (ii) the system architecture, (iii) possible run-time environments, virtual machine options and memory management support and (iv) applications. We present a systematic study of synchronization strategies, focusing on concurrent data structures. We have chosen concurrent data structures with different number of contention spots. We consider both coarse-grain and fine-grain locking strategies, as well as lock-free methods. We have investigated synchronization-aware implementations in C++, C# (.NET and Mono) and Java. Considering the machine architectures, we have studied the behavior of the implementations on both Intel's Nehalem and AMD's Bulldozer. The properties that we study are throughput and fairness under different workloads and multiprogramming execution environments. For NUMA architectures fairness is becoming as important as the typically considered throughput property. To the best of our knowledge this is the first systematic and comprehensive study of synchronization-aware implementations. This paper takes steps towards capturing a number of guiding principles and concerns for the selection of the programming environment and synchronization methods in connection to the application and the system characteristics. Daniel Cederman, Bapi Chatterjee, Nhan Nguyen Dang, Yiannis Nikolakopoulos, Marina Papatriantafilou, Philippas Tsigas |
IPDPS | 2 |
| 2012 | Understanding the Performance of Concurrent Data Structures on Graphics Processors
Daniel Cederman, Bapi Chatterjee, Philippas Tsigas |
Euro-Par | 2 |