VLDB 2026 Research / reviewers in the wild / expert
Abhirup Ghosh
dblp:143/7545
· DBLP profile ↗
11ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0003-4044-8523ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BoTTA: Benchmarking On-device Test Time Adaptation
Michal Danilowski, Soumyajit Chatterjee, Abhirup Ghosh |
SenSys | 3 |
| 2025 | E-BATS: Efficient Backpropagation-Free Test-Time Adaptation for Speech Foundation ModelsabstractSpeech Foundation Models encounter significant performance degradation when deployed in real-world scenarios involving acoustic domain shifts, such as background noise and speaker accents. Test-time adaptation (TTA) has recently emerged as a viable strategy to address such domain shifts at inference time without requiring access to source data or labels. However, existing TTA approaches, particularly those relying on backpropagation, are memory-intensive, limiting their applicability in speech tasks and resource-constrained settings. Although backpropagation-free methods offer improved efficiency, existing ones exhibit poor accuracy. This is because they are predominantly developed for vision tasks, which fundamentally differ from speech task formulations, noise characteristics, and model architecture, posing unique transferability challenges.
In this paper, we introduce E-BAT, first Efficient BAckpropagation-free TTA framework designed explicitly for speech foundation models. E-BAT achieves a balance between adaptation effectiveness and memory efficiency through three key components: (i) lightweight prompt adaptation for a forward-pass-based feature alignment, (ii) a multi-scale loss to capture both global (utterance-level) and local distribution shifts (token-level) and (iii) a test-time exponential moving average mechanism for stable adaptation across utterances. Experiments conducted on four noisy speech datasets spanning sixteen acoustic conditions demonstrate consistent improvements, with 4.1\%--13.5% accuracy gains over backpropogation-free baselines and 2.0$\times$–6.4$\times$ GPU memory savings compared to backpropogation-based methods. By enabling scalable and robust adaptation under acoustic variability, this work paves the way for developing more efficient adaptation approaches for practical speech processing systems in real-world environments. Jiaheng Dong, Hong Jia, Soumyajit Chatterjee, Abhirup Ghosh, Ting Dang |
NeurIPS | 4 |
| 2024 | In-Network Approximate and Efficient Spatiotemporal Range Queries on Moving Objects
Guang Yang 0044, Abhirup Ghosh, Thomas Heinis |
EDBT | 2 |
| 2024 | FLea: Addressing Data Scarcity and Label Skew in Federated Learning via Privacy-preserving Feature AugmentationabstractFederated Learning (FL) enables model development by leveraging data distributed across numerous edge devices without transferring local data to a central server. However, existing FL methods still face challenges when dealing with scarce and label-skewed data across devices, resulting in local model overfitting and drift, consequently hindering the performance of the global model. In response to these challenges, we propose a pioneering framework called FLea, incorporating the following key components: i) A global feature buffer that stores activation-target pairs shared from multiple clients to support local training. This design mitigates local model drift caused by the absence of certain classes; ii) A feature augmentation approach based on local and global activation mix-ups for local training. This strategy enlarges the training samples, thereby reducing the risk of local overfitting; iii) An obfuscation method to minimize the correlation between intermediate activations and the source data, enhancing the privacy of shared features. To verify the superiority of FLea, we conduct extensive experiments using a wide range of data modalities, simulating different levels of local data scarcity and label skew. The results demonstrate that FLea consistently outperforms state-of-the-art FL counterparts (among 13 of the experimented 18 settings, the improvement is over 5%) while concurrently mitigating the privacy vulnerabilities associated with shared features. Tong Xia, Abhirup Ghosh, Xinchi Qiu, Cecilia Mascolo |
KDD | 2 |
| 2023 | Cross-Device Federated Learning for Mobile Health Diagnostics: A First Study on COVID-19 DetectionabstractFederated learning (FL) aided health diagnostic models can incorporate data from a large number of personal edge devices (e.g., mobile phones) while keeping the data local to the originating devices, largely ensuring privacy. However, such a cross-device FL approach for health diagnostics still imposes many challenges due to both local data imbalance (as extreme as local data consists of a single disease class) and global data imbalance (the disease prevalence is generally low in a population). Since the federated server has no access to data distribution information, it is not trivial to solve the imbalance issue towards an unbiased model. In this paper, we propose FedLoss, a novel cross-device FL framework for health diagnostics. Here the federated server averages the models trained on edge devices according to the predictive loss on the local data, rather than using only the number of samples as weights. As the predictive loss better quantifies the data distribution at a device, FedLoss alleviates the impact of data imbalance. Through a real-world dataset on respiratory sound and symptom-based COVID-19 detection task, we validate the superiority of FedLoss. It achieves competitive COVID-19 detection performance compared to a centralised model with an AUC-ROC of 79%. It also outperforms the state-of-the-art FL baselines in sensitivity and convergence speed. Our work not only demonstrates the promise of federated COVID-19 detection but also paves the way to a plethora of mobile health model development in a privacy-preserving fashion. Tong Xia, Jing Han 0010, Abhirup Ghosh, Cecilia Mascolo |
ICASSP | 3 |
| 2023 | Modeling with Homophily Driven Heterogeneous Data in Gossip LearningabstractTraining deep learning models on data distributed and local to edge devices such as mobile phones is a prominent recent research direction. In a Gossip Learning (GL) system, each participating device maintains a model trained on its local data and iteratively aggregates it with the models from its neighbours in a communication network. While the fully distributed operation in GL comes with natural advantages over the centralized orchestration in Federated Learning (FL), its convergence becomes particularly slow when the data distribution is heterogeneous and aligns with the clustered structure of the communication network. These characteristics are pervasive across practical applications as people with similar interests (thus producing similar data) tend to create communities. This paper proposes a data-driven neighbor weighting strategy for aggregating the models: this enables faster diffusion of knowledge across the communities in the network and leads to quicker convergence. We augment the method to make it computationally efficient and fair: the devices quickly converge to the same model. We evaluate our model on real and synthetic datasets that we generate using a novel generative model for communication networks with heterogeneous data. Our exhaustive empirical evaluation verifies that our proposed method attains a faster convergence rate than the baselines. For example, the median test accuracy for a decentralized bird image classifier application reaches 81% with our proposed method within 80 rounds, whereas the baseline only reaches 46%. Abhirup Ghosh, Cecilia Mascolo |
IJCAI | 1 |
| 2022 | Publishing Asynchronous Event Times with Pufferfish PrivacyabstractPublishing data from IoT devices raises concerns of leaking sensitive information. In this paper we consider the scenario of publishing data on events with timestamps. We formulate three privacy issues, namely, whether one can tell if an event happened or not; whether one can nail down the timestamp of an event within a given time interval; and whether one can infer the relative order of any two nearby events. We show that perturbation of event timestamps or adding fake events following carefully chosen distributions can address these privacy concerns. We present a rigorous study of privately publishing discrete event timestamps with privacy guarantees under the Pufferfish privacy framework. We also conduct extensive experiments to evaluate utility of the modified time series with real world location check-in and app usage data. Our mechanisms preserve the statistical utility of event data which are suitable for aggregate queries. Jiaxin Ding 0001, Abhirup Ghosh, Rik Sarkar, Jie Gao 0001 |
DCOSS | 2 |
| 2021 | Mobility-based Individual POI Recommendation to Control the COVID-19 SpreadabstractSocietal functions have stalled during COVID-19 to reduce its spread in the population. It has been shown that visits to different venues have a large effect on spreading the virus. Hence, population-level mobility interventions like reopening selective category of venues have been proposed, for example, opening schools and offices but preventing people from visiting restaurants. These measures, although help to mitigate infection, still fail to satisfy people’s needs and hope of going back to normality. In this context, here we propose an individual level POI recommendation system that can recommend venues to users according to their preference and at the same time, can lead to as few infections as possible. The key idea behind the system is that the risk of getting infected grows with the number of unique customers that had visited the venue previously, and it is safer to visit a less crowded place during a specific time slot. We evaluate the proposed system using both theory and real check-in datasets from three cities. Based on simulation on real-world data, we present a surprising result: it is possible to recommend POIs in such a way that the total infected population reduces by up to 50% compared to that following original check-ins. This result is comparable to that when 50% of the visits are blocked, yet our method allows all check-in needs. Abhirup Ghosh, Tong Xia |
IEEE BigData | 1 |
| 2020 | Differentially Private Range Counting in Planar Graphs for Spatial SensingabstractThis paper considers the problem of privately reporting counts of events recorded by devices in different regions of the plane. Unlike previous range query methods, our approach is not limited to rectangular ranges. We devise novel hierarchical data structures to answer queries over arbitrary planar graphs. This construction relies on balanced planar separators to represent shortest paths using O(logn) number of canonical paths, where n is the number of nodes in the graph. Pre-computed sums along these canonical paths allow efficient computations of 1D counting range queries along any shortest path. We make use of differential forms together with the 1D mechanism to answer 2D queries in which a range is a union of faces in the planar graph. The methods are designed such that the range queries could be answered with differential privacy guarantee on any single event, with only a poly-logarithmic error. They also allow private range queries to be performed in a distributed setup. Theoretical and experimental results confirm that the methods are efficient and accurate on real data and incur less error than competing existing methods. Abhirup Ghosh, Jiaxin Ding 0001, Rik Sarkar, Jie Gao 0001 |
INFOCOM | 1 |
| 2018 | Topological signatures for fast mobility analysisabstractAnalytic methods can be difficult to build and costly to train for mobility data. We show that information about the topology of the space and how mobile objects navigate the obstacles can be used to extract insights about mobility at larger distance scales. The main contribution of this paper is a topological signature that maps each trajectory to a relatively low dimensional Euclidean space, so that now they are amenable to standard analytic techniques. Data mining tasks: nearest neighbor search with locality sensitive hashing, clustering, regression, etc., work more efficiently in this signature space. We define the problem of mobility prediction at different distance scales, and show that with the signatures simple k nearest neighbor based regression perform accurate prediction. Experiments on multiple real datasets show that the framework using topological signatures is accurate on all tasks, and substantially more efficient than machine learning applied to raw data. Theoretical results show that the signatures contain enough topological information to reconstruct non-self-intersecting trajectories upto homotopy type. The construction of signatures is based on a differential form that can be generated in a distributed setting using local communication, and a signature can be locally and inexpensively updated and communicated by a mobile agent. Abhirup Ghosh, Benedek Rozemberczki, Subramanian Ramamoorthy, Rik Sarkar |
SIGSPATIAL/GIS | 1 |
| 2017 | Finding Periodic Discrete Events in Noisy StreamsabstractPeriodic phenomena are ubiquitous, but detecting and predicting periodic events can be difficult in noisy environments. We describe a model of periodic events that covers both idealized and realistic scenarios characterized by multiple kinds of noise. The model incorporates false-positive events and the possibility that the underlying period and phase of the events change over time. We then describe a particle filter that can efficiently and accurately estimate the parameters of the process generating periodic events intermingled with independent noise events. The system has a small memory footprint, and, unlike alternative methods, its computational complexity is constant in the number of events that have been observed. As a result, it can be applied in low-resource settings that require real-time performance over long periods of time. In experiments on real and simulated data we find that it outperforms existing methods in accuracy and can track changes in periodicity and other characteristics in dynamic event streams. Abhirup Ghosh, Christopher G. Lucas, Rik Sarkar |
CIKM | 1 |