VLDB 2026 Research / reviewers in the wild / expert
Dhrubajyoti Ghosh
dblp:155/0356
· DBLP profile ↗
13ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Integrating Causal Inference with Graph Neural Networks for Alzheimer's Disease Analysis
Pranay Kumar Peddi, Dhrubajyoti Ghosh |
DATA (1) | 2 |
| 2026 | Equivalence and Separation Between Heard-Of and Asynchronous Message-Passing Models
Hagit Attiya, Armando Castañeda, Dhrubajyoti Ghosh, Thomas Nowak 0001 |
SIROCCO | 3 |
| 2026 | Machine Learning Based Bot Detection on X With Temporal and Semantic Feature IntegrationabstractEscalating proliferation of inorganic accounts, commonly known as bots, within the digital ecosystem represents an ongoing and multifaceted challenge to online security, trustworthiness, and user experience. These bots, often employed for the dissemination of malicious propaganda and manipulation of public opinion, wield significant influence in social media spheres with far-reaching implications for electoral processes, political campaigns, and international conflicts. Swift and accurate identification of inorganic accounts is of paramount importance in mitigating their detrimental effects. This research article focuses on the identification of such accounts and explores various effective methods for their detection through machine learning techniques. In response to the pervasive presence of bots in the contemporary digital landscape, this study extracts temporal and semantic features from tweet behaviors and proposes a bot detection framework using fundamental machine learning approaches, including support vector machines (SVMs) and k-means clustering. Furthermore, the research ranks the importance of these extracted features for each detection technique and also provides uncertainty quantification using a distribution-free method, called the conformal prediction, thereby contributing to the development of effective strategies for combating the prevalence of inorganic accounts in social media platforms. Dhrubajyoti Ghosh, William Boettcher, Rob Johnston, Soumendra Lahiri |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2025 | XCAT 3.0: A comprehensive library of personalized digital twins derived from CT scans
Lavsen Dahal, Mobina Ghojogh Nejad, Liesbeth Vancoillie, Dhrubajyoti Ghosh, Yubraj Bhandari, Fong Chi Ho, Fakrul Islam Tushar, Sheng Luo 0005, Kyle J. Lafata, Ehsan Abadi, Ehsan Samei, Joseph Y. Lo, William Paul Segars |
Medical Image Anal. | 4 |
| 2025 | Virtual lung screening trial (VLST): An in silico study inspired by the national lung screening trial for lung cancer detectionabstractstudy inspired by the National Lung Screening Trial (NLST), illustrates the potential of VITs to expedite clinical trials, minimize risks to participants, and promote optimal use of imaging technologies in healthcare. This study aimed to show that a virtual imaging trial platform could investigate some key elements of a major clinical trial, specifically the NLST, which compared Computed tomography (CT) and chest radiography (CXR) for lung cancer screening. With simulated cancerous lung nodules, a virtual patient cohort of 294 subjects was created using XCAT human models. Each virtual patient underwent both CT and CXR imaging, with deep learning models, the AI CT-Reader and AI CXR-Reader, acting as virtual readers to perform recall patients with suspicion of lung cancer. The primary outcome was the difference in diagnostic performance between CT and CXR, measured by the Area Under the Curve (AUC). The AI CT-Reader showed superior diagnostic accuracy, achieving an AUC of 0.92 (95% CI: 0.90-0.95) compared to the AI CXR-Reader's AUC of 0.72 (95% CI: 0.67-0.77). Furthermore, at the same 94% CT sensitivity reported by the NLST, the VLST specificity of 73% was similar to the NLST specificity of 73.4%. This CT performance highlights the potential of VITs to replicate certain aspects of clinical trials effectively, paving the way toward a safe and efficient method for advancing imaging-based diagnostics. Fakrul Islam Tushar, Liesbeth Vancoillie, Cindy McCabe, Amareswararao Kavuri, Lavsen Dahal, Brian P. Harrawood, Milo Fryling, Mojtaba Zarei, Saman Sotoudeh-Paima, Fong Chi Ho, Dhrubajyoti Ghosh, Michael R. Harowicz, Tina D. Tailor, Sheng Luo 0005, William Paul Segars, Ehsan Abadi, Kyle J. Lafata, Joseph Y. Lo, Ehsan Samei |
Medical Image Anal. | 11 |
| 2024 | Prism: Privacy-Preserving and Verifiable Set Computation Over Multi-Owner Secret Shared Outsourced DatabasesabstractPrivate set computation over multi-owner databases is an important problem with many applications — the most well studied of which is private set intersection (PSI). This paper proposesPrism, a secret-sharing based approach to compute private set operations (i.e., intersection and union, as well as aggregates such as count, sum, average, maximum, minimum, and median) over outsourced databases belonging to multiple owners.Prismenables data owners to pre-load the data onto non-colluding servers and exploits the additive and multiplicative properties of secret-shares to compute the above-listed operations.Prismtakes (at most) two rounds of communication between non-colluding servers (storing the secret-shares) and the querier for executing the above-mentioned operations, resulting in a very efficient implementation.Prismalso supports result verification techniques for each operation to detect malicious adversaries. Experimental results show thatPrismscales both in terms of the number of data owners and database sizes, to which prior approaches do not scale. Shantanu Sharma 0001, Yin Li 0001, Sharad Mehrotra, Nisha Panwar, Peeyush Gupta, Dhrubajyoti Ghosh |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2023 | Supporting Complex Query Time Enrichment For Analytics
Dhrubajyoti Ghosh, Peeyush Gupta, Sharad Mehrotra, Shantanu Sharma 0001 |
EDBT | 1 |
| 2022 | MIDE: Accuracy Aware Minimally Invasive Data Exploration For Decision SupportabstractThis paper studies privacy in the context of decision-support queries that classify objects as either true or false based on whether they satisfy the query. Mechanisms to ensure privacy may result in false positives and false negatives. In decision-support applications, often, false negatives have to remain bounded. Existing accuracy-aware privacy preserving techniques cannot directly be used to support such an accuracy requirement and their naive adaptations to support bounded accuracy of false negatives results in significant privacy loss depending upon distribution of data. This paper explores the concept of minimally-invasive data exploration for decision support that attempts to minimize privacy loss while supporting bounded guarantee on false negatives by adaptively adjusting privacy based on data distribution. Our experimental results show that the MIDE algorithms perform well and are robust over variations in data distributions. Sameera Ghayyur, Dhrubajyoti Ghosh, Xi He 0001, Sharad Mehrotra |
Proc. VLDB Endow. | 2 |
| 2022 | JENNER: Just-in-time Enrichment in Query ProcessingabstractEmerging domains, such as sensor-driven smart spaces and social media analytics, require incoming data to be enriched prior to its use. Enrichment often consists of machine learning (ML) functions that are too expensive/infeasible to execute at ingestion. We develop a strategy entitled Just-in-time ENrichmeNt in quERy Processing (JENNER) to support interactive analytics over data as soon as it arrives for such application context. JENNER exploits the inherent tradeoffs of cost and quality often displayed by the ML functions to progressively improve query answers during query execution. We describe how JENNER works for a large class of SPJ and aggregation queries that form the bulk of data analytics workload. Our experimental results on real datasets (IoT and Tweet) show that JENNER achieves progressive answers performing significantly better than the naive strategies of achieving progressive computation. Dhrubajyoti Ghosh, Peeyush Gupta, Sharad Mehrotra, Roberto Yus, Yasser Altowim |
Proc. VLDB Endow. | 1 |
| 2021 | PRISM: Private Verifiable Set Computation over Multi-Owner Outsourced DatabasesabstractThis paper proposes Prism, a secret sharing based approach to compute private set operations (i.e., intersection and union), as well as aggregates over outsourced databases belonging to multiple owners. Prism enables data owners to pre-load the data onto non-colluding servers and exploits the additive and multiplicative properties of secret-shares to compute the above-listed operations in (at most) two rounds of communication between the servers (storing the secret-shares) and the querier, resulting in a very efficient implementation. Also, Prism does not require communication among the servers and supports result verification techniques for each operation to detect malicious adversaries. Experimental results show that Prism scales both in terms of the number of data owners and database sizes, to which prior approaches do not scale. Yin Li 0001, Dhrubajyoti Ghosh, Peeyush Gupta, Sharad Mehrotra, Nisha Panwar, Shantanu Sharma 0001 |
SIGMOD Conference | 2 |
| 2020 | IoT Expunge: Implementing Verifiable Retention of IoT DataabstractThe growing deployment of Internet of Things (IoT) systems aims to ease the daily life of end-users by providing several value-added services. However, IoT systems may capture and store sensitive, personal data about individuals in the cloud, thereby jeopardizing user-privacy. Emerging legislation, such as California's CalOPPA and GDPR in Europe, support strong privacy laws to protect an individual's data in the cloud. One such law relates to strict enforcement of data retention policies. This paper proposes a framework, entitled IoT Expunge that allows sensor data providers to store the data in cloud platforms that will ensure enforcement of retention policies. Additionally, the cloud provider produces verifiable proofs of its adherence to the retention policies. Experimental results on a real-world smart building testbed show that IoT Expunge imposes minimal overheads to the user to verify the data against data retention policies. Nisha Panwar, Shantanu Sharma 0001, Peeyush Gupta, Dhrubajyoti Ghosh, Sharad Mehrotra, Nalini Venkatasubramanian |
CODASPY | 4 |
| 2020 | A privacy-enabled platform for COVID-19 applications: poster abstractabstractWe present our experiences in adapting and deploying TIPPERS1, a novel privacy-enabled IoT data collection and management system for smart spaces, to facilitate the monitoring of adherence to COVID-19 regulations in a university campus and a military facility. Michael August, Mamadou H. Diallo, Dhrubajyoti Ghosh, Peeyush Gupta, Christopher Graves 0002, Michael Holstrom, Pramod P. Khargonekar, Megan Kline, Sharad Mehrotra, Shantanu Sharma 0001, Nalini Venkatasubramanian, Guoxi Wang, Roberto Yus |
SenSys | 4 |
| 2019 | SCARF: a scalable data management framework for context-aware applications in smart environmentsabstractThis paper proposes a context data management framework for efficient context data collection and evaluation to support context-aware applications. The proposed framework named SCARF leverages two architectures: a server-based centralized architecture and an edge-based architecture to address the problem of scalability of context collection and evaluation. SCARF collects the context data efficiently from resource constrained mobile/edge and multiple in-situ sensors (not resource constrained) deployed in smart environments. In particular, SCARF takes a set of contextual queries from applications and optimizes the rate at which the context data needs to be collected and/or transmitted from the sensor to the server. Furthermore, SCARF decides whether the code for evaluating context should be executed on the server or on the edge. Execution on the edge will result in reduced communication between the server and the edge but can result in low performance due to resource constraints of the edge. On the other hand, evaluation of context on the server, will require SCARF to transmit sensor data from the sensor to the server but can result in high performance due to high resource availability of the server. SCARF is an adaptive framework which generates an optimal context acquisition plan and a context evaluation plan such that the overall latency is minimized. Eun-Jeong Shin, Dhrubajyoti Ghosh, Sharad Mehrotra, Nalini Venkatasubramanian |
MobiQuitous | 2 |