VLDB 2026 Research / reviewers in the wild / expert
Disha Makhija
dblp:169/9962
· DBLP profile ↗
6ranked-venue papers
2as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 62% Learning paradigms · 14% Probabilistic and Bayesian machine learning · 12% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 83% Recommender systems · 9% Web and social media mining · 8% | |
| Network and information security
1 paper |
Privacy and data protection · 100% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
federated learning |
1.3 | 2 | 2024 | A Bayesian Approach for Personalized Federated Learning in Heterogeneous Settings · NeurIPS 2024 Architecture Agnostic Federated Learning for Neural Networks · ICML 2022 |
Machine learning › Efficient and distributed learning › federated learning
personalized federated learning |
1.3 | 2 | 2024 | A Bayesian Approach for Personalized Federated Learning in Heterogeneous Settings · NeurIPS 2024 Architecture Agnostic Federated Learning for Neural Networks · ICML 2022 |
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks |
0.8 | 1 | 2024 | A Bayesian Approach for Personalized Federated Learning in Heterogeneous Settings · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.8 | 1 | 2024 | A Bayesian Approach for Personalized Federated Learning in Heterogeneous Settings · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › federated learning
heterogeneous architectures |
0.6 | 1 | 2022 | Architecture Agnostic Federated Learning for Neural Networks · ICML 2022 |
Machine learning › Efficient and distributed learning › federated learning
model-agnostic federated learning |
0.6 | 1 | 2022 | Architecture Agnostic Federated Learning for Neural Networks · ICML 2022 |
Machine learning › Learning paradigms › semi-supervised learning
graph-based semi-supervised learning |
0.4 | 1 | 2020 | Understanding the Success of Graph-based Semi-Supervised Learning using Partially Labelled Stochastic Block Model · IJCAI 2020 |
Machine learning › Learning paradigms › semi-supervised learning › graph-based semi-supervised learning
label propagation |
0.4 | 1 | 2020 | Understanding the Success of Graph-based Semi-Supervised Learning using Partially Labelled Stochastic Block Model · IJCAI 2020 |
Graph algorithms and graph theory › graph clustering › community detection
stochastic block model |
0.4 | 1 | 2020 | Understanding the Success of Graph-based Semi-Supervised Learning using Partially Labelled Stochastic Block Model · IJCAI 2020 |
Data mining › anomaly detection
fraud detection |
0.3 | 1 | 2018 | REV2: Fraudulent User Prediction in Rating Platforms · WSDM 2018 |
Data mining › structured data mining › graph mining › graph learning
graph classification |
0.3 | 1 | 2017 | ZooBP: Belief Propagation for Heterogeneous Networks · Proc. VLDB Endow. 2017 |
Privacy and data protection
differential privacy |
0.2 | 1 | 2024 | A Bayesian Approach for Personalized Federated Learning in Heterogeneous Settings · NeurIPS 2024 |
Privacy and data protection
privacy-preserving machine learning |
0.2 | 1 | 2024 | A Bayesian Approach for Personalized Federated Learning in Heterogeneous Settings · NeurIPS 2024 |
Recommender systems › rating systems
online rating systems |
0.1 | 1 | 2018 | REV2: Fraudulent User Prediction in Rating Platforms · WSDM 2018 |
Methods — techniques the papers use, named apart from their topics
differential privacy · 1.5bayesian neural network · 1.5bayesian learning · 1.5knowledge distillation · 0.6instance-level representation · 0.6closed-form inference · 0.3belief propagation · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Bayesian Approach for Personalized Federated Learning in Heterogeneous SettingsabstractFederated learning (FL), through its privacy-preserving collaborative learning approach, has significantly empowered decentralized devices. However, constraints in either data and/or computational resources among participating clients introduce several challenges in learning, including the inability to train large model architectures, heightened risks of overfitting, and more. In this work, we present a novel FL framework grounded in Bayesian learning to address these challenges. Our approach involves training personalized Bayesian models at each client tailored to the unique complexities of the clients' datasets and efficiently collaborating across these clients. By leveraging Bayesian neural networks and their uncertainty quantification capabilities, our local training procedure robustly learns from small datasets. And the novel collaboration procedure utilizing priors in the functional (output) space of the networks facilitates collaboration across models of varying sizes, enabling the framework to adapt well in heterogeneous data and computational settings. Furthermore, we present a differentially private version of the algorithm, accompanied by formal differential privacy guarantees that apply without any assumptions on the learning algorithm. Through experiments on popular FL datasets, we demonstrate that our approach outperforms strong baselines in both homogeneous and heterogeneous settings, and under strict privacy constraints. Disha Makhija, Joydeep Ghosh, Nhat Ho |
NeurIPS | 1 |
| 2022 | Architecture Agnostic Federated Learning for Neural NetworksabstractWith growing concerns regarding data privacy and rapid increase in data volume, Federated Learning (FL) has become an important learning paradigm. However, jointly learning a deep neural network model in a FL setting proves to be a non-trivial task because of the complexities associated with the neural networks, such as varied architectures across clients, permutation invariance of the neurons, and presence of non-linear transformations in each layer. This work introduces a novel framework, Federated Heterogeneous Neural Networks (FedHeNN), that allows each client to build a personalised model without enforcing a common architecture across clients. This allows each client to optimize with respect to local data and compute constraints, while still benefiting from the learnings of other (potentially more powerful) clients. The key idea of FedHeNN is to use the instance-level representations obtained from peer clients to guide the simultaneous training on each client. The extensive experimental results demonstrate that the FedHeNN framework is capable of learning better performing models on clients in both the settings of homogeneous and heterogeneous architectures across clients. Disha Makhija, Xing Han, Nhat Ho, Joydeep Ghosh |
ICML | 1 |
| 2020 | Understanding the Success of Graph-based Semi-Supervised Learning using Partially Labelled Stochastic Block ModelabstractWith the proliferation of learning scenarios with an abundance of instances, but limited amount of high-quality labels, semi-supervised learning algorithms came to prominence. Graph-based semi-supervised learning (G-SSL) algorithms, of which Label Propagation (LP) is a prominent example, are particularly well-suited for these problems. The premise of LP is the existence of homophily in the graph, but beyond that nothing is known about the efficacy of LP. In particular, there is no characterisation that connects the structural constraints, volume and quality of the labels to the accuracy of LP. In this work, we draw upon the notion of recovery from the literature on community detection, and provide guarantees on accuracy for partially-labelled graphs generated from the Partially-Labelled Stochastic Block Model (PLSBM). Extensive experiments performed on synthetic data verify the theoretical findings. Avirup Saha, Shreyas Sheshadri, Samik Datta, Niloy Ganguly, Disha Makhija, Priyank Patel |
IJCAI | 5 |
| 2018 | REV2: Fraudulent User Prediction in Rating PlatformsabstractRating platforms enable large-scale collection of user opinion about items(e.g., products or other users). However, untrustworthy users give fraudulent ratings for excessive monetary gains. In this paper, we present REV2, a system to identify such fraudulent users. We propose three interdependent intrinsic quality metrics---fairness of a user, reliability of a rating and goodness of a product. The fairness and reliability quantify the trustworthiness of a user and rating, respectively, and goodness quantifies the quality of a product. Intuitively, a user is fair if it provides reliable scores that are close to the goodness of products. We propose six axioms to establish the interdependency between the scores, and then, formulate a mutually recursive definition that satisfies these axioms. We extend the formulation to address cold start problem and incorporate behavior properties. We develop the REV2 algorithm to calculate these intrinsic quality scores for all users, ratings, and products. We show that this algorithm is guaranteed to converge and has linear time complexity. By conducting extensive experiments on five rating datasets, we show that REV2 outperforms nine existing algorithms in detecting fair and unfair users. We reported the 150 most unfair users in the Flipkart network to their review fraud investigators, and 127 users were identified as being fraudulent(84.6% accuracy). The REV2 algorithm is being deployed at Flipkart. Srijan Kumar, Bryan Hooi, Disha Makhija, Mohit Kumar 0008, Christos Faloutsos, V. S. Subrahmanian |
WSDM | 3 |
| 2017 | ZooBP: Belief Propagation for Heterogeneous NetworksabstractGiven a heterogeneous network, with nodes of different types - e.g., products, users and sellers from an online recommendation site like Amazon - and labels for a few nodes ('honest', 'suspicious', etc), can we find a closed formula for Belief Propagation (BP), exact or approximate? Can we say whether it will converge? BP, traditionally an inference algorithm for graphical models, exploits so-called "network effects" to perform graph classification tasks when labels for a subset of nodes are provided; and it has been successful in numerous settings like fraudulent entity detection in online retailers and classification in social networks. However, it does not have a closed-form nor does it provide convergence guarantees in general. We propose ZooBP, a method to perform fast BP on undirected heterogeneous graphs with provable convergence guarantees. ZooBP has the following advantages: (1) Generality : It works on heterogeneous graphs with multiple types of nodes and edges; (2) Closed-form solution: ZooBP gives a closed-form solution as well as convergence guarantees; (3) Scalability: ZooBP is linear on the graph size and is up to 600× faster than BP, running on graphs with 3.3 million edges in a few seconds. (4) Effectiveness: Applied on real data (a F lipkart e-commerce network with users, products and sellers), ZooBP identifies fraudulent users with a near-perfect precision of 92.3 % over the top 300 results. Dhivya Eswaran, Stephan Günnemann, Christos Faloutsos, Disha Makhija, Mohit Kumar 0008 |
Proc. VLDB Endow. | 4 |
| 2016 | BIRDNEST: Bayesian Inference for Ratings-Fraud DetectionabstractReview fraud is a pervasive problem in online commerce, in which fraudulent sellers write or purchase fake reviews to manipulate perception of their products and services. Fake reviews are often detected based on several signs, including 1) they occur in short bursts of time; 2) fraudulent user accounts have skewed rating distributions. However, these may both be true in any given dataset. Hence, in this paper, we propose an approach for detecting fraudulent reviews which combines these 2 approaches in a principled manner, allowing successful detection even when one of these signs is not present. To combine these 2 approaches, we formulate our Bayesian Inference for Rating Data (BIRD) model, a flexible Bayesian model of user rating behavior. Based on our model we formulate a likelihood-based suspiciousness metric, Normalized Expected Surprise Total (NEST). We propose a linear-time algorithm for performing Bayesian inference using our model and computing the metric. Experiments on real data show that BIRDNEST successfully spots review fraud in large, real-world graphs: the 50 most suspicious users of the Flipkart platform flagged by our algorithm were investigated and all identified as fraudulent by domain experts at Flipkart. Bryan Hooi, Neil Shah, Alex Beutel, Stephan Günnemann, Leman Akoglu, Mohit Kumar 0008, Disha Makhija, Christos Faloutsos |
SDM | 7 |