VLDB 2026 Research / reviewers in the wild / expert
Karthika Mohan
dblp:92/8278
· DBLP profile ↗
12ranked-venue papers
5as first author
4since 2021 · last 2026
0000-0003-0554-9433ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Probabilistic and Bayesian machine learning · 67% Knowledge representation and reasoning · 21% Trustworthy machine learning · 8% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning |
1.9 | 2 | 2026 | Discovering Linear Non-Gaussian Models for All Categories of Missing Data (Student Abstract) · AAAI 2026 Agents Robust to Distribution Shifts Learn Causal World Models Even Under Mediation · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery |
1.8 | 2 | 2026 | Discovering Linear Non-Gaussian Models for All Categories of Missing Data (Student Abstract) · AAAI 2026 Do Finetti: On Causal Effects for Exchangeable Data · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
1.7 | 4 | 2024 | Do Finetti: On Causal Effects for Exchangeable Data · NeurIPS 2024 Causal Inference with Non-IID Data using Linear Graphical Models · NeurIPS 2022 Graphical Models for Recovering Probabilistic and Causal Queries from Missing Data · NIPS 2014 |
Machine learning › Probabilistic and Bayesian machine learning
missing data |
1.2 | 2 | 2026 | Discovering Linear Non-Gaussian Models for All Categories of Missing Data (Student Abstract) · AAAI 2026 Graphical Models for Inference with Missing Data · NIPS 2013 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
linear non-gaussian acyclic model |
1.0 | 1 | 2026 | Discovering Linear Non-Gaussian Models for All Categories of Missing Data (Student Abstract) · AAAI 2026 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal model
causal world model |
0.9 | 1 | 2025 | Agents Robust to Distribution Shifts Learn Causal World Models Even Under Mediation · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift |
0.9 | 1 | 2025 | Agents Robust to Distribution Shifts Learn Causal World Models Even Under Mediation · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal effect estimation |
0.8 | 1 | 2024 | Do Finetti: On Causal Effects for Exchangeable Data · NeurIPS 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
interference |
0.6 | 1 | 2022 | Causal Inference with Non-IID Data using Linear Graphical Models · NeurIPS 2022 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning
graphical model structure learning |
0.4 | 2 | 2014 | Graphical Models for Recovering Probabilistic and Causal Queries from Missing Data · NIPS 2014 Graphical Models for Inference with Missing Data · NIPS 2013 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
partially observable markov decision process |
0.3 | 1 | 2025 | Agents Robust to Distribution Shifts Learn Causal World Models Even Under Mediation · NeurIPS 2025 |
Machine learning › Generative modeling
exchangeable data |
0.2 | 1 | 2024 | Do Finetti: On Causal Effects for Exchangeable Data · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.2 | 1 | 2022 | Causal Inference with Non-IID Data using Linear Graphical Models · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
causal inference · 1.1imputation · 1.0functional causal model · 1.0LiNGAM · 1.0optimal policy oracles · 0.9truncated factorization · 0.8spectral methods · 0.8pólya urn model · 0.8linear graphical model · 0.6causal graph · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discovering Linear Non-Gaussian Models for All Categories of Missing Data (Student Abstract)abstractCausal discovery is the task of learning causal models, encoding causal relationships, from a source of information, such as a dataset containing observational data. While many algorithms have been developed to discover causal models under varied sets of assumptions, the case in which the dataset is affected by missing data remains significantly underexplored. Naively applying standard causal discovery algorithms to listwise, test-wise, or regression-wise deleted datasets, or imputing the missing data, can introduce spurious associations between variables and bias function estimation in functional causal models. This issue arises when the data is missing at random or not at random. It ultimately invalidates the theoretical guarantees of these algorithms and prevents finding the true underlying causal model, even in the large-sample limit. An established family of causal models is the Linear Non-Gaussian Acyclic Model (LiNGAM), which assumes linear functional relationships and non-Gaussian independent noise terms. We propose a new causal discovery algorithm for LiNGAM, capable of recovering the underlying causal structure and providing unbiased estimates of the model’s parameters, even when the data is affected by MNAR missingness. Matteo Ceriscioli, Shohei Shimizu, Karthika Mohan |
AAAI | 3 |
| 2025 | Agents Robust to Distribution Shifts Learn Causal World Models Even Under MediationabstractIn this work, we prove that agents capable of adapting to distribution shifts must have learned the causal model of their environment even in the presence of mediation. This term describes situations where an agent's actions affect its environment, a dynamic common to most real-world settings. For example, a robot in an industrial plant might interact with tools, move through space, and transform products to complete its task. We introduce an algorithm for eliciting causal knowledge from robust agents using optimal policy oracles, with the flexibility to incorporate prior causal knowledge. We further demonstrate its effectiveness in mediated single-agent scenarios and multi-agent environments. We identify conditions under which the presence of a single robust agent is sufficient to recover the full causal model and derive optimal policies for other agents in the same environment. Finally, we show how to apply these results to sequential decision-making tasks modeled as Partially Observable Markov Decision Processes (POMDPs). Matteo Ceriscioli, Karthika Mohan |
NeurIPS | 2 |
| 2024 | Do Finetti: On Causal Effects for Exchangeable DataabstractWe study causal effect estimation in a setting where the data are not i.i.d.$\ $(independent and identically distributed). We focus on exchangeable data satisfying an assumption of independent causal mechanisms. Traditional causal effect estimation frameworks, e.g., relying on structural causal models and do-calculus, are typically limited to i.i.d. data and do not extend to more general exchangeable generative processes, which naturally arise in multi-environment data. To address this gap, we develop a generalized framework for exchangeable data and introduce a truncated factorization formula that facilitates both the identification and estimation of causal effects in our setting. To illustrate potential applications, we introduce a causal Pólya urn model and demonstrate how intervention propagates effects in exchangeable data settings. Finally, we develop an algorithm that performs simultaneous causal discovery and effect estimation given multi-environment data. Siyuan Guo 0003, Karthika Mohan, Ferenc Huszar, Bernhard Schölkopf |
NeurIPS | 3 |
| 2022 | Causal Inference with Non-IID Data using Linear Graphical ModelsabstractTraditional causal inference techniques assume data are independent and identically distributed (IID) and thus ignores interactions among units. However, a unit’s treatment may affect another unit's outcome (interference), a unit’s treatment may be correlated with another unit’s outcome, or a unit’s treatment and outcome may be spuriously correlated through another unit. To capture such nuances, we model the data generating process using causal graphs and conduct a systematic analysis of the bias caused by different types of interactions when computing causal effects. We derive theorems to detect and quantify the interaction bias, and derive conditions under which it is safe to ignore interactions. Put differently, we present conditions under which causal effects can be computed with negligible bias by assuming that samples are IID. Furthermore, we develop a method to eliminate bias in cases where blindly assuming IID is expected to yield a significantly biased estimate. Finally, we test the coverage and performance of our methods through simulations. Chi Zhang 0016, Karthika Mohan, Judea Pearl |
NeurIPS | 2 |
| 2019 | Causal Discovery in the Presence of Missing DataabstractMissing data are ubiquitous in many domains such as healthcare. When these data entries are not missing completely at random, the (conditional) independence relations in the observed data may be different from those in the complete data generated by the underlying causal process. Consequently, simply applying existing causal discovery methods to the observed data may lead to wrong conclusions. In this paper, we aim at developing a causal discovery method to recover the underlying causal structure from observed data that are missing under different mechanisms, including missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR). With missingness mechanisms represented by missingness graphs (m-graphs), we analyze conditions under which additional correction is needed to derive conditional independence/dependence relations in the complete data. Based on our analysis, we propose Missing Value PC (MVPC), which extends the PC algorithm to incorporate additional corrections. Our proposed MVPC is shown in theory to give asymptotically correct results even on data that are MAR or MNAR. Experimental results on both synthetic data and real healthcare applications illustrate that the proposed algorithm is able to find correct causal relations even in the general case of MNAR. Ruibo Tu, Cheng Zhang 0005, Paul Ackermann, Karthika Mohan, Hedvig Kjellström, Kun Zhang 0001 |
AISTATS | 4 |
| 2018 | Estimation with Incomplete Data: The Linear CaseabstractTraditional methods for handling incomplete data, including Multiple Imputation and Maximum Likelihood, require that the data be Missing At Random (MAR). In most cases, however, missingness in a variable depends on the underlying value of that variable. In this work, we devise model-based methods to consistently estimate mean, variance and covariance given data that are Missing Not At Random (MNAR). While previous work on MNAR data require variables to be discrete, we extend the analysis to continuous variables drawn from Gaussian distributions. We demonstrate the merits of our techniques by comparing it empirically to state of the art software packages. Karthika Mohan, Felix Thömmes, Judea Pearl |
IJCAI | 1 |
| 2015 | Efficient Algorithms for Bayesian Network Parameter Learning from Incomplete Data
Guy Van den Broeck, Karthika Mohan, Arthur Choi, Adnan Darwiche, Judea Pearl |
UAI | 2 |
| 2015 | Missing Data as a Causal and Probabilistic Problem
Ilya Shpitser, Karthika Mohan, Judea Pearl |
UAI | 2 |
| 2014 | On the Testability of Models with Missing DataabstractGraphical models that depict the process by which data are lost are helpful in recovering information from missing data. We address the question of whether any such model can be submitted to a statistical test given that the data available are corrupted by missingness. We present sufficient conditions for testability in missing data applications and note the impediments for testability when data are contaminated by missing entries. Our results strengthen the available tests for MCAR and MAR and further provide tests in the category of MNAR. Furthermore, we provide sufficient conditions to detect the existence of dependence between a variable and its missingness mechanism. We use our results to show that model sensitivity persists in almost all models typically categorized as MNAR. Karthika Mohan, Judea Pearl |
AISTATS | 1 |
| 2014 | Graphical Models for Recovering Probabilistic and Causal Queries from Missing Data
Karthika Mohan, Judea Pearl |
NIPS | 1 |
| 2013 | Graphical Models for Inference with Missing DataabstractWe address the problem of deciding whether there exists a consistent estimator of a given relation Q, when data are missing not at random. We employ a formal representation called `Missingness Graphs' to explicitly portray the causal mechanisms responsible for missingness and to encode dependencies between these mechanisms and the variables being measured. Using this representation, we define the notion of \textit{recoverability} which ensures that, for a given missingness-graph $G$ and a given query $Q$ an algorithm exists such that in the limit of large samples, it produces an estimate of $Q$ \textit{as if} no data were missing. We further present conditions that the graph should satisfy in order for recoverability to hold and devise algorithms to detect the presence of these conditions. Karthika Mohan, Judea Pearl, Jin Tian 0001 |
NIPS | 1 |
| 2010 | A post-processing scheme for malayalam using statistical sub-character language modelsabstractMost of the Indian scripts do not have any robust commercial OCRs. Many of the laboratory prototypes report reasonable results at recognition/classification stage. However, word level accuracies are still poor. It is well known that word accuracy decreases as the number of characters in a word increase. For Malayalam, the average number of characters in a word is almost twice that of English. Moreover, the number of words required to cover 80% of the Malayalam language is more than forty times that of other Indian languages such as Hindi. Hence a direct dictionary based post-processing scheme is not suitable for Malayalam. Karthika Mohan, C. V. Jawahar |
Document Analysis Systems | 1 |