VLDB 2026 Research / reviewers in the wild / expert
Osman Mian
dblp:294/0815 · also Osman Ali Mian
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
0009-0006-1112-6145ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Probabilistic and Bayesian machine learning · 94% Learning paradigms · 3% Trustworthy machine learning · 3% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery |
2.9 | 4 | 2026 | Causal Structure Learning for Dynamical Systems with Theoretical Score Analysis · AAAI 2026 Learning Causal Networks from Episodic Data · KDD 2024 Information-Theoretic Causal Discovery and Intervention Detection over Multiple Environments · AAAI 2023 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
1.6 | 2 | 2026 | Causal Structure Learning for Dynamical Systems with Theoretical Score Analysis · AAAI 2026 Inferring Cause and Effect in the Presence of Heteroscedastic Noise · ICML 2022 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
1.0 | 1 | 2026 | Causal Structure Learning for Dynamical Systems with Theoretical Score Analysis · AAAI 2026 |
Data mining
pattern mining |
1.0 | 1 | 2026 | SEQRET: Mining Rule Sets from Event Sequences · AAAI 2026 |
Data mining › pattern mining
sequential pattern mining |
1.0 | 1 | 2026 | SEQRET: Mining Rule Sets from Event Sequences · AAAI 2026 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
multi-environment causal discovery |
0.7 | 1 | 2023 | Information-Theoretic Causal Discovery and Intervention Detection over Multiple Environments · AAAI 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › probabilistic regression
heteroscedastic modeling |
0.6 | 1 | 2022 | Inferring Cause and Effect in the Presence of Heteroscedastic Noise · ICML 2022 |
Machine learning › Learning paradigms
continual learning |
0.2 | 1 | 2024 | Learning Causal Networks from Episodic Data · KDD 2024 |
Machine learning › Trustworthy machine learning › dataset bias
selection bias |
0.2 | 1 | 2024 | Learning Causal Networks from Episodic Data · KDD 2024 |
Methods — techniques the papers use, named apart from their topics
minimum description length · 3.2greedy search · 1.0gaussian process inference · 1.0information-theoretic scoring · 0.8continual learning · 0.8algorithmic model of causation · 0.7penalized regression score · 0.6dynamic programming · 0.6non-parametric multivariate regression · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SEQRET: Mining Rule Sets from Event SequencesabstractSummarizing event sequences is a key aspect of data mining. Most existing methods neglect conditional dependencies and focus on discovering sequential patterns only. In this paper, we study the problem of discovering both conditional and unconditional dependencies from event sequences. We do so by discovering rules of the form X --> Y where X and Y are sequential patterns. Rules like these are simple to understand and provide a clear description of the relation between the antecedent and the consequent. To discover succinct and non-redundant sets of rules we formalize the problem in terms of the Minimum Description Length principle. As the search space is enormous and does not exhibit helpful structure, we propose the SEQRET method to discover high-quality rule sets in practice. Through extensive empirical evaluation we show that unlike the state of the art, SEQRET ably recovers the ground truth on synthetic datasets and finds useful rules from real datasets. Aleena Siji, Joscha Cüppers, Osman Mian, Jilles Vreeken |
AAAI | 3 |
| 2026 | Causal Structure Learning for Dynamical Systems with Theoretical Score AnalysisabstractReal world systems evolve in continuous-time according to their underlying causal relationships, yet their dynamics are often unknown. Existing approaches to learning such dynamics typically either discretize time ---leading to poor performance on irregularly sampled data--- or ignore the underlying causality. We propose CADYT, a novel method for causal discovery on dynamical systems addressing both these challenges. In contrast to state-of-the-art causal discovery methods that model the problem using discrete-time Dynamic Bayesian networks, our formulation is grounded in Difference-based causal models, which allow milder assumptions for modeling the continuous nature of the system. CADYT leverages exact Gaussian Process inference for modeling the continuous-time dynamics which is more aligned with the underlying dynamical process. We propose a practical instantiation that identifies the causal structure via a greedy search guided by the Algorithmic Markov Condition and Minimum Description Length principle. Our experiments show that CADYT outperforms state-of-the-art methods on both regularly and irregularly-sampled data, discovering causal networks closer to the true underlying dynamics. Nicholas Tagliapietra, Katharina Ensinger, Christoph Zimmer, Osman Mian |
AAAI | 4 |
| 2024 | Learning Causal Networks from Episodic DataabstractIn numerous real-world domains, spanning from environmental monitoring to long-term medical studies, observations do not arrive in a single batch but rather over time in episodes. This challenges the traditional assumption in causal discovery of a single, observational dataset, not only because each episode may be a biased sample of the population but also because multiple episodes could differ in the causal interactions underlying the observed variables. We address these issues using notions of context switches and episodic selection bias, and introduce a framework for causal modeling of episodic data. We show under which conditions we can apply information-theoretic scoring criteria for causal discovery while preserving consistency. To in practice discover the causal model progressively over time, we propose the CONTINENT algorithm which, taking inspiration from continual learning, discovers the causal model in an online fashion without having to re-learn the model upon arrival of each new episode. Our experiments over a variety of settings including selection bias, unknown interventions, and network changes showcase that CONTINENT works well in practice and outperforms the baselines by a clear margin. Osman Mian, Sarah Mameche, Jilles Vreeken |
KDD | 1 |
| 2023 | Information-Theoretic Causal Discovery and Intervention Detection over Multiple EnvironmentsabstractGiven multiple datasets over a fixed set of random variables, each collected from a different environment, we are interested in discovering the shared underlying causal network and the local interventions per environment, without assuming prior knowledge on which datasets are observational or interventional, and without assuming the shape of the causal dependencies. We formalize this problem using the Algorithmic Model of Causation, instantiate a consistent score via the Minimum Description Length principle, and show under which conditions the network and interventions are identifiable. To efficiently discover causal networks and intervention targets in practice, we introduce the ORION algorithm, which through extensive experiments we show outperforms the state of the art in causal inference over multiple environments. Osman Mian, Michael Kamp, Jilles Vreeken |
AAAI | 1 |
| 2023 | Nothing but Regrets - Privacy-Preserving Federated Causal DiscoveryabstractIn critical applications, causal models are the prime choice for their trustworthiness and explainability. If data is inherently distributed and privacy-sensitive, federated learning allows for collaboratively training a joint model. Existing approaches for federated causal discovery share locally discovered causal model in every iteration, therewith not only revealing local structure but also leading to very high communication costs. Instead, we propose an approach for privacy-preserving federated causal discovery by distributed min-max regret optimization. We prove that max-regret is a consistent scoring criterion that can be used within the well-known Greedy Equivalence Search to discover causal networks in a federated setting and is provably privacy-preserving at the same time. Through extensive experiments, we show that our approach reliably discovers causal networks without ever looking at local data and beats the state of the art both in terms of the quality of discovered causal networks as well as communication efficiency. Osman Mian, David Kaltenpoth, Michael Kamp, Jilles Vreeken |
AISTATS | 1 |
| 2022 | Inferring Cause and Effect in the Presence of Heteroscedastic NoiseabstractWe study the problem of identifying cause and effect over two univariate continuous variables $X$ and $Y$ from a sample of their joint distribution. Our focus lies on the setting when the variance of the noise may be dependent on the cause. We propose to partition the domain of the cause into multiple segments where the noise indeed is dependent. To this end, we minimize a scale-invariant, penalized regression score, finding the optimal partitioning using dynamic programming. We show under which conditions this allows us to identify the causal direction for the linear setting with heteroscedastic noise, for the non-linear setting with homoscedastic noise, as well as empirically confirm that these results generalize to the non-linear and heteroscedastic case. Altogether, the ability to model heteroscedasticity translates into an improved performance in telling cause from effect on a wide range of synthetic and real-world datasets. Sascha Xu, Osman Mian, Alexander Marx 0001, Jilles Vreeken |
ICML | 2 |
| 2021 | Discovering Fully Oriented Causal NetworksabstractWe study the problem of inferring causal graphs from observational data. We are particularly interested in discovering graphs where all edges are oriented, as opposed to the partially directed graph that the state of the art discover. To this end, we base our approach on the algorithmic Markov condition. Unlike the statistical Markov condition, it uniquely identifies the true causal network as the one that provides the simplest— as measured in Kolmogorov complexity—factorization of the joint distribution. Although Kolmogorov complexity is not computable, we can approximate it from above via the Minimum Description Length principle, which allows us to define a consistent and computable score based on non-parametric multivariate regression. To efficiently discover causal networks in practice, we introduce the GLOBE algorithm, which greedily adds, removes, and orients edges such that it minimizes the overall cost. Through an extensive set of experiments, we show GLOBE performs very well in practice, beating the state of the art by a margin. Osman Mian, Alexander Marx 0001, Jilles Vreeken |
AAAI | 1 |