Osman Mian

dblp:294/0815 · also Osman Ali Mian · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
0009-0006-1112-6145ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Probabilistic and Bayesian machine learning · 94% Learning paradigms · 3% Trustworthy machine learning · 3%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
2.942026
Causal Structure Learning for Dynamical Systems with Theoretical Score Analysis · AAAI 2026
Learning Causal Networks from Episodic Data · KDD 2024
Information-Theoretic Causal Discovery and Intervention Detection over Multiple Environments · AAAI 2023
Machine learning › Probabilistic and Bayesian machine learning
causal inference
1.622026
Causal Structure Learning for Dynamical Systems with Theoretical Score Analysis · AAAI 2026
Inferring Cause and Effect in the Presence of Heteroscedastic Noise · ICML 2022
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
1.012026
Causal Structure Learning for Dynamical Systems with Theoretical Score Analysis · AAAI 2026
Data mining
pattern mining
1.012026
SEQRET: Mining Rule Sets from Event Sequences · AAAI 2026
Data mining › pattern mining
sequential pattern mining
1.012026
SEQRET: Mining Rule Sets from Event Sequences · AAAI 2026
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
multi-environment causal discovery
0.712023
Information-Theoretic Causal Discovery and Intervention Detection over Multiple Environments · AAAI 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › probabilistic regression
heteroscedastic modeling
0.612022
Inferring Cause and Effect in the Presence of Heteroscedastic Noise · ICML 2022
Machine learning › Learning paradigms
continual learning
0.212024
Learning Causal Networks from Episodic Data · KDD 2024
Machine learning › Trustworthy machine learning › dataset bias
selection bias
0.212024
Learning Causal Networks from Episodic Data · KDD 2024

Methods — techniques the papers use, named apart from their topics

minimum description length · 3.2greedy search · 1.0gaussian process inference · 1.0information-theoretic scoring · 0.8continual learning · 0.8algorithmic model of causation · 0.7penalized regression score · 0.6dynamic programming · 0.6non-parametric multivariate regression · 0.5
YearPublicationVenuePosition
2026 SEQRET: Mining Rule Sets from Event Sequences
abstract
Summarizing event sequences is a key aspect of data mining. Most existing methods neglect conditional dependencies and focus on discovering sequential patterns only. In this paper, we study the problem of discovering both conditional and unconditional dependencies from event sequences. We do so by discovering rules of the form X --> Y where X and Y are sequential patterns. Rules like these are simple to understand and provide a clear description of the relation between the antecedent and the consequent. To discover succinct and non-redundant sets of rules we formalize the problem in terms of the Minimum Description Length principle. As the search space is enormous and does not exhibit helpful structure, we propose the SEQRET method to discover high-quality rule sets in practice. Through extensive empirical evaluation we show that unlike the state of the art, SEQRET ably recovers the ground truth on synthetic datasets and finds useful rules from real datasets.
Aleena Siji, Joscha Cüppers, Osman Mian, Jilles Vreeken
AAAI3
2026 Causal Structure Learning for Dynamical Systems with Theoretical Score Analysis
abstract
Real world systems evolve in continuous-time according to their underlying causal relationships, yet their dynamics are often unknown. Existing approaches to learning such dynamics typically either discretize time ---leading to poor performance on irregularly sampled data--- or ignore the underlying causality. We propose CADYT, a novel method for causal discovery on dynamical systems addressing both these challenges. In contrast to state-of-the-art causal discovery methods that model the problem using discrete-time Dynamic Bayesian networks, our formulation is grounded in Difference-based causal models, which allow milder assumptions for modeling the continuous nature of the system. CADYT leverages exact Gaussian Process inference for modeling the continuous-time dynamics which is more aligned with the underlying dynamical process. We propose a practical instantiation that identifies the causal structure via a greedy search guided by the Algorithmic Markov Condition and Minimum Description Length principle. Our experiments show that CADYT outperforms state-of-the-art methods on both regularly and irregularly-sampled data, discovering causal networks closer to the true underlying dynamics.
Nicholas Tagliapietra, Katharina Ensinger, Christoph Zimmer, Osman Mian
AAAI4
2024 Learning Causal Networks from Episodic Data
abstract
In numerous real-world domains, spanning from environmental monitoring to long-term medical studies, observations do not arrive in a single batch but rather over time in episodes. This challenges the traditional assumption in causal discovery of a single, observational dataset, not only because each episode may be a biased sample of the population but also because multiple episodes could differ in the causal interactions underlying the observed variables. We address these issues using notions of context switches and episodic selection bias, and introduce a framework for causal modeling of episodic data. We show under which conditions we can apply information-theoretic scoring criteria for causal discovery while preserving consistency. To in practice discover the causal model progressively over time, we propose the CONTINENT algorithm which, taking inspiration from continual learning, discovers the causal model in an online fashion without having to re-learn the model upon arrival of each new episode. Our experiments over a variety of settings including selection bias, unknown interventions, and network changes showcase that CONTINENT works well in practice and outperforms the baselines by a clear margin.
Osman Mian, Sarah Mameche, Jilles Vreeken
KDD1
2023 Information-Theoretic Causal Discovery and Intervention Detection over Multiple Environments
abstract
Given multiple datasets over a fixed set of random variables, each collected from a different environment, we are interested in discovering the shared underlying causal network and the local interventions per environment, without assuming prior knowledge on which datasets are observational or interventional, and without assuming the shape of the causal dependencies. We formalize this problem using the Algorithmic Model of Causation, instantiate a consistent score via the Minimum Description Length principle, and show under which conditions the network and interventions are identifiable. To efficiently discover causal networks and intervention targets in practice, we introduce the ORION algorithm, which through extensive experiments we show outperforms the state of the art in causal inference over multiple environments.
Osman Mian, Michael Kamp, Jilles Vreeken
AAAI1
2023 Nothing but Regrets - Privacy-Preserving Federated Causal Discovery
abstract
In critical applications, causal models are the prime choice for their trustworthiness and explainability. If data is inherently distributed and privacy-sensitive, federated learning allows for collaboratively training a joint model. Existing approaches for federated causal discovery share locally discovered causal model in every iteration, therewith not only revealing local structure but also leading to very high communication costs. Instead, we propose an approach for privacy-preserving federated causal discovery by distributed min-max regret optimization. We prove that max-regret is a consistent scoring criterion that can be used within the well-known Greedy Equivalence Search to discover causal networks in a federated setting and is provably privacy-preserving at the same time. Through extensive experiments, we show that our approach reliably discovers causal networks without ever looking at local data and beats the state of the art both in terms of the quality of discovered causal networks as well as communication efficiency.
Osman Mian, David Kaltenpoth, Michael Kamp, Jilles Vreeken
AISTATS1
2022 Inferring Cause and Effect in the Presence of Heteroscedastic Noise
abstract
We study the problem of identifying cause and effect over two univariate continuous variables $X$ and $Y$ from a sample of their joint distribution. Our focus lies on the setting when the variance of the noise may be dependent on the cause. We propose to partition the domain of the cause into multiple segments where the noise indeed is dependent. To this end, we minimize a scale-invariant, penalized regression score, finding the optimal partitioning using dynamic programming. We show under which conditions this allows us to identify the causal direction for the linear setting with heteroscedastic noise, for the non-linear setting with homoscedastic noise, as well as empirically confirm that these results generalize to the non-linear and heteroscedastic case. Altogether, the ability to model heteroscedasticity translates into an improved performance in telling cause from effect on a wide range of synthetic and real-world datasets.
Sascha Xu, Osman Mian, Alexander Marx 0001, Jilles Vreeken
ICML2
2021 Discovering Fully Oriented Causal Networks
abstract
We study the problem of inferring causal graphs from observational data. We are particularly interested in discovering graphs where all edges are oriented, as opposed to the partially directed graph that the state of the art discover. To this end, we base our approach on the algorithmic Markov condition. Unlike the statistical Markov condition, it uniquely identifies the true causal network as the one that provides the simplest— as measured in Kolmogorov complexity—factorization of the joint distribution. Although Kolmogorov complexity is not computable, we can approximate it from above via the Minimum Description Length principle, which allows us to define a consistent and computable score based on non-parametric multivariate regression. To efficiently discover causal networks in practice, we introduce the GLOBE algorithm, which greedily adds, removes, and orients edges such that it minimizes the overall cost. Through an extensive set of experiments, we show GLOBE performs very well in practice, beating the state of the art by a margin.
Osman Mian, Alexander Marx 0001, Jilles Vreeken
AAAI1