EDBT 2026 Demo / reviewers in the wild / expert
Kailash Budhathoki
dblp:167/3244
· DBLP profile ↗
15ranked-venue papers
10as first author
8since 2021 · last 2025
0000-0002-5255-8642ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 9 · 8 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Probabilistic and Bayesian machine learning · 40% Knowledge representation and reasoning · 35% Efficient and distributed learning · 20% | |
| Databases, data mining, and information retrieval
4 papers |
Data mining · 100% | |
| Theoretical computer science
3 papers |
Information theory · 67% Mathematical optimization · 33% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 100% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
causal inference |
0.9 | 3 | 2018 | Accurate Causal Inference on Discrete Data · ICDM 2018 MDL for Causal Inference on Discrete Data · ICDM 2017 Causal Inference by Compression · ICDM 2016 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
causal graphical model |
0.8 | 1 | 2024 | DoWhy-GCM: An Extension of DoWhy for Causal Inference in Graphical Causal Models · J. Mach. Learn. Res. 2024 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.8 | 1 | 2024 | DoWhy-GCM: An Extension of DoWhy for Causal Inference in Graphical Causal Models · J. Mach. Learn. Res. 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | Inference Optimization of Foundation Models on AI Accelerators · KDD 2024 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
root cause analysis |
0.8 | 1 | 2024 | DoWhy-GCM: An Extension of DoWhy for Causal Inference in Graphical Causal Models · J. Mach. Learn. Res. 2024 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.8 | 1 | 2024 | Inference Optimization of Foundation Models on AI Accelerators · KDD 2024 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
transformer inference accelerator |
0.8 | 1 | 2024 | Inference Optimization of Foundation Models on AI Accelerators · KDD 2024 |
Mathematical optimization
causal inference |
0.6 | 2 | 2018 | Accurate Causal Inference on Discrete Data · ICDM 2018 Causal Inference by Compression · ICDM 2016 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning |
0.6 | 1 | 2022 | Causal structure-based root cause analysis of outliers · ICML 2022 |
Data mining
anomaly detection |
0.6 | 1 | 2022 | Causal structure-based root cause analysis of outliers · ICML 2022 |
Data mining › anomaly detection
outlier explanation |
0.6 | 1 | 2022 | Causal structure-based root cause analysis of outliers · ICML 2022 |
Data mining › causal inference › causal modeling
causal discovery |
0.5 | 2 | 2017 | MDL for Causal Inference on Discrete Data · ICDM 2017 Causal Inference by Compression · ICDM 2016 |
Information theory
minimum description length |
0.5 | 2 | 2017 | MDL for Causal Inference on Discrete Data · ICDM 2017 Causal Inference by Compression · ICDM 2016 |
Information theory › information measures › entropy
shannon entropy |
0.3 | 1 | 2018 | Accurate Causal Inference on Discrete Data · ICDM 2018 |
Information theory › algorithmic information theory
stochastic complexity |
0.3 | 1 | 2017 | MDL for Causal Inference on Discrete Data · ICDM 2017 |
Machine learning › Deep learning architectures and training
transformer |
0.2 | 1 | 2024 | Inference Optimization of Foundation Models on AI Accelerators · KDD 2024 |
Methods — techniques the papers use, named apart from their topics
causal graph · 1.9quantization · 1.5attention computation optimization · 1.5functional causal model · 1.1minimum description length · 1.1kolmogorov complexity · 1.1causal mechanism fitting · 0.8shannon entropy · 0.7additive noise model · 0.7compression · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Block-Diagonal LoRA for Eliminating Communication Overhead in Tensor Parallel LoRA ServingabstractWhen serving a single base LLM with several different LoRA adapters simultaneously, the adapters cannot simply be merged with the base model’s weights as the adapter swapping would create overhead and requests using different adapters could not be batched. Rather, the LoRA computations have to be separated from the base LLM computations, and in a multi-device setup the LoRA adapters can be sharded in a way that is well aligned with the base model’s tensor parallel execution, as proposed in S-LoRA. However, the S-LoRA sharding strategy encounters some communication overhead, which may be small in theory, but can be large in practice. In this paper, we propose to constrain certain LoRA factors to be block-diagonal, which allows for an alternative way of sharding LoRA adapters that does not require any additional communication for the LoRA computations. We demonstrate in extensive experiments that our block-diagonal LoRA approach is similarly parameter efficient as standard LoRA (i.e., for a similar number of parameters it achieves similar downstream performance) and that it leads to significant end-to-end speed-up over S-LoRA. For example, when serving on eight A100 GPUs, we observe up to 1.79x (1.23x) end-to-end speed-up with 0.87x (1.74x) the number of adapter parameters for Llama-3.1-70B, and up to 1.63x (1.3x) end-to-end speed-up with 0.86x (1.73x) the number of adapter parameters for Llama-3.1-8B. Xinyu Wang 0062, Jonas M. Kübler, Kailash Budhathoki, Yida Wang 0003, Matthäus Kleindessner |
NeurIPS | 3 |
| 2024 | Quantifying intrinsic causal contributions via structure preserving interventionsabstractWe propose a notion of causal influence that describes the ‘intrinsic’ part of the contribution of a node on a target node in a DAG. By recursively writing each node as a function of the upstream noise terms, we separate the intrinsic information added by each node from the one obtained from its ancestors. To interpret the intrinsic information as a causal contribution, we consider ‘structure-preserving interventions’ that randomize each node in a way that mimics the usual dependence on the parents and does not perturb the observed joint distribution. To get a measure that is invariant across arbitrary orderings of nodes we use Shapley based symmetrization and show that it reduces in the linear case to simple ANOVA after resolving the target node into noise variables. We describe our contribution analysis for variance and entropy, but contributions for other target metrics can be defined analogously. Dominik Janzing, Patrick Blöbaum, Atalanti-Anastasia Mastakouri, Philipp Michael Faller, Lenon Minorics, Kailash Budhathoki |
AISTATS | 6 |
| 2024 | Inference Optimization of Foundation Models on AI AcceleratorsabstractPowerful foundation models, including large language models (LLMs), with Transformer architectures have ushered in a new era of Generative AI across various industries. Industry and research community have witnessed a large number of new applications, based on those foundation models. Such applications include question and answer, customer services, image and video generation, and code completions, among others. However, as the number of model parameters reaches to hundreds of billions, their deployment incurs prohibitive inference costs and high latency in real-world scenarios. As a result, the demand for cost-effective and fast inference using AI accelerators is ever more higher. To this end, our tutorial offers a comprehensive discussion on complementary inference optimization techniques using AI accelerators. Beginning with an overview of basic Transformer architectures and deep learning system frameworks, we deep dive into system optimization techniques for fast and memory-efficient attention computations and discuss how they can be implemented efficiently on AI accelerators. Next, we describe architectural elements that are key for fast transformer inference. Finally, we examine various model compression and fast decoding strategies in the same context. Youngsuk Park, Kailash Budhathoki, Liangfu Chen, Jonas M. Kübler, Jiaji Huang, Matthäus Kleindessner, Jun Huan, Volkan Cevher, Yida Wang 0003, George Karypis |
KDD | 2 |
| 2024 | DoWhy-GCM: An Extension of DoWhy for Causal Inference in Graphical Causal ModelsabstractWe present DoWhy-GCM, an extension of the DoWhy Python library, which leverages graphical causal models. Unlike existing causality libraries, which mainly focus on effect estimation, DoWhy-GCM addresses diverse causal queries, such as identifying the root causes of outliers and distributional changes, attributing causal influences to the data generating process of each node, or diagnosis of causal structures. With DoWhy-GCM, users typically specify cause-effect relations via a causal graph, fit causal mechanisms, and pose causal queries---all with just a few lines of code. The general documentation is available at https://www.pywhy.org/dowhy and the DoWhy-GCM specific code at https://github.com/py-why/dowhy/tree/main/dowhy/gcm. Patrick Blöbaum, Peter Götz, Kailash Budhathoki, Atalanti-Anastasia Mastakouri, Dominik Janzing |
J. Mach. Learn. Res. | 3 |
| 2023 | Evaluating the Fairness of Discriminative Foundation Models in Computer VisionabstractWe propose a novel taxonomy for bias evaluation of discriminative foundation models, such as Contrastive Language-Pretraining (CLIP), that are used for labeling tasks. We then systematically evaluate existing methods for mitigating bias in these models with respect to our taxonomy. Specifically, we evaluate OpenAI’s CLIP and OpenCLIP models for key applications, such as zero-shot classification, image retrieval and image captioning. We categorize desired behaviors based around three axes: (i) if the task concerns humans; (ii) how subjective the task is (i.e., how likely it is that people from a diverse range of backgrounds would agree on a labeling); and (iii) the intended purpose of the task and if fairness is better served by impartiality (i.e., making decisions independent of the protected attributes) or representation (i.e., making decisions to maximize diversity). Finally, we provide quantitative fairness evaluations for both binary-valued and multi-valued protected attributes over ten diverse datasets. We find that fair PCA, a post-processing method for fair representations, works very well for debiasing in most of the aforementioned tasks while incurring only minor loss of performance. However, different debiasing approaches vary in their effectiveness depending on the task. Hence, one should choose the debiasing approach depending on the specific use case. Junaid Ali 0001, Matthäus Kleindessner, Florian Wenzel, Kailash Budhathoki, Volkan Cevher, Chris Russell 0001 |
AIES | 4 |
| 2022 | Causal structure-based root cause analysis of outliersabstractCurrent techniques for explaining outliers cannot tell what caused the outliers. We present a formal method to identify "root causes" of outliers, amongst variables. The method requires a causal graph of the variables along with the functional causal model. It quantifies the contribution of each variable to the target outlier score, which explains to what extent each variable is a "root cause" of the target outlier. We study the empirical performance of the method through simulations and present a real-world case study identifying "root causes" of extreme river flows. Kailash Budhathoki, Lenon Minorics, Patrick Blöbaum, Dominik Janzing |
ICML | 1 |
| 2021 | Why did the distribution change?abstractWe describe a formal approach based on graphical causal models to identify the "root causes" of the change in the probability distribution of variables. After factorizing the joint distribution into conditional distributions of each variable, given its parents (the "causal mechanisms"), we attribute the change to changes of these causal mechanisms. This attribution analysis accounts for the fact that mechanisms often change independently and sometimes only some of them change. Through simulations, we study the performance of our distribution change attribution proposal. We then present a real-world case study identifying the drivers of the difference in the income distribution between men and women. Kailash Budhathoki, Dominik Janzing, Patrick Blöbaum, Hoiyi Ng |
AISTATS | 1 |
| 2021 | Discovering Reliable Causal RulesabstractWe study the problem of deriving policies, or rules, that when enacted on a complex system, cause a desired outcome. Absent the ability to perform controlled experiments, such rules have to be inferred from past observations of the system's behaviour. This is a challenging problem for two reasons: First, observational effects are often unrepresentative of the underlying causal effect because they are skewed by the presence of confounding factors. Second, naive empirical estimations of a rule's effect have a high variance, and, hence, their maximisation can lead to random results. To address these issues, first we measure the causal effect of a rule from observational data---adjusting for the effect of potential confounders. Importantly, we provide a graphical criteria under which causal rule discovery is possible. Moreover, to discover reliable causal rules from a sample, we propose a conservative and consistent estimator of the causal effect, and derive an efficient and exact algorithm that maximises the estimator. On synthetic data, the proposed estimator converges faster to the ground truth than the naive estimator and recovers relevant causal rules even at small sample sizes. Extensive experiments on a variety of real-world datasets show that the proposed algorithm is efficient and discovers meaningful rules. Kailash Budhathoki, Mario Boley, Jilles Vreeken |
SDM | 1 |
| 2018 | Accurate Causal Inference on Discrete DataabstractAdditive Noise Models (ANMs) provide a theoretically sound approach to inferring the most likely causal direction between pairs of random variables given only a sample from their joint distribution. The key assumption is that the effect is a function of the cause, with additive noise that is independent of the cause. In many cases ANMs are identifiable. Their performance, however, hinges on the chosen dependence measure, the assumption we make on the true distribution. In this paper we propose to use Shannon entropy to measure the dependence within an ANM, which gives us a general approach by which we do not have to assume a true distribution, nor have to perform explicit significance tests during optimization. The information-theoretic formulation gives us a general, efficient, identifiable, and, as the experiments show, highly accurate method for causal inference on pairs of discrete variables-achieving (near) 100% accuracy on both synthetic and real data. Kailash Budhathoki, Jilles Vreeken |
ICDM | 1 |
| 2018 | Causal Inference on Event SequencesabstractGiven two discrete valued time series—that is, event sequences—of length n can we tell whether they are causally related? That is, can we tell whether xn causes yn, whether yn causes xn? Can we do so without having to make assumptions on the distribution of these time series, or about the lag of the causal effect? And, importantly for practical application, can we do so accurately and efficiently? These are exactly the questions we answer in this paper. We propose a causal inference framework for event sequences based on information theory. We build upon the well-known notion of Granger causality, and define causality in terms of compression. We infer that xn is likely a cause of yn if yn can be (much) better sequentially compressed given the past of both yn and xn, than for the other way around. To compress the data we use the notion of sequential normalized maximal likelihood, which means we use minimax optimal codes with respect to a parametric family of distributions. To show this works in practice, we propose CUTE, a linear time method for inferring the causal direction between two event sequences. Empirical evaluation shows that CUTE works well in practice, is much more robust than transfer entropy, and ably reconstructs the ground truth on river flow and spike train data. Kailash Budhathoki, Jilles Vreeken |
SDM | 1 |
| 2018 | Origo: causal inference by compressionabstractCausal inference from observational data is one of the most fundamental problems in science. In general, the task is to tell whether it is more likely that $$X$$ caused $$Y$$ , or vice versa, given only data over their joint distribution. In this paper we propose a general inference framework based on Kolmogorov complexity, as well as a practical and computable instantiation based on the Minimum Description Length principle. Simply put, we propose causal inference by compression. That is, we infer that $$X$$ is a likely cause of $$Y$$ if we can better compress the data by first encoding $$X$$ , and then encoding $$Y$$ given $$X$$ , than in the other direction. To show this works in practice, we propose Origo, an efficient method for inferring the causal direction from binary data. Origo employs the lossless Pack compressor and searches for that set of decision trees that encodes the data most succinctly. Importantly, it works directly on the data and does not require assumptions about neither distributions nor the type of causal relations. To evaluate Origo in practice, we provide extensive experiments on synthetic, benchmark, and real-world data, including three case studies. Altogether, the experiments show that Origo reliably infers the correct causal direction on a wide range of settings. Kailash Budhathoki, Jilles Vreeken |
Knowl. Inf. Syst. | 1 |
| 2017 | MDL for Causal Inference on Discrete DataabstractThe algorithmic Markov condition states that the most likely causal direction between two random variables X and Y can be identified as the direction with the lowest Kolmogorov complexity. This notion is very powerful as it can detect any causal dependency that can be explained by a physical process. However, due to the halting problem, it is also not computable. In this paper we propose an computable instantiation that provably maintains the key aspects of the ideal. We propose to approximate Kolmogorov complexity via the Minimum Description Length (MDL) principle, using a score that is mini-max optimal with regard to the model class under consideration. This means that even in an adversarial setting, the score degrades gracefully, and we are still maximally able to detect dependencies between the marginal and the conditional distribution. As a proof of concept, we propose CISC, a linear-time algorithm for causal inference by stochastic complexity, for pairs of univariate discrete variables. Experiments show that CISC is highly accurate on synthetic, benchmark, as well as real-world data, outperforming the state of the art by a margin, and scales extremely well with regard to sample and domain sizes. Kailash Budhathoki, Jilles Vreeken |
ICDM | 1 |
| 2017 | Correlation by CompressionabstractDiscovering correlated variables is one of the core problems in data analysis. Many measures for correlation have been proposed, yet it is surprisingly ill-defined in general. That is, most, if not all, measures make very strong assumptions on the data distribution or type of dependency they can detect. In this work, we provide a general theory on correlation, without making any such assumptions. Simply put, we propose correlation by compression. To this end, we propose two correlation measures based on solid information theoretic foundations, i.e. Kolmogorov complexity. The proposed correlation measures possess interesting properties desirable for any sensible correlation measure. However, Kolmogorov complexity is not computable, and hence we propose practical and computable instantiations based on the Minimum Description Length (MDL) principle. In practice, we can apply the proposed measures on any type of data by instantiating them with any lossless real-world compressors that reward pairwise dependencies. Extensive experiments show that the correlation measures works well in practice, have high statistical power, and find meaningful correlations on binary data, while they are easily extendible to other data types. Kailash Budhathoki, Jilles Vreeken |
SDM | 1 |
| 2016 | Causal Inference by CompressionabstractCausal inference is one of the fundamental problems in science. In recent years, several methods have been proposed for discovering causal structure from observational data. These methods, however, focus specifically on numeric data, and are not applicable on nominal or binary data. In this work, we focus on causal inference for binary data. Simply put, we propose causal inference by compression. To this end we propose an inference framework based on solid information theoretic foundations, i.e. Kolmogorov complexity. However, Kolmogorov complexity is not computable, and hence we propose a practical and computable instantiation based on the Minimum Description Length (MDL) principle. To apply the framework in practice, we propose ORIGO, an efficient method for inferring the causal direction from binary data. ORIGO employs the lossless PACK compressor, works directly on the data and does not require assumptions about neither distributions nor the type of causal relations. Extensive evaluation on synthetic, benchmark, and real-world data shows that ORIGO discovers meaningful causal relations, and outperforms state-of-the-art methods by a wide margin. Kailash Budhathoki, Jilles Vreeken |
ICDM | 1 |
| 2015 | The Difference and the Norm - Characterising Similarities and Differences Between Databases
Kailash Budhathoki, Jilles Vreeken |
ECML/PKDD (2) | 1 |