Jonas M. Kübler

dblp:241/9893 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-8634-7990ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Learning theory · 52% Efficient and distributed learning · 21% Probabilistic and Bayesian machine learning · 9%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 100%
Theoretical computer science
1 paper
Quantum computing and quantum information · 100%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
hypothesis testing
1.022022
AutoML Two-Sample Test · NeurIPS 2022
Learning Kernel Tests Without Data Splitting · NeurIPS 2020
Machine learning › Efficient and distributed learning
model compression
0.812024
Inference Optimization of Foundation Models on AI Accelerators · KDD 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.812024
Inference Optimization of Foundation Models on AI Accelerators · KDD 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
transformer inference accelerator
0.812024
Inference Optimization of Foundation Models on AI Accelerators · KDD 2024
Machine learning › Efficient and distributed learning
automated machine learning
0.612022
AutoML Two-Sample Test · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.612022
Causal Inference Through the Structural Causal Marginal Problem · ICML 2022
Machine learning › Learning theory › hypothesis testing
two-sample testing
0.612022
AutoML Two-Sample Test · NeurIPS 2022
Machine learning › Learning theory
inductive bias
0.512021
The Inductive Bias of Quantum Kernels · NeurIPS 2021
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.512021
The Inductive Bias of Quantum Kernels · NeurIPS 2021
Quantum computing and quantum information
quantum machine learning
0.512021
The Inductive Bias of Quantum Kernels · NeurIPS 2021
Machine learning › Learning theory › hypothesis testing › two-sample testing
kernel two-sample test
0.412020
Learning Kernel Tests Without Data Splitting · NeurIPS 2020
Machine learning › Learning theory › probability metric › integral probability metric
maximum mean discrepancy
0.412020
Learning Kernel Tests Without Data Splitting · NeurIPS 2020
Machine learning › Deep learning architectures and training
transformer
0.212024
Inference Optimization of Foundation Models on AI Accelerators · KDD 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
counterfactual reasoning
0.212022
Causal Inference Through the Structural Causal Marginal Problem · ICML 2022
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.212022
AutoML Two-Sample Test · NeurIPS 2022
Machine learning › Learning theory
generalization bounds
0.112021
The Inductive Bias of Quantum Kernels · NeurIPS 2021
Machine learning › Learning theory › hypothesis testing
selective inference
0.112020
Learning Kernel Tests Without Data Splitting · NeurIPS 2020

Methods — techniques the papers use, named apart from their topics

quantization · 1.5attention computation optimization · 1.5reproducing kernel hilbert space · 1.0quantum kernel · 1.0witness function · 0.6structural causal model · 0.6response function · 0.6mean discrepancy · 0.6AutoML · 0.6data splitting · 0.4
YearPublicationVenuePosition
2025 Block-Diagonal LoRA for Eliminating Communication Overhead in Tensor Parallel LoRA Serving
abstract
When serving a single base LLM with several different LoRA adapters simultaneously, the adapters cannot simply be merged with the base model’s weights as the adapter swapping would create overhead and requests using different adapters could not be batched. Rather, the LoRA computations have to be separated from the base LLM computations, and in a multi-device setup the LoRA adapters can be sharded in a way that is well aligned with the base model’s tensor parallel execution, as proposed in S-LoRA. However, the S-LoRA sharding strategy encounters some communication overhead, which may be small in theory, but can be large in practice. In this paper, we propose to constrain certain LoRA factors to be block-diagonal, which allows for an alternative way of sharding LoRA adapters that does not require any additional communication for the LoRA computations. We demonstrate in extensive experiments that our block-diagonal LoRA approach is similarly parameter efficient as standard LoRA (i.e., for a similar number of parameters it achieves similar downstream performance) and that it leads to significant end-to-end speed-up over S-LoRA. For example, when serving on eight A100 GPUs, we observe up to 1.79x (1.23x) end-to-end speed-up with 0.87x (1.74x) the number of adapter parameters for Llama-3.1-70B, and up to 1.63x (1.3x) end-to-end speed-up with 0.86x (1.73x) the number of adapter parameters for Llama-3.1-8B.
Xinyu Wang 0062, Jonas M. Kübler, Kailash Budhathoki, Yida Wang 0003, Matthäus Kleindessner
NeurIPS2
2024 Inference Optimization of Foundation Models on AI Accelerators
abstract
Powerful foundation models, including large language models (LLMs), with Transformer architectures have ushered in a new era of Generative AI across various industries. Industry and research community have witnessed a large number of new applications, based on those foundation models. Such applications include question and answer, customer services, image and video generation, and code completions, among others. However, as the number of model parameters reaches to hundreds of billions, their deployment incurs prohibitive inference costs and high latency in real-world scenarios. As a result, the demand for cost-effective and fast inference using AI accelerators is ever more higher. To this end, our tutorial offers a comprehensive discussion on complementary inference optimization techniques using AI accelerators. Beginning with an overview of basic Transformer architectures and deep learning system frameworks, we deep dive into system optimization techniques for fast and memory-efficient attention computations and discuss how they can be implemented efficiently on AI accelerators. Next, we describe architectural elements that are key for fast transformer inference. Finally, we examine various model compression and fast decoding strategies in the same context.
Youngsuk Park, Kailash Budhathoki, Liangfu Chen, Jonas M. Kübler, Jiaji Huang, Matthäus Kleindessner, Jun Huan, Volkan Cevher, Yida Wang 0003, George Karypis
KDD4
2022 A Witness Two-Sample Test
abstract
The Maximum Mean Discrepancy (MMD) has been the state-of-the-art nonparametric test for tackling the two-sample problem. Its statistic is given by the difference in expectations of the witness function, a real-valued function defined as a weighted sum of kernel evaluations on a set of basis points. Typically the kernel is optimized on a training set, and hypothesis testing is performed on a separate test set to avoid overfitting (i.e., control type-I error). That is, the test set is used to simultaneously estimate the expectations and define the basis points, while the training set only serves to select the kernel and is discarded. In this work, we propose to use the training data to also define the weights and the basis points for better data efficiency. We show that 1) the new test is consistent and has a well-controlled type-I error; 2) the optimal witness function is given by a precision-weighted mean in the reproducing kernel Hilbert space associated with the kernel; and 3) the test power of the proposed test is comparable or exceeds that of the MMD and other modern tests, as verified empirically on challenging synthetic and real problems (e.g., Higgs data).
Jonas M. Kübler, Wittawat Jitkrittum, Bernhard Schölkopf, Krikamol Muandet
AISTATS1
2022 Causal Inference Through the Structural Causal Marginal Problem
abstract
We introduce an approach to counterfactual inference based on merging information from multiple datasets. We consider a causal reformulation of the statistical marginal problem: given a collection of marginal structural causal models (SCMs) over distinct but overlapping sets of variables, determine the set of joint SCMs that are counterfactually consistent with the marginal ones. We formalise this approach for categorical SCMs using the response function formulation and show that it reduces the space of allowed marginal and joint SCMs. Our work thus highlights a new mode of falsifiability through additional variables, in contrast to the statistical one via additional data.
Luigi Gresele, Julius von Kügelgen, Jonas M. Kübler, Elke Kirschbaum, Bernhard Schölkopf, Dominik Janzing
ICML3
2022 AutoML Two-Sample Test
abstract
Two-sample tests are important in statistics and machine learning, both as tools for scientific discovery as well as to detect distribution shifts.This led to the development of many sophisticated test procedures going beyond the standard supervised learning frameworks, whose usage can require specialized knowledge about two-sample testing. We use a simple test that takes the mean discrepancy of a witness function as the test statistic and prove that minimizing a squared loss leads to a witness with optimal testing power. This allows us to leverage recent advancements in AutoML. Without any user input about the problems at hand, and using the same method for all our experiments, our AutoML two-sample test achieves competitive performance on a diverse distribution shift benchmark as well as on challenging two-sample testing problems.
Jonas M. Kübler, Vincent Stimper, Simon Buchholz, Krikamol Muandet, Bernhard Schölkopf
NeurIPS1
2021 The Inductive Bias of Quantum Kernels
abstract
It has been hypothesized that quantum computers may lend themselves well to applications in machine learning. In the present work, we analyze function classes defined via quantum kernels. Quantum computers offer the possibility to efficiently compute inner products of exponentially large density operators that are classically hard to compute. However, having an exponentially large feature space renders the problem of generalization hard. Furthermore, being able to evaluate inner products in high dimensional spaces efficiently by itself does not guarantee a quantum advantage, as already classically tractable kernels can correspond to high- or infinite-dimensional reproducing kernel Hilbert spaces (RKHS). We analyze the spectral properties of quantum kernels and find that we can expect an advantage if their RKHS is low dimensional and contains functions that are hard to compute classically. If the target function is known to lie in this class, this implies a quantum advantage, as the quantum computer can encode this inductive bias, whereas there is no classically efficient way to constrain the function class in the same way. However, we show that finding suitable quantum kernels is not easy because the kernel evaluation might require exponentially many measurements. In conclusion, our message is a somewhat sobering one: we conjecture that quantum machine learning models can offer speed-ups only if we manage to encode knowledge about the problem at hand into quantum circuits, while encoding the same bias into a classical model would be hard. These situations may plausibly occur when learning on data generated by a quantum process, however, they appear to be harder to come by for classical datasets.
Jonas M. Kübler, Simon Buchholz, Bernhard Schölkopf
NeurIPS1
2020 Learning Kernel Tests Without Data Splitting
abstract
Modern large-scale kernel-based tests such as maximum mean discrepancy (MMD) and kernelized Stein discrepancy (KSD) optimize kernel hyperparameters on a held-out sample via data splitting to obtain the most powerful test statistics. While data splitting results in a tractable null distribution, it suffers from a reduction in test power due to smaller test sample size. Inspired by the selective inference framework, we propose an approach that enables learning the hyperparameters and testing on the full sample without data splitting. Our approach can correctly calibrate the test in the presence of such dependency, and yield a test threshold in closed form. At the same significance level, our approach’s test power is empirically larger than that of the data-splitting approach, regardless of its split proportion.
Jonas M. Kübler, Wittawat Jitkrittum, Bernhard Schölkopf, Krikamol Muandet
NeurIPS1
2020 Kernel Conditional Moment Test via Maximum Moment Restriction
abstract
We propose a new family of specification tests called kernel conditional moment (KCM) tests. Our tests are built on a novel representation of conditional moment restrictions in a reproducing kernel Hilbert space (RKHS) called conditional moment embedding (CMME). After transforming the conditional moment restrictions into a continuum of unconditional counterparts, the test statistic is defined as the maximum moment restriction (MMR) within the unit ball of the RKHS. We show that the MMR not only fully characterizes the original conditional moment restrictions, leading to consistency in both hypothesis testing and parameter estimation, but also has an analytic expression that is easy to compute as well as closed-form asymptotic distributions. Our empirical studies show that the KCM test has a promising finite-sample performance compared to existing tests.
Krikamol Muandet, Wittawat Jitkrittum, Jonas M. Kübler
UAI3