Carl-Johann Simon-Gabriel

dblp:163/2039 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
Learning theory · 25% Trustworthy machine learning · 23% Kernel, tree and ensemble methods · 16%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Kernel, tree and ensemble methods
kernel methods
2.042024
Targeted Separation and Convergence with Kernel Discrepancies · J. Mach. Learn. Res. 2024
Metrizing Weak Convergence with Maximum Mean Discrepancies · J. Mach. Learn. Res. 2023
Kernel Distribution Embeddings: Universal Kernels, Characteristic Kernels and Kernel Metrics on Distributions · J. Mach. Learn. Res. 2018
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
1.632024
Robust NAS under adversarial training: benchmark, theory, and beyond · ICLR 2024
PopSkipJump: Decision-Based Attack for Probabilistic Classifiers · ICML 2021
First-Order Adversarial Vulnerability of Neural Networks and Input Dimension · ICML 2019
Machine learning › Learning theory › probability metric › integral probability metric
maximum mean discrepancy
1.422024
Targeted Separation and Convergence with Kernel Discrepancies · J. Mach. Learn. Res. 2024
Metrizing Weak Convergence with Maximum Mean Discrepancies · J. Mach. Learn. Res. 2023
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
1.322023
Bridging the Gap to Real-World Object-Centric Learning · ICLR 2023
Object-Centric Multiple Object Tracking · ICCV 2023
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.812024
Robust NAS under adversarial training: benchmark, theory, and beyond · ICLR 2024
Machine learning › Learning theory › generalization
generalization theory
0.812024
Robust NAS under adversarial training: benchmark, theory, and beyond · ICLR 2024
Machine learning › Learning theory
hypothesis testing
0.812024
Targeted Separation and Convergence with Kernel Discrepancies · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning › divergence measure
kernel stein discrepancy
0.812024
Targeted Separation and Convergence with Kernel Discrepancies · J. Mach. Learn. Res. 2024
Machine learning › Learning theory › hypothesis testing › two-sample testing
kernel two-sample test
0.812024
Targeted Separation and Convergence with Kernel Discrepancies · J. Mach. Learn. Res. 2024
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.812024
Robust NAS under adversarial training: benchmark, theory, and beyond · ICLR 2024
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
0.812024
Robust NAS under adversarial training: benchmark, theory, and beyond · ICLR 2024
Computer vision › Video understanding and tracking
multi-object tracking
0.712023
Object-Centric Multiple Object Tracking · ICCV 2023
Natural language and speech › Information extraction and text analysis › open vocabulary learning
open-vocabulary recognition
0.712023
Unsupervised Open-Vocabulary Object Localization in Videos · ICCV 2023
Machine learning › Learning theory
probability metric
0.712023
Metrizing Weak Convergence with Maximum Mean Discrepancies · J. Mach. Learn. Res. 2023
Computer vision › Segmentation and scene understanding
scene understanding
0.712023
Bridging the Gap to Real-World Object-Centric Learning · ICLR 2023
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel mean embedding
0.622018
Kernel Distribution Embeddings: Universal Kernels, Characteristic Kernels and Kernel Metrics on Distributions · J. Mach. Learn. Res. 2018
Consistent Kernel Mean Estimation for Functions of Random Variables · NIPS 2016
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.612022
Assaying Out-Of-Distribution Generalization in Transfer Learning · NeurIPS 2022
Machine learning › Trustworthy machine learning
robustness
0.612022
Assaying Out-Of-Distribution Generalization in Transfer Learning · NeurIPS 2022
Machine learning › Trustworthy machine learning › robustness › adversarial attack
hard-label black-box attack
0.512021
PopSkipJump: Decision-Based Attack for Probabilistic Classifiers · ICML 2021
Machine learning › Trustworthy machine learning › adversarial machine learning
adversarial vulnerability
0.412019
First-Order Adversarial Vulnerability of Neural Networks and Input Dimension · ICML 2019
Machine learning › Kernel, tree and ensemble methods › kernel methods
characteristic kernels
0.312018
Kernel Distribution Embeddings: Universal Kernels, Characteristic Kernels and Kernel Metrics on Distributions · J. Mach. Learn. Res. 2018
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting
0.312017
AdaGAN: Boosting Generative Models · NIPS 2017
Machine learning › Generative modeling
generative adversarial network
0.312017
AdaGAN: Boosting Generative Models · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model
0.312017
AdaGAN: Boosting Generative Models · NIPS 2017
Machine learning › Generative modeling › generative adversarial network › GAN training
mode collapse
0.312017
AdaGAN: Boosting Generative Models · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning
probabilistic programming
0.212016
Consistent Kernel Mean Estimation for Functions of Random Variables · NIPS 2016
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › particle-based variational inference
stein variational gradient descent
0.212024
Targeted Separation and Convergence with Kernel Discrepancies · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.212024
Targeted Separation and Convergence with Kernel Discrepancies · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.212015
Removing systematic errors for exoplanet search via latent causes · ICML 2015
Machine learning › Trustworthy machine learning › causal machine learning
confounder removal
0.212015
Removing systematic errors for exoplanet search via latent causes · ICML 2015

Methods — techniques the papers use, named apart from their topics

slot attention · 1.3self-supervised learning · 1.3reproducing kernel hilbert space · 1.0weak convergence · 0.8reproducing kernel hilbert space embedding · 0.8neural tangent kernel · 0.8multi-objective adversarial training · 0.8bochner embedding · 0.8expectation-maximization · 0.7CLIP · 0.7regularization · 0.4gradient norm analysis · 0.4half-sibling regression · 0.2additive noise models · 0.2
YearPublicationVenuePosition
2024 Robust NAS under adversarial training: benchmark, theory, and beyond
abstract
Recent developments in neural architecture search (NAS) emphasize the significance of considering robust architectures against malicious data. However, there is a notable absence of benchmark evaluations and theoretical guarantees for searching these robust architectures, especially when adversarial training is considered. In this work, we aim to address these two challenges, making twofold contributions. First, we release a comprehensive data set that encompasses both clean accuracy and robust accuracy for a vast array of adversarially trained networks from the NAS-Bench-201 search space on image datasets. Then, leveraging the neural tangent kernel (NTK) tool from deep learning theory, we establish a generalization theory for searching architecture in terms of clean accuracy and robust accuracy under multi-objective adversarial training. We firmly believe that our benchmark and theoretical insights will significantly benefit the NAS community through reliable reproducibility, efficient assessment, and theoretical foundation, particularly in the pursuit of robust architectures.
Yongtao Wu, Fanghui Liu 0001, Carl-Johann Simon-Gabriel, Grigorios Chrysos 0002, Volkan Cevher
ICLR3
2024 Targeted Separation and Convergence with Kernel Discrepancies
abstract
Maximum mean discrepancies (MMDs) like the kernel Stein discrepancy (KSD) have grown central to a wide range of applications, including hypothesis testing, sampler selection, distribution approximation, and variational inference. In each setting, these kernel-based discrepancy measures are required to $(i)$ separate a target $\mathrm{P}$ from other probability measures or even $(ii)$ control weak convergence to $\mathrm{P}$. In this article we derive new sufficient and necessary conditions to ensure $(i)$ and $(ii)$. For MMDs on separable metric spaces, we characterize those kernels that separate Bochner embeddable measures and introduce simple conditions for separating all measures with unbounded kernels and for controlling convergence with bounded kernels. We use these results on $\mathbb{R}^d$ to substantially broaden the known conditions for KSD separation and convergence control and to develop the first KSDs known to exactly metrize weak convergence to $\mathrm{P}$. Along the way, we highlight the implications of our results for hypothesis testing, measuring and improving sample quality, and sampling with Stein variational gradient descent.
Alessandro Barp, Carl-Johann Simon-Gabriel, Mark A. Girolami, Lester Mackey
J. Mach. Learn. Res.2
2023 Unsupervised Open-Vocabulary Object Localization in Videos
abstract
In this paper, we show that recent advances in video representation learning and pre-trained vision-language models allow for substantial improvements in self-supervised video object localization. We propose a method that first localizes objects in videos via a slot attention approach and then assigns text to the obtained slots. The latter is achieved by an unsupervised way to read localized semantic information from the pre-trained CLIP model. The resulting video object localization is entirely unsupervised apart from the implicit annotation contained in CLIP, and it is effectively the first unsupervised approach that yields good results on regular video benchmarks.
Zechen Bai, Tianjun Xiao, Dominik Zietlow, Max Horn, Carl-Johann Simon-Gabriel, Zheng Shou 0001, Francesco Locatello, Bernt Schiele, Thomas Brox, Zheng Zhang 0001, Yanwei Fu 0001, Tong He 0002
ICCV7
2023 Object-Centric Multiple Object Tracking
abstract
Unsupervised object-centric learning methods allow the partitioning of scenes into entities without additional localization information and are excellent candidates for reducing the annotation burden of multiple-object tracking (MOT) pipelines. Unfortunately, they lack two key properties: objects are often split into parts and are not consistently tracked over time. In fact, state-of-the-art models achieve pixel-level accuracy and temporal consistency by relying on supervised object detection with additional ID labels for the association through time. This paper proposes a video object-centric model for MOT. It consists of an index-merge module that adapts the object-centric slots into detection outputs and an object memory module that builds complete object prototypes to handle occlusions. Benefited from object-centric learning, we only require sparse detection labels (0%-6.25%) for object localization and feature binding. Relying on our self-supervised Expectation-Maximization-inspired loss for object association, our approach requires no ID labels. Our experiments significantly narrow the gap between the existing object-centric model and the fully supervised state-of-the-art and outperform several unsupervised trackers. Code is available at https://github.com/amazon-science/object-centric-multiple-object-tracking.
Max Horn, Yizhuo Ding, Tong He 0002, Zechen Bai, Dominik Zietlow, Carl-Johann Simon-Gabriel, Bing Shuai, Zhuowen Tu, Thomas Brox, Bernt Schiele, Yanwei Fu 0001, Francesco Locatello, Zheng Zhang 0001, Tianjun Xiao
ICCV8
2023 Bridging the Gap to Real-World Object-Centric Learning
Maximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Zietlow, Tianjun Xiao, Carl-Johann Simon-Gabriel, Tong He 0002, Zheng Zhang 0001, Bernhard Schölkopf, Thomas Brox, Francesco Locatello
ICLR6
2023 Metrizing Weak Convergence with Maximum Mean Discrepancies
abstract
This paper characterizes the maximum mean discrepancies (MMD) that metrize the weak convergence of probability measures for a wide class of kernels. More precisely, we prove that, on a locally compact, non-compact, Hausdorff space, the MMD of a bounded continuous Borel measurable kernel $k$, whose RKHS-functions vanish at infinity (i.e., $H_k \subset C_0$), metrizes the weak convergence of probability measures if and only if $k$ is continuous and integrally strictly positive definite ($\int$s.p.d.) over all signed, finite, regular Borel measures. We also correct a prior result of Simon-Gabriel and Schölkopf (JMLR 2018, Thm. 12) by showing that there exist both bounded continuous $\int$s.p.d. kernels that do not metrize weak convergence and bounded continuous non-$\int$s.p.d. kernels that do metrize it.
Carl-Johann Simon-Gabriel, Alessandro Barp, Bernhard Schölkopf, Lester Mackey
J. Mach. Learn. Res.1
2022 Assaying Out-Of-Distribution Generalization in Transfer Learning
abstract
Since out-of-distribution generalization is a generally ill-posed problem, various proxy targets (e.g., calibration, adversarial robustness, algorithmic corruptions, invariance across shifts) were studied across different research programs resulting in different recommendations. While sharing the same aspirational goal, these approaches have never been tested under the same experimental conditions on real data. In this paper, we take a unified view of previous work, highlighting message discrepancies that we address empirically, and providing recommendations on how to measure the robustness of a model and how to improve it. To this end, we collect 172 publicly available dataset pairs for training and out-of-distribution evaluation of accuracy, calibration error, adversarial attacks, environment invariance, and synthetic corruptions. We fine-tune over 31k networks, from nine different architectures in the many- and few-shot setting. Our findings confirm that in- and out-of-distribution accuracies tend to increase jointly, but show that their relation is largely dataset-dependent, and in general more nuanced and more complex than posited by previous, smaller scale studies.
Florian Wenzel, Andrea Dittadi, Peter V. Gehler, Carl-Johann Simon-Gabriel, Max Horn, Dominik Zietlow, David Kernert, Chris Russell 0001, Thomas Brox, Bernt Schiele, Bernhard Schölkopf, Francesco Locatello
NeurIPS4
2021 PopSkipJump: Decision-Based Attack for Probabilistic Classifiers
abstract
Most current classifiers are vulnerable to adversarial examples, small input perturbations that change the classification output. Many existing attack algorithms cover various settings, from white-box to black-box classifiers, but usually assume that the answers are deterministic and often fail when they are not. We therefore propose a new adversarial decision-based attack specifically designed for classifiers with probabilistic outputs. It is based on the HopSkipJump attack by Chen et al. (2019), a strong and query efficient decision-based attack originally designed for deterministic classifiers. Our P(robabilisticH)opSkipJump attack adapts its amount of queries to maintain HopSkipJump’s original output quality across various noise levels, while converging to its query efficiency as the noise level decreases. We test our attack on various noise models, including state-of-the-art off-the-shelf randomized defenses, and show that they offer almost no extra robustness to decision-based attacks. Code is available at https://github.com/cjsg/PopSkipJump.
Carl-Johann Simon-Gabriel, Noman Ahmed Sheikh, Andreas Krause 0001
ICML1
2019 First-Order Adversarial Vulnerability of Neural Networks and Input Dimension
abstract
Over the past few years, neural networks were proven vulnerable to adversarial images: targeted but imperceptible image perturbations lead to drastically different predictions. We show that adversarial vulnerability increases with the gradients of the training objective when viewed as a function of the inputs. Surprisingly, vulnerability does not depend on network topology: for many standard network architectures, we prove that at initialization, the L1-norm of these gradients grows as the square root of the input dimension, leaving the networks increasingly vulnerable with growing image size. We empirically show that this dimension-dependence persists after either usual or robust training, but gets attenuated with higher regularization.
Carl-Johann Simon-Gabriel, Yann Ollivier, Léon Bottou, Bernhard Schölkopf, David Lopez-Paz
ICML1
2018 Kernel Distribution Embeddings: Universal Kernels, Characteristic Kernels and Kernel Metrics on Distributions
abstract
Kernel mean embeddings have become a popular tool in machine learning. They map probability measures to functions in a reproducing kernel Hilbert space. The distance between two mapped measures defines a semi-distance over the probability measures known as the maximum mean discrepancy (MMD). Its properties depend on the underlying kernel and have been linked to three fundamental concepts of the kernel literature: universal, characteristic and strictly positive definite kernels. The contributions of this paper are three-fold. First, by slightly extending the usual definitions of universal, characteristic and strictly positive definite kernels, we show that these three concepts are essentially equivalent. Second, we give the first complete characterization of those kernels whose associated MMD-distance metrizes the weak convergence of probability measures. Third, we show that kernel mean embeddings can be extended from probability measures to generalized measures called Schwartz-distributions and analyze a few properties of these distribution embeddings.
Carl-Johann Simon-Gabriel, Bernhard Schölkopf
J. Mach. Learn. Res.1
2017 AdaGAN: Boosting Generative Models
abstract
Generative Adversarial Networks (GAN) are an effective method for training generative models of complex data such as natural images. However, they are notoriously hard to train and can suffer from the problem of missing modes where the model is not able to produce examples in certain regions of the space. We propose an iterative procedure, called AdaGAN, where at every step we add a new component into a mixture model by running a GAN algorithm on a re-weighted sample. This is inspired by boosting algorithms, where many potentially weak individual predictors are greedily aggregated to form a strong composite predictor. We prove analytically that such an incremental procedure leads to convergence to the true distribution in a finite number of steps if each step is optimal, and convergence at an exponential rate otherwise. We also illustrate experimentally that this procedure addresses the problem of missing modes.
Ilya O. Tolstikhin, Sylvain Gelly, Olivier Bousquet, Carl-Johann Simon-Gabriel, Bernhard Schölkopf
NIPS4
2016 Consistent Kernel Mean Estimation for Functions of Random Variables
abstract
We provide a theoretical foundation for non-parametric estimation of functions of random variables using kernel mean embeddings. We show that for any continuous function f, consistent estimators of the mean embedding of a random variable X lead to consistent estimators of the mean embedding of f(X). For Matern kernels and sufficiently smooth functions we also provide rates of convergence. Our results extend to functions of multiple random variables. If the variables are dependent, we require an estimator of the mean embedding of their joint distribution as a starting point; if they are independent, it is sufficient to have separate estimators of the mean embeddings of their marginal distributions. In either case, our results cover both mean embeddings based on i.i.d. samples as well as "reduced set" expansions in terms of dependent expansion points. The latter serves as a justification for using such expansions to limit memory resources when applying the approach as a basis for probabilistic programming.
Carl-Johann Simon-Gabriel, Adam Scibior, Ilya O. Tolstikhin, Bernhard Schölkopf
NIPS1
2015 Removing systematic errors for exoplanet search via latent causes
abstract
We describe a method for removing the effect of confounders in order to reconstruct a latent quantity of interest. The method, referred to as half-sibling regression, is inspired by recent work in causal inference using additive noise models. We provide a theoretical justification and illustrate the potential of the method in a challenging astronomy application.
Bernhard Schölkopf, David W. Hogg, Dun Wang, Daniel Foreman-Mackey, Dominik Janzing, Carl-Johann Simon-Gabriel, Jonas Peters
ICML6