Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Pan Kessel

dblp:238/1381 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
8since 2021 · last 2025
0009-0006-5951-7188ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Trustworthy machine learning · 30% Generative modeling · 18% Optimization for machine learning · 14%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Theoretical computer science
1 paper
Quantum computing and quantum information · 100%

Topics — the 27 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.632024
Diffeomorphic Counterfactuals With Generative Models · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Fairwashing explanations with off-manifold detergent · ICML 2020
Explanations can be manipulated and geometry is to blame · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
conditional density estimation
0.912025
Generative property enhancer: implicit guided generation through conditional density estimation · NeurIPS 2025
Machine learning › Deep learning architectures and training
equivariant neural network
0.912025
Equivariant Neural Tangent Kernels · ICML 2025
Machine learning › Deep learning architectures and training › equivariant neural network
group convolutional network
0.912025
Equivariant Neural Tangent Kernels · ICML 2025
Machine learning › Generative modeling › diffusion model › controllable generation
guided generation
0.912025
Generative property enhancer: implicit guided generation through conditional density estimation · NeurIPS 2025
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
0.912025
Equivariant Neural Tangent Kernels · ICML 2025
Bioinformatics and computational biology
protein design
0.912025
Generative property enhancer: implicit guided generation through conditional density estimation · NeurIPS 2025
Bioinformatics and computational biology › protein design
protein fitness optimization
0.912025
Generative property enhancer: implicit guided generation through conditional density estimation · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability › explanation evaluation
explanation robustness
0.822020
Fairwashing explanations with off-manifold detergent · ICML 2020
Explanations can be manipulated and geometry is to blame · NeurIPS 2019
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation
0.812024
Diffeomorphic Counterfactuals With Generative Models · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Kernel, tree and ensemble methods › ensemble learning
deep ensembles
0.812024
Emergent Equivariance in Deep Ensembles · ICML 2024
Machine learning › Optimization for machine learning
gradient estimator
0.812024
Fast and unified path gradient estimators for normalizing flows · ICLR 2024
Machine learning › Generative modeling
normalizing flow
0.812024
Fast and unified path gradient estimators for normalizing flows · ICLR 2024
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
reparameterization gradient
0.812024
Fast and unified path gradient estimators for normalizing flows · ICLR 2024
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
acquisition function
0.712023
Physics-Informed Bayesian Optimization of Variational Quantum Circuits · NeurIPS 2023
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.712023
Physics-Informed Bayesian Optimization of Variational Quantum Circuits · NeurIPS 2023
Quantum computing and quantum information
quantum algorithms
0.712023
Physics-Informed Bayesian Optimization of Variational Quantum Circuits · NeurIPS 2023
Quantum computing and quantum information › quantum algorithms
variational quantum eigensolver
0.712023
Physics-Informed Bayesian Optimization of Variational Quantum Circuits · NeurIPS 2023
Machine learning › Generative modeling › normalizing flow
continuous normalizing flow
0.612022
Path-Gradient Estimators for Continuous Normalizing Flows · ICML 2022
Machine learning › Trustworthy machine learning › interpretability
attribution methods
0.412020
Fairwashing explanations with off-manifold detergent · ICML 2020
Machine learning › Trustworthy machine learning
fairness
0.412020
Fairwashing explanations with off-manifold detergent · ICML 2020
Machine learning › Trustworthy machine learning › fairness
fairwashing
0.412020
Fairwashing explanations with off-manifold detergent · ICML 2020
Security and privacy of machine learning
adversarial attack
0.412019
Explanations can be manipulated and geometry is to blame · NeurIPS 2019
Machine learning › Deep learning architectures and training
data augmentation
0.312025
Equivariant Neural Tangent Kernels · ICML 2025
Machine learning › Representation and self-supervised learning
equivariance
0.312025
Equivariant Neural Tangent Kernels · ICML 2025
Machine learning › Generative modeling
generative model
0.212024
Diffeomorphic Counterfactuals With Generative Models · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Generative modeling
maximum likelihood learning
0.212024
Fast and unified path gradient estimators for normalizing flows · ICLR 2024

Methods — techniques the papers use, named apart from their topics

synthetic data augmentation · 1.7iterative training · 1.7conditional density estimation · 1.7kernel regression · 0.9group convolution · 0.9path gradient estimation · 0.8neural tangent kernel · 0.8maximum likelihood · 0.8diffeomorphic coordinate transformation · 0.8data augmentation · 0.8bayesian optimization · 0.7VQE-kernel · 0.7EMICoRe · 0.7perturbation analysis · 0.4
YearPublicationVenuePosition
2025 Equivariant Neural Tangent Kernels
abstract
Little is known about the training dynamics of equivariant neural networks, in particular how it compares to data augmented training of their non-equivariant counterparts. Recently, neural tangent kernels (NTKs) have emerged as a powerful tool to analytically study the training dynamics of wide neural networks. In this work, we take an important step towards a theoretical understanding of training dynamics of equivariant models by deriving neural tangent kernels for a broad class of equivariant architectures based on group convolutions. As a demonstration of the capabilities of our framework, we show an interesting relationship between data augmentation and group convolutional networks. Specifically, we prove that they share the same expected prediction over initializations at all training times and even off the data manifold. In this sense, they have the same training dynamics. We demonstrate in numerical experiments that this still holds approximately for finite-width ensembles. By implementing equivariant NTKs for roto-translations in the plane ($G=C_{n}\ltimes\mathbb{R}^{2}$) and 3d rotations ($G=\mathrm{SO}(3)$), we show that equivariant NTKs outperform their non-equivariant counterparts as kernel predictors for histological image classification and quantum mechanical property prediction.
Philipp Misof, Pan Kessel, Jan E. Gerken
ICML2
2025 Generative property enhancer: implicit guided generation through conditional density estimation
abstract
Generative modeling is increasingly important for data-driven computational design. Conventional approaches pair a generative model with a discriminative model to select or guide samples toward optimized designs. Yet discriminative models often struggle in data-scarce settings, common in scientific applications, and are unreliable in the tails of the distribution where optimal designs typically lie. We introduce generative property enhancer (GPE), an approach that implicitly guides generation by matching samples with lower property values to higher-value ones. Formulated as conditional density estimation, our framework defines a target distribution with improved properties, compelling the generative model to produce enhanced, diverse designs without auxiliary predictors. GPE is simple, scalable, end-to-end, modality-agnostic, and integrates seamlessly with diverse generative model architectures and losses. We demonstrate competitive empirical results on standard _in silico_ offline (non-sequential) protein fitness optimization benchmarks. Finally, we propose iterative training on a combination of limited real data and self-generated synthetic data, enabling extrapolation beyond the original property ranges.
Pedro O. Pinheiro, Pan Kessel, Aya Abdelsalam Ismail, Sai Pooja Mahajan, Kyunghyun Cho, Saeed Saremi, Natasa Tagasovska
NeurIPS2
2024 Fast and unified path gradient estimators for normalizing flows
abstract
Recent work shows that path gradient estimators for normalizing flows have lower variance compared to standard estimators, resulting in improved training. However, they are often prohibitively more expensive from a computational point of view and cannot be applied to maximum likelihood training in a scalable manner, which severely hinders their widespread adoption. In this work, we overcome these crucial limitations. Specifically, we propose a fast path gradient estimator which works for all normalizing flow architectures of practical relevance for sampling from an unnormalized target distribution. We then show that this estimator can also be applied to maximum likelihood training and empirically establish its superior performance for several natural sciences applications.
Lorenz Vaitl, Ludwig Winkler, Lorenz Richter, Pan Kessel
ICLR4
2024 Emergent Equivariance in Deep Ensembles
abstract
We show that deep ensembles become equivariant for all inputs and at all training times by simply using data augmentation. Crucially, equivariance holds off-manifold and for any architecture in the infinite width limit. The equivariance is emergent in the sense that predictions of individual ensemble members are not equivariant but their collective prediction is. Neural tangent kernel theory is used to derive this result and we verify our theoretical insights using detailed numerical experiments.
Jan E. Gerken, Pan Kessel
ICML2
2024 Diffeomorphic Counterfactuals With Generative Models
abstract
Counterfactuals can explain classification decisions of neural networks in a human interpretable way. We propose a simple but effective method to generate such counterfactuals. More specifically, we perform a suitable diffeomorphic coordinate transformation and then perform gradient ascent in these coordinates to find counterfactuals which are classified with great confidence as a specified target class. We propose two methods to leverage generative models to construct such suitable coordinate systems that are either exactly or approximately diffeomorphic. We analyze the generation process theoretically using Riemannian differential geometry and validate the quality of the generated counterfactuals using various qualitative and quantitative measures.
Ann-Kathrin Dombrowski, Jan E. Gerken, Klaus-Robert Müller, Pan Kessel
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Physics-Informed Bayesian Optimization of Variational Quantum Circuits
abstract
In this paper, we propose a novel and powerful method to harness Bayesian optimization for variational quantum eigensolvers (VQEs) - a hybrid quantum-classical protocol used to approximate the ground state of a quantum Hamiltonian. Specifically, we derive a *VQE-kernel* which incorporates important prior information about quantum circuits: the kernel feature map of the VQE-kernel exactly matches the known functional form of the VQE's objective function and thereby significantly reduces the posterior uncertainty. Moreover, we propose a novel acquisition function for Bayesian optimization called \emph{Expected Maximum Improvement over Confident Regions} (EMICoRe) which can actively exploit the inductive bias of the VQE-kernel by treating regions with low predictive uncertainty as indirectly "observed". As a result, observations at as few as three points in the search domain are sufficient to determine the complete objective function along an entire one-dimensional subspace of the optimization landscape. Our numerical experiments demonstrate that our approach improves over state-of-the-art baselines.
Kim Nicoli, Christopher J. Anders, Lena Funcke, Tobias Hartung, Karl Jansen, Stefan Kühn, Klaus-Robert Müller, Paolo Stornati, Pan Kessel, Shinichi Nakajima
NeurIPS9
2022 Path-Gradient Estimators for Continuous Normalizing Flows
abstract
Recent work has established a path-gradient estimator for simple variational Gaussian distributions and has argued that the path-gradient is particularly beneficial in the regime in which the variational distribution approaches the exact target distribution. In many applications, this regime can however not be reached by a simple Gaussian variational distribution. In this work, we overcome this crucial limitation by proposing a path-gradient estimator for the considerably more expressive variational family of continuous normalizing flows. We outline an efficient algorithm to calculate this estimator and establish its superior performance empirically.
Lorenz Vaitl, Kim Nicoli, Shinichi Nakajima, Pan Kessel
ICML4
2022 Towards robust explanations for deep neural networks
abstract
Explanation methods shed light on the decision process of black-box classifiers such as deep neural networks. But their usefulness can be compromised because they are susceptible to manipulations. With this work, we aim to enhance the resilience of explanations. We develop a unified theoretical framework for deriving bounds on the maximal manipulability of a model. Based on these theoretical insights, we present three different techniques to boost robustness against manipulation: training with weight decay, smoothing activation functions, and minimizing the Hessian of the network. Our experimental results confirm the effectiveness of these approaches.
Ann-Kathrin Dombrowski, Christopher J. Anders, Klaus-Robert Müller, Pan Kessel
Pattern Recognit.4
2020 Fairwashing explanations with off-manifold detergent
abstract
Explanation methods promise to make black-box classifiers more transparent. As a result, it is hoped that they can act as proof for a sensible, fair and trustworthy decision-making process of the algorithm and thereby increase its acceptance by the end-users. In this paper, we show both theoretically and experimentally that these hopes are presently unfounded. Specifically, we show that, for any classifier $g$, one can always construct another classifier $\tilde{g}$ which has the same behavior on the data (same train, validation, and test error) but has arbitrarily manipulated explanation maps. We derive this statement theoretically using differential geometry and demonstrate it experimentally for various explanation methods, architectures, and datasets. Motivated by our theoretical insights, we then propose a modification of existing explanation methods which makes them significantly more robust.
Christopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller, Pan Kessel
ICML5
2019 Explanations can be manipulated and geometry is to blame
abstract
Explanation methods aim to make neural networks more trustworthy and interpretable. In this paper, we demonstrate a property of explanation methods which is disconcerting for both of these purposes. Namely, we show that explanations can be manipulated arbitrarily by applying visually hardly perceptible perturbations to the input that keep the network's output approximately constant. We establish theoretically that this phenomenon can be related to certain geometrical properties of neural networks. This allows us to derive an upper bound on the susceptibility of explanations to manipulations. Based on this result, we propose effective mechanisms to enhance the robustness of explanations.
Ann-Kathrin Dombrowski, Maximilian Alber, Christopher J. Anders, Marcel Ackermann 0001, Klaus-Robert Müller, Pan Kessel
NeurIPS6