Jan Hendrik Metzen

dblp:93/1712 · DBLP profile ↗
← Back
28ranked-venue papers
10as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 10 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 Training-free diffusion for controlling illumination conditions in images
Xiaoyan Xing, Vincent Tao Hu, Jan Hendrik Metzen, Konrad Groh, Sezer Karaoglu, Theo Gevers
Comput. Vis. Image Underst.3
2024 Label-Free Neural Semantic Image Synthesis
Kevin Alexander Laube, Jan Hendrik Metzen, Shin-I Cheng, Julio Borges, Anna Khoreva
ECCV (53)4
2024 SciOL and MuLMS-Img: Introducing A Large-Scale Multimodal Scientific Dataset and Models for Image-Text Tasks in the Scientific Domain
abstract
In scientific publications, a substantial part of the information is expressed via figures containing images and diagrams. Hence, the retrieval of relevant figures given a natural language query is an important real-world task. However, due to the lack of training and evaluation data, most existing approaches are either limited to one modality or focus on non-scientific domains, making their application to scientific publications challenging.In this paper, we address this gap by introducing two novel datasets: (1) SciOL, the largest openly-licensed pre-training corpus for multimodal models in the scientific domain, covering multiple sciences including materials science, physics, and computer science, and (2) MuLMS-Img, a high-quality dataset in the materials science domain, manually annotated for various image-text tasks. Our experiments show that pre-training large-scale vision-language models on SciOL increases performance considerably across a broad variety of image-text tasks including figure type classification, optical character recognition, captioning, and figure retrieval. Using MuLMS-Img, we show that integrating text-based features extracted via a fine-tuned model for a specific domain can boost cross-modal scientific figure retrieval performance by up to 50%.
Tim Tarsi, Heike Adel, Jan Hendrik Metzen, Matteo Finco, Annemarie Friedrich
WACV3
2023 Identification of Systematic Errors of Image Classifiers on Rare Subgroups
abstract
Despite excellent average-case performance of many image classifiers, their performance can substantially deteriorate on semantically coherent subgroups of the data that were under-represented in the training data. These systematic errors can impact both fairness for demographic minority groups as well as robustness and safety under domain shift. A major challenge is to identify such subgroups with subpar performance when the subgroups are not annotated and their occurrence is very rare. We leverage recent advances in text-to-image models and search in the space of textual descriptions of subgroups ("prompts") for sub-groups where the target model has low performance on the prompt-conditioned synthesized data. To tackle the exponentially growing number of subgroups, we employ combinatorial testing. We denote this procedure as PromptAttack as it can be interpreted as an adversarial attack in a prompt space. We study subgroup coverage and identifiability with PromptAttack in a controlled setting and find that it identifies systematic errors with high accuracy. Thereupon, we apply PromptAttack to ImageNet classifiers and identify novel systematic errors on rare subgroups.
Jan Hendrik Metzen, Robin Hutmacher, N. Grace Hua, Valentyn Boreiko
ICCV1
2023 Certified Defences Against Adversarial Patch Attacks on Semantic Segmentation
Maksym Yatsura, Kaspar Sakmann, N. Grace Hua, Matthias Hein 0001, Jan Hendrik Metzen
ICLR5
2023 Neural Architecture Search for Dense Prediction Tasks in Computer Vision
abstract
Abstract The success of deep learning in recent years has lead to a rising demand for neural network architecture engineering. As a consequence, neural architecture search (NAS), which aims at automatically designing neural network architectures in a data-driven manner rather than manually, has evolved as a popular field of research. With the advent of weight sharing strategies across architectures, NAS has become applicable to a much wider range of problems. In particular, there are now many publications for dense prediction tasks in computer vision that require pixel-level predictions, such as semantic segmentation or object detection. These tasks come with novel challenges, such as higher memory footprints due to high-resolution data, learning multi-scale representations, longer training times, and more complex and larger neural architectures. In this manuscript, we provide an overview of NAS for dense prediction tasks by elaborating on these novel challenges and surveying ways to address them to ease future research and application of existing methods to novel problems.
Rohit Mohan, Thomas Elsken, Arber Zela, Jan Hendrik Metzen, Benedikt Staffler, Thomas Brox, Abhinav Valada, Frank Hutter
Int. J. Comput. Vis.4
2022 Give Me Your Attention: Dot-Product Attention Considered Harmful for Adversarial Patch Robustness
abstract
Neural architectures based on attention such as vision transformers are revolutionizing image recognition. Their main benefit is that attention allows reasoning about all parts of a scene jointly. In this paper, we show how the global reasoning of (scaled) dot-product attention can be the source of a major vulnerability when confronted with adversarial patch attacks. We provide a theoretical understanding of this vulnerability and relate it to an adversary's ability to misdirect the attention of all queries to a single key token under the control of the adversarial patch. We propose novel adversarial objectives for crafting adversarial patches which target this vulnerability explicitly. We show the effectiveness of the proposed patch attacks on popular image classification (ViTs and DeiTs) and object detection models (DETR). We find that adversarial patches occupying 0.5% of the input can lead to robust accuracies as low as 0% for ViT on ImageNet, and reduce the mAP of DETR on MS COCO to less than 3%.
Giulio Lovisotto, Nicole Finnie, Mauricio Munoz, Chaithanya Kumar Mummadi, Jan Hendrik Metzen
CVPR5
2021 Efficient Certified Defenses Against Patch Attacks on Image Classifiers
Jan Hendrik Metzen, Maksym Yatsura
ICLR1
2021 Does enhanced shape bias improve neural network robustness to common corruptions?
Chaithanya Kumar Mummadi, Ranjitha Subramaniam, Robin Hutmacher, Julien Vitay, Volker Fischer 0003, Jan Hendrik Metzen
ICLR6
2021 Meta-Learning the Search Distribution of Black-Box Random Search Based Adversarial Attacks
abstract
Adversarial attacks based on randomized search schemes have obtained state-of-the-art results in black-box robustness evaluation recently. However, as we demonstrate in this work, their efficiency in different query budget regimes depends on manual design and heuristic tuning of the underlying proposal distributions. We study how this issue can be addressed by adapting the proposal distribution online based on the information obtained during the attack. We consider Square Attack, which is a state-of-the-art score-based black-box attack, and demonstrate how its performance can be improved by a learned controller that adjusts the parameters of the proposal distribution online during the attack. We train the controller using gradient-based end-to-end training on a CIFAR10 model with white box access. We demonstrate that plugging the learned controller into the attack consistently improves its black-box robustness estimate in different query regimes by up to 20% for a wide range of different models with black-box access. We further show that the learned adaptation principle transfers well to the other data distributions such as CIFAR100 or ImageNet and to the targeted attack setting.
Maksym Yatsura, Jan Hendrik Metzen, Matthias Hein 0001
NeurIPS2
2020 Meta-Learning of Neural Architectures for Few-Shot Learning
abstract
The recent progress in neural architecture search (NAS) has allowed scaling the automated design of neural architectures to real-world domains, such as object detection and semantic segmentation. However, one prerequisite for the application of NAS are large amounts of labeled data and compute resources. This renders its application challenging in few-shot learning scenarios, where many related tasks need to be learned, each with limited amounts of data and compute time. Thus, few-shot learning is typically done with a fixed neural architecture. To improve upon this, we propose MetaNAS, the first method which fully integrates NAS with gradient-based meta-learning. MetaNAS optimizes a meta-architecture along with the meta-weights during meta-training. During meta-testing, architectures can be adapted to a novel task with a few steps of the task optimizer, that is: task adaptation becomes computationally cheap and requires only little data per task. Moreover, MetaNAS is agnostic in that it can be used with arbitrary model-agnostic meta-learning algorithms and arbitrary gradient-based NAS methods. Empirical results on standard few-shot classification benchmarks show that MetaNAS with a combination of DARTS and REPTILE yields state-of-the-art results.
Thomas Elsken, Benedikt Staffler, Jan Hendrik Metzen, Frank Hutter
CVPR3
2019 Defending Against Universal Perturbations With Shared Adversarial Training
abstract
Classifiers such as deep neural networks have been shown to be vulnerable against adversarial perturbations on problems with high-dimensional input space. While adversarial training improves the robustness of image classifiers against such adversarial perturbations, it leaves them sensitive to perturbations on a non-negligible fraction of the inputs. In this work, we show that adversarial training is more effective in preventing universal perturbations, where the same perturbation needs to fool a classifier on many inputs. Moreover, we investigate the trade-off between robustness against universal perturbations and performance on unperturbed data and propose an extension of adversarial training that handles this trade-off more gracefully. We present results for image classification and semantic segmentation to showcase that universal perturbations that fool a model hardened with adversarial training become clearly perceptible and show patterns of the target scene.
Chaithanya Kumar Mummadi, Thomas Brox, Jan Hendrik Metzen
ICCV3
2019 Efficient Multi-Objective Neural Architecture Search via Lamarckian Evolution
Thomas Elsken, Jan Hendrik Metzen, Frank Hutter
ICLR (Poster)2
2019 Neural Architecture Search: A Survey
abstract
Deep Learning has enabled remarkable progress over the last years on a variety of tasks, such as image recognition, speech recognition, and machine translation. One crucial aspect for this progress are novel neural architectures. Currently employed architectures have mostly been developed manually by human experts, which is a time-consuming and error-prone process. Because of this, there is growing interest in automated \emph{neural architecture search} methods. We provide an overview of existing work in this field of research and categorize them according to three dimensions: search space, search strategy, and performance estimation strategy.
Thomas Elsken, Jan Hendrik Metzen, Frank Hutter
J. Mach. Learn. Res.2
2018 Scaling provable adversarial defenses
abstract
Recent work has developed methods for learning deep network classifiers that are \emph{provably} robust to norm-bounded adversarial perturbation; however, these methods are currently only possible for relatively small feedforward networks. In this paper, in an effort to scale these approaches to substantially larger models, we extend previous work in three main directly. First, we present a technique for extending these training procedures to much more general networks, with skip connections (such as ResNets) and general nonlinearities; the approach is fully modular, and can be implemented automatically analogously to automatic differentiation. Second, in the specific case of $\ell_\infty$ adversarial perturbations and networks with ReLU nonlinearities, we adopt a nonlinear random projection for training, which scales \emph{linearly} in the number of hidden units (previous approached scaled quadratically). Third, we show how to further improve robust error through cascade models. On both MNIST and CIFAR data sets, we train classifiers that improve substantially on the state of the art in provable robust adversarial error bounds: from 5.8% to 3.1% on MNIST (with $\ell_\infty$ perturbations of $\epsilon=0.1$), and from 80% to 36.4% on CIFAR (with $\ell_\infty$ perturbations of $\epsilon=2/255$).
Eric Wong 0001, Frank R. Schmidt, Jan Hendrik Metzen, J. Zico Kolter
NeurIPS3
2017 Universal Adversarial Perturbations Against Semantic Image Segmentation
abstract
While deep learning is remarkably successful on perceptual tasks, it was also shown to be vulnerable to adversarial perturbations of the input. These perturbations denote noise added to the input that was generated specifically to fool the system while being quasi-imperceptible for humans. More severely, there even exist universal perturbations that are input-agnostic but fool the network on the majority of inputs. While recent work has focused on image classification, this work proposes attacks against semantic image segmentation: we present an approach for generating (universal) adversarial perturbations that make the network yield a desired target segmentation as output. We show empirically that there exist barely perceptible universal noise patterns which result in nearly the same predicted segmentation for arbitrary inputs. Furthermore, we also show the existence of universal noise which removes a target class (e.g., all pedestrians) from the segmentation while leaving the segmentation mostly unchanged otherwise.
Jan Hendrik Metzen, Chaithanya Kumar Mummadi, Thomas Brox, Volker Fischer 0003
ICCV1
2017 On Detecting Adversarial Perturbations
Jan Hendrik Metzen, Tim Genewein, Volker Fischer 0003, Bastian Bischoff
ICLR (Poster)1
2016 Minimum Regret Search for Single- and Multi-Task Optimization
abstract
We propose minimum regret search (MRS), a novel acquisition function for Bayesian optimization. MRS bears similarities with information-theoretic approaches such as entropy search (ES). However, while ES aims in each query at maximizing the information gain with respect to the global maximum, MRS aims at minimizing the expected simple regret of its ultimate recommendation for the optimum. While empirically ES and MRS perform similar in most of the cases, MRS produces fewer outliers with high simple regret than ES. We provide empirical results both for a synthetic single-task optimization problem as well as for a simulated multi-task robotic control problem.
Jan Hendrik Metzen
ICML1
2014 Velocity-Based Multiple Change-Point Inference for Unsupervised Segmentation of Human Movement Behavior
abstract
In order to transfer complex human behavior to a robot, segmentation methods are needed which are able to detect central movement patterns that can be combined to generate a wide range of behaviors. We propose an algorithm that segments human movements into behavior building blocks in a fully automatic way, called velocity-based Multiple Change-point Inference (vMCI). Based on characteristic bell-shaped velocity patterns that can be found in point-to-point arm movements, the algorithm infers segment borders using Bayesian inference. Different segment lengths and variations in the movement execution can be handled. Moreover, the number of segments the movement is composed of need not be known in advance. Several experiments are performed on synthetic and motion capturing data of human movements to compare vMCI with other techniques for unsupervised segmentation. The results show that vMCI is able to detect segment borders even in noisy data and in demonstrations with smooth transitions between segments.
Lisa Senger, Martin Schröer, Jan Hendrik Metzen, Elsa Andrea Kirchner
ICPR3
2014 Active contextual policy search
Alexander Fabisch, Jan Hendrik Metzen
J. Mach. Learn. Res.2
2013 Learning Graph-Based Representations for Continuous Reinforcement Learning Domains
Jan Hendrik Metzen
ECML/PKDD (1)1
2009 Matching of anatomical tree structures for registration of medical images
Jan Hendrik Metzen, Tim Kröger, Andrea Schenk, Stephan Zidowitz, Heinz-Otto Peitgen, Xiaoyi Jiang 0001
Image Vis. Comput.1
2008 Accelerating neuroevolutionary methods using a Kalman filter
abstract
In recent years, neuroevolutionary methods have shown great promise in solving learning tasks, especially in domains that are stochastic, partially observable, and noisy. In this paper, we show how the Kalman filter can be exploited (1) to efficiently find an optimal solution (i. e. reducing the number of evaluations needed to find the solution), (2) to find solutions that are robust against noise, and (3) to recover or reconstruct missing state variables, traditionally known as state estimation in control engineering community. Our algorithm has been tested on the double pole balancing without velocities benchmark, and has achieved significantly better results on this benchmark than the published results of other algorithms to date.
Yohannes Kassahun, Jose de Gea, Mark Edgington, Jan Hendrik Metzen, Frank Kirchner
GECCO4
2008 Towards efficient online reinforcement learning using neuroevolution
abstract
For many complex Reinforcement Learning (RL) problems with large and continuous state spaces, neuroevolution has achieved promising results. This is especially true when there is noise in sensor and/or actuator signals. These results have mainly been obtained in offline learning settings, where the training and the evaluation phases of the systems are separated. In contrast, for online RL tasks, the actual performance of a system matters during its learning phase. In these tasks, neuroevolutionary systems are often impaired by their purely exploratory nature, meaning that they usually do not use (i.e. exploit) their knowledge of a single individual's performance to improve performance during learning. In this paper we describe modifications that significantly improve the online performance of the neuroevolutionary method Evolutionary Acquisition of Neural Topologies and discuss the results obtained in the Mountain Car benchmark.
Jan Hendrik Metzen, Frank Kirchner, Mark Edgington, Yohannes Kassahun
GECCO1
2008 Evolving Neural Networks for Online Reinforcement Learning
Jan Hendrik Metzen, Mark Edgington, Yohannes Kassahun, Frank Kirchner
PPSN1
2008 Learning Walking Patterns for Kinematically Complex Robots Using Evolution Strategies
Malte Römmermann, Mark Edgington, Jan Hendrik Metzen, Jose de Gea, Yohannes Kassahun, Frank Kirchner
PPSN3
2007 A common genetic encoding for both direct and indirect encodings of networks
abstract
In this paper we present a Common Genetic Encoding (CGE) for networks that can be applied to both direct and indirect encoding methods. As a direct encoding method, CGE allows the implicit evaluation of an encoded phenotype without the need to decode the phenotype from the genotype. On the other hand, one can easily decode the structure of a phenotype network, since its topology is implicitly encoded in the genotype's gene-order. Furthermore, we illustrate how CGE can be used for the indirect encoding of networks. CGE has useful properties that makes it suitable for evolving neural networks. A formal definition of the encoding is given, and some of the important properties of the encoding are proven such as its closure under mutation operators, its completeness in representing any phenotype network, and the existence of an algorithm that can evaluate any given phenotype without running into an infinite loop.
Yohannes Kassahun, Mark Edgington, Jan Hendrik Metzen, Gerald Sommer, Frank Kirchner
GECCO3
2007 Performance evaluation of EANT in the robocup keepaway benchmark
abstract
Several methods have been proposed for solving reinforcement learning (RL) problems. In addition to temporal difference (TD) methods, evolutionary algorithms (EA) are among the most promising approaches. The relative performance of these approaches in certain subdomains of the general RL problem remains an open question at this time. In addition to theoretical analysis, benchmarks are one of the most important tools for comparing different RL methods in certain problem domains. A recently proposed RL benchmark problem is the Keepaway benchmark, which is based on the RoboCup Soccer Simulator. This benchmark is one of the most challenging multiagent learning problems because its state-space is continuous and high dimensional, and both the sensors and actuators are noisy. In this paper we analyze the performance of the neuroevolutionary approach called evolutionary acquisition of neural topologies (EANT) in the Keepaway benchmark, and compare the results obtained using EANT with the results of other algorithms tested on the same benchmark.
Jan Hendrik Metzen, Mark Edgington, Yohannes Kassahun, Frank Kirchner
ICMLA1