EDBT 2026 Demo / reviewers in the wild / expert
Felix A. Wichmann
dblp:42/5049
· DBLP profile ↗
15ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-2592-634XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Trustworthy machine learning · 55% Learning theory · 18% Image recognition and object detection · 14% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 50% Image and video processing · 50% |
Topics — the 15 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
robustness |
2.7 | 5 | 2025 | Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of Classifiers · NeurIPS 2025 Trivial or Impossible --- dichotomous data difficulty masks model differences (on ImageNet and beyond) · ICLR 2022 Partial success in closing the gap between human and machine vision · NeurIPS 2021 |
Machine learning › Learning theory › statistical estimation › confidence set construction
confidence intervals |
0.9 | 1 | 2025 | Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of Classifiers · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift |
0.8 | 2 | 2021 | Partial success in closing the gap between human and machine vision · NeurIPS 2021 Generalisation in humans and deep neural networks · NeurIPS 2018 |
Machine learning › Learning theory › generalization error
generalization gap |
0.6 | 1 | 2022 | Trivial or Impossible --- dichotomous data difficulty masks model differences (on ImageNet and beyond) · ICLR 2022 |
Computer vision › Image recognition and object detection
image classification |
0.6 | 1 | 2022 | Trivial or Impossible --- dichotomous data difficulty masks model differences (on ImageNet and beyond) · ICLR 2022 |
Machine learning › Trustworthy machine learning
interpretability |
0.4 | 2 | 2020 | Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency · NeurIPS 2020 Machine Learning Applied to Perception: Decision Images for Gender Classification · NIPS 2004 |
Natural language and speech › Language models and text generation › large language model evaluation
human-model comparison |
0.4 | 1 | 2020 | Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.4 | 1 | 2019 | Perceiving the arrow of time in autoregressive motion · NeurIPS 2019 |
Computer vision › Image recognition and object detection
shape bias |
0.4 | 1 | 2019 | ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness · ICLR 2019 |
Machine learning › Trustworthy machine learning › robustness › corruption robustness
image distortion robustness |
0.3 | 1 | 2018 | Generalisation in humans and deep neural networks · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning › bayesian decision theory
ideal observer model |
0.1 | 1 | 2019 | Perceiving the arrow of time in autoregressive motion · NeurIPS 2019 |
Computer vision › Image recognition and object detection
saliency prediction |
0.1 | 1 | 2006 | A Nonparametric Approach to Bottom-Up Visual Saliency · NIPS 2006 |
Image and video processing › saliency detection
bottom-up saliency |
0.1 | 1 | 2006 | A Nonparametric Approach to Bottom-Up Visual Saliency · NIPS 2006 |
Visualization and visual analytics
visual saliency |
0.1 | 1 | 2006 | A Nonparametric Approach to Bottom-Up Visual Saliency · NIPS 2006 |
Computer vision › Face, body and person analysis › facial attribute analysis
gender recognition |
0.0 | 1 | 2004 | Machine Learning Applied to Perception: Decision Images for Gender Classification · NIPS 2004 |
Methods — techniques the papers use, named apart from their topics
psychophysical experiment · 0.9significance testing · 0.9bootstrapping · 0.9self-supervised learning · 0.5adversarial training · 0.5CLIP · 0.5trial-by-trial analysis · 0.4convolutional neural network · 0.4bayesian inference · 0.4additive noise model · 0.4nonparametric learning · 0.1center-surround filter · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of ClassifiersabstractBenchmarking models is a key factor for the rapid progress in machine learning (ML) research. Thus, further progress depends on improving benchmarking metrics. A standard metric to measure the behavioral alignment between ML models and human observers is error consistency (EC). EC allows for more fine-grained comparisons of behavior than other metrics such as accuracy, and has been used in the influential Brain-Score benchmark to rank different DNNs by their behavioral consistency with humans. Previously, EC values have been reported without confidence intervals. However, empirically measured EC values are typically noisy - thus, without confidence intervals, valid benchmarking conclusions are problematic. Here we improve on standard EC in two ways: First, we show how to obtain confidence intervals for EC using a bootstrapping technique, allowing us to derive significance tests for EC. Second, we propose a new computational model relating the EC between two classifiers to the implicit probability that one of them copies responses from the other. This view of EC allows us to give practical guidance to scientists regarding the number of trials required for sufficiently powerful, conclusive experiments.
Finally, we use our methodology to revisit popular NeuroAI-results. We find that while the general trend of behavioral differences between humans and machines holds up to scrutiny, many reported differences between deep vision models are statistically insignificant. Our methodology enables researchers to design adequately powered experiments that can reliably detect behavioral differences between models, providing a foundation for more rigorous benchmarking of behavioral alignment. Thomas Klein, Sascha Meyen, Wieland Brendel, Felix A. Wichmann, Kristof Meding |
NeurIPS | 4 |
| 2022 | Trivial or Impossible --- dichotomous data difficulty masks model differences (on ImageNet and beyond)
Kristof Meding, Luca M. Schulze Buschoff, Robert Geirhos, Felix A. Wichmann |
ICLR | 4 |
| 2021 | Partial success in closing the gap between human and machine visionabstractA few years ago, the first CNN surpassed human performance on ImageNet. However, it soon became clear that machines lack robustness on more challenging test cases, a major obstacle towards deploying machines "in the wild" and towards obtaining better computational models of human visual perception. Here we ask: Are we making progress in closing the gap between human and machine vision? To answer this question, we tested human observers on a broad range of out-of-distribution (OOD) datasets, recording 85,120 psychophysical trials across 90 participants. We then investigated a range of promising machine learning developments that crucially deviate from standard supervised CNNs along three axes: objective function (self-supervised, adversarially trained, CLIP language-image training), architecture (e.g. vision transformers), and dataset size (ranging from 1M to 1B).Our findings are threefold. (1.) The longstanding distortion robustness gap between humans and CNNs is closing, with the best models now exceeding human feedforward performance on most of the investigated OOD datasets. (2.) There is still a substantial image-level consistency gap, meaning that humans make different errors than models. In contrast, most models systematically agree in their categorisation errors, even substantially different ones like contrastive self-supervised vs. standard supervised models. (3.) In many cases, human-to-model consistency improves when training dataset size is increased by one to three orders of magnitude. Our results give reason for cautious optimism: While there is still much room for improvement, the behavioural difference between human and machine vision is narrowing. In order to measure future progress, 17 OOD datasets with image-level human behavioural data and evaluation code are provided as a toolbox and benchmark at: https://github.com/bethgelab/model-vs-human/ Robert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer, Matthias Bethge, Felix A. Wichmann, Wieland Brendel |
NeurIPS | 6 |
| 2020 | Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistencyabstractA central problem in cognitive science and behavioural neuroscience as well as in machine learning and artificial intelligence research is to ascertain whether two or more decision makers---be they brains or algorithms---use the same strategy. Accuracy alone cannot distinguish between strategies: two systems may achieve similar accuracy with very different strategies. The need to differentiate beyond accuracy is particularly pressing if two systems are at or near ceiling performance, like Convolutional Neural Networks (CNNs) and humans on visual object recognition. Here we introduce trial-by-trial error consistency, a quantitative analysis for measuring whether two decision making systems systematically make errors on the same inputs. Making consistent errors on a trial-by-trial basis is a necessary condition if we want to ascertain similar processing strategies between decision makers. Our analysis is applicable to compare algorithms with algorithms, humans with humans, and algorithms with humans. When applying error consistency to visual object recognition we obtain three main findings: (1.) Irrespective of architecture, CNNs are remarkably consistent with one another. (2.) The consistency between CNNs and human observers, however, is little above what can be expected by chance alone---indicating that humans and CNNs are likely implementing very different strategies. (3.) CORnet-S, a recurrent model termed the "current best model of the primate ventral visual stream", fails to capture essential characteristics of human behavioural data and behaves essentially like a standard purely feedforward ResNet-50 in our analysis; highlighting that certain behavioural failure cases are not limited to feedforward models. Taken together, error consistency analysis suggests that the strategies used by human and machine vision are still very different---but we envision our general-purpose error consistency analysis to serve as a fruitful tool for quantifying future progress. Robert Geirhos, Kristof Meding, Felix A. Wichmann |
NeurIPS | 3 |
| 2019 | ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, Wieland Brendel |
ICLR | 5 |
| 2019 | Perceiving the arrow of time in autoregressive motionabstractUnderstanding the principles of causal inference in the visual system has a long history at least since the seminal studies by Albert Michotte. Many cognitive and machine learning scientists believe that intelligent behavior requires agents to possess causal models of the world. Recent ML algorithms exploit the dependence structure of additive noise terms for inferring causal structures from observational data, e.g. to detect the direction of time series; the arrow of time. This raises the question whether the subtle asymmetries between the time directions can also be perceived by humans. Here we show that human observers can indeed discriminate forward and backward autoregressive motion with non-Gaussian additive independent noise, i.e. they appear sensitive to subtle asymmetries between the time directions. We employ a so-called frozen noise paradigm enabling us to compare human performance with four different algorithms on a trial-by-trial basis: A causal inference algorithm exploiting the dependence structure of additive noise terms, a neurally inspired network, a Bayesian ideal observer model as well as a simple heuristic. Our results suggest that all human observers use similar cues or strategies to solve the arrow of time motion discrimination task, but the human algorithm is significantly different from the three machine algorithms we compared it to. In fact, our simple heuristic appears most similar to our human observers. Kristof Meding, Dominik Janzing, Bernhard Schölkopf, Felix A. Wichmann |
NeurIPS | 4 |
| 2019 | Neural Signatures of Motor Skill in the Resting BrainabstractStroke-induced disturbances of large-scale cortical networks are known to be associated with the extent of motor deficits. We argue that identifying brain networks representative of motor behavior in the resting brain would provide significant insights for current neurorehabilitation approaches. Particularly, we aim to investigate the global configuration of brain rhythms and their relation to motor skill, instead of learning performance as broadly studied. We empirically approach this problem by conducting a three-dimensional physical space visuomotor learning experiment during electroencephalographic (EEG) data recordings with thirty-seven healthy participants. We demonstrate that across-subjects variations in average movement smoothness as the quantified measure of subjects' motor skills can be predicted from the global configuration of resting-state EEG alpha-rhythms (8-14 Hz) recorded prior to the experiment. Importantly, this neural signature of motor skill was found to be orthogonal to (independent of) task-as well as to learning-related changes in alpha-rhythms, which we interpret as an organizing principle of the brain. We argue that disturbances of such configurations in the brain may contribute to motor deficits in stroke, and that reconfiguring stroke patients' brain rhythms by neurofeedback may enhance post-stroke neurorehabilitation. Ozan Özdenizci, Timm Meyer, Felix A. Wichmann, Jan Peters 0001, Bernhard Schölkopf, Müjdat Çetin, Moritz Grosse-Wentrup |
SMC | 3 |
| 2018 | Generalisation in humans and deep neural networksabstractWe compare the robustness of humans and current convolutional deep neural networks (DNNs) on object recognition under twelve different types of image degradations. First, using three well known DNNs (ResNet-152, VGG-19, GoogLeNet) we find the human visual system to be more robust to nearly all of the tested image manipulations, and we observe progressively diverging classification error-patterns between humans and DNNs when the signal gets weaker. Secondly, we show that DNNs trained directly on distorted images consistently surpass human performance on the exact distortion types they were trained on, yet they display extremely poor generalisation abilities when tested on other distortion types. For example, training on salt-and-pepper noise does not imply robustness on uniform white noise and vice versa. Thus, changes in the noise distribution between training and testing constitutes a crucial challenge to deep learning vision systems that can be systematically addressed in a lifelong machine learning approach. Our new dataset consisting of 83K carefully measured human psychophysical trials provide a useful reference for lifelong robustness against image degradations set by the human visual system. Robert Geirhos, Carlos R. Medina Temme, Jonas Rauber, Heiko H. Schütt, Matthias Bethge, Felix A. Wichmann |
NeurIPS | 6 |
| 2013 | How Sensitive Is the Human Visual System to the Local Statistics of Natural Images?abstractA key hypothesis in sensory system neuroscience is that sensory representations are adapted to the statistical regularities in sensory signals and thereby incorporate knowledge about the outside world. Supporting this hypothesis, several probabilistic models of local natural image regularities have been proposed that reproduce neural response properties. Although many such physiological links have been made, these models have not been linked directly to visual sensitivity. Previous psychophysical studies of sensitivity to natural image regularities focus on global perception of large images, but much less is known about sensitivity to local natural image regularities. We present a new paradigm for controlled psychophysical studies of local natural image regularities and compare how well such models capture perceptually relevant image content. To produce stimuli with precise statistics, we start with a set of patches cut from natural images and alter their content to generate a matched set whose joint statistics are equally likely under a probabilistic natural image model. The task is forced choice to discriminate natural patches from model patches. The results show that human observers can learn to discriminate the higher-order regularities in natural images from those of model samples after very few exposures and that no current model is perfect for patches as small as 5 by 5 pixels or larger. Discrimination performance was accurately predicted by model likelihood, an information theoretic measure of model efficacy, indicating that the visual system possesses a surprisingly detailed knowledge of natural image higher-order correlations, much more so than current image models. We also perform three cue identification experiments to interpret how model features correspond to perceptually relevant image features. Holly E. Gerhard, Felix A. Wichmann, Matthias Bethge |
PLoS Comput. Biol. | 2 |
| 2012 | A New Perceptual Bias Reveals Suboptimal Population Decoding of Sensory ResponsesabstractSeveral studies have reported optimal population decoding of sensory responses in two-alternative visual discrimination tasks. Such decoding involves integrating noisy neural responses into a more reliable representation of the likelihood that the stimuli under consideration evoked the observed responses. Importantly, an ideal observer must be able to evaluate likelihood with high precision and only consider the likelihood of the two relevant stimuli involved in the discrimination task. We report a new perceptual bias suggesting that observers read out the likelihood representation with remarkably low precision when discriminating grating spatial frequencies. Using spectrally filtered noise, we induced an asymmetry in the likelihood function of spatial frequency. This manipulation mainly affects the likelihood of spatial frequencies that are irrelevant to the task at hand. Nevertheless, we find a significant shift in perceived grating frequency, indicating that observers evaluate likelihoods of a broad range of irrelevant frequencies and discard prior knowledge of stimulus alternatives when performing two-alternative discrimination. Tom Putzeys, Matthias Bethge, Felix A. Wichmann, Johan Wagemans, Robbe L. T. Goris |
PLoS Comput. Biol. | 3 |
| 2006 | A Nonparametric Approach to Bottom-Up Visual SaliencyabstractThis paper addresses the bottom-up influence of local image information on human eye movements. Most existing computational models use a set of biologically plausible linear filters, e.g., Gabor or Difference-of-Gaussians filters as a front-end, the outputs of which are nonlinearly combined into a real number that indicates visual saliency. Unfortunately, this requires many design parameters such as the number, type, and size of the front-end filters, as well as the choice of nonlinearities, weighting and normalization schemes etc., for which biological plausibility cannot always be justified. As a result, these parameters have to be chosen in a more or less ad hoc way. Here, we propose to learn a visual saliency model directly from human eye movement data. The model is rather simplistic and essentially parameter-free, and therefore contrasts recent developments in the field that usually aim at higher prediction rates at the cost of additional parameters and increasing model complexity. Experimental results show that--despite the lack of any biological prior knowledge--our model performs comparably to existing approaches, and in fact learns image features that resemble findings from several previous studies. In particular, its maximally excitatory stimuli have center-surround structure, similar to receptive fields in the early human visual system. Wolf Kienzle, Felix A. Wichmann, Bernhard Schölkopf, Matthias O. Franz |
NIPS | 2 |
| 2006 | Inducing Metric Violations in Human Similarity JudgementsabstractAttempting to model human categorization and similarity judgements is both a very interesting but also an exceedingly difficult challenge. Some of the difficulty arises because of conflicting evidence whether human categorization and similarity judgements should or should not be modelled as to operate on a mental representation that is essentially metric. Intuitively, this has a strong appeal as it would allow (dis)similarity to be represented geometrically as distance in some internal space. Here we show how a single stimulus, carefully constructed in a psychophysical experiment, introduces l2 violations in what used to be an internal similarity space that could be adequately modelled as Euclidean. We term this one influential data point a conflictual judgement. We present an algorithm of how to analyse such data and how to identify the crucial point. Thus there may not be a strict dichotomy between either a metric or a non-metric internal space but rather degrees to which potentially large subsets of stimuli are represented metrically with a small subset causing a global violation of metricity. Julian Laub, Jakob H. Macke, Klaus-Robert Müller, Felix A. Wichmann |
NIPS | 4 |
| 2006 | Classification of Faces in Man and MachineabstractWe attempt to shed light on the algorithms humans use to classify images of human faces according to their gender. For this, a novel methodology combining human psychophysics and machine learning is introduced. We proceed as follows. First, we apply principal component analysis (PCA) on the pixel information of the face stimuli. We then obtain a data set composed of these PCA eigenvectors combined with the subjects' gender estimates of the corresponding stimuli. Second, we model the gender classification process on this data set using a separating hyperplane (SH) between both classes. This SH is computed using algorithms from machine learning: the support vector machine (SVM), the relevance vector machine, the prototype classifier, and the K-means classifier. The classification behavior of humans and machines is then analyzed in three steps. First, the classification errors of humans and machines are compared for the various classifiers, and we also assess how well machines can recreate the subjects' internal decision boundary by studying the training errors of the machines. Second, we study the correlations between the rank-order of the subjects' responses to each stimulus-the gender estimate with its reaction time and confidence rating-and the rank-order of the distance of these stimuli to the SH. Finally, we attempt to compare the metric of the representations used by humans and machines for classification by relating the subjects' gender estimate of each stimulus and the distance of this stimulus to the SH. While we show that the classification error alone is not a sufficient selection criterion between the different algorithms humans might use to classify face stimuli, the distance of these stimuli to the SH is shown to capture essentials of the internal decision space of humans. Furthermore, algorithms such as the prototype classifier using stimuli in the center of the classes are shown to be less adapted to model human classification behavior than algorithms such as the SVM based on stimuli close to the boundary between the classes. Arnulf B. A. Graf, Felix A. Wichmann, Heinrich H. Bülthoff, Bernhard Schölkopf |
Neural Comput. | 2 |
| 2004 | Machine Learning Applied to Perception: Decision Images for Gender ClassificationabstractWe study gender discrimination of human faces using a combination of psychophysical classification and discrimination experiments together with methods from machine learning. We reduce the dimensionality of a set of face images using principal component analysis, and then train a set of linear classifiers on this reduced representation (linear support vec- tor machines (SVMs), relevance vector machines (RVMs), Fisher linear discriminant (FLD), and prototype (prot) classifiers) using human clas- sification data. Because we combine a linear preprocessor with linear classifiers, the entire system acts as a linear classifier, allowing us to visu- alise the decision-image corresponding to the normal vector of the separ- ating hyperplanes (SH) of each classifier. We predict that the female-to- maleness transition along the normal vector for classifiers closely mim- icking human classification (SVM and RVM [1]) should be faster than the transition along any other direction. A psychophysical discrimina- tion experiment using the decision images as stimuli is consistent with this prediction. Felix A. Wichmann, Arnulf B. A. Graf, Eero P. Simoncelli, Heinrich H. Bülthoff, Bernhard Schölkopf |
NIPS | 1 |
| 2003 | Insights from Machine Learning Applied to Human Visual ClassificationabstractWe attempt to understand visual classification in humans using both psy- chophysical and machine learning techniques. Frontal views of human faces were used for a gender classification task. Human subjects classi- fied the faces and their gender judgment, reaction time and confidence rating were recorded. Several hyperplane learning algorithms were used on the same classification task using the Principal Components of the texture and shape representation of the faces. The classification perfor- mance of the learning algorithms was estimated using the face database with the true gender of the faces as labels, and also with the gender es- timated by the subjects. We then correlated the human responses to the distance of the stimuli to the separating hyperplane of the learning algo- rithms. Our results suggest that human classification can be modeled by some hyperplane algorithms in the feature space we used. For classifica- tion, the brain needs more processing for stimuli close to that hyperplane than for those further away. Arnulf B. A. Graf, Felix A. Wichmann |
NIPS | 2 |