Thomas Unterthiner

dblp:50/9446 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
7since 2021 · last 2024
0000-0001-5361-3087ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
15 papers
Deep learning architectures and training · 21% Image recognition and object detection · 16% Representation and self-supervised learning · 15%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 30 heaviest of 45, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › transformer
vision transformer
1.132021
Do Vision Transformers See Like Convolutional Neural Networks? · NeurIPS 2021
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale · ICLR 2021
Understanding Robustness of Transformers for Image Classification · ICCV 2021
Computer vision › Image recognition and object detection
image classification
1.022021
MLP-Mixer: An all-MLP Architecture for Vision · NeurIPS 2021
Understanding Robustness of Transformers for Image Classification · ICCV 2021
Machine learning › Generative modeling
generative adversarial network
0.932018
First Order Generative Adversarial Networks · ICML 2018
Coulomb GANs: Provably Optimal Nash Equilibria via Potential Fields · ICLR (Poster) 2018
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium · NIPS 2017
Machine learning › Trustworthy machine learning
calibration
0.812024
Set Learning for Accurate and Calibrated Models · ICLR 2024
Computer vision › 3D vision › geometric deep learning
set learning
0.812024
Set Learning for Accurate and Calibrated Models · ICLR 2024
Machine learning › Deep learning architectures and training › architecture learning
network growth
0.612022
GradMax: Growing Neural Networks using Gradient Information · ICLR 2022
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.522017
Rectified factor networks for biclustering of omics data · Bioinform. 2017
Rectified Factor Networks · NIPS 2015
Machine learning › Deep learning architectures and training
convolutional neural network
0.512021
Do Vision Transformers See Like Convolutional Neural Networks? · NeurIPS 2021
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.512021
Differentiable Patch Selection for Image Recognition · CVPR 2021
Computer vision › Image recognition and object detection › visual recognition
high-resolution image recognition
0.512021
Differentiable Patch Selection for Image Recognition · CVPR 2021
Machine learning › Efficient and distributed learning
inference efficiency
0.512021
Differentiable Patch Selection for Image Recognition · CVPR 2021
Machine learning › Efficient and distributed learning › adaptive computation
input-adaptive computation
0.512021
Differentiable Patch Selection for Image Recognition · CVPR 2021
Machine learning › Efficient and distributed learning
model compression
0.512021
Differentiable Patch Selection for Image Recognition · CVPR 2021
Machine learning › Deep learning architectures and training › feedforward neural network
multilayer perceptron
0.512021
MLP-Mixer: An all-MLP Architecture for Vision · NeurIPS 2021
Machine learning › Efficient and distributed learning › data selection
patch selection
0.512021
Differentiable Patch Selection for Image Recognition · CVPR 2021
Machine learning › Representation and self-supervised learning
representation analysis
0.512021
Do Vision Transformers See Like Convolutional Neural Networks? · NeurIPS 2021
Machine learning › Representation and self-supervised learning › representation analysis
representation comparison
0.512021
Do Vision Transformers See Like Convolutional Neural Networks? · NeurIPS 2021
Machine learning › Trustworthy machine learning
robustness
0.512021
Understanding Robustness of Transformers for Image Classification · ICCV 2021
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
0.412020
Object-Centric Learning with Slot Attention · NeurIPS 2020
Computer vision › Image recognition and object detection
object discovery
0.412020
Object-Centric Learning with Slot Attention · NeurIPS 2020
Machine learning › Representation and self-supervised learning › representation learning › object-centric representation learning
slot attention
0.412020
Object-Centric Learning with Slot Attention · NeurIPS 2020
Computer vision › Image recognition and object detection › object discovery
unsupervised object discovery
0.412020
Object-Centric Learning with Slot Attention · NeurIPS 2020
Machine learning › Reinforcement learning › value-based reinforcement learning
q-value estimation
0.412019
RUDDER: Return Decomposition for Delayed Rewards · NeurIPS 2019
Machine learning › Reinforcement learning › reward design › reward shaping
reward redistribution
0.412019
RUDDER: Return Decomposition for Delayed Rewards · NeurIPS 2019
Machine learning › Reinforcement learning › reward design
reward shaping
0.412019
RUDDER: Return Decomposition for Delayed Rewards · NeurIPS 2019
Machine learning › Reinforcement learning
value-based reinforcement learning
0.412019
RUDDER: Return Decomposition for Delayed Rewards · NeurIPS 2019
Knowledge, reasoning and agents › Multi-agent systems › game theory
nash equilibrium
0.312018
Coulomb GANs: Provably Optimal Nash Equilibria via Potential Fields · ICLR (Poster) 2018
Machine learning › Deep learning architectures and training
transformer
0.322021
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale · ICLR 2021
Understanding Robustness of Transformers for Image Classification · ICCV 2021
Machine learning › Deep learning architectures and training › normalization
activation normalization
0.312017
Self-Normalizing Neural Networks · NIPS 2017
Machine learning › Deep learning architectures and training
feedforward neural network
0.312017
Self-Normalizing Neural Networks · NIPS 2017

Methods — techniques the papers use, named apart from their topics

posterior regularization · 0.8alternating minimization · 0.8odd-k-out learning · 0.8cross-entropy minimization · 0.8gradient information · 0.6self-attention · 0.5large-scale pretraining · 0.5empirical robustness study · 0.5differentiable top-k · 0.5backpropagation · 0.5potential fields · 0.3coulomb potential · 0.3rectified factor networks · 0.3pre-filtering · 0.2molecule kernels · 0.2
YearPublicationVenuePosition
2024 Set Learning for Accurate and Calibrated Models
abstract
Model overconfidence and poor calibration are common in machine learning and difficult to account for when applying standard empirical risk minimization. In this work, we propose a novel method to alleviate these problems that we call odd-$k$-out learning (OKO), which minimizes the cross-entropy error for sets rather than for single examples. This naturally allows the model to capture correlations across data examples and achieves both better accuracy and calibration, especially in limited training data and class-imbalanced regimes. Perhaps surprisingly, OKO often yields better calibration even when training with hard labels and dropping any additional calibration parameter tuning, such as temperature scaling. We demonstrate this in extensive experimental analyses and provide a mathematical theory to interpret our findings. We emphasize that OKO is a general framework that can be easily adapted to many settings and a trained model can be applied to single examples at inference time, without significant run-time overhead or architecture changes.
Lukas Muttenthaler, Robert A. Vandermeulen, Qiuyi Zhang 0001, Thomas Unterthiner, Klaus-Robert Müller
ICLR4
2022 GradMax: Growing Neural Networks using Gradient Information
Utku Evci, Bart van Merrienboer, Thomas Unterthiner, Fabian Pedregosa, Max Vladymyrov
ICLR3
2021 Differentiable Patch Selection for Image Recognition
abstract
Neural Networks require large amounts of memory and compute to process high resolution images, even when only a small part of the image is actually informative for the task at hand. We propose a method based on a differentiable Top-K operator to select the most relevant parts of the input to efficiently process high resolution images. Our method may be interfaced with any downstream neural network, is able to aggregate information from different patches in a flexible way, and allows the whole model to be trained end-to-end using backpropagation. We show results for traffic sign recognition, inter-patch relationship reasoning, and fine-grained recognition without using object/part bounding box annotations during training.
Jean-Baptiste Cordonnier, Aravindh Mahendran, Alexey Dosovitskiy, Dirk Weissenborn, Jakob Uszkoreit, Thomas Unterthiner
CVPR6
2021 Understanding Robustness of Transformers for Image Classification
abstract
Deep Convolutional Neural Networks (CNNs) have long been the architecture of choice for computer vision tasks. Recently, Transformer-based architectures like Vision Transformer (ViT) have matched or even surpassed ResNets for image classification. However, details of the Transformer architecture –such as the use of non-overlapping patches– lead one to wonder whether these networks are as robust. In this paper, we perform an extensive study of a variety of different measures of robustness of ViT models and compare the findings to ResNet baselines. We investigate robustness to input perturbations as well as robustness to model perturbations. We find that when pre-trained with a sufficient amount of data, ViT models are at least as robust as the ResNet counterparts on a broad range of perturbations. We also find that Transformers are robust to the removal of almost any single layer, and that while activations from later layers are highly correlated with each other, they nevertheless play an important role in classification.
Srinadh Bhojanapalli, Ayan Chakrabarti, Daniel Glasner, Daliang Li, Thomas Unterthiner, Andreas Veit
ICCV5
2021 An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov 0003, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani 0001, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby
ICLR6
2021 Do Vision Transformers See Like Convolutional Neural Networks?
abstract
Convolutional neural networks (CNNs) have so far been the de-facto model for visual data. Recent work has shown that (Vision) Transformer models (ViT) can achieve comparable or even superior performance on image classification tasks. This raises a central question: how are Vision Transformers solving these tasks? Are they acting like convolutional networks, or learning entirely different visual representations? Analyzing the internal representation structure of ViTs and CNNs on image classification benchmarks, we find striking differences between the two architectures, such as ViT having more uniform representations across all layers. We explore how these differences arise, finding crucial roles played by self-attention, which enables early aggregation of global information, and ViT residual connections, which strongly propagate features from lower to higher layers. We study the ramifications for spatial localization, demonstrating ViTs successfully preserve input spatial information, with noticeable effects from different classification methods. Finally, we study the effect of (pretraining) dataset scale on intermediate features and transfer learning, and conclude with a discussion on connections to new architectures such as the MLP-Mixer.
Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, Alexey Dosovitskiy
NeurIPS2
2021 MLP-Mixer: An all-MLP Architecture for Vision
abstract
Convolutional Neural Networks (CNNs) are the go-to model for computer vision. Recently, attention-based networks, such as the Vision Transformer, have also become popular. In this paper we show that while convolutions and attention are both sufficient for good performance, neither of them are necessary. We present MLP-Mixer, an architecture based exclusively on multi-layer perceptrons (MLPs). MLP-Mixer contains two types of layers: one with MLPs applied independently to image patches (i.e. "mixing" the per-location features), and one with MLPs applied across patches (i.e. "mixing" spatial information). When trained on large datasets, or with modern regularization schemes, MLP-Mixer attains competitive scores on image classification benchmarks, with pre-training and inference cost comparable to state-of-the-art models. We hope that these results spark further research beyond the realms of well established CNNs and Transformers.
Ilya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov 0003, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner 0001, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, Alexey Dosovitskiy
NeurIPS6
2020 Object-Centric Learning with Slot Attention
abstract
Learning object-centric representations of complex scenes is a promising step towards enabling efficient abstract reasoning from low-level perceptual features. Yet, most deep learning approaches learn distributed representations that do not capture the compositional properties of natural scenes. In this paper, we present the Slot Attention module, an architectural component that interfaces with perceptual representations such as the output of a convolutional neural network and produces a set of task-dependent abstract representations which we call slots. These slots are exchangeable and can bind to any object in the input by specializing through a competitive procedure over multiple rounds of attention. We empirically demonstrate that Slot Attention can extract object-centric representations that enable generalization to unseen compositions when trained on unsupervised object discovery and supervised property prediction tasks.
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, Thomas Kipf
NeurIPS3
2019 RUDDER: Return Decomposition for Delayed Rewards
abstract
We propose RUDDER, a novel reinforcement learning approach for delayed rewards in finite Markov decision processes (MDPs). In MDPs the Q-values are equal to the expected immediate reward plus the expected future rewards. The latter are related to bias problems in temporal difference (TD) learning and to high variance problems in Monte Carlo (MC) learning. Both problems are even more severe when rewards are delayed. RUDDER aims at making the expected future rewards zero, which simplifies Q-value estimation to computing the mean of the immediate reward. We propose the following two new concepts to push the expected future rewards toward zero. (i) Reward redistribution that leads to return-equivalent decision processes with the same optimal policies and, when optimal, zero expected future rewards. (ii) Return decomposition via contribution analysis which transforms the reinforcement learning task into a regression task at which deep learning excels. On artificial tasks with delayed rewards, RUDDER is significantly faster than MC and exponentially faster than Monte Carlo Tree Search (MCTS), TD(λ), and reward shaping approaches. At Atari games, RUDDER on top of a Proximal Policy Optimization (PPO) baseline improves the scores, which is most prominent at games with delayed rewards.
Jose A. Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Unterthiner, Johannes Brandstetter, Sepp Hochreiter
NeurIPS4
2018 Coulomb GANs: Provably Optimal Nash Equilibria via Potential Fields
Thomas Unterthiner, Bernhard Nessler, Calvin Seward, Günter Klambauer, Martin Heusel, Hubert Ramsauer, Sepp Hochreiter
ICLR (Poster)1
2018 First Order Generative Adversarial Networks
abstract
GANs excel at learning high dimensional distributions, but they can update generator parameters in directions that do not correspond to the steepest descent direction of the objective. Prominent examples of problematic update directions include those used in both Goodfellow’s original GAN and the WGAN-GP. To formally describe an optimal update direction, we introduce a theoretical framework which allows the derivation of requirements on both the divergence and corresponding method for determining an update direction, with these requirements guaranteeing unbiased mini-batch updates in the direction of steepest descent. We propose a novel divergence which approximates the Wasserstein distance while regularizing the critic’s first order information. Together with an accompanying update direction, this divergence fulfills the requirements for unbiased steepest descent updates. We verify our method, the First Order GAN, with image generation on CelebA, LSUN and CIFAR-10 and set a new state of the art on the One Billion Word language generation task.
Calvin Seward, Thomas Unterthiner, Urs Bergmann, Nikolay Jetchev, Sepp Hochreiter
ICML2
2017 GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
abstract
Generative Adversarial Networks (GANs) excel at creating realistic images with complex models for which maximum likelihood is infeasible. However, the convergence of GAN training has still not been proved. We propose a two time-scale update rule (TTUR) for training GANs with stochastic gradient descent on arbitrary GAN loss functions. TTUR has an individual learning rate for both the discriminator and the generator. Using the theory of stochastic approximation, we prove that the TTUR converges under mild assumptions to a stationary local Nash equilibrium. The convergence carries over to the popular Adam optimization, for which we prove that it follows the dynamics of a heavy ball with friction and thus prefers flat minima in the objective landscape. For the evaluation of the performance of GANs at image generation, we introduce the `Fréchet Inception Distance'' (FID) which captures the similarity of generated images to real ones better than the Inception Score. In experiments, TTUR improves learning for DCGANs and Improved Wasserstein GANs (WGAN-GP) outperforming conventional GAN training on CelebA, CIFAR-10, SVHN, LSUN Bedrooms, and the One Billion Word Benchmark.
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, Sepp Hochreiter
NIPS3
2017 Self-Normalizing Neural Networks
abstract
Deep Learning has revolutionized vision via convolutional neural networks (CNNs) and natural language processing via recurrent neural networks (RNNs). However, success stories of Deep Learning with standard feed-forward neural networks (FNNs) are rare. FNNs that perform well are typically shallow and, therefore cannot exploit many levels of abstract representations. We introduce self-normalizing neural networks (SNNs) to enable high-level abstract representations. While batch normalization requires explicit normalization, neuron activations of SNNs automatically converge towards zero mean and unit variance. The activation function of SNNs are "scaled exponential linear units" (SELUs), which induce self-normalizing properties. Using the Banach fixed-point theorem, we prove that activations close to zero mean and unit variance that are propagated through many network layers will converge towards zero mean and unit variance -- even under the presence of noise and perturbations. This convergence property of SNNs allows to (1) train deep networks with many layers, (2) employ strong regularization, and (3) to make learning highly robust. Furthermore, for activations not close to unit variance, we prove an upper and lower bound on the variance, thus, vanishing and exploding gradients are impossible. We compared SNNs on (a) 121 tasks from the UCI machine learning repository, on (b) drug discovery benchmarks, and on (c) astronomy tasks with standard FNNs and other machine learning methods such as random forests and support vector machines. For FNNs we considered (i) ReLU networks without normalization, (ii) batch normalization, (iii) layer normalization, (iv) weight normalization, (v) highway networks, (vi) residual networks. SNNs significantly outperformed all competing FNN methods at 121 UCI tasks, outperformed all competing methods at the Tox21 dataset, and set a new record at an astronomy data set. The winning SNN architectures are often very deep.
Günter Klambauer, Thomas Unterthiner, Sepp Hochreiter
NIPS2
2017 Rectified factor networks for biclustering of omics data
abstract
MOTIVATION: Biclustering has become a major tool for analyzing large datasets given as matrix of samples times features and has been successfully applied in life sciences and e-commerce for drug design and recommender systems, respectively. actor nalysis for cluster cquisition (FABIA), one of the most successful biclustering methods, is a generative model that represents each bicluster by two sparse membership vectors: one for the samples and one for the features. However, FABIA is restricted to about 20 code units because of the high computational complexity of computing the posterior. Furthermore, code units are sometimes insufficiently decorrelated and sample membership is difficult to determine. We propose to use the recently introduced unsupervised Deep Learning approach Rectified Factor Networks (RFNs) to overcome the drawbacks of existing biclustering methods. RFNs efficiently construct very sparse, non-linear, high-dimensional representations of the input via their posterior means. RFN learning is a generalized alternating minimization algorithm based on the posterior regularization method which enforces non-negative and normalized posterior means. Each code unit represents a bicluster, where samples for which the code unit is active belong to the bicluster and features that have activating weights to the code unit belong to the bicluster. RESULTS: On 400 benchmark datasets and on three gene expression datasets with known clusters, RFN outperformed 13 other biclustering methods including FABIA. On data of the 1000 Genomes Project, RFN could identify DNA segments which indicate, that interbreeding with other hominins starting already before ancestors of modern humans left Africa. AVAILABILITY AND IMPLEMENTATION: https://github.com/bioinf-jku/librfn. CONTACT: [email protected] or [email protected].
Djork-Arné Clevert, Thomas Unterthiner, Gundula Povysil, Sepp Hochreiter
Bioinform.2
2015 Rectified Factor Networks
abstract
We propose rectified factor networks (RFNs) to efficiently construct very sparse, non-linear, high-dimensional representations of the input. RFN models identify rare and small events, have a low interference between code units, have a small reconstruction error, and explain the data covariance structure. RFN learning is a generalized alternating minimization algorithm derived from the posterior regularization method which enforces non-negative and normalized posterior means. We proof convergence and correctness of the RFN learning algorithm.On benchmarks, RFNs are compared to other unsupervised methods like autoencoders, RBMs, factor analysis, ICA, and PCA. In contrast to previous sparse coding methods, RFNs yield sparser codes, capture the data's covariance structure more precisely, and have a significantly smaller reconstruction error. We test RFNs as pretraining technique of deep networks on different vision datasets, where RFNs were superior to RBMs and autoencoders. On gene expression data from two pharmaceutical drug discovery studies, RFNs detected small and rare gene modules that revealed highly relevant new biological insights which were so far missed by other unsupervised methods.RFN package for GPU/CPU is available at http://www.bioinf.jku.at/software/rfn.
Djork-Arné Clevert, Thomas Unterthiner, Sepp Hochreiter
NIPS3
2015 Rchemcpp: a web service for structural analoging in ChEMBL, Drugbank and the Connectivity Map
abstract
UNLABELLED: We have developed Rchempp, a web service that identifies structurally similar compounds (structural analogs) in large-scale molecule databases. The service allows compounds to be queried in the widely used ChEMBL, DrugBank and the Connectivity Map databases. Rchemcpp utilizes the best performing similarity functions, i.e. molecule kernels, as measures for structural similarity. Molecule kernels have proven superior performance over other similarity measures and are currently excelling at machine learning challenges. To considerably reduce computational time, and thereby make it feasible as a web service, a novel efficient prefiltering strategy has been developed, which maintains the sensitivity of the method. By exploiting information contained in public databases, the web service facilitates many applications crucial for the drug development process, such as prioritizing compounds after screening or reducing adverse side effects during late phases. Rchemcpp was used in the DeepTox pipeline that has won the Tox21 Data Challenge and is frequently used by researchers in pharmaceutical companies. AVAILABILITY AND IMPLEMENTATION: The web service and the R package are freely available via http://shiny.bioinf.jku.at/Analoging/ and via Bioconductor. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Günter Klambauer, Martin Wischenbart, Michael Mahr, Thomas Unterthiner, Sepp Hochreiter
Bioinform.4
2011 Detection of viral sequence fragments of HIV-1 subfamilies yet unknown
abstract
BACKGROUND: Methods of determining whether or not any particular HIV-1 sequence stems - completely or in part - from some unknown HIV-1 subtype are important for the design of vaccines and molecular detection systems, as well as for epidemiological monitoring. Nevertheless, a single algorithm only, the Branching Index (BI), has been developed for this task so far. Moving along the genome of a query sequence in a sliding window, the BI computes a ratio quantifying how closely the query sequence clusters with a subtype clade. In its current version, however, the BI does not provide predicted boundaries of unknown fragments. RESULTS: We have developed Unknown Subtype Finder (USF), an algorithm based on a probabilistic model, which automatically determines which parts of an input sequence originate from a subtype yet unknown. The underlying model is based on a simple profile hidden Markov model (pHMM) for each known subtype and an additional pHMM for an unknown subtype. The emission probabilities of the latter are estimated using the emission frequencies of the known subtypes by means of a (position-wise) probabilistic model for the emergence of new subtypes. We have applied USF to SIV and HIV-1 sequences formerly classified as having emerged from an unknown subtype. Moreover, we have evaluated its performance on artificial HIV-1 recombinants and non-recombinant HIV-1 sequences. The results have been compared with the corresponding results of the BI. CONCLUSIONS: Our results demonstrate that USF is suitable for detecting segments in HIV-1 sequences stemming from yet unknown subtypes. Comparing USF with the BI shows that our algorithm performs as good as the BI or better.
Thomas Unterthiner, Anne-Kathrin Schultz, Jan Bulla, Burkhard Morgenstern, Mario Stanke, Ingo Bulla
BMC Bioinform.1