VLDB 2026 Research / reviewers in the wild / expert
Jakob Uszkoreit
dblp:87/4805
· DBLP profile ↗
27ranked-venue papers
2as first author
4since 2021 · last 2022
0000-0001-5066-7530ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
23 papers |
Deep learning architectures and training · 25% Image recognition and object detection · 13% Generative modeling · 13% | |
| Computer graphics and multimedia
3 papers |
Audio and music processing · 45% Image and video processing · 39% Visual content generation and editing · 16% |
Topics — the 30 heaviest of 51, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
autoregressive model |
1.1 | 3 | 2020 | Scaling Autoregressive Video Models · ICLR 2020 Music Transformer: Generating Music with Long-Term Structure · ICLR (Poster) 2019 Image Transformer · ICML 2018 |
Machine learning › Deep learning architectures and training
transformer |
1.0 | 4 | 2021 | Universal Transformers · ICLR (Poster) 2019 Image Transformer · ICML 2018 An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale · ICLR 2021 |
Natural language and speech › Machine translation
neural machine translation |
0.7 | 3 | 2018 | The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation · ACL (1) 2018 Attention is All you Need · NIPS 2017 Fast Decoding in Sequence Models Using Discrete Latent Variables · ICML 2018 |
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
non-autoregressive generation |
0.7 | 2 | 2019 | Insertion Transformer: Flexible Sequence Generation via Insertion Operations · ICML 2019 Blockwise Parallel Decoding for Deep Autoregressive Models · NeurIPS 2018 |
Natural language and speech › Language models and text generation › decoding › decoding strategy
parallel decoding |
0.7 | 2 | 2018 | Blockwise Parallel Decoding for Deep Autoregressive Models · NeurIPS 2018 Fast Decoding in Sequence Models Using Discrete Latent Variables · ICML 2018 |
Computer vision › 3D vision
novel view synthesis |
0.6 | 1 | 2022 | Scene Representation Transformer: Geometry-Free Novel View Synthesis Through Set-Latent Scene Representations · CVPR 2022 |
Computer vision › 3D vision › 3d scene modeling
scene representation |
0.6 | 1 | 2022 | Scene Representation Transformer: Geometry-Free Novel View Synthesis Through Set-Latent Scene Representations · CVPR 2022 |
Computer vision › Image recognition and object detection › image classification
fine-grained image classification |
0.5 | 1 | 2021 | Differentiable Patch Selection for Image Recognition · CVPR 2021 |
Computer vision › Image recognition and object detection › visual recognition
high-resolution image recognition |
0.5 | 1 | 2021 | Differentiable Patch Selection for Image Recognition · CVPR 2021 |
Computer vision › Image recognition and object detection
image classification |
0.5 | 1 | 2021 | MLP-Mixer: An all-MLP Architecture for Vision · NeurIPS 2021 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.5 | 1 | 2021 | Differentiable Patch Selection for Image Recognition · CVPR 2021 |
Machine learning › Efficient and distributed learning › adaptive computation
input-adaptive computation |
0.5 | 1 | 2021 | Differentiable Patch Selection for Image Recognition · CVPR 2021 |
Machine learning › Efficient and distributed learning
model compression |
0.5 | 1 | 2021 | Differentiable Patch Selection for Image Recognition · CVPR 2021 |
Machine learning › Deep learning architectures and training › feedforward neural network
multilayer perceptron |
0.5 | 1 | 2021 | MLP-Mixer: An all-MLP Architecture for Vision · NeurIPS 2021 |
Machine learning › Efficient and distributed learning › data selection
patch selection |
0.5 | 1 | 2021 | Differentiable Patch Selection for Image Recognition · CVPR 2021 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.5 | 1 | 2021 | An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale · ICLR 2021 |
Machine learning › Generative modeling › video generation
autoregressive video generation |
0.4 | 1 | 2020 | Scaling Autoregressive Video Models · ICLR 2020 |
Natural language and speech › Machine translation › neural machine translation
non-autoregressive machine translation |
0.4 | 1 | 2020 | An Empirical Study of Generation Order for Machine Translation · EMNLP (1) 2020 |
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning |
0.4 | 1 | 2020 | Object-Centric Learning with Slot Attention · NeurIPS 2020 |
Computer vision › Image recognition and object detection
object discovery |
0.4 | 1 | 2020 | Object-Centric Learning with Slot Attention · NeurIPS 2020 |
Machine learning › Representation and self-supervised learning › representation learning › object-centric representation learning
slot attention |
0.4 | 1 | 2020 | Object-Centric Learning with Slot Attention · NeurIPS 2020 |
Computer vision › Image recognition and object detection › object discovery
unsupervised object discovery |
0.4 | 1 | 2020 | Object-Centric Learning with Slot Attention · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation |
0.4 | 1 | 2019 | Insertion Transformer: Flexible Sequence Generation via Insertion Operations · ICML 2019 |
Machine learning › Deep learning architectures and training › transformer › recurrent transformer
universal transformer |
0.4 | 1 | 2019 | Universal Transformers · ICLR (Poster) 2019 |
Audio and music processing
music generation |
0.4 | 1 | 2019 | Music Transformer: Generating Music with Long-Term Structure · ICLR (Poster) 2019 |
Natural language and speech › Machine translation
statistical machine translation |
0.4 | 4 | 2011 | Inducing Sentence Structure from Parallel Corpora for Reordering · EMNLP 2011 'Poetic' Statistical Machine Translation: Rhyme and Meter · EMNLP 2010 Lattice-based Minimum Error Rate Training for Statistical Machine Translation · EMNLP 2008 |
Natural language and speech › Language models and text generation › decoding
autoregressive decoding |
0.3 | 1 | 2018 | Blockwise Parallel Decoding for Deep Autoregressive Models · NeurIPS 2018 |
Natural language and speech › Language models and text generation › large language model inference
decoding efficiency |
0.3 | 1 | 2018 | Blockwise Parallel Decoding for Deep Autoregressive Models · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
discrete latent variable |
0.3 | 1 | 2018 | Fast Decoding in Sequence Models Using Discrete Latent Variables · ICML 2018 |
Machine learning › Generative modeling
image generation |
0.3 | 1 | 2018 | Image Transformer · ICML 2018 |
Methods — techniques the papers use, named apart from their topics
self-attention · 1.9autoregressive modeling · 1.9transformer · 1.1attention mechanism · 0.7vision transformer · 0.6light field rendering · 0.6feed-forward inference · 0.6large-scale pretraining · 0.5differentiable top-k · 0.5backpropagation · 0.5scaling laws · 0.4relative positional encoding · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Scene Representation Transformer: Geometry-Free Novel View Synthesis Through Set-Latent Scene RepresentationsabstractA classical problem in computer vision is to infer a 3D scene representation from few images that can be used to render novel views at interactive rates. Previous work focuses on reconstructing pre-defined 3D representations, e.g. textured meshes, or implicit representations, e.g. radiance fields, and often requires input images with precise camera poses and long processing times for each novel scene. In this work, we propose the Scene Representation Transformer (SRT), a method which processes posed or unposed RGB images of a new area, infers a “set-latent scene representation ”, and synthesises novel views, all in a single feed-forward pass. To calculate the scene representation, we propose a generalization of the Vision Transformer to sets of images, enabling global information integration, and hence 3D reasoning. An efficient decoder transformer parameterizes the light field by attending into the scene representation to render novel views. Learning is supervised end-to-end by minimizing a novel-view reconstruction error. We show that this method outperforms recent baselines in terms of PSNR and speed on synthetic datasets, including a new dataset created for the paper. Further, we demonstrate that SRT scales to support interactive visualization and semantic segmentation of real-world outdoor environments using Street View imagery. Mehdi S. M. Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann, Klaus Greff, Noha Radwan, Suhani Vora, Mario Lucic, Daniel Duckworth, Alexey Dosovitskiy, Jakob Uszkoreit, Thomas A. Funkhouser, Andrea Tagliasacchi |
CVPR | 11 |
| 2021 | Differentiable Patch Selection for Image RecognitionabstractNeural Networks require large amounts of memory and compute to process high resolution images, even when only a small part of the image is actually informative for the task at hand. We propose a method based on a differentiable Top-K operator to select the most relevant parts of the input to efficiently process high resolution images. Our method may be interfaced with any downstream neural network, is able to aggregate information from different patches in a flexible way, and allows the whole model to be trained end-to-end using backpropagation. We show results for traffic sign recognition, inter-patch relationship reasoning, and fine-grained recognition without using object/part bounding box annotations during training. Jean-Baptiste Cordonnier, Aravindh Mahendran, Alexey Dosovitskiy, Dirk Weissenborn, Jakob Uszkoreit, Thomas Unterthiner |
CVPR | 5 |
| 2021 | An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov 0003, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani 0001, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby |
ICLR | 11 |
| 2021 | MLP-Mixer: An all-MLP Architecture for VisionabstractConvolutional Neural Networks (CNNs) are the go-to model for computer vision. Recently, attention-based networks, such as the Vision Transformer, have also become popular. In this paper we show that while convolutions and attention are both sufficient for good performance, neither of them are necessary. We present MLP-Mixer, an architecture based exclusively on multi-layer perceptrons (MLPs). MLP-Mixer contains two types of layers: one with MLPs applied independently to image patches (i.e. "mixing" the per-location features), and one with MLPs applied across patches (i.e. "mixing" spatial information). When trained on large datasets, or with modern regularization schemes, MLP-Mixer attains competitive scores on image classification benchmarks, with pre-training and inference cost comparable to state-of-the-art models. We hope that these results spark further research beyond the realms of well established CNNs and Transformers. Ilya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov 0003, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner 0001, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, Alexey Dosovitskiy |
NeurIPS | 10 |
| 2020 | An Empirical Study of Generation Order for Machine TranslationabstractIn this work, we present an empirical study of generation order for machine translation.Building on recent advances in insertion-based modeling, we first introduce a soft orderreward framework that enables us to train models to follow arbitrary oracle generation policies.We then make use of this framework to explore a large variety of generation orders, including uninformed orders, locationbased orders, frequency-based orders, contentbased orders, and model-based orders.Curiously, we find that for the WMT'14 English → German and WMT'18 English → Chinese translation tasks, order does not have a substantial impact on output quality.Moreover, for English → German, we even discover that unintuitive orderings such as alphabetical and shortest-first can match the performance of a standard Transformer, suggesting that traditional left-to-right generation may not be necessary to achieve high performance. Mitchell Stern, Jamie Kiros, Jakob Uszkoreit |
EMNLP (1) | 4 |
| 2020 | Scaling Autoregressive Video Models
Dirk Weissenborn, Oscar Täckström, Jakob Uszkoreit |
ICLR | 3 |
| 2020 | Object-Centric Learning with Slot AttentionabstractLearning object-centric representations of complex scenes is a promising step towards enabling efficient abstract reasoning from low-level perceptual features. Yet, most deep learning approaches learn distributed representations that do not capture the compositional properties of natural scenes. In this paper, we present the Slot Attention module, an architectural component that interfaces with perceptual representations such as the output of a convolutional neural network and produces a set of task-dependent abstract representations which we call slots. These slots are exchangeable and can bind to any object in the input by specializing through a competitive procedure over multiple rounds of attention. We empirically demonstrate that Slot Attention can extract object-centric representations that enable generalization to unseen compositions when trained on unsupervised object discovery and supervised property prediction tasks. Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, Thomas Kipf |
NeurIPS | 6 |
| 2019 | Universal Transformers
Mostafa Dehghani 0001, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, Lukasz Kaiser |
ICLR (Poster) | 4 |
| 2019 | Music Transformer: Generating Music with Long-Term Structure
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M. Dai, Matthew Hoffman 0001, Monica Dinculescu, Douglas Eck |
ICLR (Poster) | 3 |
| 2019 | Insertion Transformer: Flexible Sequence Generation via Insertion OperationsabstractWe present the Insertion Transformer, an iterative, partially autoregressive model for sequence generation based on insertion operations. Unlike typical autoregressive models which rely on a fixed, often left-to-right ordering of the output, our approach accommodates arbitrary orderings by allowing for tokens to be inserted anywhere in the sequence during decoding. This flexibility confers a number of advantages: for instance, not only can our model be trained to follow specific orderings such as left-to-right generation or a binary tree traversal, but it can also be trained to maximize entropy over all valid insertions for robustness. In addition, our model seamlessly accommodates both fully autoregressive generation (one insertion at a time) and partially autoregressive generation (simultaneous insertions at multiple locations). We validate our approach by analyzing its performance on the WMT 2014 English-German machine translation task under various settings for training and decoding. We find that the Insertion Transformer outperforms many prior non-autoregressive approaches to translation at comparable or better levels of parallelism, and successfully recovers the performance of the original Transformer while requiring only logarithmically many iterations during decoding. Mitchell Stern, Jamie Kiros, Jakob Uszkoreit |
ICML | 4 |
| 2019 | Natural Questions: a Benchmark for Question Answering ResearchabstractWe present the Natural Questions corpus, a question answering data set. Questions consist of real anonymized, aggregated queries issued to the Google search engine. An annotator is presented with a question along with a Wikipedia page from the top 5 search results, and annotates a long answer (typically a paragraph) and a short answer (one or more entities) if present on the page, or marks null if no long/short answer is present. The public release consists of 307,373 training examples with single annotations; 7,830 examples with 5-way annotations for development data; and a further 7,842 examples with 5-way annotated sequestered as test data. We present experiments validating quality of the data. We also describe analysis of 25-way annotations on 302 examples, giving insights into human variability on the annotation task. We introduce robust metrics for the purposes of evaluating question answering systems; demonstrate high human upper bounds on these metrics; and establish baseline results using competitive methods drawn from related literature. Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins 0001, Ankur P. Parikh, Christopher Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc V. Le, Slav Petrov |
Trans. Assoc. Comput. Linguistics | 16 |
| 2018 | The Best of Both Worlds: Combining Recent Advances in Neural Machine TranslationabstractMia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, Macduff Hughes. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George F. Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Macduff Hughes |
ACL (1) | 12 |
| 2018 | Fast Decoding in Sequence Models Using Discrete Latent VariablesabstractAutoregressive sequence models based on deep neural networks, such as RNNs, Wavenet and Transformer are the state-of-the-art on many tasks. However, they lack parallelism and are thus slow for long sequences. RNNs lack parallelism both during training and decoding, while architectures like WaveNet and Transformer are much more parallel during training, but still lack parallelism during decoding. We present a method to extend sequence models using discrete latent variables that makes decoding much more parallel. The main idea behind this approach is to first autoencode the target sequence into a shorter discrete latent sequence, which is generated autoregressively, and finally decode the full sequence from this shorter latent sequence in a parallel manner. To this end, we introduce a new method for constructing discrete latent variables and compare it with previously introduced methods. Finally, we verify that our model works on the task of neural machine translation, where our models are an order of magnitude faster than comparable autoregressive models and, while lower in BLEU than purely autoregressive models, better than previously proposed non-autogregressive translation. Lukasz Kaiser, Samy Bengio, Aurko Roy, Ashish Vaswani, Niki Parmar, Jakob Uszkoreit, Noam Shazeer |
ICML | 6 |
| 2018 | Image TransformerabstractImage generation has been successfully cast as an autoregressive sequence generation or transformation problem. Recent work has shown that self-attention is an effective way of modeling textual sequences. In this work, we generalize a recently proposed model architecture based on self-attention, the Transformer, to a sequence modeling formulation of image generation with a tractable likelihood. By restricting the self-attention mechanism to attend to local neighborhoods we significantly increase the size of images the model can process in practice, despite maintaining significantly larger receptive fields per layer than typical convolutional neural networks. While conceptually simple, our generative models significantly outperform the current state of the art in image generation on ImageNet, improving the best published negative log-likelihood on ImageNet from 3.83 to 3.77. We also present results on image super-resolution with a large magnification ratio, applying an encoder-decoder configuration of our architecture. In a human evaluation study, we find that images generated by our super-resolution model fool human observers three times more often than the previous state of the art. Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, Dustin Tran |
ICML | 3 |
| 2018 | Blockwise Parallel Decoding for Deep Autoregressive ModelsabstractDeep autoregressive sequence-to-sequence models have demonstrated impressive performance across a wide variety of tasks in recent years. While common architecture classes such as recurrent, convolutional, and self-attention networks make different trade-offs between the amount of computation needed per layer and the length of the critical path at training time, generation still remains an inherently sequential process. To overcome this limitation, we propose a novel blockwise parallel decoding scheme in which we make predictions for multiple time steps in parallel then back off to the longest prefix validated by a scoring model. This allows for substantial theoretical improvements in generation speed when applied to architectures that can process output sequences in parallel. We verify our approach empirically through a series of experiments using state-of-the-art self-attention models for machine translation and image super-resolution, achieving iteration reductions of up to 2x over a baseline greedy decoder with no loss in quality, or up to 7x in exchange for a slight decrease in performance. In terms of wall-clock time, our fastest models exhibit real-time speedups of up to 4x over standard greedy decoding. Mitchell Stern, Noam Shazeer, Jakob Uszkoreit |
NeurIPS | 3 |
| 2017 | Coarse-to-Fine Question Answering for Long DocumentsabstractEunsol Choi, Daniel Hewlett, Jakob Uszkoreit, Illia Polosukhin, Alexandre Lacoste, Jonathan Berant. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Eunsol Choi, Daniel Hewlett, Jakob Uszkoreit, Illia Polosukhin, Alexandre Lacoste, Jonathan Berant |
ACL (1) | 3 |
| 2017 | Attention is All you NeedabstractThe dominant sequence transduction models are based on complex recurrent orconvolutional neural networks in an encoder and decoder configuration. The best performing such models also connect the encoder and decoder through an attentionm echanisms. We propose a novel, simple network architecture based solely onan attention mechanism, dispensing with recurrence and convolutions entirely.Experiments on two machine translation tasks show these models to be superiorin quality while being more parallelizable and requiring significantly less timeto train. Our single model with 165 million parameters, achieves 27.5 BLEU onEnglish-to-German translation, improving over the existing best ensemble result by over 1 BLEU. On English-to-French translation, we outperform the previoussingle state-of-the-art with model by 0.7 BLEU, achieving a BLEU score of 41.1. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin |
NIPS | 4 |
| 2016 | A Decomposable Attention Model for Natural Language InferenceabstractWe propose a simple neural architecture for natural language inference.Our approach uses attention to decompose the problem into subproblems that can be solved separately, thus making it trivially parallelizable.On the Stanford Natural Language Inference (SNLI) dataset, we obtain state-of-the-art results with almost an order of magnitude fewer parameters than previous work and without relying on any word-order information.Adding intra-sentence attention that takes a minimum amount of order into account yields further improvements. Ankur P. Parikh, Oscar Täckström, Dipanjan Das 0001, Jakob Uszkoreit |
EMNLP | 4 |
| 2013 | Language-Independent Discriminative Parsing of Temporal Expressions
Gabor Angeli, Jakob Uszkoreit |
ACL (1) | 2 |
| 2012 | Cross-lingual Word Clusters for Direct Transfer of Linguistic Structure
Oscar Täckström, Ryan T. McDonald, Jakob Uszkoreit |
HLT-NAACL | 3 |
| 2011 | Inducing Sentence Structure from Parallel Corpora for Reordering
John DeNero, Jakob Uszkoreit |
EMNLP | 2 |
| 2011 | Watermarking the Outputs of Structured Prediction with an application in Statistical Machine Translation
Ashish Venugopal, Jakob Uszkoreit, David Talbot, Franz Josef Och, Juri Ganitkevitch |
EMNLP | 2 |
| 2010 | Large Scale Parallel Document Mining for Machine Translation
Jakob Uszkoreit, Jay Ponte, Ashok C. Popat, Moshe Dubiner |
COLING | 1 |
| 2010 | 'Poetic' Statistical Machine Translation: Rhyme and Meter
Dmitriy Genzel, Jakob Uszkoreit, Franz Josef Och |
EMNLP | 2 |
| 2009 | Creating a High-Quality Machine Translation System for a Low-Resource Language: Yiddish
Dmitriy Genzel, Klaus Macherey, Jakob Uszkoreit |
MTSummit | 3 |
| 2008 | Distributed Word Clustering for Large Scale Class-Based Language Modeling in Machine Translation
Jakob Uszkoreit, Thorsten Brants |
ACL | 1 |
| 2008 | Lattice-based Minimum Error Rate Training for Statistical Machine Translation
Wolfgang Macherey, Franz Josef Och, Ignacio Thayer, Jakob Uszkoreit |
EMNLP | 4 |