Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jean-Baptiste Cordonnier

dblp:227/3062 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
3since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Deep learning architectures and training · 37% Efficient and distributed learning · 29% Image recognition and object detection · 19%
Databases, data mining, and information retrieval
1 paper
Graph data management · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
1.432021
Attention is not all you need: pure attention loses rank doubly exponentially with depth · ICML 2021
Group Equivariant Stand-Alone Self-Attention For Vision · ICLR 2021
On the Relationship between Self-Attention and Convolutional Layers · ICLR 2020
Machine learning › Deep learning architectures and training
transformer
0.922021
Attention is not all you need: pure attention loses rank doubly exponentially with depth · ICML 2021
On the Relationship between Self-Attention and Convolutional Layers · ICLR 2020
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.512021
Differentiable Patch Selection for Image Recognition · CVPR 2021
Computer vision › Image recognition and object detection › visual recognition
high-resolution image recognition
0.512021
Differentiable Patch Selection for Image Recognition · CVPR 2021
Machine learning › Learning theory
inductive bias
0.512021
Attention is not all you need: pure attention loses rank doubly exponentially with depth · ICML 2021
Machine learning › Efficient and distributed learning
inference efficiency
0.512021
Differentiable Patch Selection for Image Recognition · CVPR 2021
Machine learning › Efficient and distributed learning › adaptive computation
input-adaptive computation
0.512021
Differentiable Patch Selection for Image Recognition · CVPR 2021
Machine learning › Efficient and distributed learning
model compression
0.512021
Differentiable Patch Selection for Image Recognition · CVPR 2021
Machine learning › Efficient and distributed learning › data selection
patch selection
0.512021
Differentiable Patch Selection for Image Recognition · CVPR 2021
Machine learning › Deep learning architectures and training
convolutional neural network
0.412020
On the Relationship between Self-Attention and Convolutional Layers · ICLR 2020
Machine learning › Graph learning
graph neural network
0.412019
Extrapolating Paths with Graph Neural Networks · IJCAI 2019
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
0.312018
Sparsified SGD with Memory · NeurIPS 2018
Machine learning › Optimization for machine learning
stochastic gradient descent
0.312018
Sparsified SGD with Memory · NeurIPS 2018
Machine learning › Deep learning architectures and training
skip connections
0.112021
Attention is not all you need: pure attention loses rank doubly exponentially with depth · ICML 2021

Methods — techniques the papers use, named apart from their topics

graph neural network · 0.8rank analysis · 0.5path decomposition · 0.5differentiable top-k · 0.5backpropagation · 0.5sparsification · 0.3quantization · 0.3error compensation · 0.3
YearPublicationVenuePosition
2021 Differentiable Patch Selection for Image Recognition
abstract
Neural Networks require large amounts of memory and compute to process high resolution images, even when only a small part of the image is actually informative for the task at hand. We propose a method based on a differentiable Top-K operator to select the most relevant parts of the input to efficiently process high resolution images. Our method may be interfaced with any downstream neural network, is able to aggregate information from different patches in a flexible way, and allows the whole model to be trained end-to-end using backpropagation. We show results for traffic sign recognition, inter-patch relationship reasoning, and fine-grained recognition without using object/part bounding box annotations during training.
Jean-Baptiste Cordonnier, Aravindh Mahendran, Alexey Dosovitskiy, Dirk Weissenborn, Jakob Uszkoreit, Thomas Unterthiner
CVPR1
2021 Group Equivariant Stand-Alone Self-Attention For Vision
David W. Romero, Jean-Baptiste Cordonnier
ICLR2
2021 Attention is not all you need: pure attention loses rank doubly exponentially with depth
abstract
Attention-based architectures have become ubiquitous in machine learning. Yet, our understanding of the reasons for their effectiveness remains limited. This work proposes a new way to understand self-attention networks: we show that their output can be decomposed into a sum of smaller terms—or paths—each involving the operation of a sequence of attention heads across layers. Using this path decomposition, we prove that self-attention possesses a strong inductive bias towards "token uniformity". Specifically, without skip connections or multi-layer perceptrons (MLPs), the output converges doubly exponentially to a rank-1 matrix. On the other hand, skip connections and MLPs stop the output from degeneration. Our experiments verify the convergence results on standard transformer architectures.
Yihe Dong, Jean-Baptiste Cordonnier, Andreas Loukas
ICML2
2020 On the Relationship between Self-Attention and Convolutional Layers
Jean-Baptiste Cordonnier, Andreas Loukas, Martin Jaggi
ICLR1
2019 Extrapolating Paths with Graph Neural Networks
abstract
We consider the problem of path inference: given a path prefix, i.e., a partially observed sequence of nodes in a graph, we want to predict which nodes are in the missing suffix. In particular, we focus on natural paths occurring as a by-product of the interaction of an agent with a network---a driver on the transportation network, an information seeker in Wikipedia, or a client in an online shop. Our interest is sparked by the realization that, in contrast to shortest-path problems, natural paths are usually not optimal in any graph-theoretic sense, but might still follow predictable patterns. Our main contribution is a graph neural network called Gretel. Conditioned on a path prefix, this network can efficiently extrapolate path suffixes, evaluate path likelihood, and sample from the future path distribution. Our experiments with GPS traces on a road network and user-navigation paths in Wikipedia confirm that Gretel is able to adapt to graphs with very different properties, while also comparing favorably to previous solutions.
Jean-Baptiste Cordonnier, Andreas Loukas
IJCAI1
2018 Sparsified SGD with Memory
abstract
Huge scale machine learning problems are nowadays tackled by distributed optimization algorithms, i.e. algorithms that leverage the compute power of many devices for training. The communication overhead is a key bottleneck that hinders perfect scalability. Various recent works proposed to use quantization or sparsification techniques to reduce the amount of data that needs to be communicated, for instance by only sending the most significant entries of the stochastic gradient (top-k sparsification). Whilst such schemes showed very promising performance in practice, they have eluded theoretical analysis so far. In this work we analyze Stochastic Gradient Descent (SGD) with k-sparsification or compression (for instance top-k or random-k) and show that this scheme converges at the same rate as vanilla SGD when equipped with error compensation (keeping track of accumulated errors in memory). That is, communication can be reduced by a factor of the dimension of the problem (sometimes even more) whilst still converging at the same rate. We present numerical experiments to illustrate the theoretical findings and the good scalability for distributed applications.
Sebastian U. Stich, Jean-Baptiste Cordonnier, Martin Jaggi
NeurIPS2