EDBT 2026 Demo / reviewers in the wild / expert
Arthur Mensch
dblp:156/2229
· DBLP profile ↗
14ranked-venue papers
6as first author
6since 2021 · last 2022
0000-0001-7866-0461ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Deep learning architectures and training · 21% Language models and text generation · 16% Learning theory · 11% | |
| Theoretical computer science
3 papers |
Algorithmic game theory and mechanism design · 60% Mathematical optimization · 40% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 50% Recommender systems · 22% Data mining · 22% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Medical and health informatics · 100% |
Topics — the 30 heaviest of 38, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
scaling laws |
1.1 | 2 | 2022 | An empirical analysis of compute-optimal large language model training · NeurIPS 2022 Unified Scaling Laws for Routed Language Models · ICML 2022 |
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training |
0.6 | 1 | 2022 | An empirical analysis of compute-optimal large language model training · NeurIPS 2022 |
Machine learning › Transfer learning and domain adaptation › few-shot learning
cross-modal few-shot learning |
0.6 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Natural language and speech › Language models and text generation
in-context learning |
0.6 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Natural language and speech › Question answering and dialogue systems
knowledge-intensive tasks |
0.6 | 1 | 2022 | Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.6 | 1 | 2022 | Unified Scaling Laws for Routed Language Models · ICML 2022 |
Natural language and speech › Language models and text generation
multimodal language model |
0.6 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Natural language and speech › Language models and text generation
retrieval-augmented language models |
0.6 | 1 | 2022 | Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022 |
Computer vision › Vision and language
vision-language model |
0.6 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Information retrieval
document retrieval |
0.6 | 1 | 2022 | Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022 |
Medical and health informatics › neuroimaging
neuroimaging analysis |
0.5 | 2 | 2017 | Learning Neural Representations of Human Cognition across Many fMRI Studies · NIPS 2017 Learning brain regions via large-scale online structured sparse dictionary learning · NIPS 2016 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2020 | A mean-field analysis of two-player zero-sum games · NeurIPS 2020 |
Knowledge, reasoning and agents › Multi-agent systems › equilibrium computation
nash equilibrium computation |
0.4 | 1 | 2020 | Extra-gradient with player sampling for faster convergence in n-player games · ICML 2020 |
Machine learning › Learning theory › online learning
online estimation |
0.4 | 1 | 2020 | Online Sinkhorn: Optimal Transport distances from sample streams · NeurIPS 2020 |
Mathematical optimization › optimal transport
entropic optimal transport |
0.4 | 1 | 2020 | Online Sinkhorn: Optimal Transport distances from sample streams · NeurIPS 2020 |
Algorithmic game theory and mechanism design
equilibrium computation |
0.4 | 1 | 2020 | Extra-gradient with player sampling for faster convergence in n-player games · ICML 2020 |
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts › nash equilibrium
mixed nash equilibrium |
0.4 | 1 | 2020 | A mean-field analysis of two-player zero-sum games · NeurIPS 2020 |
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts
nash equilibrium |
0.4 | 1 | 2020 | A mean-field analysis of two-player zero-sum games · NeurIPS 2020 |
Mathematical optimization
optimal transport |
0.4 | 1 | 2020 | Online Sinkhorn: Optimal Transport distances from sample streams · NeurIPS 2020 |
Machine learning › Learning theory
distribution learning |
0.4 | 1 | 2019 | Geometric Losses for Distributional Learning · ICML 2019 |
Machine learning › Learning theory
loss function |
0.4 | 1 | 2019 | Geometric Losses for Distributional Learning · ICML 2019 |
Machine learning › Optimization for machine learning
optimal transport |
0.4 | 1 | 2019 | Geometric Losses for Distributional Learning · ICML 2019 |
Machine learning › Representation and self-supervised learning › representation learning › joint representation learning
multi-task representation learning |
0.3 | 1 | 2017 | Learning Neural Representations of Human Cognition across Many fMRI Studies · NIPS 2017 |
Medical and health informatics › neuroimaging
fMRI decoding |
0.3 | 1 | 2017 | Learning Neural Representations of Human Cognition across Many fMRI Studies · NIPS 2017 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning |
0.2 | 1 | 2016 | Learning brain regions via large-scale online structured sparse dictionary learning · NIPS 2016 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding |
0.2 | 1 | 2016 | Learning brain regions via large-scale online structured sparse dictionary learning · NIPS 2016 |
Data mining › representation learning
dictionary learning |
0.2 | 1 | 2016 | Dictionary Learning for Massive Matrix Factorization · ICML 2016 |
Recommender systems › collaborative filtering
matrix factorization |
0.2 | 1 | 2016 | Dictionary Learning for Massive Matrix Factorization · ICML 2016 |
Computer vision › Vision and language
visual question answering |
0.2 | 1 | 2022 | Flamingo: a Visual Language Model for Few-Shot Learning · NeurIPS 2022 |
Machine learning › Optimization for machine learning
convex optimization |
0.1 | 1 | 2019 | Geometric Losses for Distributional Learning · ICML 2019 |
Methods — techniques the papers use, named apart from their topics
differentiable encoder · 1.1chunked cross-attention · 1.1player sampling · 0.9extragradient · 0.9scaling law estimation · 0.6pretrained vision encoder · 0.6pre-trained language model · 0.6power-law scaling · 0.6interleaved multimodal pretraining · 0.6effective parameter count · 0.6randomized method · 0.5online optimization · 0.5coordinate descent · 0.5variance reduction · 0.4sinkhorn algorithm · 0.4sample streaming · 0.4mirror descent · 0.4mean-field analysis · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Improving Language Models by Retrieving from Trillions of TokensabstractWe enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a 2 trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance to GPT-3 and Jurassic-1 on the Pile, despite using 25{\texttimes} fewer parameters. After fine-tuning, RETRO performance translates to downstream knowledge-intensive tasks such as question answering. RETRO combines a frozen Bert retriever, a differentiable encoder and a chunked cross-attention mechanism to predict tokens based on an order of magnitude more data than what is typically consumed during training. We typically train RETRO from scratch, yet can also rapidly RETROfit pre-trained transformers with retrieval and still achieve good performance. Our work opens up new avenues for improving language models through explicit memory at unprecedented scale. Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche 0002, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Albin Cassirer, Andrew Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, Laurent Sifre |
ICML | 2 |
| 2022 | Unified Scaling Laws for Routed Language ModelsabstractThe performance of a language model has been shown to be effectively modeled as a power-law in its parameter count. Here we study the scaling behaviors of Routing Networks: architectures that conditionally use only a subset of their parameters while processing an input. For these models, parameter count and computational requirement form two independent axes along which an increase leads to better performance. In this work we derive and justify scaling laws defined on these two variables which generalize those known for standard language models and describe the performance of a wide range of routing architectures trained via three different techniques. Afterwards we provide two applications of these laws: first deriving an Effective Parameter Count along which all models scale at the same rate, and then using the scaling coefficients to give a quantitative comparison of the three routing techniques considered. Our analysis derives from an extensive evaluation of Routing Networks across five orders of magnitude of size, including models with hundreds of experts and hundreds of billions of parameters. Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake A. Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche 0002, Eliza Rutherford, Tom Hennigan, Matthew J. Johnson 0002, Albin Cassirer, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Marc'Aurelio Ranzato, Jack W. Rae, Erich Elsen, Koray Kavukcuoglu, Karen Simonyan |
ICML | 4 |
| 2022 | Flamingo: a Visual Language Model for Few-Shot LearningabstractBuilding models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. We propose key architectural innovations to: (i) bridge powerful pretrained vision-only and language-only models, (ii) handle sequences of arbitrarily interleaved visual and textual data, and (iii) seamlessly ingest images or videos as inputs. Thanks to their flexibility, Flamingo models can be trained on large-scale multimodal web corpora containing arbitrarily interleaved text and images, which is key to endow them with in-context few-shot learning capabilities. We perform a thorough evaluation of our models, exploring and measuring their ability to rapidly adapt to a variety of image and video tasks. These include open-ended tasks such as visual question-answering, where the model is prompted with a question which it has to answer, captioning tasks, which evaluate the ability to describe a scene or an event, and close-ended tasks such as multiple-choice visual question-answering. For tasks lying anywhere on this spectrum, a single Flamingo model can achieve a new state of the art with few-shot learning, simply by prompting the model with task-specific examples. On numerous benchmarks, Flamingo outperforms models fine-tuned on thousands of times more task-specific data. Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, Karen Simonyan |
NeurIPS | 8 |
| 2022 | An empirical analysis of compute-optimal large language model trainingabstractWe investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4$\times$ more data. Chinchilla uniformly and significantly outperformsGopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, a 7% improvement over Gopher. Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche 0002, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack W. Rae, Laurent Sifre |
NeurIPS | 3 |
| 2021 | Differentiable Divergences Between Time SeriesabstractComputing the discrepancy between time series of variable sizes is notoriously challenging. While dynamic time warping (DTW) is popularly used for this purpose, it is not differentiable everywhere and is known to lead to bad local optima when used as a “loss”. Soft-DTW addresses these issues, but it is not a positive definite divergence: due to the bias introduced by entropic regularization, it can be negative and it is not minimized when the time series are equal. We propose in this paper a new divergence, dubbed soft-DTW divergence, which aims to correct these issues. We study its properties; in particular, under conditions on the ground cost, we show that it is a valid divergence: it is non-negative and minimized if and only if the two time series are equal. We also propose a new “sharp” variant by further removing entropic bias. We showcase our divergences on time series averaging and demonstrate significant accuracy improvements compared to both DTW and soft-DTW on 84 time series classification datasets. Mathieu Blondel, Arthur Mensch, Jean-Philippe Vert |
AISTATS | 2 |
| 2021 | Extracting representations of cognition across neuroimaging studies improves brain decodingabstractCognitive brain imaging is accumulating datasets about the neural substrate of many different mental processes. Yet, most studies are based on few subjects and have low statistical power. Analyzing data across studies could bring more statistical power; yet the current brain-imaging analytic framework cannot be used at scale as it requires casting all cognitive tasks in a unified theoretical framework. We introduce a new methodology to analyze brain responses across tasks without a joint model of the psychological processes. The method boosts statistical power in small studies with specific cognitive focus by analyzing them jointly with large studies that probe less focal mental processes. Our approach improves decoding performance for 80% of 35 widely-different functional-imaging studies. It finds commonalities across tasks in a data-driven way, via common brain representations that predict mental processes. These are brain networks tuned to psychological manipulations. They outline interpretable and plausible brain structures. The extracted networks have been made available; they can be readily reused in new neuro-imaging studies. We provide a multi-study decoding tool to adapt to new data. Arthur Mensch, Julien Mairal, Bertrand Thirion, Gaël Varoquaux |
PLoS Comput. Biol. | 1 |
| 2020 | Extra-gradient with player sampling for faster convergence in n-player gamesabstractData-driven modeling increasingly requires to find a Nash equilibrium in multi-player games, e.g. when training GANs. In this paper, we analyse a new extra-gradient method for Nash equilibrium finding, that performs gradient extrapolations and updates on a random subset of players at each iteration. This approach provably exhibits a better rate of convergence than full extra-gradient for non-smooth convex games with noisy gradient oracle. We propose an additional variance reduction mechanism to obtain speed-ups in smooth convex games. Our approach makes extrapolation amenable to massive multiplayer settings, and brings empirical speed-ups, in particular when using a heuristic cyclic sampling scheme. Most importantly, it allows to train faster and better GANs and mixtures of GANs. Samy Jelassi, Carles Domingo-Enrich, Damien Scieur, Arthur Mensch, Joan Bruna |
ICML | 4 |
| 2020 | A mean-field analysis of two-player zero-sum gamesabstractFinding Nash equilibria in two-player zero-sum continuous games is a central problem in machine learning, e.g. for training both GANs and robust models. The existence of pure Nash equilibria requires strong conditions which are not typically met in practice. Mixed Nash equilibria exist in greater generality and may be found using mirror descent. Yet this approach does not scale to high dimensions. To address this limitation, we parametrize mixed strategies as mixtures of particles, whose positions and weights are updated using gradient descent-ascent. We study this dynamics as an interacting gradient flow over measure spaces endowed with the Wasserstein-Fisher-Rao metric. We establish global convergence to an approximate equilibrium for the related Langevin gradient-ascent dynamic. We prove a law of large numbers that relates particle dynamics to mean-field dynamics. Our method identifies mixed equilibria in high dimensions and is demonstrably effective for training mixtures of GANs. Carles Domingo-Enrich, Samy Jelassi, Arthur Mensch, Grant M. Rotskoff, Joan Bruna |
NeurIPS | 3 |
| 2020 | Online Sinkhorn: Optimal Transport distances from sample streamsabstractOptimal Transport (OT) distances are now routinely used as loss functions in ML tasks. Yet, computing OT distances between arbitrary (i.e. not necessarily discrete) probability distributions remains an open problem. This paper introduces a new online estimator of entropy-regularized OT distances between two such arbitrary distributions. It uses streams of samples from both distributions to iteratively enrich a non-parametric representation of the transportation plan. Compared to the classic Sinkhorn algorithm, our method leverages new samples at each iteration, which enables a consistent estimation of the true regularized OT distance. We provide a theoretical analysis of the convergence of the online Sinkhorn algorithm, showing a nearly-1/n asymptotic sample complexity for the iterate sequence. We validate our method on synthetic 1-d to 10-d data and on real 3-d shape data. Arthur Mensch, Gabriel Peyré |
NeurIPS | 1 |
| 2019 | Geometric Losses for Distributional LearningabstractBuilding upon recent advances in entropy-regularized optimal transport, and upon Fenchel duality between measures and continuous functions, we propose a generalization of the logistic loss that incorporates a metric or cost between classes. Unlike previous attempts to use optimal transport distances for learning, our loss results in unconstrained convex objective functions, supports infinite (or very large) class spaces, and naturally defines a geometric generalization of the softmax operator. The geometric properties of this loss make it suitable for predicting sparse and singular distributions, for instance supported on curves or hyper-surfaces. We study the theoretical properties of our loss and showcase its effectiveness on two applications: ordinal regression and drawing generation. Arthur Mensch, Mathieu Blondel, Gabriel Peyré |
ICML | 1 |
| 2018 | Differentiable Dynamic Programming for Structured Prediction and AttentionabstractDynamic programming (DP) solves a variety of structured combinatorial problems by iteratively breaking them down into smaller subproblems. In spite of their versatility, many DP algorithms are non-differentiable, which hampers their use as a layer in neural networks trained by backpropagation. To address this issue, we propose to smooth the max operator in the dynamic programming recursion, using a strongly convex regularizer. This allows to relax both the optimal value and solution of the original combinatorial problem, and turns a broad class of DP algorithms into differentiable operators. Theoretically, we provide a new probabilistic perspective on backpropagating through these DP operators, and relate them to inference in graphical models. We derive two particular instantiations of our framework, a smoothed Viterbi algorithm for sequence prediction and a smoothed DTW algorithm for time-series alignment. We showcase these instantiations on structured prediction (audio-to-score alignment, NER) and on structured and sparse attention for translation. Arthur Mensch, Mathieu Blondel |
ICML | 1 |
| 2017 | Learning Neural Representations of Human Cognition across Many fMRI StudiesabstractCognitive neuroscience is enjoying rapid increase in extensive public brain-imaging datasets. It opens the door to large-scale statistical models. Finding a unified perspective for all available data calls for scalable and automated solutions to an old challenge: how to aggregate heterogeneous information on brain function into a universal cognitive system that relates mental operations/cognitive processes/psychological tasks to brain networks? We cast this challenge in a machine-learning approach to predict conditions from statistical brain maps across different studies. For this, we leverage multi-task learning and multi-scale dimension reduction to learn low-dimensional representations of brain images that carry cognitive information and can be robustly associated with psychological stimuli. Our multi-dataset classification model achieves the best prediction performance on several large reference datasets, compared to models without cognitive-aware low-dimension representations; it brings a substantial performance boost to the analysis of small datasets, and can be introspected to identify universal template cognitive concepts. Arthur Mensch, Julien Mairal, Danilo Bzdok, Bertrand Thirion, Gaël Varoquaux |
NIPS | 1 |
| 2016 | Dictionary Learning for Massive Matrix FactorizationabstractSparse matrix factorization is a popular tool to obtain interpretable data decompositions, which are also effective to perform data completion or denoising. Its applicability to large datasets has been addressed with online and randomized methods, that reduce the complexity in one of the matrix dimension, but not in both of them. In this paper, we tackle very large matrices in both dimensions. We propose a new factorization method that scales gracefully to terabyte-scale datasets. Those could not be processed by previous algorithms in a reasonable amount of time. We demonstrate the efficiency of our approach on massive functional Magnetic Resonance Imaging (fMRI) data, and on matrix completion problems for recommender systems, where we obtain significant speed-ups compared to state-of-the art coordinate descent methods. Arthur Mensch, Julien Mairal, Bertrand Thirion, Gaël Varoquaux |
ICML | 1 |
| 2016 | Learning brain regions via large-scale online structured sparse dictionary learningabstractWe propose a multivariate online dictionary-learning method for obtaining decompositions of brain images with structured and sparse components (aka atoms). Sparsity is to be understood in the usual sense: the dictionary atoms are constrained to contain mostly zeros. This is imposed via an $\ell_1$-norm constraint. By "structured", we mean that the atoms are piece-wise smooth and compact, thus making up blobs, as opposed to scattered patterns of activation. We propose to use a Sobolev (Laplacian) penalty to impose this type of structure. Combining the two penalties, we obtain decompositions that properly delineate brain structures from functional images. This non-trivially extends the online dictionary-learning work of Mairal et al. (2010), at the price of only a factor of 2 or 3 on the overall running time. Just like the Mairal et al. (2010) reference method, the online nature of our proposed algorithm allows it to scale to arbitrarily sized datasets. Experiments on brain data show that our proposed method extracts structured and denoised dictionaries that are more intepretable and better capture inter-subject variability in small medium, and large-scale regimes alike, compared to state-of-the-art models. Elvis Dohmatob, Arthur Mensch, Gaël Varoquaux, Bertrand Thirion |
NIPS | 2 |