Angelos Katharopoulos

dblp:188/1159 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Deep learning architectures and training · 35% Efficient and distributed learning · 24% Representation and self-supervised learning · 9%
Computer graphics and multimedia
2 papers
Geometric modeling and processing · 87% Multimedia analysis and retrieval · 13%

Topics — the 27 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
mixture of experts
1.722025
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging · ICML 2025
No Need to Talk: Asynchronous Mixture of Language Models · ICLR 2025
Machine learning › Efficient and distributed learning › distributed training
asynchronous training
0.912025
No Need to Talk: Asynchronous Mixture of Language Models · ICLR 2025
Machine learning › Efficient and distributed learning
distributed training
0.912025
No Need to Talk: Asynchronous Mixture of Language Models · ICLR 2025
Natural language and speech › Language models and text generation
model routing
0.912025
No Need to Talk: Asynchronous Mixture of Language Models · ICLR 2025
Machine learning › Transfer learning and domain adaptation › model adaptation
model specialization
0.912025
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging · ICML 2025
Machine learning › Efficient and distributed learning
parameter averaging
0.912025
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging · ICML 2025
Computer vision › Vision and language › vision-language model
CLIP
0.712023
Masked Autoencoding Does Not Help Natural Language Supervision at Scale · CVPR 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder
0.712023
Masked Autoencoding Does Not Help Natural Language Supervision at Scale · CVPR 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling
0.712023
Masked Autoencoding Does Not Help Natural Language Supervision at Scale · CVPR 2023
Computer vision › 3D vision
3d shape reconstruction
0.512021
Neural Parts: Learning Expressive 3D Shape Abstractions With Invertible Neural Networks · CVPR 2021
Geometric modeling and processing
shape decomposition
0.512021
Neural Parts: Learning Expressive 3D Shape Abstractions With Invertible Neural Networks · CVPR 2021
Machine learning › Efficient and distributed learning › attention efficiency
attention approximation
0.412020
Fast Transformers with Clustered Attention · NeurIPS 2020
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
autoregressive generation
0.412020
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention · ICML 2020
Machine learning › Deep learning architectures and training › attention mechanism › efficient attention
clustered attention
0.412020
Fast Transformers with Clustered Attention · NeurIPS 2020
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention
0.412020
Fast Transformers with Clustered Attention · NeurIPS 2020
Machine learning › Deep learning architectures and training › transformer
efficient transformer
0.412020
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention · ICML 2020
Machine learning › Deep learning architectures and training › attention mechanism › efficient attention
linear attention
0.412020
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention · ICML 2020
Machine learning › Efficient and distributed learning
model compression
0.412020
Fast Transformers with Clustered Attention · NeurIPS 2020
Machine learning › Deep learning architectures and training
transformer
0.412020
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention · ICML 2020
Machine learning › Deep learning architectures and training › attention mechanism
transformer attention
0.412020
Fast Transformers with Clustered Attention · NeurIPS 2020
Computer vision › Image recognition and object detection
image classification
0.412019
Processing Megapixel Images with Deep Attention-Sampling Models · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
importance sampling
0.312018
Not All Samples Are Created Equal: Deep Learning with Importance Sampling · ICML 2018
Machine learning › Optimization for machine learning
stochastic gradient descent
0.312018
Not All Samples Are Created Equal: Deep Learning with Importance Sampling · ICML 2018
Machine learning › Optimization for machine learning
variance reduction
0.312018
Not All Samples Are Created Equal: Deep Learning with Importance Sampling · ICML 2018
Natural language and speech › Information extraction and text analysis
topic model
0.212016
Fast Supervised LDA for Discovering Micro-Events in Large-Scale Video Datasets · ACM Multimedia 2016
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.212023
Masked Autoencoding Does Not Help Natural Language Supervision at Scale · CVPR 2023
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction
0.112021
Neural Parts: Learning Expressive 3D Shape Abstractions With Invertible Neural Networks · CVPR 2021

Methods — techniques the papers use, named apart from their topics

invertible neural network · 1.0implicit surface · 1.0homeomorphic mapping · 1.0routing · 0.9parameter averaging · 0.9mixture of experts · 0.9language modeling · 0.9backpropagation · 0.9masked autoencoder · 0.7contrastive learning · 0.7variational inference · 0.2supervised LDA · 0.2
YearPublicationVenuePosition
2025 No Need to Talk: Asynchronous Mixture of Language Models
abstract
We introduce SMALLTALK LM, an innovative method for training a mixture of language models in an almost asynchronous manner. Each model of the mixture specializes in distinct parts of the data distribution, without the need of high-bandwidth communication between the nodes training each model. At inference, a lightweight router directs a given sequence to a single expert, according to a short prefix. This inference scheme naturally uses a fraction of the parameters from the overall mixture model. Unlike prior works on asynchronous LLM training, our routing method does not rely on full corpus clustering or access to metadata, making it more suitable for real-world applications. Our experiments on language modeling demonstrate that SMALLTALK LM achieves significantly lower perplexity than dense model baselines for the same total training FLOPs and an almost identical inference cost. Finally, in our downstream evaluations we outperform the dense baseline on 75% of the tasks.
Anastasiia Filippova, Angelos Katharopoulos, David Grangier, Ronan Collobert
ICLR2
2025 Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
abstract
Machine learning models are routinely trained on a mixture of different data domains. Different domain weights yield very different downstream performances. We propose the Soup-of-Experts, a novel architecture that can instantiate a model at test time for any domain weights with minimal computational cost and without re-training the model. Our architecture consists of a bank of expert parameters, which are linearly combined to instantiate one model. We learn the linear combination coefficients as a function of the input domain weights. To train this architecture, we sample random domain weights, instantiate the corresponding model, and backprop through one batch of data sampled with these domain weights. We demonstrate how our approach obtains small specialized models on several language modeling tasks quickly. Soup-of-Experts are particularly appealing when one needs to ship many different specialist models quickly under a size constraint.
Pierre Ablin, Angelos Katharopoulos, Skyler Seto, David Grangier
ICML2
2023 Masked Autoencoding Does Not Help Natural Language Supervision at Scale
abstract
Self supervision and natural language supervision have emerged as two exciting ways to train general purpose image encoders which excel at a variety of downstream tasks. Recent works such as M3AE [31] and SLIP [63] have suggested that these approaches can be effectively combined, but most notably their results use small100M samples) that is commonly used for these approaches. Here we investigate whether a similar approach can be effective when trained with a much larger amount of data. We find that a combination of two state of the art approaches: masked autoencoders, MAE [37] and contrastive language image pretraining, CLIP [68] provides a benefit over CLIP when trained on a corpus of 11.3M image-text pairs, but little to no benefit (as evaluated on a suite of common vision tasks) over CLIP when trained on a large corpus of 1.4B images. Our work provides some much needed clarity into the effectiveness (or lack thereof) of self supervision for large-scale image-text training.
Floris Weers, Vaishaal Shankar, Angelos Katharopoulos, Yinfei Yang, Tom Gunter
CVPR3
2021 Neural Parts: Learning Expressive 3D Shape Abstractions With Invertible Neural Networks
abstract
Impressive progress in 3D shape extraction led to representations that can capture object geometries with high fidelity. In parallel, primitive-based methods seek to represent objects as semantically consistent part arrangements. However, due to the simplicity of existing primitive representations, these methods fail to accurately reconstruct 3D shapes using a small number of primitives/parts. We address the trade-off between reconstruction quality and number of parts with Neural Parts, a novel 3D primitive representation that defines primitives using an Invertible Neural Network (INN) which implements homeomorphic mappings between a sphere and the target object. The INN allows us to compute the inverse mapping of the homeomorphism, which in turn, enables the efficient computation of both the implicit surface function of a primitive and its mesh, without any additional post-processing. Our model learns to parse 3D objects into semantically consistent part arrangements without any part-level supervision. Evaluations on ShapeNet, D-FAUST and FreiHAND demonstrate that our primitives can capture complex geometries and thus simultaneously achieve geometrically accurate as well as interpretable reconstructions using an order of magnitude fewer primitives than state-of-the-art shape abstraction methods.
Despoina Paschalidou, Angelos Katharopoulos, Andreas Geiger 0001, Sanja Fidler
CVPR2
2020 Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
abstract
Transformers achieve remarkable performance in several tasks but due to their quadratic complexity, with respect to the input’s length, they are prohibitively slow for very long sequences. To address this limitation, we express the self-attention as a linear dot-product of kernel feature maps and make use of the associativity property of matrix products to reduce the complexity from $\bigO{N^2}$ to $\bigO{N}$, where $N$ is the sequence length. We show that this formulation permits an iterative implementation that dramatically accelerates autoregressive transformers and reveals their relationship to recurrent neural networks. Our \emph{Linear Transformers} achieve similar performance to vanilla Transformers and they are up to 4000x faster on autoregressive prediction of very long sequences.
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas 0002, François Fleuret
ICML1
2020 Fast Transformers with Clustered Attention
abstract
Transformers have been proven a successful model for a variety of tasks in sequence modeling. However, computing the attention matrix, which is their key component, has quadratic complexity with respect to the sequence length, thus making them prohibitively expensive for large sequences. To address this, we propose clustered attention, which instead of computing the attention for every query, groups queries into clusters and computes attention just for the centroids. To further improve this approximation, we use the computed clusters to identify the keys with the highest attention per query and compute the exact key/query dot products. This results in a model with linear complexity with respect to the sequence length for a fixed number of clusters. We evaluate our approach on two automatic speech recognition datasets and show that our model consistently outperforms vanilla transformers for a given computational budget. Finally, we demonstrate that our model can approximate arbitrarily complex attention distributions with a minimal number of clusters by approximating a pretrained BERT model on GLUE and SQuAD benchmarks with only 25 clusters and no loss in performance.
Apoorv Vyas, Angelos Katharopoulos, François Fleuret
NeurIPS2
2019 Processing Megapixel Images with Deep Attention-Sampling Models
abstract
Existing deep architectures cannot operate on very large signals such as megapixel images due to computational and memory constraints. To tackle this limitation, we propose a fully differentiable end-to-end trainable model that samples and processes only a fraction of the full resolution input image. The locations to process are sampled from an attention distribution computed from a low resolution view of the input. We refer to our method as attention sampling and it can process images of several megapixels with a standard single GPU setup. We show that sampling from the attention distribution results in an unbiased estimator of the full model with minimal variance, and we derive an unbiased estimator of the gradient that we use to train our model end-to-end with a normal SGD procedure. This new method is evaluated on three classification tasks, where we show that it allows to reduce computation and memory footprint by an order of magnitude for the same accuracy as classical architectures. We also show the consistency of the sampling that indeed focuses on informative parts of the input images.
Angelos Katharopoulos, François Fleuret
ICML1
2018 Not All Samples Are Created Equal: Deep Learning with Importance Sampling
abstract
Deep Neural Network training spends most of the computation on examples that are properly handled, and could be ignored. We propose to mitigate this phenomenon with a principled importance sampling scheme that focuses computation on "informative" examples, and reduces the variance of the stochastic gradients during training. Our contribution is twofold: first, we derive a tractable upper bound to the per-sample gradient norm, and second we derive an estimator of the variance reduction achieved with importance sampling, which enables us to switch it on when it will result in an actual speedup. The resulting scheme can be used by changing a few lines of code in a standard SGD procedure, and we demonstrate experimentally on image classification, CNN fine-tuning, and RNN training, that for a fixed wall-clock time budget, it provides a reduction of the train losses of up to an order of magnitude and a relative improvement of test errors between 5% and 17%.
Angelos Katharopoulos, François Fleuret
ICML1
2016 Fast Supervised LDA for Discovering Micro-Events in Large-Scale Video Datasets
abstract
This paper introduces fsLDA, a fast variational inference method for supervised LDA, which overcomes the computational limitations of the original supervised LDA and enables its application in large-scale video datasets. In addition to its scalability, our method also overcomes the drawbacks of standard, unsupervised LDA for video, including its focus on dominant but often irrelevant video information (e.g. background, camera motion). As a result, experiments in the UCF11 and UCF101 datasets show that our method consistently outperforms unsupervised LDA in every metric. Furthermore, analysis shows that class-relevant topics of fsLDA lead to sparse video representations and encapsulate high-level information corresponding to parts of video events, which we denote "micro-events".
Angelos Katharopoulos, Despoina Paschalidou, Christos Diou, Anastasios Delopoulos
ACM Multimedia1