Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Pietro Astolfi

dblp:208/4543 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-5192-9608ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Vision and language · 32% Generative modeling · 26% Deep learning architectures and training · 20%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
latent diffusion model
1.622025
Boosting Latent Diffusion with Perceptual Objectives · ICLR 2025
On improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models · NeurIPS 2024
Computer vision › Vision and language › vision-language model
CLIP
0.912025
Object-centric binding in Contrastive Language-Image Pretraining · NeurIPS 2025
Computer vision › Vision and language › compositionality
compositional understanding
0.912025
Object-centric binding in Contrastive Language-Image Pretraining · NeurIPS 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
X-Sample Contrastive Loss: Improving Contrastive Learning with Sample Similarity Graphs · ICLR 2025
Machine learning › Deep learning architectures and training › loss function design
perceptual loss
0.912025
Boosting Latent Diffusion with Perceptual Objectives · ICLR 2025
Machine learning › Deep learning architectures and training
scaling laws
0.912025
Improving the Scaling Laws of Synthetic Data with Deliberate Practice · ICML 2025
Machine learning › Generative modeling
synthetic data generation
0.912025
Improving the Scaling Laws of Synthetic Data with Deliberate Practice · ICML 2025
Computer vision › Vision and language › image captioning
dense captioning
0.812024
A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions · CVPR 2024
Machine learning › Generative modeling
diffusion model
0.812024
On improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models · NeurIPS 2024
Machine learning › Deep learning architectures and training
pre-training strategy
0.812024
On improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models · NeurIPS 2024
Computer vision › Vision and language
vision-language model
0.812024
A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions · CVPR 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › clustering-based representation learning
self-supervised clustering
0.712023
Semi-supervised learning made simple with self-supervised clustering · CVPR 2023
Machine learning › Learning paradigms › semi-supervised learning
semi-supervised clustering
0.712023
Semi-supervised learning made simple with self-supervised clustering · CVPR 2023
Machine learning › Learning paradigms
semi-supervised learning
0.712023
Semi-supervised learning made simple with self-supervised clustering · CVPR 2023

Methods — techniques the papers use, named apart from their topics

slot attention · 0.9scene graph · 0.9sample similarity graphs · 0.9latent perceptual loss · 0.9hard-negative augmentation · 0.9flow matching · 0.9deliberate practice · 0.9data pruning · 0.9contrastive loss · 0.9DDPM · 0.9
YearPublicationVenuePosition
2025 Boosting Latent Diffusion with Perceptual Objectives
abstract
Latent diffusion models (LDMs) power state-of-the-art high-resolution generative image models. LDMs learn the data distribution in the latent space of an autoencoder (AE) and produce images by mapping the generated latents into RGB image space using the AE decoder. While this approach allows for efficient model training and sampling, it induces a disconnect between the training of the diffusion model and the decoder, resulting in a loss of detail in the generated images. To remediate this disconnect, we propose to leverage the internal features of the decoder to define a latent perceptual loss (LPL). This loss encourages the models to create sharper and more realistic images. Our loss can be seamlessly integrated with common autoencoders used in latent diffusion models, and can be applied to different generative modeling paradigms such as DDPM with epsilon and velocity prediction, as well as flow matching. Extensive experiments with models trained on three datasets at 256 and 512 resolution show improved quantitative -- with boosts between 6% and 20% in FID -- and qualitative results when using our perceptual loss.
Tariq Berrada, Pietro Astolfi, Melissa Hall, Marton Havasi, Yohann Benchetrit, Adriana Romero-Soriano, Karteek Alahari, Michal Drozdzal, Jakob Verbeek
ICLR2
2025 X-Sample Contrastive Loss: Improving Contrastive Learning with Sample Similarity Graphs
Vlad Sobal, Mark Ibrahim, Randall Balestriero, Vivien Cabannes, Diane Bouchacourt, Pietro Astolfi, Kyunghyun Cho, Yann LeCun
ICLR6
2025 Improving the Scaling Laws of Synthetic Data with Deliberate Practice
abstract
Inspired by the principle of deliberate practice in human learning, we propose Deliberate Practice for Synthetic Data Generation (DP), a novel framework that improves sample efficiency through dynamic synthetic data generation. Prior work has shown that scaling synthetic data is inherently challenging, as naively adding new data leads to diminishing returns. To address this, pruning has been identified as a key mechanism for improving scaling, enabling models to focus on the most informative synthetic samples. Rather than generating a large dataset and pruning it afterward, DP efficiently approximates the direct generation of informative samples. We theoretically show how training on challenging, informative examples improves scaling laws and empirically validate that DP achieves better scaling performance with significantly fewer training samples and iterations. On ImageNet-100, DP generates 3.4x fewer samples and requires six times fewer iterations, while on ImageNet-1k, it generates 8x fewer samples with a 30% reduction in iterations, all while achieving superior performance compared to prior work.
Reyhane Askari Hemmat, Mohammad Pezeshki, Elvis Dohmatob, Florian Bordes, Pietro Astolfi, Melissa Hall, Jakob Verbeek, Michal Drozdzal, Adriana Romero-Soriano
ICML5
2025 Object-centric binding in Contrastive Language-Image Pretraining
abstract
Recent advances in vision language models (VLM) have been driven by contrastive models such as CLIP, which learn to associate visual information with their corresponding text descriptions. However, these models have limitations in understanding complex compositional scenes involving multiple objects and their spatial relationships. To address these challenges, we propose a novel approach that diverges from commonly used strategies that rely on the design of finegrained hard-negative augmentations. Instead, our work focuses on integrating inductive biases into the pretraining of CLIP-like models to improve their compositional understanding. To that end, we introduce a binding module that connects a scene graph, derived from a text description, with a slot-structured image representation, facilitating a structured similarity assessment between the two modalities. We also leverage relationships as text-conditioned visual constraints, thereby capturing the intricate interactions between objects and their contextual relationships more effectively. Our resulting model not only enhances the performance of CLIP-based models in multi-object compositional understanding but also paves the way towards more accurate and sample-efficient image-text matching of complex scenes.
Rim Assouel, Pietro Astolfi, Florian Bordes, Michal Drozdzal, Adriana Romero-Soriano
NeurIPS2
2024 A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions
abstract
Curation methods for massive vision-language datasets trade off between dataset size and quality. However, even the highest quality of available curated captions are far too short to capture the rich visual detail in an image. To show the value of dense and highly-aligned image-text pairs, we collect the Densely Captioned Images (DCI) dataset, containing 7805 natural images human-annotated with mask-aligned descriptions averaging above 1000 words each. With precise and reliable captions associated with specific parts of an image, we can evaluate vision-language models' (VLMs) understanding of image content with a novel task that matches each caption with its corresponding subcrop. As current models are often limited to 77 text tokens, we also introduce a summarized version (sDCI) in which each caption length is limited. We show that modern techniques that make progress on standard benchmarks do not correspond with significant improvement on our sDCI based benchmark. Lastly, we finetune CLIP using sDCI and show significant improvements over the baseline despite a small training set. By releasing the first human annotated dense image captioning dataset, we hope to enable the development of new benchmarks or finetuning recipes for the next generation of VLMs to come.
Jack Urbanek, Florian Bordes, Pietro Astolfi, Mary Williamson, Vasu Sharma, Adriana Romero-Soriano
CVPR3
2024 On improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models
abstract
Large-scale training of latent diffusion models (LDMs) has enabled unprecedented quality in image generation. However, large-scale end-to-end training of these models is computationally costly, and hence most research focuses either on finetuning pretrained models or experiments at smaller scales. In this work we aim to improve the training efficiency and performance of LDMs with the goal of scaling to larger datasets and higher resolutions. We focus our study on two points that are critical for good performance and efficient training: (i) the mechanisms used for semantic level (\eg a text prompt, or class name) and low-level (crop size, random flip, \etc) conditioning of the model, and (ii) pre-training strategies to transfer representations learned on smaller and lower-resolution datasets to larger ones. The main contributions of our work are the following: we present systematic experimental study of these points, we propose a novel conditioning mechanism that disentangles semantic and low-level conditioning, we obtain state-of-the-art performance on CC12M for text-to-image at 512 resolution.
Tariq Berrada, Pietro Astolfi, Melissa Hall, Reyhane Askari Hemmat, Yohann Benchetrit, Marton Havasi, Matthew J. Muckley, Karteek Alahari, Adriana Romero-Soriano, Jakob Verbeek, Michal Drozdzal
NeurIPS2
2023 Semi-supervised learning made simple with self-supervised clustering
abstract
Self-supervised learning models have been shown to learn rich visual representations without requiring human annotations. However, in many real-world scenarios, labels are partially available, motivating a recent line of work on semi-supervised methods inspired by self-supervised principles. In this paper, we propose a conceptually simple yet empirically powerful approach to turn clustering-based self-supervised methods such as SwAV or DINO into semi-supervised learners. More precisely, we introduce a multi-task framework merging a supervised objective using ground-truth labels and a self-supervised objective relying on clustering assignments with a single cross-entropy loss. This approach may be interpreted as imposing the cluster centroids to be class prototypes. Despite its simplicity, we provide empirical evidence that our approach is highly effective and achieves state-of-the-art performance on CI-FAR100 and ImageNet.
Enrico Fini, Pietro Astolfi, Karteek Alahari, Xavier Alameda-Pineda, Julien Mairal, Moin Nabi, Elisa Ricci 0001
CVPR2
2023 Supervised tractogram filtering using Geometric Deep Learning
Pietro Astolfi, Ruben S. Verhagen, Laurent Petit, Emanuele Olivetti, Silvio Sarubbo, Jonathan Masci, Davide Boscaini, Paolo Avesani
Medical Image Anal.1
2020 Clustered Dynamic Graph CNN for Biometric 3D Hand Shape Recognition
abstract
The research in biometric recognition using hand shape has been somewhat stagnating in the last decade. Meanwhile, computer vision and machine learning have experienced a paradigm shift with the renaissance of deep learning, which has set the new state-of-the-art in many related fields. Inspired by successful applications of deep learning for other biometric modalities, we propose a novel approach to 3D hand shape recognition from RGB-D data based on geometric deep learning techniques. We show how to train our model on synthetic data and retain the performance on real samples during test time. To evaluate our method, we provide a new dataset NNHand RGB- D of short video sequences and show encouraging performance compared to diverse baselines on the new data, as well as current benchmark dataset HKPolyU. Moreover, the new dataset opens door to many new research directions in hand shape recognition.
Jan Svoboda, Pietro Astolfi, Davide Boscaini, Jonathan Masci, Michael M. Bronstein
IJCB2
2020 Tractogram Filtering of Anatomically Non-plausible Fibers with Geometric Deep Learning
Pietro Astolfi, Ruben S. Verhagen, Laurent Petit, Emanuele Olivetti, Jonathan Masci, Davide Boscaini, Paolo Avesani
MICCAI (7)1