Federico Baldassarre

dblp:211/6876 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0001-8152-767XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 32% Representation and self-supervised learning · 32% Trustworthy machine learning · 32%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
cross-modal alignment
0.912025
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
vision foundation model
0.912025
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment · CVPR 2025
Bioinformatics and computational biology › protein structure prediction
model quality assessment
0.512021
GraphQA: protein model quality assessment using graph convolutional networks · Bioinform. 2021
Bioinformatics and computational biology
protein structure prediction
0.512021
GraphQA: protein model quality assessment using graph convolutional networks · Bioinform. 2021
Machine learning › Trustworthy machine learning › interpretability
explanation-based learning
0.412020
Explanation-Based Weakly-Supervised Learning of Visual Relations with Graph Networks · ECCV (28) 2020
Machine learning › Trustworthy machine learning
interpretability
0.412020
Explanation-Based Weakly-Supervised Learning of Visual Relations with Graph Networks · ECCV (28) 2020
Machine learning › Graph learning › graph neural network
graph convolutional network
0.112021
GraphQA: protein model quality assessment using graph convolutional networks · Bioinform. 2021

Methods — techniques the papers use, named apart from their topics

graph convolutional network · 1.0text encoder alignment · 0.9lit training · 0.9contrastive learning · 0.9weakly supervised learning · 0.4graph neural network · 0.4
YearPublicationVenuePosition
2025 DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment
abstract
Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models such as CLIP [67], self-supervised visual features are not readily aligned with language, hindering their adoption in open-vocabulary tasks. Our method, named dino.txt, unlocks this new ability for DINOv2 [63], a widely used self-supervised visual encoder. We build upon the LiT training strategy [97], which trains a text encoder to align with a frozen vision model but leads to unsatisfactory results on dense tasks. We propose several key ingredients to improve performance on both global and dense tasks, such as concatenating the [CLS] token with the patch average to train the alignment and curating data using both text and image modalities. With these, we successfully train a CLIP-like model with only a fraction of the computational cost compared to CLIP while achieving state-of-the-art results in zero-shot classification and open-vocabulary semantic segmentation.
Cijo Jose, Théo Moutakanni, Dahyun Kang, Federico Baldassarre, Timothée Darcet, Hu Xu 0001, Daniel Li 0006, Marc Szafraniec, Michaël Ramamonjisoa, Maxime Oquab, Oriane Siméoni, Huy V. Vo, Patrick Labatut, Piotr Bojanowski
CVPR4
2023 Variable Rate Allocation for Vector-Quantized Autoencoders
abstract
Vector-quantized autoencoders have recently gained interest in image compression, generation and self-supervised learning. However, as a neural compression method, they lack the possibility to allocate a variable number of bits to each image location, e.g. according to the semantic content or local saliency. In this paper, we address this limitation in a simple yet effective way. We adopt a product quantizer (PQ) that produces a set of discrete codes for each image patch rather than a single index. This PQ-autoencoder is trained end-to-end with a structured dropout that selectively masks a variable number of codes at each location. These mechanisms force the decoder to reconstruct the original image based on partial information and allow us to control the local rate. The resulting model can compress images on a wide range of operating points of the rate-distortion curve and can be paired with any external method for saliency estimation to control the compression rate at a local level. We demonstrate the effectiveness of our approach on the popular Kodak and ImageNet datasets by measuring both distortion and perceptual quality metrics.
Federico Baldassarre, Alaaeldin El-Nouby, Hervé Jégou
ICASSP1
2022 Quantitative Metrics for Evaluating Explanations of Video DeepFake Detectors
Federico Baldassarre, Quentin Debard, Gonzalo Fiz Pontiveros, Tri Kurniawan Wijaya
BMVC1
2022 Learnable Masked Tokens for Improved Transferability of Self-supervised Vision Transformers
Federico Baldassarre, Hossein Azizpour
ECML/PKDD (3)2
2021 GraphQA: protein model quality assessment using graph convolutional networks
abstract
MOTIVATION: Proteins are ubiquitous molecules whose function in biological processes is determined by their 3D structure. Experimental identification of a protein's structure can be time-consuming, prohibitively expensive and not always possible. Alternatively, protein folding can be modeled using computational methods, which however are not guaranteed to always produce optimal results. GraphQA is a graph-based method to estimate the quality of protein models, that possesses favorable properties such as representation learning, explicit modeling of both sequential and 3D structure, geometric invariance and computational efficiency. RESULTS: GraphQA performs similarly to state-of-the-art methods despite using a relatively low number of input features. In addition, the graph network structure provides an improvement over the architecture used in ProQ4 operating on the same input features. Finally, the individual contributions of GraphQA components are carefully evaluated. AVAILABILITY AND IMPLEMENTATION: PyTorch implementation, datasets, experiments and link to an evaluation server are available through this GitHub repository: github.com/baldassarreFe/graphqa. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Federico Baldassarre, David Menéndez Hurtado, Arne Elofsson, Hossein Azizpour
Bioinform.1
2020 Explanation-Based Weakly-Supervised Learning of Visual Relations with Graph Networks
Federico Baldassarre, Kevin Smith 0001, Josephine Sullivan, Hossein Azizpour
ECCV (28)1