Simone Ricci

dblp:02/5746 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0001-9838-6076ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 31% Image recognition and object detection · 25% Language models and text generation · 25%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
image classification
1.012026
Mitigating Negative Flips via Margin Preserving Training · AAAI 2026
Natural language and speech › Language models and text generation › knowledge editing
negative flip reduction
1.012026
Mitigating Negative Flips via Margin Preserving Training · AAAI 2026
Machine learning › Representation and self-supervised learning
compatible representations
0.812024
Stationary Representations: Optimally Approximating Compatibility and Implications for Improved Model Replacements · CVPR 2024
Multimedia analysis and retrieval
cross-modal retrieval
0.812024
Learning Backward Compatible Representations · ACM Multimedia 2024

Methods — techniques the papers use, named apart from their topics

fine-tuning · 1.5embedding-based retrieval · 1.5d-simplex fixed classifier · 1.5backward-compatible learning · 1.5margin calibration · 1.0knowledge distillation · 1.0focal loss · 1.0
YearPublicationVenuePosition
2026 Mitigating Negative Flips via Margin Preserving Training
abstract
Minimizing inconsistencies across successive versions of an AI system is as crucial as reducing the overall error. In image classification, such inconsistencies manifest as negative flips, where an updated model misclassifies test samples that were previously classified correctly. This issue becomes increasingly pronounced as the number of training classes grows over time, since adding new categories reduces the margin of each class and may introduce conflicting patterns that undermine their learning process, thereby degrading performance on the original subset. To mitigate negative flips, we propose a novel approach that preserves the margins of the original model while learning an improved one. Our method encourages a larger relative margin between the previously learned and newly introduced classes by introducing an explicit margin-calibration term on the logits. However, overly constraining the logit margin for the new classes can significantly degrade their accuracy compared to a new independently trained model. To address this, we integrate a double-source focal distillation loss with the previous model and a new independently trained model, learning an appropriate decision margin from both old and new data, even under a logit margin calibration. Extensive experiments on image classification benchmarks demonstrate that our approach consistently reduces the negative flip rate with high overall accuracy.
Simone Ricci, Niccolò Biondi, Federico Pernici, Alberto Del Bimbo
AAAI1
2025 Learning Compatible Representations
Alberto Del Bimbo, Niccolò Biondi, Simone Ricci, Federico Pernici
ICPRAM3
2025 λ-Orthogonality Regularization for Compatible Representation Learning
Simone Ricci, Niccolò Biondi, Federico Pernici, Ioannis Patras, Alberto Del Bimbo
NeurIPS1
2024 Stationary Representations: Optimally Approximating Compatibility and Implications for Improved Model Replacements
abstract
Learning compatible representations enables the interchangeable use of semantic features as models are updated over time. This is particularly relevant in search and retrieval systems where it is crucial to avoid reprocessing of the gallery images with the updated model. While recent research has shown promising empirical evidence, there is still a lack of comprehensive theoretical understanding about learning compatible representations. In this paper, we demonstrate that the stationary representations learned by the d-Simplex fixed classifier optimally approximate compatibility representation according to the two inequality constraints of its formal definition. This not only establishes a solid foundation for future works in this line of research but also presents implications that can be exploited in practical learning scenarios. An exemplary application is the nowstandard practice of downloading and fine-tuning new pretrained models. Specifically, we show the strengths and critical issues of stationary representations in the case in which a model undergoing sequential fine-tuning is asynchronously replaced by downloading a better-performing model pretrained elsewhere. Such a representation enables seamless delivery of retrieval service (i.e., no reprocessing of gallery images) and offers improved performance without operational disruptions during model replacement. Code available at: https://github.com/miccunifi/iamcl2r.
Niccolò Biondi, Federico Pernici, Simone Ricci, Alberto Del Bimbo
CVPR3
2024 Learning Backward Compatible Representations
abstract
In today's multimedia-rich environment, the rapid growth of data poses significant challenges for developing efficient multi-modal retrieval systems essential for retrieving text, images, audio, and video. As data expands, newer, scalable, and high-performance retrieval systems are increasingly necessary. Embedding-based deep neural networks (DNNs) have become key solutions, transforming high-dimensional data into lower-dimensional embeddings for easy comparison and retrieval. However, updating DNNs changes the internal feature representations, necessitating the extraction of new feature vectors for all gallery data, which is costly, especially with gallery sets comprising billions of data. Learning backward-compatible representations addresses this by allowing new representation to be matched with old gallery data without recalculating features. This tutorial aims to equip participants with the knowledge and tools to apply backward-compatible representations, enhancing multimedia retrieval systems' efficiency and scalability. Participants will learn the importance of compatible representations, basic methods and techniques, and explore challenging open questions that are becoming increasingly relevant to multimedia and cross-modal retrieval.
Niccolò Biondi, Simone Ricci, Federico Pernici, Alberto Del Bimbo
ACM Multimedia2
2023 Meta-learning Advisor Networks for Long-tail and Noisy Labels in Social Image Classification
abstract
Deep neural networks (DNNs)for social image classification are prone to performance reduction and overfitting when trained on datasets plagued by noisy or imbalanced labels. Weight loss methods tend to ignore the influence of noisy or frequent category examples during the training, resulting in a reduction of final accuracy and, in the presence of extreme noise, even a failure of the learning process. A new advisor network is introduced to address both imbalance and noise problems, and is able to pilot learning of a main network by adjusting the visual features and the gradient with a meta-learning strategy. In a curriculum learning fashion, the impact of redundant data is reduced while recognizable noisy label images are downplayed or redirected.Meta Feature Re-Weighting (MFRW)andMeta Equalization Softmax (MES)methods are introduced to let the main network focus only on the information in an image deemed relevant by the advisor network and to adjust the training gradient to reduce the adverse effects of frequent or noisy categories. The proposed method is first tested on synthetic versions of CIFAR10 and CIFAR100, and then on the more realistic ImageNet-LT, Places-LT, and Clothing1M datasets, reporting state-of-the-art results.
Simone Ricci, Tiberio Uricchio, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.1