VLDB 2026 Research / reviewers in the wild / expert
Federico Pernici
dblp:79/2114
· DBLP profile ↗
31ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0001-7036-6655ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 19 · 5 first-author · 5 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Negative Flips via Margin Preserving TrainingabstractMinimizing inconsistencies across successive versions of an AI system is as crucial as reducing the overall error. In image classification, such inconsistencies manifest as negative flips, where an updated model misclassifies test samples that were previously classified correctly. This issue becomes increasingly pronounced as the number of training classes grows over time, since adding new categories reduces the margin of each class and may introduce conflicting patterns that undermine their learning process, thereby degrading performance on the original subset. To mitigate negative flips, we propose a novel approach that preserves the margins of the original model while learning an improved one. Our method encourages a larger relative margin between the previously learned and newly introduced classes by introducing an explicit margin-calibration term on the logits. However, overly constraining the logit margin for the new classes can significantly degrade their accuracy compared to a new independently trained model. To address this, we integrate a double-source focal distillation loss with the previous model and a new independently trained model, learning an appropriate decision margin from both old and new data, even under a logit margin calibration. Extensive experiments on image classification benchmarks demonstrate that our approach consistently reduces the negative flip rate with high overall accuracy. Simone Ricci, Niccolò Biondi, Federico Pernici, Alberto Del Bimbo |
AAAI | 3 |
| 2025 | Learning Compatible Representations
Alberto Del Bimbo, Niccolò Biondi, Simone Ricci, Federico Pernici |
ICPRAM | 4 |
| 2025 | λ-Orthogonality Regularization for Compatible Representation Learning
Simone Ricci, Niccolò Biondi, Federico Pernici, Ioannis Patras, Alberto Del Bimbo |
NeurIPS | 3 |
| 2024 | Stationary Representations: Optimally Approximating Compatibility and Implications for Improved Model ReplacementsabstractLearning compatible representations enables the interchangeable use of semantic features as models are updated over time. This is particularly relevant in search and retrieval systems where it is crucial to avoid reprocessing of the gallery images with the updated model. While recent research has shown promising empirical evidence, there is still a lack of comprehensive theoretical understanding about learning compatible representations. In this paper, we demonstrate that the stationary representations learned by the d-Simplex fixed classifier optimally approximate compatibility representation according to the two inequality constraints of its formal definition. This not only establishes a solid foundation for future works in this line of research but also presents implications that can be exploited in practical learning scenarios. An exemplary application is the nowstandard practice of downloading and fine-tuning new pretrained models. Specifically, we show the strengths and critical issues of stationary representations in the case in which a model undergoing sequential fine-tuning is asynchronously replaced by downloading a better-performing model pretrained elsewhere. Such a representation enables seamless delivery of retrieval service (i.e., no reprocessing of gallery images) and offers improved performance without operational disruptions during model replacement. Code available at: https://github.com/miccunifi/iamcl2r. Niccolò Biondi, Federico Pernici, Simone Ricci, Alberto Del Bimbo |
CVPR | 2 |
| 2024 | Learning Backward Compatible RepresentationsabstractIn today's multimedia-rich environment, the rapid growth of data poses significant challenges for developing efficient multi-modal retrieval systems essential for retrieving text, images, audio, and video. As data expands, newer, scalable, and high-performance retrieval systems are increasingly necessary. Embedding-based deep neural networks (DNNs) have become key solutions, transforming high-dimensional data into lower-dimensional embeddings for easy comparison and retrieval. However, updating DNNs changes the internal feature representations, necessitating the extraction of new feature vectors for all gallery data, which is costly, especially with gallery sets comprising billions of data. Learning backward-compatible representations addresses this by allowing new representation to be matched with old gallery data without recalculating features. This tutorial aims to equip participants with the knowledge and tools to apply backward-compatible representations, enhancing multimedia retrieval systems' efficiency and scalability. Participants will learn the importance of compatible representations, basic methods and techniques, and explore challenging open questions that are becoming increasingly relevant to multimedia and cross-modal retrieval. Niccolò Biondi, Simone Ricci, Federico Pernici, Alberto Del Bimbo |
ACM Multimedia | 3 |
| 2023 | CoReS: Compatible Representations via StationarityabstractCompatible features enable the direct comparison of old and new learned features allowing to use them interchangeably over time. In visual search systems, this eliminates the need to extract new features from the gallery-set when the representation model is upgraded with novel data. This has a big value in real applications as re-indexing the gallery-set can be computationally expensive when the gallery-set is large, or even infeasible due to privacy or other concerns of the application. In this paper, we propose CoReS, a new training procedure to learn representations that are compatible with those previously learned, grounding on the stationarity of the features as provided by fixed classifiers based on polytopes. With this solution, classes are maximally separated in the representation space and maintain their spatial configuration stationary as new classes are added, so that there is no need to learn any mappings between representations nor to impose pairwise training with the previously learned model. We demonstrate that our training procedure largely outperforms the current state of the art and is particularly effective in the case of multiple upgrades of the training-set, which is the typical case in real applications. Niccolò Biondi, Federico Pernici, Matteo Bruni, Alberto Del Bimbo |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Regular Polytope NetworksabstractNeural networks are widely used as a model for classification in a large variety of tasks. Typically, a learnable transformation (i.e., the classifier) is placed at the end of such models returning a value for each class used for classification. This transformation plays an important role in determining how the generated features change during the learning process. In this work, we argue that this transformation not only can be fixed (i.e., set as nontrainable) with no loss of accuracy and with a reduction in memory usage, but it can also be used to learn stationary and maximally separated embeddings. We show that the stationarity of the embedding and its maximal separated representation can be theoretically justified by setting the weights of the fixed classifier to values taken from the coordinate vertices of the three regular polytopes available in [Formula: see text], namely, the d -Simplex, the d -Cube, and the d -Orthoplex. These regular polytopes have the maximal amount of symmetry that can be exploited to generate stationary features angularly centered around their corresponding fixed weights. Our approach improves and broadens the concept of a fixed classifier, recently proposed by Hoffer et al., to a larger class of fixed classifier models. Experimental results confirm the theoretical analysis, the generalization capability, the faster convergence, and the improved performance of the proposed method. Code will be publicly available. Federico Pernici, Matteo Bruni, Claudio Baecchi, Alberto Del Bimbo |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | CL2R: Compatible Lifelong Learning RepresentationsabstractIn this article, we propose a method to partially mimic natural intelligence for the problem of lifelong learning representations that are compatible. We take the perspective of a learning agent that is interested in recognizing object instances in an open dynamic universe in a way in which any update to its internal feature representation does not render the features in the gallery unusable for visual search. We refer to this learning problem as Compatible Lifelong Learning Representations (CL 2 R), as it considers compatible representation learning within the lifelong learning paradigm. We identify stationarity as the property that the feature representation is required to hold to achieve compatibility and propose a novel training procedure that encourages local and global stationarity on the learned representation. Due to stationarity, the statistical properties of the learned features do not change over time, making them interoperable with previously learned features. Extensive experiments on standard benchmark datasets show that our CL 2 R training procedure outperforms alternative baselines and state-of-the-art methods. We also provide novel metrics to specifically evaluate compatible representation learning under catastrophic forgetting in various sequential learning tasks. Code is available at https://github.com/NiccoBiondi/CompatibleLifelongRepresentation . Niccolò Biondi, Federico Pernici, Matteo Bruni, Daniele Mugnai, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Fine-Grained Adversarial Semi-Supervised LearningabstractIn this article, we exploit Semi-Supervised Learning ( SSL ) to increase the amount of training data to improve the performance of Fine-Grained Visual Categorization ( FGVC ). This problem has not been investigated in the past in spite of prohibitive annotation costs that FGVC requires. Our approach leverages unlabeled data with an adversarial optimization strategy in which the internal features representation is obtained with a second-order pooling model. This combination allows one to back-propagate the information of the parts, represented by second-order pooling, onto unlabeled data in an adversarial training setting. We demonstrate the effectiveness of the combined use by conducting experiments on six state-of-the-art fine-grained datasets, which include Aircrafts, Stanford Cars, CUB-200-2011, Oxford Flowers, Stanford Dogs, and the recent Semi-Supervised iNaturalist-Aves. Experimental results clearly show that our proposed method has better performance than the only previous approach that examined this problem; it also obtained higher classification accuracy with respect to the supervised learning methods with which we compared. Daniele Mugnai, Federico Pernici, Francesco Turchini, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Temporal Binary Representation for Event-Based Action RecognitionabstractIn this paper we present an event aggregation strategy to convert the output of an event camera into frames processable by traditional Computer Vision algorithms. The proposed method first generates sequences of intermediate binary representations, which are then losslessly transformed into a compact format by simply applying a binary-to-decimal conversion. This strategy allows us to encode temporal information directly into pixel values, which are then interpreted by deep learning models. We apply our strategy, called Temporal Binary Representation, to the task of Gesture Recognition, obtaining state of the art results on the popular DVS128 Gesture Dataset. To underline the effectiveness of the proposed method compared to existing ones, we also collect an extension of the dataset under more challenging conditions on which to perform experiments. Simone Undri Innocenti, Federico Becattini, Federico Pernici, Alberto Del Bimbo |
ICPR | 3 |
| 2020 | Class-incremental Learning with Pre-allocated Fixed ClassifiersabstractIn class-incremental learning, a learning agent faces a stream of data with the goal of learning new classes while not forgetting previous ones. Neural networks are known to suffer under this setting, as they forget previously acquired knowledge. To address this problem, effective methods exploit past data stored in an episodic memory while expanding the final classifier nodes to accommodate the new classes. In this work, we substitute the expanding classifier with a novel fixed classifier in which a number of pre-allocated output nodes are subject to the classification loss right from the beginning of the learning phase. Contrarily to the standard expanding classifier, this allows: (a) the output nodes of future unseen classes to firstly see negative samples since the beginning of learning together with the positive samples that incrementally arrive; (b) to learn features that do not change their geometric configuration as novel classes are incorporated in the learning model. Experiments with public datasets show that the proposed approach is as effective as the expanding classifier while exhibiting novel intriguing properties of the internal feature representation that are otherwise not-existent. Our ablation study on pre-allocating a large number of classes further validates the approach. Federico Pernici, Matteo Bruni, Claudio Baecchi, Francesco Turchini, Alberto Del Bimbo |
ICPR | 1 |
| 2020 | Self-supervised on-line cumulative learning from video streams
Federico Pernici, Matteo Bruni, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 1 |
| 2019 | Incremental Learning of People Identities
Federico Bartoli, Federico Pernici, Matteo Bruni, Alberto Del Bimbo |
CIARP | 2 |
| 2018 | Memory Based Online Learning of Deep Representations From Video StreamsabstractWe present a novel online unsupervised method for face identity learning from video streams. The method exploits deep face descriptors together with a memory based learning mechanism that takes advantage of the temporal coherence of visual data. Specifically, we introduce a discriminative descriptor matching solution based on Reverse Nearest Neighbour and a forgetting strategy that detect redundant descriptors and discard them appropriately while time progresses. It is shown that the proposed learning procedure is asymptotically stable and can be effectively used in relevant applications like multiple face identification and tracking from unconstrained video streams. Experimental results show that the proposed method achieves comparable results in the task of multiple face tracking and better performance in face identification with offline approaches exploiting future information. Code will be publicly available. Federico Pernici, Federico Bartoli, Matteo Bruni, Alberto Del Bimbo |
CVPR | 1 |
| 2016 | Continuous localization and mapping of a pan-tilt-zoom camera for wide area tracking
Giuseppe Lisanti, Iacopo Masi, Federico Pernici, Alberto Del Bimbo |
Mach. Vis. Appl. | 3 |
| 2015 | Non-myopic information theoretic sensor management of a single pan-tilt-zoom camera for multiple object detection and tracking
Pietro Salvagnini, Federico Pernici, Marco Cristani, Giuseppe Lisanti, Alberto Del Bimbo, Vittorio Murino |
Comput. Vis. Image Underst. | 2 |
| 2014 | Information theoretic sensor management for multi-target tracking with a single pan-tilt-zoom cameraabstractAutomatic multiple target tracking with pan-tilt-zoom (PTZ) cameras is a hard task, with few approaches in the literature, most of them proposing simplistic scenarios. In this paper, we present a PTZ camera management framework which lies on information theoretic principles: at each time step, the next camera pose (pan, tilt, focal length) is chosen, according to a policy which ensures maximum information gain. The formulation takes into account occlusions, physical extension of targets, realistic pedestrian detectors and the mechanical constraints of the camera. Convincing comparative results on synthetic data, realistic simulations and the implementation on a real video surveillance camera validate the effectiveness of the proposed method. Pietro Salvagnini, Federico Pernici, Marco Cristani, Giuseppe Lisanti, Iacopo Masi, Alberto Del Bimbo, Vittorio Murino |
WACV | 2 |
| 2014 | Object Tracking by Oversampling Local FeaturesabstractIn this paper, we present the ALIEN tracking method that exploits oversampling of local invariant representations to build a robust object/context discriminative classifier. To this end, we use multiple instances of scale invariant local features weakly aligned along the object template. This allows taking into account the 3D shape deviations from planarity and their interactions with shadows, occlusions, and sensor quantization for which no invariant representations can be defined. A non-parametric learning algorithm based on the transitive matching property discriminates the object from the context and prevents improper object template updating during occlusion. We show that our learning rule has asymptotic stability under mild conditions and confirms the drift-free capability of the method in long-term tracking. A real-time implementation of the ALIEN tracker has been evaluated in comparison with the state-of-the-art tracking systems on an extensive set of publicly available video sequences that represent most of the critical conditions occurring in real tracking environments. We have reported superior or equal performance in most of the cases and verified tracking with no drift in very long video sequences. Federico Pernici, Alberto Del Bimbo |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Continuous recovery for real time pan tilt zoom localization and mappingabstractWe propose a method for real time recovering from tracking failure in monocular localization and mapping with a Pan Tilt Zoom camera (PTZ). The method automatically detects and seamlessly recovers from tracking failure while preserving map integrity. By extending recent advances in the PTZ localization and mapping, the system can quickly and continuously resume tracking failures by determining the best way to task two different localization modalities. The tradeoff involved when choosing between the two modalities is captured by maximizing the information expected to be extracted from the scene map. This is especially helpful in four main viewing condition: blurred frames, weak textured scene, not up to date map and occlusions due to sensor quantization or moving objects. Extensive tests show that the resulting system is able to recover from several different failures while zooming-in weak textured scene, all in real time. Alberto Del Bimbo, Giuseppe Lisanti, Iacopo Masi, Federico Pernici |
AVSS | 4 |
| 2010 | Sensor Fusion for Cooperative Head LocalizationabstractIn modern video surveillance systems, pan-tilt-zoom (PTZ) cameras certainly have the potential to allow the coverage of wide areas with a much smaller number of sensors, compared to the common approach of fixed camera networks. This paper describes a general framework that aims at exploiting the capabilities of modern PTZ cameras in order to acquire high resolution images of body parts, such as the head, from the observation of pedestrians moving in a wide outdoor area. The framework allows to organize the sensors in a network with arbitrary topology, and to establish pairwise master-slave relationship between them. In this way a slave camera can be steered to acquire imagery of a target keeping into account both target and zooming uncertainties. Experiments show good performance in localizing target's head, independently from the zooming factor of the slave camera. Alberto Del Bimbo, Fabrizio Dini, Giuseppe Lisanti, Federico Pernici |
ICPR | 4 |
| 2010 | Person Detection Using Temporal and Geometric Context with a Pan Tilt Zoom CameraabstractIn this paper we present a system that integrates automatic camera geometry estimation and object detection from a Pan Tilt Zoom camera. We estimate camera pose with respect to a world scene plane in real-time and perform human detection exploiting the relative space-time context. Using camera self-localization, 2D object detections are clustered in a 3D world coordinate frame. Target scale inference is further exploited to reduce the number of false alarms and to increase also the detection rate in the final non-maximum suppression stage. Our integrated system applied on real-world data shows superior performance with respect to the standard detector used. Alberto Del Bimbo, Giuseppe Lisanti, Iacopo Masi, Federico Pernici |
ICPR | 4 |
| 2010 | Exploiting distinctive visual landmark maps in pan-tilt-zoom camera networks
Alberto Del Bimbo, Fabrizio Dini, Giuseppe Lisanti, Federico Pernici |
Comput. Vis. Image Underst. | 4 |
| 2009 | Arneb: a rich internet application for ground truth annotation of videosabstractIn this technical demonstration we show the current version of Arneb, a web-based system for manual annotation of videos, developed within the EU VidiVideo project. This tool has been developed with the aim of creating ground truth annotations, that can be used for training and evaluating automatic video annotation systems. Annotations can be exported to MPEG-7 and OWL ontologies. The system has been developed according to the Rich Internet Application paradigm, allowing collaborative web-based annotation. Thomas M. Alisi, Marco Bertini 0001, Gianpaolo D'Amico, Alberto Del Bimbo, Andrea Ferracani, Federico Pernici, Giuseppe Serra 0001 |
ACM Multimedia | 6 |
| 2009 | Sirio: an ontology-based web search engine for videosabstractIn this technical demonstration we show a web video search engine based on ontologies, the Sirio system, that has been developed within the EU VidiVideo project. The goal of the system is to provide a search engine for videos for both technical and non-technical users. In fact, the system has different interfaces that permit different query modalities: free-text, natural language, graphical composition of concepts using boolean and temporal relations and query by visual example. In addition, the ontology structure is exploited to encode semantic relations between concepts permitting, for example, to expand queries to synonyms and concept specializations. Thomas M. Alisi, Marco Bertini 0001, Gianpaolo D'Amico, Alberto Del Bimbo, Andrea Ferracani, Federico Pernici, Giuseppe Serra 0001 |
ACM Multimedia | 6 |
| 2008 | Uncalibrated Framework for On-line Camera Cooperation to Acquire Human Head Imagery in Wide AreasabstractThis paper considers the problem of estimating on-line the time-variant transformation relating a person's feet position in the image of a first, fixed camera, to his head position in the image of a second, pan-tilt-zoom camera. The transformation allows to acquire high-resolution images by steering the PTZ camera at targets detected in a fixed camera view. Assuming a planar scene and modeling humans as vertical segments, we present the development of an uncalibrated framework which does not require any 3D known location to be specified, and it allows to take into account both zooming camera and target uncertainties. Results show good performances in slave camera target head localization, degrading when the high zoom factor causes a lack of feature points in the slave camera. Alberto Del Bimbo, Fabrizio Dini, Andrea Grifoni, Federico Pernici |
AVSS | 4 |
| 2007 | Accurate self-calibration of two cameras by observations of a moving person on a ground planeabstractA calibration algorithm of two cameras using observations of a moving person is presented. Similar methods have been proposed for self-calibration with a single camera, but internal parameter estimation is only limited to the focal length. Recently it has been demonstrated that principal point supposed in the center of the image causes inaccuracy of all estimated parameters. Our method exploits two cameras, using image points of head and foot locations of a moving person, to determine for both cameras the focal length and the principal point. Moreover with the increasing number of cameras there is a demand of procedures to determine their relative placements. In this paper we also describe a method to find the relative position and orientation of two cameras: the rotation matrix and the translation vector which describe the rigid motion between the coordinate frames fixed in two cameras. Results in synthetic and real scenes are presented to evaluate the performance of the proposed method. Tsuhan Chen, Alberto Del Bimbo, Federico Pernici, Giuseppe Serra 0001 |
AVSS | 3 |
| 2006 | Learning Foveal Sensing Strategies in Unconstrained Surveillance EnvironmentsabstractIn this paper we report on techniques for automatically learning foveal sensing strategies for an active pan-tiltzoom camera. The approach uses reinforcement learning to discover foveal actions maximizing the performance of visual detectors, that are in turn assumed to be highly correlated with the task at hand. In our case, the main goal is to recognize people, hence a frontal face detection module is employed. The system uses reinforcement learning to learn if, when and how to foveate on a subject, based on its previous experience in terms or successful actions in similar situations. An action is successful if it leads to a correct face detection in the high resolution images obtained when the subject is zoomed in. In contrast with existing methods, the proposed approach obviates the need for camera calibration and camera performance modeling. Also, the method does not rely on active tracking of targets. Experimental results show how the system is capable of learning foveation strategies without requiring extensive a priori information or environmental models. Results also illustrate how the system effectively learns a strategy that allows the camera to foveate only in situations where successful detection is highly likely. Andrew D. Bagdanov, Alberto Del Bimbo, Walter Nunziati, Federico Pernici |
AVSS | 4 |
| 2006 | Towards on-line saccade planning for high-resolution image sensing
Alberto Del Bimbo, Federico Pernici |
Pattern Recognit. Lett. | 2 |
| 2005 | Metric 3D Reconstruction and Texture Acquisition of Surfaces of Revolution from a Single Uncalibrated ViewabstractImage analysis and computer vision can be effectively employed to recover the three-dimensional structure of imaged objects, together with their surface properties. In this paper, we address the problem of metric reconstruction and texture acquisition from a single uncalibrated view of a surface of revolution (SOR). Geometric constraints induced in the image by the symmetry properties of the SOR structure are exploited to perform self-calibration of a natural camera, 3D metric reconstruction, and texture acquisition. By exploiting the analogy with the geometry of single axis motion, we demonstrate that the imaged apparent contour and the visible segments of two imaged cross sections in a single SOR view provide enough information for these tasks. Original contributions of the paper are: single view self-calibration and reconstruction based on planar rectification, previously developed for planar surfaces, has been extended to deal also with the SOR class of curved surfaces; self-calibration is obtained by estimating both camera focal length (one parameter) and principal point (two parameters) from three independent linear constraints for the SOR fixed entities; the invariant-based description of the SOR scaling function has been extended from affine to perspective projection. The solution proposed exploits both the geometric and topological properties of the transformation that relates the apparent contour to the SOR scaling function. Therefore, with this method, a metric localization of the SOR occluded parts can be made, so as to cope with them correctly. For the reconstruction of textured SORs, texture acquisition is performed without requiring the estimation of external camera calibration parameters, but only using internal camera parameters obtained from self-calibration. Carlo Colombo, Alberto Del Bimbo, Federico Pernici |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | Image Mosaicing from Uncalibrated Views of a Surface of RevolutionabstractWe present a novel approach to obtain a mosaic image for the surface texture content of a surface of revolution (SOR) from a collection of uncalibrated views. The SOR scene constraint is used to calibrate each view and align the corresponding pictorial content into a global representation. Metric surface properties are extracted from each view by exploiting special properties of the imaged SOR geometry expressed in terms of homologies. Image alignment is achieved by projecting imaged surface elements onto a reference plane, and then registering them according to a translational motion model. This work extends previous research on calibrated scenes of right circular cylinders to the more general case of uncalibrated SOR scenes. Experimental results with images taken from the web demonstrate the effectiveness and the general applicability of the approach. 1 Introduction and Related Work Image mosaicing consists in merging collections of images having a partially overlapping content. The process can be decomposed into three main steps. First, the transformations Carlo Colombo, Alberto Del Bimbo, Federico Pernici |
BMVC | 3 |
| 2002 | Shape reconstruction from a single photograph for 3D object retrieval and visualizationabstractWe describe a geometric approach for reconstructing 3D textured graphical models of surfaces of revolution (SOR) from a single uncalibrated view. Metric reconstruction of 3D shape is complemented with the extraction of flattened 2D texture, so as to support visual retrieval from 2D/3D cues and to generate realistic 3D visualization models. The approach developed is quite simple, yet accurate and robust; its applications range from the preservation, analysis and classification of cultural heritage, to advanced graphics and multimedia. Carlo Colombo, Alberto Del Bimbo, Federico Pernici |
ICME (1) | 3 |