Claudio Baecchi

dblp:155/3120 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0001-8294-4539ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Real-time GenAI Solutions for Video Streaming in Low-bandwidth Settings
abstract
Surveillance systems such as bodycams and drones often operate under bandwidth constraints that limit video quality and degrade both human monitoring and AI-based analytics. Traditional compression techniques introduce artifacts that obscure critical details, especially in high-motion scenarios. We present a generative AI-powered video compression framework developed by Small Pixels, a spin-off of the University of Florence, designed to deliver Full HD video at significantly reduced bitrates. The system combines edge-side preprocessing for compression resilience with real-time receiver-side super-resolution, enabling up to 50% bandwidth savings while preserving perceptual quality and detection accuracy. Objective evaluations on EgoSeg and VisDrone datasets show +6.2 VMAF improvement and stable YOLOv11 detection performance with 30% less bitrate. Live trials in Singapore, within the Singapore Hatch-X Global Innovation Program, validated real-time operation with minimal latency, demonstrating clearer faces and motion in challenging conditions. The solution integrates seamlessly into existing infrastructures without hardware upgrades, offering a practical path to reliable, high-quality video streaming in bandwidth-limited environments.
Claudio Baecchi, Matteo Bruni, Fabio Clabot, Marco Bertini 0001
ACM Multimedia1
2023 Fashion recommendation based on style and social events
Federico Becattini, Lavinia De Divitiis, Claudio Baecchi, Alberto Del Bimbo
Multim. Tools Appl.3
2023 Disentangling Features for Fashion Recommendation
abstract
Online stores have become fundamental for the fashion industry, revolving around recommendation systems to suggest appropriate items to customers. Such recommendations often suffer from a lack of diversity and propose items that are similar to previous purchases of a user. Recently, a novel kind of approach based on Memory Augmented Neural Networks (MANNs) has been proposed, aimed at recommending a variety of garments to create an outfit by complementing a given fashion item. In this article we address the task of compatible garment recommendation developing a MANN architecture by taking into account the co-occurrence of clothing attributes, such as shape and color, to compose an outfit. To this end we obtain disentangled representations of fashion items and store them in external memory modules, used to guide recommendations at inference time. We show that our disentangled representations are able to achieve significantly better performance compared to the state of the art and also provide interpretable latent spaces, giving a qualitative explanation of the recommendations.
Lavinia De Divitiis, Federico Becattini, Claudio Baecchi, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Regular Polytope Networks
abstract
Neural networks are widely used as a model for classification in a large variety of tasks. Typically, a learnable transformation (i.e., the classifier) is placed at the end of such models returning a value for each class used for classification. This transformation plays an important role in determining how the generated features change during the learning process. In this work, we argue that this transformation not only can be fixed (i.e., set as nontrainable) with no loss of accuracy and with a reduction in memory usage, but it can also be used to learn stationary and maximally separated embeddings. We show that the stationarity of the embedding and its maximal separated representation can be theoretically justified by setting the weights of the fixed classifier to values taken from the coordinate vertices of the three regular polytopes available in [Formula: see text], namely, the d -Simplex, the d -Cube, and the d -Orthoplex. These regular polytopes have the maximal amount of symmetry that can be exploited to generate stationary features angularly centered around their corresponding fixed weights. Our approach improves and broadens the concept of a fixed classifier, recently proposed by Hoffer et al., to a larger class of fixed classifier models. Experimental results confirm the theoretical analysis, the generalization capability, the faster convergence, and the improved performance of the proposed method. Code will be publicly available.
Federico Pernici, Matteo Bruni, Claudio Baecchi, Alberto Del Bimbo
IEEE Trans. Neural Networks Learn. Syst.3
2021 Style-Based Outfit Recommendation
abstract
In this paper we propose a garment recommendation system that leverages emotive color information to give recommendations that adhere to a desired style. We leverage previous work by Shigenobu Kobayashi on how specific color combinations, that pertain to certain pre-defined styles, are able to convey specific emotions in human beings. Leveraging this information, we extend the classic general garment recommendation to a style-driven one, where the user can adapt the suggestions to a specific style that may be more appropriate for a specific social event. Here, first we train a generalized style classifier based on Kobayashi's color triplets, then we lever-age a recent memory network-based garment recommendation system to perform suggestions of bottom garments (e.g. skirts, trousers, etc.) given a user-defined top (e.g. a shirt, T-shirt, etc.). Suggestions are then processed to maintain only the ones that, according to our classifier, are coherent with the user defined style. Experiments show that our system is able to generalise on Kobayashi's color styles and that the recommendation system is able to propose garments that are in line with the user desire while also introducing diversity in the proposed garments.
Lavinia De Divitiis, Federico Becattini, Claudio Baecchi, Alberto Del Bimbo
CBMI3
2021 PLM-IPE: A Pixel-Landmark Mutual Enhanced Framework for Implicit Preference Estimation
abstract
In this paper, we are interested in understanding how customers perceive fashion recommendations, in particular when observing a proposed combination of garments to compose an outfit. Automatically understanding how a suggested item is perceived, without any kind of active engagement, is in fact an essential block to achieve interactive applications. We propose a pixel-landmark mutual enhanced framework for implicit preference estimation, named PLM-IPE, which is capable of inferring the user’s implicit preferences exploiting visual cues, without any active or conscious engagement. PLM-IPE consists of three key modules: pixel-based estimator, landmark-based estimator and mutual learning based optimization. The former two modules work on capturing the implicit reaction of the user from the pixel level and landmark level, respectively. The last module serves to transfer knowledge between the two parallel estimators. Towards evaluation, we collected a real-world dataset, named SentiGarment, which contains 3,345 facial reaction videos paired with suggested outfits and human labeled reaction scores. Extensive experiments show the superiority of our model over state-of-the-art approaches.
Federico Becattini, Xuemeng Song, Claudio Baecchi, Shi-Ting Fang, Claudio Ferrari, Liqiang Nie, Alberto Del Bimbo
MMAsia3
2020 Class-incremental Learning with Pre-allocated Fixed Classifiers
abstract
In class-incremental learning, a learning agent faces a stream of data with the goal of learning new classes while not forgetting previous ones. Neural networks are known to suffer under this setting, as they forget previously acquired knowledge. To address this problem, effective methods exploit past data stored in an episodic memory while expanding the final classifier nodes to accommodate the new classes. In this work, we substitute the expanding classifier with a novel fixed classifier in which a number of pre-allocated output nodes are subject to the classification loss right from the beginning of the learning phase. Contrarily to the standard expanding classifier, this allows: (a) the output nodes of future unseen classes to firstly see negative samples since the beginning of learning together with the positive samples that incrementally arrive; (b) to learn features that do not change their geometric configuration as novel classes are incorporated in the learning model. Experiments with public datasets show that the proposed approach is as effective as the expanding classifier while exhibiting novel intriguing properties of the internal feature representation that are otherwise not-existent. Our ablation study on pre-allocating a large number of classes further validates the approach.
Federico Pernici, Matteo Bruni, Claudio Baecchi, Francesco Turchini, Alberto Del Bimbo
ICPR3
2020 Automatic Interest Recognition from Posture and Behaviour
abstract
In the last years, the clothing industry has attracted a lot of interest from researchers. Increasing research efforts have been devoted into giving the buyer a way to improve the shopping experience by suggesting meaningful items to purchase. These efforts result in works aiming at suggesting good matches for clothes, but seem to lack one important aspect: understanding the user's interest. In fact, to suggest something it is first necessary to collect the user's personal interests, or something about his or her previous purchases. Without this information, no personalized suggestion can be made. User interest understanding allows to recognize if a user is showing interest in a product he or she is looking at, acquiring precious information that can be later leveraged. Usually user interest is associated to facial expressions, but these are known to be easily falsifiable. Moreover, when privacy is a concern, faces are often impossible to exploit. To address all these aspects, we propose an automatic system that aims to recognize the user's interest towards a garment by just looking at body posture and behaviour. To train and evaluate our system we create a body pose interest dataset, named BodyInterest, which consists of 30 users looking at garments for a total of approximately 6 hours of videos. Extensive evaluations show the effectiveness of our proposed method.
Wolmer Bigi, Claudio Baecchi, Alberto Del Bimbo
ACM Multimedia2
2017 Deep Sentiment Features of Context and Faces for Affective Video Analysis
abstract
Given the huge quantity of hours of video available on video sharing platforms such as YouTube, Vimeo, etc. development of automatic tools that help users find videos that fit their interests has attracted the attention of both scientific and industrial communities. So far the majority of the works have addressed semantic analysis, to identify objects, scenes and events depicted in videos, but more recently affective analysis of videos has started to gain more attention. In this work we investigate the use of sentiment driven features to classify the induced sentiment of a video, i.e. the sentiment reaction of the user. Instead of using standard computer vision features such as CNN features or SIFT features trained to recognize objects and scenes, we exploit sentiment related features such as the ones provided by Deep-SentiBank, and features extracted from models that exploit deep networks trained on face expressions. We experiment on two recently introduced datasets: LIRIS-ACCEDE and MEDIAEVAL-2015, that provide sentiment annotations of a large set of short videos. We show that our approach not only outperforms the current state-of-the-art in terms of valence and arousal classification accuracy, but it also uses a smaller number of features, requiring thus less video processing.
Claudio Baecchi, Tiberio Uricchio, Marco Bertini 0001, Alberto Del Bimbo
ICMR1
2017 Outdoor Object Recognition for Smart Audio Guides
abstract
We present a smart audio guide that adapts itself to the environment the user is navigating into. The system builds automatically a point of interest database exploiting Wikipedia and Google APIs as source. We rely on a computer vision system, to overcome the likely sensor limitations, and determine with high accuracy if the user is facing a certain landmark or if he is not facing any. Thanks to this the guide presents audio description at the most appropriate moment without any user intervention, using text-to-speech augmenting the experience.
Claudio Baecchi, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo
ACM Multimedia1
2017 Deep Artwork Detection and Retrieval for Automatic Context-Aware Audio Guides
abstract
In this article, we address the problem of creating a smart audio guide that adapts to the actions and interests of museum visitors. As an autonomous agent, our guide perceives the context and is able to interact with users in an appropriate fashion. To do so, it understands what the visitor is looking at, if the visitor is moving inside the museum hall, or if he or she is talking with a friend. The guide performs automatic recognition of artworks, and it provides configurable interface features to improve the user experience and the fruition of multimedia materials through semi-automatic interaction. Our smart audio guide is backed by a computer vision system capable of working in real time on a mobile device, coupled with audio and motion sensors. We propose the use of a compact Convolutional Neural Network (CNN) that performs object classification and localization. Using the same CNN features computed for these tasks, we perform also robust artwork recognition. To improve the recognition accuracy, we perform additional video processing using shape-based filtering, artwork tracking, and temporal filtering. The system has been deployed on an NVIDIA Jetson TK1 and a NVIDIA Shield Tablet K1 and tested in a real-world environment (Bargello Museum of Florence).
Lorenzo Seidenari, Claudio Baecchi, Tiberio Uricchio, Andrea Ferracani, Marco Bertini 0001, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.2
2016 A multimodal feature learning approach for sentiment analysis of social network multimedia
Claudio Baecchi, Tiberio Uricchio, Marco Bertini 0001, Alberto Del Bimbo
Multim. Tools Appl.1
2014 Fisher Vectors over Random Density Forests for Object Recognition
abstract
In this paper we describe a Fisher vector encoding of images over Random Density Forests. Random Density Forests (RDFs) are an unsupervised variation of Random Decision Forests for density estimation. In this work we train RDFs by splitting at each node in order to minimize the Gaussian differential entropy of each split. We use this as generative model of image patch features and derive the Fisher vector representation using the RDF as the underlying model. Our approach is computationally efficient, reducing the amount of Gaussian derivatives to compute, and allows more flexibility in the feature density modelling. We evaluate our approach on the PASCAL VOC 2007 dataset showing that our approach, that only uses linear classifiers, improves over bag of visual words and is comparable to the traditional Fisher vector encoding over Gaussian Mixture Models for density estimation.
Claudio Baecchi, Francesco Turchini, Lorenzo Seidenari, Andrew D. Bagdanov, Alberto Del Bimbo
ICPR1