EDBT 2026 Demo / reviewers in the wild / expert
Federico Becattini
dblp:183/0099
· DBLP profile ↗
33ranked-venue papers
13as first author
28since 2021 · last 2026
0000-0003-2537-2700ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 10 first-author · 17 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 12 since 2021Computer networks · 7 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spatio-temporal transformers for action unit classification with event cameras
Luca Cultrera, Federico Becattini, Lorenzo Berlincioni, Claudio Ferrari, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 2 |
| 2026 | Prediction of wine quality by monitoring climatic data and weather anomalies
Federico Becattini, Andrea Ferracani, Giuseppe Becchi, Alberto Del Bimbo |
Multim. Tools Appl. | 1 |
| 2025 | FRED: The Florence RGB-Event Drone DatasetabstractSmall, fast, and lightweight drones present significant challenges for traditional RGB cameras due to their limitations in capturing fast-moving objects, especially under challenging lighting conditions. Event cameras offer an ideal solution, providing high temporal definition and dynamic range, yet existing benchmarks often lack fine temporal resolution or drone-specific motion patterns, hindering progress in these areas. This paper introduces the Florence RGB-Event Drone dataset (FRED), a novel multimodal dataset specifically designed for drone detection, tracking, and trajectory forecasting, combining RGB video and event streams. FRED features more than 7 hours of densely annotated drone trajectories, using 5 different drone models and including challenging scenarios such as rain and adverse lighting conditions. We provide detailed evaluation protocols and standard metrics for each task, facilitating reproducible benchmarking. The authors hope FRED will advance research in high-speed drone perception and multimodal spatiotemporal understanding. Gabriele Magrini, Niccolò Marini, Federico Becattini, Lorenzo Berlincioni, Niccolò Biondi, Pietro Pala, Alberto Del Bimbo |
ACM Multimedia | 3 |
| 2025 | Immunizing Images from Text to Image Editing via Adversarial Cross-AttentionabstractRecent advances in text-based image editing have enabled fine-grained manipulation of visual content guided by natural language. However, such methods are susceptible to adversarial attacks. In this work, we propose a novel attack that targets the visual component of editing methods. We introduce Attention Attack, which disrupts the cross-attention between a textual prompt and the visual representation of the image by using an automatically generated caption of the source image as a proxy for the edit prompt. This breaks the alignment between the contents of the image and their textual description, without requiring knowledge of the editing method or the editing prompt. Reflecting on the reliability of existing metrics for immunization success, we propose two novel evaluation strategies: Caption Similarity, which quantifies semantic consistency between original and adversarial edits, and semantic Intersection over Union (IoU), which measures spatial layout disruption via segmentation masks. Experiments conducted on the TEDBench++ benchmark demonstrate that our attack significantly degrades editing performance while remaining imperceptible. Matteo Trippodo, Federico Becattini, Lorenzo Seidenari |
ACM Multimedia | 2 |
| 2025 | 3D Pose Nowcasting: Forecast the future to improve the presentabstractTechnologies to enable safe and effective collaboration and coexistence between humans and robots have gained significant importance in the last few years. A critical component useful for realizing this collaborative paradigm is the understanding of human and robot 3D poses using non-invasive systems. Therefore, in this paper, we propose a novel vision-based system leveraging depth data to accurately establish the 3D locations of skeleton joints. Specifically, we introduce the concept of Pose Nowcasting, denoting the capability of the proposed system to enhance its current pose estimation accuracy by jointly learning to forecast future poses. The experimental evaluation is conducted on two different datasets, providing accurate and real-time performance and confirming the validity of the proposed method on both the robotic and human scenarios. • We introduce the novel task of 3D Pose Nowcasting. • Our Pose Nowcasting system is based on both 3D Pose Estimation and Forecasting. • We show that knowledge about pose forecasting improves the accuracy of pose estimation. • We apply the proposed system both to human and robots. • Result on different dataset show state-of-the-art performance and robustness. Alessandro Simoni, Francesco Marchetti, Guido Borghi, Federico Becattini, Lorenzo Seidenari, Roberto Vezzani, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 4 |
| 2025 | Neuromorphic face analysis: A surveyabstractNeuromorphic sensors, also known as event cameras, are a class of imaging devices mimicking the function of biological visual systems. Unlike traditional frame-based cameras, which capture fixed images at discrete intervals, neuromorphic sensors continuously generate events that represent changes in light intensity or motion in the visual field with high temporal resolution and low latency. These properties have proven to be interesting in modeling human faces, both from an effectiveness and a privacy-preserving point of view. Neuromorphic face analysis however is still a raw and unstructured field of research, with several attempts at addressing different tasks with no clear standard or benchmark. This survey paper presents a comprehensive overview of capabilities, challenges and emerging applications in the domain of neuromorphic face analysis, to outline promising directions and open issues. After discussing the fundamental working principles of neuromorphic vision and presenting an in-depth overview of the related research, we explore the current state of available data, data representations, emerging challenges, and limitations that require further investigation. This paper aims to highlight the recent progress in this evolving field to provide researchers an all-encompassing analysis of the state of the art along with its problems and shortcomings. • Neuromorphic sensors are becoming more common in human-face related vision tasks. • As a new niche field it lacks overview papers helping researchers orient themselves. • We compare the relevant literature to classical methods and highlights pros and cons. • Several tasks and recently published open dataset are presented and compared. • Future research directions are provided after analyzing the relevant shortcomings. Federico Becattini, Lorenzo Berlincioni, Luca Cultrera, Alberto Del Bimbo |
Pattern Recognit. Lett. | 1 |
| 2025 | Spike-TBR: A noise resilient neuromorphic event representationabstractEvent cameras offer significant advantages over traditional frame-based sensors, including higher temporal resolution, lower latency and dynamic range. However, efficiently converting event streams into formats compatible with standard computer vision pipelines remains a challenging problem, particularly in the presence of noise. In this paper, we propose Spike-TBR, a novel event-based encoding strategy based on Temporal Binary Representation (TBR), addressing its vulnerability to noise by integrating spiking neurons. Spike-TBR combines the frame-based advantages of TBR with the noise-filtering capabilities of spiking neural networks, creating a more robust representation of event streams. We evaluate four variants of Spike-TBR, each using different spiking neurons, across multiple datasets, demonstrating superior performance in noise-affected scenarios while improving the results on clean data. Our method bridges the gap between spike-based and frame-based processing, offering a simple noise-resilient solution for event-driven vision applications. • We present Spike-TBR an enhanced version of Temporal Binary Representation. • By using spiking neurons we make the frame-based representation resilient to noise. • We obtain state-of-the-art results on four different neuromorphic datasets. Gabriele Magrini, Federico Becattini, Luca Cultrera, Lorenzo Berlincioni, Pietro Pala, Alberto Del Bimbo |
Pattern Recognit. Lett. | 2 |
| 2025 | Interactive Garment Recommendation with User in the LoopabstractRecommending fashion items often leverages rich user profiles and makes targeted suggestions based on past history and previous purchases. In this paper, we work under the assumption that no prior knowledge is given about a user. We propose to build a user profile on the fly by integrating user reactions as we recommend complementary items to compose an outfit. We present a reinforcement learning agent capable of suggesting appropriate garments and ingesting user feedback so to improve its recommendations and maximize user satisfaction. To train such a model, we resort to a proxy model to be able to simulate having user feedback in the training loop. We experiment on the IQON3000 fashion dataset and we find that a reinforcement learning-based agent becomes capable of improving its recommendations by taking into account personal preferences. Furthermore, such task demonstrated to be hard for non-reinforcement models, that cannot exploit exploration during training. Federico Becattini, Xiaolin Chen 0001, Andrea Puccia, Haokun Wen, Xuemeng Song, Liqiang Nie, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Continual Neural Computation
Matteo Tiezzi, Simone Marullo, Federico Becattini, Stefano Melacci |
ECML/PKDD (2) | 3 |
| 2024 | SMEMO: Social Memory for Trajectory ForecastingabstractEffective modeling of human interactions is of utmost importance when forecasting behaviors such as future trajectories. Each individual, with its motion, influences surrounding agents since everyone obeys to social non-written rules such as collision avoidance or group following. In this paper we model such interactions, which constantly evolve through time, by looking at the problem from an algorithmic point of view, i.e., as a data manipulation task. We present a neural network based on an end-to-end trainable working memory, which acts as an external storage where information about each agent can be continuously written, updated and recalled. We show that our method is capable of learning explainable cause-effect relationships between motions of different agents, obtaining state-of-the-art results on multiple trajectory forecasting datasets. Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | FLODCAST: Flow and depth forecasting via multimodal recurrent architecturesabstractForecasting motion and spatial positions of objects is of fundamental importance, especially in safety-critical settings such as autonomous driving. In this work, we address the issue by forecasting two different modalities that carry complementary information, namely optical flow and depth. To this end we propose FLODCAST a flow and depth forecasting model that leverages a multitask recurrent architecture, trained to jointly forecast both modalities at once. We stress the importance of training using flows and depth maps together, demonstrating that both tasks improve when the model is informed of the other modality. We train the proposed model to also perform predictions for several timesteps in the future. This provides better supervision and leads to more precise predictions, retaining the capability of the model to yield outputs autoregressively for any future time horizon. We test our model on the challenging Cityscapes dataset, obtaining state of the art results for both flow and depth forecasting. Thanks to the high quality of the generated flows, we also report benefits on the downstream task of segmentation forecasting, injecting our predictions in a flow-based mask-warping framework. Andrea Ciamarra, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo |
Pattern Recognit. | 2 |
| 2024 | Head Pose Estimation Patterns as Deepfake DetectorsabstractThe capacity to create “fake” videos has recently raised concerns about the reliability of multimedia content. Identifying between true and false information is a critical step toward resolving this problem. On this issue, several algorithms utilizing deep learning and facial landmarks have yielded intriguing results. Facial landmarks are traits that are solely tied to the subject’s head posture. Based on this observation, we study how Head Pose Estimation (HPE) patterns may be utilized to detect deepfakes in this work. The HPE patterns studied are based on FSA-Net, SynergyNet, and WSM, which are among the most performant approaches on the state-of-the-art. Finally, using a machine learning technique based on K-Nearest Neighbor and Dynamic Time Warping, their temporal patterns are categorized as authentic or false. We also offer a set of experiments for examining the feasibility of using deep learning techniques on such patterns. The findings reveal that the ability to recognize a deepfake video utilizing an HPE pattern is dependent on the HPE methodology. On the contrary, performance is less dependent on the performance of the utilized HPE technique. Experiments are carried out on the FaceForensics++ dataset that presents both identity swap and expression swap examples. The findings show that FSA-Net is an effective feature extraction method for determining whether a pattern belongs to a deepfake or not. The approach is also robust in comparison to deepfake videos created using various methods or for different goals. In the mean the method obtain 86% of accuracy on the identity swap task and 86.5% of accuracy on the expression swap. These findings offer up various possibilities and future directions for solving the deepfake detection problem using specialized HPE approaches, which are also known to be fast and reliable. Federico Becattini, Carmen Bisogni, Vincenzo Loia, Chiara Pero, Fei Hao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Deep Variational Learning for 360° Adaptive StreamingabstractPrediction of head movements in immersive media is key to designing efficient streaming systems able to focus the bandwidth budget on visible areas of the content. However, most of the numerous proposals made to predict user head motion in 360° images and videos do not explicitly consider a prominent characteristic of the head motion data: its intrinsic uncertainty. In this article, we present an approach to generate multiple plausible futures of head motion in 360° videos, given a common past trajectory. To our knowledge, this is the first work that considers the problem of multiple head motion prediction for 360° video streaming. We introduce our discrete variational multiple sequence (DVMS) learning framework, which builds on deep latent variable models. We design a training procedure to obtain a flexible, lightweight stochastic prediction model compatible with sequence-to-sequence neural architectures. Experimental results on four different datasets show that DVMS outperforms competitors adapted from the self-driving domain by up to 41% on prediction horizons up to 5 s, at lower computational and memory costs. To understand how the learned features account for the motion uncertainty, we analyze the structure of the learned latent space and connect it with the physical properties of the trajectories. We also introduce a method to estimate the likelihood of each generated trajectory, enabling the integration of DVMS in a streaming system. We hence deploy an extensive evaluation of the interest of our DVMS proposal for a streaming system. To do so, we first introduce a new Python-based 360° streaming simulator that we make available to the community. On real-world user, video, and networking data, we show that predicting multiple trajectories yields higher fairness between the traces, the gains for 20–30% of the users reaching up to 10% in visual quality for the best number K of trajectories to generate. Quentin Guimard, Lucile Sassatelli, Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Fashion recommendation based on style and social events
Federico Becattini, Lavinia De Divitiis, Claudio Baecchi, Alberto Del Bimbo |
Multim. Tools Appl. | 1 |
| 2023 | Multiple Trajectory Prediction of Moving Agents With Memory Augmented NetworksabstractPedestrians and drivers are expected to safely navigate complex urban environments along with several non cooperating agents. Autonomous vehicles will soon replicate this capability. Each agent acquires a representation of the world from an egocentric perspective and must make decisions ensuring safety for itself and others. This requires to predict motion patterns of observed agents for a far enough future. In this paper we propose MANTRA, a model that exploits memory augmented networks to effectively predict multiple trajectories of other agents, observed from an egocentric perspective. Our model stores observations in memory and uses trained controllers to write meaningful pattern encodings and read trajectories that are most likely to occur in future. We show that our method is able to natively perform multi-modal trajectory prediction obtaining state-of-the art results on four datasets. Moreover, thanks to the non-parametric nature of the memory module, we show how once trained our system can continuously improve by ingesting novel patterns. Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Attribute disentanglement with gradient reversal for interactive fashion retrievalabstractInteractive fashion search is gaining more and more interest thanks to the rapid diffusion of online retailers. It allows users to browse fashion items and perform attribute manipulations, modifying parts or details of given garments. To successfully model and analyze garments at such a fine-grained level, it is necessary to obtain attribute-wise representations, separating information relative to different characteristics. In this work we propose an attribute disentanglement method based on attribute classifiers and the usage of gradient reversal layers. This combination allows us to learn attribute-specific features, removing unwanted details from each representation. We test the effectiveness of our learned features in a fashion attribute manipulation task, obtaining state of the art results. Furthermore, to favor training stability we present a novel loss balancing approach, preventing reversed losses to diverge during the optimization process. Giovanna Scaramuzzino, Federico Becattini, Alberto Del Bimbo |
Pattern Recognit. Lett. | 2 |
| 2023 | VISCOUNTH: A Large-scale Multilingual Visual Question Answering Dataset for Cultural HeritageabstractVisual question answering has recently been settled as a fundamental multi-modal reasoning task of artificial intelligence that allows users to get information about visual content by asking questions in natural language. In the cultural heritage domain, this task can contribute to assisting visitors in museums and cultural sites, thus increasing engagement. However, the development of visual question answering models for cultural heritage is prevented by the lack of suitable large-scale datasets. To meet this demand, we built a large-scale heterogeneous and multilingual (Italian and English) dataset for cultural heritage that comprises approximately 500K Italian cultural assets and 6.5M question-answer pairs. We propose a novel formulation of the task that requires reasoning over both the visual content and an associated natural language description, and present baselines for this task. Results show that the current state of the art is reasonably effective but still far from satisfactory; therefore, further research in this area is recommended. Nonetheless, we also present a holistic baseline to address visual and contextual questions and foster future research on the topic. Federico Becattini, Pietro Bongini, Luana Bulla, Alberto Del Bimbo, Ludovica Marinucci, Misael Mongiovì, Valentina Presutti |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Disentangling Features for Fashion RecommendationabstractOnline stores have become fundamental for the fashion industry, revolving around recommendation systems to suggest appropriate items to customers. Such recommendations often suffer from a lack of diversity and propose items that are similar to previous purchases of a user. Recently, a novel kind of approach based on Memory Augmented Neural Networks (MANNs) has been proposed, aimed at recommending a variety of garments to create an outfit by complementing a given fashion item. In this article we address the task of compatible garment recommendation developing a MANN architecture by taking into account the co-occurrence of clothing attributes, such as shape and color, to compose an outfit. To this end we obtain disentangled representations of fashion items and store them in external memory modules, used to guide recommendations at inference time. We show that our disentangled representations are able to achieve significantly better performance compared to the state of the art and also provide interpretable latent spaces, giving a qualitative explanation of the recommendations. Lavinia De Divitiis, Federico Becattini, Claudio Baecchi, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | (Compress and Restore)N: A Robust Defense Against Adversarial Attacks on Image ClassificationabstractModern image classification approaches often rely on deep neural networks, which have shown pronounced weakness to adversarial examples: images corrupted with specifically designed yet imperceptible noise that causes the network to misclassify. In this article, we propose a conceptually simple yet robust solution to tackle adversarial attacks on image classification. Our defense works by first applying a JPEG compression with a random quality factor; compression artifacts are subsequently removed by means of a generative model Artifact Restoration GAN. The process can be iterated ensuring the image is not degraded and hence the classification not compromised. We train different AR-GANs for different compression factors, so that we can change its parameters dynamically at each iteration depending on the current compression, making the gradient approximation difficult. We experiment with our defense against three white-box and two black-box attacks, with a particular focus on the state-of-the-art BPDA attack. Our method does not require any adversarial training, and is independent of both the classifier and the attack. Experiments demonstrate that dynamically changing the AR-GAN parameters is of fundamental importance to obtain significant robustness. Claudio Ferrari, Federico Becattini, Leonardo Galteri, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Online Deep Clustering with Video Track ConsistencyabstractSeveral unsupervised and self-supervised approaches have been developed in recent years to learn visual features from large-scale unlabeled datasets. Their main drawback however is that these methods are hardly able to recognize visual features of the same object if it is simply rotated or the perspective of the camera changes. To overcome this limitation and at the same time exploit a useful source of supervision, we take into account video object tracks. Following the intuition that two patches in a track should have similar visual representations in a learned feature space, we adopt an unsupervised clustering-based approach and constrain such representations to be labeled as the same category since they likely belong to the same object or object part. Experimental results on two downstream tasks on different datasets demonstrate the effectiveness of our Online Deep Clustering with Video Track Consistency (ODCT) approach compared to prior work, which did not leverage temporal information. In addition we show that exploiting an unsupervised class-agnostic, yet noisy, track generator yields to better accuracy compared to relying on costly and precise track annotations. Alessandra Alfani, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo |
ICPR | 2 |
| 2022 | Memory NetworksabstractMemory Networks are models equipped with a storage component where information can generally be written and successively retrieved for any purpose. Simple forms of memory networks like the popular recurrent neural networks (RNN), LSTMs or GRUs, have limited storage capabilities and for specific tasks. In contrast, recent works, starting from Memory Augmented Neural Networks, overcome storage and computational limitations with the addition of a controller network with an external element-wise addressable memory. This tutorial aims at providing an overview of such memory-based techniques and their applications in multimedia. It will cover an explanation of the basic concepts behind recurrent neural networks and will then delve into the advanced details of memory augmented neural networks, their structure and how such models can be trained. We target a broad audience, from beginners to experienced researchers, offering an in-depth introduction to an important crop of literature which is starting to gain interest in the multimedia, computer vision and natural language processing communities. Federico Becattini, Tiberio Uricchio |
ACM Multimedia | 1 |
| 2022 | MCFR'22: 1st Workshop on Multimedia Computing towards Fashion RecommendationabstractWith the proliferation of online shopping, fashion recommendation, which aims to provide suitable suggestions to support the consumer's purchase in e-commerce platforms, has gained increasing research attention from both academia and industry. Although existing efforts have achieved great progress, they focus on the visual modality, lacking the exploration of other modalities of items, e.g., the textual descriptions and attributes of items. Accordingly, this workshop targets calling for a coordinated effort to promote the multimedia computing towards fashion recommendation. This workshop will showcase the innovative methodologies and ideas on new yet challenging research problems, including (not limited to) fashion recommendation, interactive fashion recommendation, interactive garment retrieval, and outfit compatibility modeling. Xuemeng Song, Jingjing Chen 0001, Federico Becattini, Weili Guan, Yibing Zhan, Tat-Seng Chua |
ACM Multimedia | 3 |
| 2022 | Deep variational learning for multiple trajectory prediction of 360° head movementsabstractPrediction of head movements in immersive media is key to design efficient streaming systems able to focus the bandwidth budget on visible areas of the content. Numerous proposals have therefore been made in the recent years to predict 360° images and videos. However, the performance of these models is limited by a main characteristic of the head motion data: its intrinsic uncertainty. In this article, we present an approach to generate multiple plausible futures of head motion in 360° videos, given a common past trajectory. Our method provides likelihood estimates of every predicted trajectory, enabling direct integration in streaming optimization. To the best of our knowledge, this is the first work that considers the problem of multiple head motion prediction for 360° video streaming. We first quantify this uncertainty from the data. We then introduce our discrete variational multiple sequence (DVMS) learning framework, which builds on deep latent variable models. We design a training procedure to obtain a flexible and lightweight stochastic prediction model compatible with sequence-to-sequence recurrent neural architectures. Experimental results on 3 different datasets show that our method DVMS outperforms competitors adapted from the self-driving domain by up to 37% on prediction horizons up to 5 sec., at lower computational and memory costs. Finally, we design a method to estimate the respective likelihoods of the multiple predicted trajectories, by exploiting the stationarity of the distribution of the prediction error over the latent space. Experimental results on 3 datasets show the quality of these estimates, and how they depend on the video category. Quentin Guimard, Lucile Sassatelli, Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo |
MMSys | 4 |
| 2022 | Events in crowded places: A smart service management
Federico Becattini, Andrea Ferracani, Giuseppe Becchi, Alberto Del Bimbo |
Pattern Recognit. Lett. | 1 |
| 2022 | Understanding Human Reactions Looking at Facial Microexpressions With an Event CameraabstractWith the establishment ofIndustry 4.0, machines are now required to interact with workers. By observing biometrics they can assess if humans are authorized, or mentally and physically fit to work. Understanding body language, makes human–machine interaction more natural, secure, and effective. Nonetheless, traditional cameras have limitations; low frame rate and dynamic range hinder a comprehensive human understanding. This poses a challenge, since faces undergo frequent instantaneous microexpressions. In addition, this is privacy-sensitive information that must be protected. We propose to model expressions with event cameras, bio-inspired vision sensors that have found application within the Industry 4.0 scope. They capture motion at millisecond rates and work under challenging conditions like low illumination and highly dynamic scenes. Such cameras are also privacy-preserving, making them extremely interesting for industry. We show that using event cameras, we can understand human reactions by only observing facial expressions. Comparison with red-green-blue (RGB)-based modeling demonstrates improved effectiveness and robustness. Federico Becattini, Federico Palai, Alberto Del Bimbo |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Style-Based Outfit RecommendationabstractIn this paper we propose a garment recommendation system that leverages emotive color information to give recommendations that adhere to a desired style. We leverage previous work by Shigenobu Kobayashi on how specific color combinations, that pertain to certain pre-defined styles, are able to convey specific emotions in human beings. Leveraging this information, we extend the classic general garment recommendation to a style-driven one, where the user can adapt the suggestions to a specific style that may be more appropriate for a specific social event. Here, first we train a generalized style classifier based on Kobayashi's color triplets, then we lever-age a recent memory network-based garment recommendation system to perform suggestions of bottom garments (e.g. skirts, trousers, etc.) given a user-defined top (e.g. a shirt, T-shirt, etc.). Suggestions are then processed to maintain only the ones that, according to our classifier, are coherent with the user defined style. Experiments show that our system is able to generalise on Kobayashi's color styles and that the recommendation system is able to propose garments that are in line with the user desire while also introducing diversity in the proposed garments. Lavinia De Divitiis, Federico Becattini, Claudio Baecchi, Alberto Del Bimbo |
CBMI | 2 |
| 2021 | PLM-IPE: A Pixel-Landmark Mutual Enhanced Framework for Implicit Preference EstimationabstractIn this paper, we are interested in understanding how customers perceive fashion recommendations, in particular when observing a proposed combination of garments to compose an outfit. Automatically understanding how a suggested item is perceived, without any kind of active engagement, is in fact an essential block to achieve interactive applications. We propose a pixel-landmark mutual enhanced framework for implicit preference estimation, named PLM-IPE, which is capable of inferring the user’s implicit preferences exploiting visual cues, without any active or conscious engagement. PLM-IPE consists of three key modules: pixel-based estimator, landmark-based estimator and mutual learning based optimization. The former two modules work on capturing the implicit reaction of the user from the pixel level and landmark level, respectively. The last module serves to transfer knowledge between the two parallel estimators. Towards evaluation, we collected a real-world dataset, named SentiGarment, which contains 3,345 facial reaction videos paired with suggested outfits and human labeled reaction scores. Extensive experiments show the superiority of our model over state-of-the-art approaches. Federico Becattini, Xuemeng Song, Claudio Baecchi, Shi-Ting Fang, Claudio Ferrari, Liqiang Nie, Alberto Del Bimbo |
MMAsia | 1 |
| 2021 | Am I Done? Predicting Action Progress in VideosabstractIn this article, we deal with the problem of predicting action progress in videos. We argue that this is an extremely important task, since it can be valuable for a wide range of interaction applications. To this end, we introduce a novel approach, named ProgressNet, capable of predicting when an action takes place in a video, where it is located within the frames, and how far it has progressed during its execution. To provide a general definition of action progress, we ground our work in the linguistics literature, borrowing terms and concepts to understand which actions can be the subject of progress estimation. As a result, we define a categorization of actions and their phases. Motivated by the recent success obtained from the interaction of Convolutional and Recurrent Neural Networks, our model is based on a combination of the Faster R-CNN framework, to make framewise predictions, and LSTM networks, to estimate action progress through time. After introducing two evaluation protocols for the task at hand, we demonstrate the capability of our model to effectively predict action progress on the UCF-101 and J-HMDB datasets. Federico Becattini, Tiberio Uricchio, Lorenzo Seidenari, Lamberto Ballan, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | MANTRA: Memory Augmented Networks for Multiple Trajectory PredictionabstractAutonomous vehicles are expected to drive in complex scenarios with several independent non cooperating agents. Path planning for safely navigating in such environments can not just rely on perceiving present location and motion of other agents. It requires instead to predict such variables in a far enough future. In this paper we address the problem of multimodal trajectory prediction exploiting a Memory Augmented Neural Network. Our method learns past and future trajectory embeddings using recurrent neural networks and exploits an associative external memory to store and retrieve such embeddings. Trajectory prediction is then performed by decoding in-memory future encodings conditioned with the observed past. We incorporate scene knowledge in the decoding state by learning a CNN on top of semantic scene maps. Memory growth is limited by learning a writing controller based on the predictive capability of existing embeddings. We show that our method is able to natively perform multi-modal trajectory prediction obtaining state-of-the art results on three datasets. Moreover, thanks to the non-parametric nature of the memory module, we show how once trained our system can continuously improve by ingesting novel patterns. Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo |
CVPR | 2 |
| 2020 | Multiple Future Prediction Leveraging Synthetic TrajectoriesabstractTrajectory prediction is an important task, especially in autonomous driving. The ability to forecast the position of other moving agents can yield to an effective planning, ensuring safety for the autonomous vehicle as well for the observed entities. In this work we propose a data driven approach based on Markov Chains to generate synthetic trajectories, which are useful for training a multiple future trajectory predictor. The advantages are twofold: on the one hand synthetic samples can be used to augment existing datasets and train more effective predictors; on the other hand, it allows to generate samples with multiple ground truths, corresponding to diverse equally likely outcomes of the observed trajectory. We define a trajectory prediction model and a loss that explicitly address the multimodality of the problem and we show that combining synthetic and real data leads to prediction improvements, obtaining state of the art results. Lorenzo Berlincioni, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo |
ICPR | 2 |
| 2020 | Temporal Binary Representation for Event-Based Action RecognitionabstractIn this paper we present an event aggregation strategy to convert the output of an event camera into frames processable by traditional Computer Vision algorithms. The proposed method first generates sequences of intermediate binary representations, which are then losslessly transformed into a compact format by simply applying a binary-to-decimal conversion. This strategy allows us to encode temporal information directly into pixel values, which are then interpreted by deep learning models. We apply our strategy, called Temporal Binary Representation, to the task of Gesture Recognition, obtaining state of the art results on the popular DVS128 Gesture Dataset. To underline the effectiveness of the proposed method compared to existing ones, we also collect an extension of the dataset under more challenging conditions on which to perform experiments. Simone Undri Innocenti, Federico Becattini, Federico Pernici, Alberto Del Bimbo |
ICPR | 2 |
| 2019 | NeuronUnityIntegration2.0. A Unity Based Application for Motion Capture and Gesture RecognitionabstractNeuronUnityIntgration2.0 (demo video is avilable at http://tiny.cc/u1lz6y) is a plugin for Unity which provides gesture recognition functionalities through the Perception Neuron motion capture suit. The system offers a recording mode, which guides the user through the collection of a dataset of gestures, and a recognition mode, capable of detecting the recorded actions in real time. Gestures are recognized by training Support Vector Machines directly within our plugin. We demonstrate the effectiveness of our application through an experimental evaluation on a newly collected dataset. Furthermore, external applications can exploit NeuronUnityIntgration2.0's recognition capabilities thanks to a set of exposed API. Federico Becattini, Andrea Ferracani, Filippo Principi, Marioemanuele Ghianni, Alberto Del Bimbo |
ACM Multimedia | 1 |
| 2017 | Indexing quantized ensembles of exemplar-SVMs with rejecting taxonomies
Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo |
Multim. Tools Appl. | 1 |