Lorenzo Seidenari

dblp:63/8108 · DBLP profile ↗
← Back
49ranked-venue papers
5as first author
15since 2021 · last 2025
0000-0003-4816-0268ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 21 · 2 first-author · 7 since 2021Computer networks · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Immunizing Images from Text to Image Editing via Adversarial Cross-Attention
abstract
Recent advances in text-based image editing have enabled fine-grained manipulation of visual content guided by natural language. However, such methods are susceptible to adversarial attacks. In this work, we propose a novel attack that targets the visual component of editing methods. We introduce Attention Attack, which disrupts the cross-attention between a textual prompt and the visual representation of the image by using an automatically generated caption of the source image as a proxy for the edit prompt. This breaks the alignment between the contents of the image and their textual description, without requiring knowledge of the editing method or the editing prompt. Reflecting on the reliability of existing metrics for immunization success, we propose two novel evaluation strategies: Caption Similarity, which quantifies semantic consistency between original and adversarial edits, and semantic Intersection over Union (IoU), which measures spatial layout disruption via segmentation masks. Experiments conducted on the TEDBench++ benchmark demonstrate that our attack significantly degrades editing performance while remaining imperceptible.
Matteo Trippodo, Federico Becattini, Lorenzo Seidenari
ACM Multimedia3
2025 3D Pose Nowcasting: Forecast the future to improve the present
abstract
Technologies to enable safe and effective collaboration and coexistence between humans and robots have gained significant importance in the last few years. A critical component useful for realizing this collaborative paradigm is the understanding of human and robot 3D poses using non-invasive systems. Therefore, in this paper, we propose a novel vision-based system leveraging depth data to accurately establish the 3D locations of skeleton joints. Specifically, we introduce the concept of Pose Nowcasting, denoting the capability of the proposed system to enhance its current pose estimation accuracy by jointly learning to forecast future poses. The experimental evaluation is conducted on two different datasets, providing accurate and real-time performance and confirming the validity of the proposed method on both the robotic and human scenarios. • We introduce the novel task of 3D Pose Nowcasting. • Our Pose Nowcasting system is based on both 3D Pose Estimation and Forecasting. • We show that knowledge about pose forecasting improves the accuracy of pose estimation. • We apply the proposed system both to human and robots. • Result on different dataset show state-of-the-art performance and robustness.
Alessandro Simoni, Francesco Marchetti, Guido Borghi, Federico Becattini, Lorenzo Seidenari, Roberto Vezzani, Alberto Del Bimbo
Comput. Vis. Image Underst.5
2025 A fine-tuning approach based on spatio-temporal features for few-shot video object detection
abstract
This paper describes a new Fine-Tuning approach for Few-Shot object detection in Videos that exploits spatio-temporal information to boost detection precision. Despite the progress made in the single image domain in recent years, the few-shot video object detection problem remains almost unexplored. A few-shot detector must quickly adapt to a new domain with a limited number of annotations per category. Therefore, it is not possible to include videos in the training set, hindering the spatio-temporal learning process. We propose augmenting each training image with synthetic frames to train the spatio-temporal module of our method. This module employs attention mechanisms to mine relationships between proposals across frames, effectively leveraging spatio-temporal information. A spatio-temporal double head then localizes objects in the current frame while classifying them using both context from nearby frames and information from the current frame. Finally, the predicted scores are fed into a long-term object-linking method that generates object tubes across the video. By optimizing the classification score based on these tubes, our approach ensures spatio-temporal consistency. Classification is the primary challenge in few-shot object detection. Our results show that spatio-temporal information helps to mitigate this issue, paving the way for future research in this direction. FTFSVid achieves 41.9 AP50 on the Few-Shot Video Object Detection (FSVOD-500) and 42.9 AP50 on the Few-Shot YouTube Video (FSYTV-40) dataset, surpassing our spatial baseline by 4.3 and 2.5 points. Additionally, FTFSVid outperforms previous few-shot video object detectors by 3.2 points on FSVOD-500 and 14.5 points on FSYTV-40, setting a new state-of-the-art.
Daniel Cores, Lorenzo Seidenari, Alberto Del Bimbo, Víctor M. Brea 0001, Manuel Mucientes
Eng. Appl. Artif. Intell.2
2025 Foundation Forecasting in IoE Networks: When Generative AI Meets Programmable Edge Nodes
abstract
Artificial intelligence (AI)-native edge networks are promising solutions to seamlessly integrate AI into modern network architectures, promoting intelligent and hyper-flexible behavior in self-adaptation and reconfiguration of hybrid networks. AI-native edge nodes have to promptly react to any change in network conditions and to be linked to Internet of Everything (IoE) deployed in etherogeneous communication domains. This paper deals with a next-generation programmable edge node capable of being employed in different domains in a unified and flexible manner. With the aim to achieve a low re-configuration overhead in such herogeneous contexts, this paper proposes the integration of a generative-AI module within an edge node exploiting foundation models for efficient general-purpose time-series prediction, without involving overhead and costs due to models trained from scratch and overcoming data scarcity. This permits to manage IoE networks deployed in both homogeneous and heterogeneous domains, i.e., aqua, ground and air, autonomously, without the need for adjustments from the outside. As foundation models, we focused on Chronos and TimesFM, in both the zero-shot and fine-tuning learning paradigms. Finally, performance results in terms of prediction accuracy, training, and inference time are provided and compared with those achieved by the state-of-the-art recurrent neural networks trained from scratch and baseline alternatives. The obtained results corroborate the potential of foundation models as key enablers of native AI networks, achieving good accuracy despite the absence of a training phase (i.e., in zero-shot mode) or with limited training (fine-tuning), w.r.t. alternatives.
Francesco Marchetti, Benedetta Picano, Lorenzo Seidenari, Romano Fantacci
IEEE Internet Things J.3
2024 SMEMO: Social Memory for Trajectory Forecasting
abstract
Effective modeling of human interactions is of utmost importance when forecasting behaviors such as future trajectories. Each individual, with its motion, influences surrounding agents since everyone obeys to social non-written rules such as collision avoidance or group following. In this paper we model such interactions, which constantly evolve through time, by looking at the problem from an algorithmic point of view, i.e., as a data manipulation task. We present a neural network based on an end-to-end trainable working memory, which acts as an external storage where information about each agent can be continuously written, updated and recalled. We show that our method is capable of learning explainable cause-effect relationships between motions of different agents, obtaining state-of-the-art results on multiple trajectory forecasting datasets.
Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 FLODCAST: Flow and depth forecasting via multimodal recurrent architectures
abstract
Forecasting motion and spatial positions of objects is of fundamental importance, especially in safety-critical settings such as autonomous driving. In this work, we address the issue by forecasting two different modalities that carry complementary information, namely optical flow and depth. To this end we propose FLODCAST a flow and depth forecasting model that leverages a multitask recurrent architecture, trained to jointly forecast both modalities at once. We stress the importance of training using flows and depth maps together, demonstrating that both tasks improve when the model is informed of the other modality. We train the proposed model to also perform predictions for several timesteps in the future. This provides better supervision and leads to more precise predictions, retaining the capability of the model to yield outputs autoregressively for any future time horizon. We test our model on the challenging Cityscapes dataset, obtaining state of the art results for both flow and depth forecasting. Thanks to the high quality of the generated flows, we also report benefits on the downstream task of segmentation forecasting, injecting our predictions in a flow-based mask-warping framework.
Andrea Ciamarra, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
Pattern Recognit.3
2024 Deep Variational Learning for 360° Adaptive Streaming
abstract
Prediction of head movements in immersive media is key to designing efficient streaming systems able to focus the bandwidth budget on visible areas of the content. However, most of the numerous proposals made to predict user head motion in 360° images and videos do not explicitly consider a prominent characteristic of the head motion data: its intrinsic uncertainty. In this article, we present an approach to generate multiple plausible futures of head motion in 360° videos, given a common past trajectory. To our knowledge, this is the first work that considers the problem of multiple head motion prediction for 360° video streaming. We introduce our discrete variational multiple sequence (DVMS) learning framework, which builds on deep latent variable models. We design a training procedure to obtain a flexible, lightweight stochastic prediction model compatible with sequence-to-sequence neural architectures. Experimental results on four different datasets show that DVMS outperforms competitors adapted from the self-driving domain by up to 41% on prediction horizons up to 5 s, at lower computational and memory costs. To understand how the learned features account for the motion uncertainty, we analyze the structure of the learned latent space and connect it with the physical properties of the trajectories. We also introduce a method to estimate the likelihood of each generated trajectory, enabling the integration of DVMS in a streaming system. We hence deploy an extensive evaluation of the interest of our DVMS proposal for a streaming system. To do so, we first introduce a new Python-based 360° streaming simulator that we make available to the community. On real-world user, video, and networking data, we show that predicting multiple trajectories yields higher fairness between the traces, the gains for 20–30% of the users reaching up to 10% in visual quality for the best number K of trajectories to generate.
Quentin Guimard, Lucile Sassatelli, Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Downsampling GAN for Small Object Data Augmentation
Daniel Cores, Víctor M. Brea 0001, Manuel Mucientes, Lorenzo Seidenari, Alberto Del Bimbo
CAIP (1)4
2023 Multiple Trajectory Prediction of Moving Agents With Memory Augmented Networks
abstract
Pedestrians and drivers are expected to safely navigate complex urban environments along with several non cooperating agents. Autonomous vehicles will soon replicate this capability. Each agent acquires a representation of the world from an egocentric perspective and must make decisions ensuring safety for itself and others. This requires to predict motion patterns of observed agents for a far enough future. In this paper we propose MANTRA, a model that exploits memory augmented networks to effectively predict multiple trajectories of other agents, observed from an egocentric perspective. Our model stores observations in memory and uses trained controllers to write meaningful pattern encodings and read trajectories that are most likely to occur in future. We show that our method is able to natively perform multi-modal trajectory prediction obtaining state-of-the art results on four datasets. Moreover, thanks to the non-parametric nature of the memory module, we show how once trained our system can continuously improve by ingesting novel patterns.
Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 A full data augmentation pipeline for small object detection based on generative adversarial networks
abstract
Object detection accuracy on small objects, i.e., objects under 32 × 32 pixels, lags behind that of large ones. To address this issue, innovative architectures have been designed and new datasets have been released. Still, the number of small objects in many datasets does not suffice for training. The advent of the generative adversarial networks (GANs) opens up a new data augmentation possibility for training architectures without the costly task of annotating huge datasets for small objects. In this paper, we propose a full pipeline for data augmentation for small object detection which combines a GAN-based object generator with techniques of object segmentation, image inpainting, and image blending to achieve high-quality synthetic data. The main component of our pipeline is DS-GAN, a novel GAN-based architecture that generates realistic small objects from larger ones. Experimental results show that our overall data augmentation method improves the performance of state-of-the-art models up to 11.9% [email protected] on UAVDT and by 4.7% [email protected] on iSAID, both for the small objects subset and for a scenario where the number of training instances is limited.
Brais Bosquet, Daniel Cores, Lorenzo Seidenari, Víctor M. Brea 0001, Manuel Mucientes, Alberto Del Bimbo
Pattern Recognit.3
2022 Online Deep Clustering with Video Track Consistency
abstract
Several unsupervised and self-supervised approaches have been developed in recent years to learn visual features from large-scale unlabeled datasets. Their main drawback however is that these methods are hardly able to recognize visual features of the same object if it is simply rotated or the perspective of the camera changes. To overcome this limitation and at the same time exploit a useful source of supervision, we take into account video object tracks. Following the intuition that two patches in a track should have similar visual representations in a learned feature space, we adopt an unsupervised clustering-based approach and constrain such representations to be labeled as the same category since they likely belong to the same object or object part. Experimental results on two downstream tasks on different datasets demonstrate the effectiveness of our Online Deep Clustering with Video Track Consistency (ODCT) approach compared to prior work, which did not leverage temporal information. In addition we show that exploiting an unsupervised class-agnostic, yet noisy, track generator yields to better accuracy compared to relying on costly and precise track annotations.
Alessandra Alfani, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
ICPR3
2022 Deep variational learning for multiple trajectory prediction of 360° head movements
abstract
Prediction of head movements in immersive media is key to design efficient streaming systems able to focus the bandwidth budget on visible areas of the content. Numerous proposals have therefore been made in the recent years to predict 360° images and videos. However, the performance of these models is limited by a main characteristic of the head motion data: its intrinsic uncertainty. In this article, we present an approach to generate multiple plausible futures of head motion in 360° videos, given a common past trajectory. Our method provides likelihood estimates of every predicted trajectory, enabling direct integration in streaming optimization. To the best of our knowledge, this is the first work that considers the problem of multiple head motion prediction for 360° video streaming. We first quantify this uncertainty from the data. We then introduce our discrete variational multiple sequence (DVMS) learning framework, which builds on deep latent variable models. We design a training procedure to obtain a flexible and lightweight stochastic prediction model compatible with sequence-to-sequence recurrent neural architectures. Experimental results on 3 different datasets show that our method DVMS outperforms competitors adapted from the self-driving domain by up to 37% on prediction horizons up to 5 sec., at lower computational and memory costs. Finally, we design a method to estimate the respective likelihoods of the multiple predicted trajectories, by exploiting the stationarity of the distribution of the prediction error over the latent space. Experimental results on 3 datasets show the quality of these estimates, and how they depend on the video category.
Quentin Guimard, Lucile Sassatelli, Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
MMSys5
2022 LANBIQUE: LANguage-based Blind Image QUality Evaluation
abstract
Image quality assessment is often performed with deep networks that are fine-tuned to regress a human provided quality score of a given image. Usually, this approach may lack generalization capabilities and, while being highly precise on similar image distribution, it may yield lower correlation on unseen distortions. In particular, they show poor performances, whereas images corrupted by noise, blur, or compression have been restored by generative models. As a matter of fact, evaluation of these generative models is often performed providing anecdotal results to the reader. In the case of image enhancement and restoration, reference images are usually available. Nevertheless, using signal based metrics often leads to counterintuitive results: Highly natural crisp images may obtain worse scores than blurry ones. However, blind reference image assessment may rank images reconstructed with GANs higher than the original undistorted images. To avoid time-consuming human-based image assessment, semantic computer vision tasks may be exploited instead. In this article, we advocate the use of language generation tasks to evaluate the quality of restored images. We refer to our assessment approach as LANguage-based Blind Image QUality Evaluation (LANBIQUE). We show experimentally that image captioning, used as a downstream task, may serve as a method to score image quality, independently of the distortion process that affects the data. Captioning scores are better aligned with human rankings with respect to classic signal based or No-reference image quality metrics. We show insights on how the corruption, by artefacts, of local image structure may steer image captions in the wrong direction.
Leonardo Galteri, Lorenzo Seidenari, Pietro Bongini, Marco Bertini 0001, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Language Based Image Quality Assessment
abstract
Evaluation of generative models, in the visual domain, is often performed providing anecdotal results to the reader. In the case of image enhancement, reference images are usually available. Nonetheless, using signal based metrics often leads to counterintuitive results: highly natural crisp images may obtain worse scores than blurry ones. On the other hand, blind reference image assessment may rank images reconstructed with GANs higher than the original undistorted images. To avoid time consuming human based image assessment, semantic computer vision tasks may be exploited instead [9, 25, 33]. In this paper we advocate the use of language generation tasks to evaluate the quality of restored images. We show experimentally that image captioning, used as a downstream task, may serve as a method to score image quality. Captioning scores are better aligned with human rankings with respect to signal based metrics or no-reference image quality metrics. We show insights on how the corruption, by artifacts, of local image structure may steer image captions in the wrong direction.
Lorenzo Seidenari, Leonardo Galteri, Pietro Bongini, Marco Bertini 0001, Alberto Del Bimbo
MMAsia1
2021 Am I Done? Predicting Action Progress in Videos
abstract
In this article, we deal with the problem of predicting action progress in videos. We argue that this is an extremely important task, since it can be valuable for a wide range of interaction applications. To this end, we introduce a novel approach, named ProgressNet, capable of predicting when an action takes place in a video, where it is located within the frames, and how far it has progressed during its execution. To provide a general definition of action progress, we ground our work in the linguistics literature, borrowing terms and concepts to understand which actions can be the subject of progress estimation. As a result, we define a categorization of actions and their phases. Motivated by the recent success obtained from the interaction of Convolutional and Recurrent Neural Networks, our model is based on a combination of the Faster R-CNN framework, to make framewise predictions, and LSTM networks, to estimate action progress through time. After introducing two evaluation protocols for the task at hand, we demonstrate the capability of our model to effectively predict action progress on the UCF-101 and J-HMDB datasets.
Federico Becattini, Tiberio Uricchio, Lorenzo Seidenari, Lamberto Ballan, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.3
2020 MANTRA: Memory Augmented Networks for Multiple Trajectory Prediction
abstract
Autonomous vehicles are expected to drive in complex scenarios with several independent non cooperating agents. Path planning for safely navigating in such environments can not just rely on perceiving present location and motion of other agents. It requires instead to predict such variables in a far enough future. In this paper we address the problem of multimodal trajectory prediction exploiting a Memory Augmented Neural Network. Our method learns past and future trajectory embeddings using recurrent neural networks and exploits an associative external memory to store and retrieve such embeddings. Trajectory prediction is then performed by decoding in-memory future encodings conditioned with the observed past. We incorporate scene knowledge in the decoding state by learning a CNN on top of semantic scene maps. Memory growth is limited by learning a writing controller based on the predictive capability of existing embeddings. We show that our method is able to natively perform multi-modal trajectory prediction obtaining state-of-the art results on three datasets. Moreover, thanks to the non-parametric nature of the memory module, we show how once trained our system can continuously improve by ingesting novel patterns.
Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
CVPR3
2020 Multiple Future Prediction Leveraging Synthetic Trajectories
abstract
Trajectory prediction is an important task, especially in autonomous driving. The ability to forecast the position of other moving agents can yield to an effective planning, ensuring safety for the autonomous vehicle as well for the observed entities. In this work we propose a data driven approach based on Markov Chains to generate synthetic trajectories, which are useful for training a multiple future trajectory predictor. The advantages are twofold: on the one hand synthetic samples can be used to augment existing datasets and train more effective predictors; on the other hand, it allows to generate samples with multiple ground truths, corresponding to diverse equally likely outcomes of the observed trajectory. We define a trajectory prediction model and a loss that explicitly address the multimodality of the problem and we show that combining synthetic and real data leads to prediction improvements, obtaining state of the art results.
Lorenzo Berlincioni, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
ICPR3
2020 Learning Group Activities from Skeletons without Individual Action Labels
abstract
To understand human behavior we must not just recognize individual actions but model possibly complex group activity and interactions. Hierarchical models obtain the best results in group activity recognition but require fine grained individual action annotations at the actor level. In this paper we show that using only skeletal data we can train a state-of-the art end-to-end system using only group activity labels at the sequence level. Our experiments show that models trained without individual action supervision perform poorly. On the other hand we show that pseudo-labels can be computed from any pre-trained feature extractor with comparable final performance. Finally our carefully designed lean pose only architecture shows highly competitive results versus more complex multimodal approaches even in the self-supervised variant.
Fabio Zappardino, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo
ICPR3
2020 Increasing Video Perceptual Quality with GANs and Semantic Coding
abstract
We have seen a rise in video based user communication in the last year, unfortunately fueled by the spread of COVID-19 disease. Efficient low-latency delay of transmission of video is a challenging problem which must also deal with the segmented nature of network infrastructure not always allowing a high throughput. Lossy video compression is a basic requirement to enable such technology widely. While this may compromise the quality of the streamed video there are recent deep learning based solutions to restore quality of a lossy compressed video.
Leonardo Galteri, Marco Bertini 0001, Lorenzo Seidenari, Tiberio Uricchio, Alberto Del Bimbo
ACM Multimedia3
2019 Towards Real-Time Image Enhancement GANs
Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo
CAIP (1)2
2019 Fast Video Quality Enhancement using GANs
abstract
Video compression algorithms result in a reduction of image quality, because of their lossy approach to reduce the required bandwidth. This affects commercial streaming services such as Netflix, or Amazon Prime Video, but affects also video conferencing and video surveillance systems. In all these cases it is possible to improve the video quality, both for human view and for automatic video analysis, without changing the compression pipeline, through a post-processing that eliminates the visual artifacts created by the compression algorithms. Generative Adversarial Networks have obtained extremely high quality results in image enhancement tasks; however, to obtain such results large generators are usually employed, resulting in high computational costs and processing time. In this work we present an architecture that can be used to reduce the computational cost and that has been implemented on mobile devices. A possible application is to improve video conferencing, or live streaming. In these cases there is no original uncompressed video stream available. Therefore, we report results using no-reference video quality metric showing high naturalness and quality even for efficient networks.
Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo
ACM Multimedia2
2019 Enhanced skeleton and face 3D data for person re-identification from depth cameras
Pietro Pala, Lorenzo Seidenari, Stefano Berretti, Alberto Del Bimbo
Comput. Graph.2
2019 Real-time demographic profiling from face imagery with Fisher vectors
Lorenzo Seidenari, Alessandro Rozza, Alberto Del Bimbo
Mach. Vis. Appl.1
2019 Deep Universal Generative Adversarial Compression Artifact Removal
abstract
Image compression is a need that arises in many circumstances. Unfortunately, whenever a lossy compression algorithm is used, artifacts will manifest. Image artifacts, caused by compression tend to eliminate higher frequency details and, in certain cases, may add noise or small image structures. There are two main drawbacks of this phenomenon. First, images appear much less pleasant to the human eye. Second, computer vision algorithms, such as object detectors, may be hindered and their performance reduced. Removing such artifacts means recovering the original image from a perturbed version of it. This means that one ideally should invert the compression process through a complicated nonlinear image transformation. We propose an image transformation approach based on a feedforward fully convolutional residual network model. We show that this model can be optimized either traditionally, directly optimizing an image similarity loss (SSIM), or using a generative adversarial approach (GAN). Our GAN is able to produce images with more photorealistic details than SSIM-based networks. We describe a novel training procedure based on subpatches and devise a novel testing protocol to evaluate restored images quantitatively. We show that our approach can be used as a preprocessing step for different computer vision tasks in case images are degraded by compression to a point that state-of-the art algorithms fail. In this case, our GAN-based approach obtains better performance than MSE or SSIM trained networks. Different from previously proposed approaches, we are able to remove artifacts generated at any QF by inferring the image quality directly from data.
Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo
IEEE Trans. Multim.2
2018 Video Compression for Object Detection Algorithms
abstract
Video compression algorithms have been designed aiming at pleasing human viewers, and are driven by video quality metrics that are designed to account for the capabilities of the human visual system. However, thanks to the advances in computer vision systems more and more videos are going to be watched by algorithms, e.g. implementing video surveillance systems or performing automatic video tagging. This paper describes an adaptive video coding approach for computer vision-based systems. We show how to control the quality of video compression so that automatic object detectors can still process the resulting video, improving their detection performance, by preserving the elements of the scene that are more likely to contain meaningful content. Our approach is based on computation of saliency maps exploiting a fast objectness measure. The computational efficiency of this approach makes it usable in a real-time video coding pipeline. Experiments show that our technique outperforms standard H.265 in speed and coding efficiency, and can be applied to different types of video domains, from surveillance to web videos.
Leonardo Galteri, Marco Bertini 0001, Lorenzo Seidenari, Alberto Del Bimbo
ICPR3
2017 Deep Generative Adversarial Compression Artifact Removal
abstract
Compression artifacts arise in images whenever a lossy compression algorithm is applied. These artifacts eliminate details present in the original image, or add noise and small structures; because of these effects they make images less pleasant for the human eye, and may also lead to decreased performance of computer vision algorithms such as object detectors. To eliminate such artifacts, when decompressing an image, it is required to recover the original image from a disturbed version. To this end, we present a feed-forward fully convolutional residual network model trained using a generative adversarial framework. To provide a baseline, we show that our model can be also trained optimizing the Structural Similarity (SSIM), which is a better loss with respect to the simpler Mean Squared Error (MSE). Our GAN is able to produce images with more photorealistic details than MSE or SSIM based networks. Moreover we show that our approach can be used as a pre-processing step for object detection in case images are degraded by compression to a point that state-of-the art detectors fail. In this task, our GAN method obtains better performance than MSE or SSIM trained networks.
Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo
ICCV2
2017 PACE: Prediction-based Annotation for Crowded Environments
abstract
We present a new tool we have developed to ease the annotation of crowded environments, typical of visual surveillance datasets. Our tool is developed using HTML5 and Javascript and has two back-ends. A PHP based back-end implement the persistence using a relational database and manage the dynamic creation of pages and the authentication procedure. A python based REST server implement all the computer vision facilities to assist annotators. Our tool allows collaborative annotation of person identity, group membership, location, gaze and occluded parts. PACE supports multiple cameras and if calibration is provided the geometry is used to improve computer vision based assistance. We detail the whole interface comprising an administrative view that ease the setup of the system.
Federico Bartoli, Giuseppe Lisanti, Lorenzo Seidenari, Alberto Del Bimbo
ICMR3
2017 Outdoor Object Recognition for Smart Audio Guides
abstract
We present a smart audio guide that adapts itself to the environment the user is navigating into. The system builds automatically a point of interest database exploiting Wikipedia and Google APIs as source. We rely on a computer vision system, to overcome the likely sensor limitations, and determine with high accuracy if the user is facing a certain landmark or if he is not facing any. Thanks to this the guide presents audio description at the most appropriate moment without any user intervention, using text-to-speech augmenting the experience.
Claudio Baecchi, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo
ACM Multimedia3
2017 Understanding and localizing activities from correspondences of clustered trajectories
Francesco Turchini, Lorenzo Seidenari, Alberto Del Bimbo
Comput. Vis. Image Underst.2
2017 Indexing quantized ensembles of exemplar-SVMs with rejecting taxonomies
Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
Multim. Tools Appl.2
2017 Automatic image annotation via label transfer in the semantic space
Tiberio Uricchio, Lamberto Ballan, Lorenzo Seidenari, Alberto Del Bimbo
Pattern Recognit.3
2017 Spatio-Temporal Closed-Loop Object Detection
abstract
Object detection is one of the most important tasks of computer vision. It is usually performed by evaluating a subset of the possible locations of an image, that are more likely to contain the object of interest. Exhaustive approaches have now been superseded by object proposal methods. The interplay of detectors and proposal algorithms has not been fully analyzed and exploited up to now, although this is a very relevant problem for object detection in video sequences. We propose to connect, in a closed-loop, detectors and object proposal generator functions exploiting the ordered and continuous nature of video sequences. Different from tracking we only require a previous frame to improve both proposal and detection: no prediction based on local motion is performed, thus avoiding tracking errors. We obtain three to four points of improvement in mAP and a detection time that is lower than Faster Regions with CNN features (R-CNN), which is the fastest Convolutional Neural Network (CNN) based generic object detector known at the moment.
Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo
IEEE Trans. Image Process.2
2017 Deep Artwork Detection and Retrieval for Automatic Context-Aware Audio Guides
abstract
In this article, we address the problem of creating a smart audio guide that adapts to the actions and interests of museum visitors. As an autonomous agent, our guide perceives the context and is able to interact with users in an appropriate fashion. To do so, it understands what the visitor is looking at, if the visitor is moving inside the museum hall, or if he or she is talking with a friend. The guide performs automatic recognition of artworks, and it provides configurable interface features to improve the user experience and the fruition of multimedia materials through semi-automatic interaction. Our smart audio guide is backed by a computer vision system capable of working in real time on a mobile device, coupled with audio and motion sensors. We propose the use of a compact Convolutional Neural Network (CNN) that performs object classification and localization. Using the same CNN features computed for these tasks, we perform also robust artwork recognition. To improve the recognition accuracy, we perform additional video processing using shape-based filtering, artwork tracking, and temporal filtering. The system has been deployed on an NVIDIA Jetson TK1 and a NVIDIA Shield Tablet K1 and tested in a real-world environment (Bargello Museum of Florence).
Lorenzo Seidenari, Claudio Baecchi, Tiberio Uricchio, Andrea Ferracani, Marco Bertini 0001, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.1
2016 User interest profiling using tracking-free coarse gaze estimation
abstract
Understanding where people attention focuses is a challenging and extremely valuable task that can be solved using computer vision technologies. In this paper we address this problem on surveillance-like scenarios, where head and body imagery are usually low resolution. We propose a method to profile the attention of people moving in a known space. We exploit coarse gaze estimation and a novel model based on optical flow to improve attention prediction without the need of a tracker. Removing the tracker dependency makes the method applicable also on highly crowded scenarios. The proposed method is able to obtain comparable performance with respect to state of the art solutions in terms of Mean Average Angular Error (MAAE) on the TownCentre dataset. We also test our approach on the publicly available MuseumVisitors dataset showing an improvement both in terms of MAAE and in terms of accuracy in the estimation of visitors' profile.
Federico Bartoli, Giuseppe Lisanti, Lorenzo Seidenari, Alberto Del Bimbo
ICPR3
2016 Do Textual Descriptions Help Action Recognition?
abstract
We present a novel method to improve action recognition by leveraging a set of captioned videos. By learning linear projections to map videos and text onto a common space, our approach shows that improved results on unseen videos can be obtained. We also propose a novel structure preserving loss that further ameliorates the quality of the projections. We tested our method on the challenging, realistic, Hollywood2 action recognition dataset where a considerable gain in performance is obtained. We show that the gain is proportional to the number of training samples used to learn the projections.
Matteo Bruni, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo
ACM Multimedia3
2016 Real-time Wearable Computer Vision System for Improved Museum Experience
abstract
The goal of this work is to implement a real-time computer vision system that can run on wearable devices to perform object classification and artwork recognition, to improve the experience of a museum visit through understanding the interests of users. Object classification helps to understand the context of the visit, e.g. differentiating when a visitor is talking with people, or just wandering through the museum, or if he is looking at an exhibit that interests him. Artwork recognition allows to provide automatically information of the observed item or to create a user profile based on what and how long a user has observed artworks.
Giovanni Taverriti, Stefano Lombini, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo
ACM Multimedia3
2015 WATTS: a Web Annotation Tool for Surveillance Scenarios
abstract
In this paper, we present a web based annotation tool we developed allowing creating collaboratively a detailed ground truth for datasets related to visual surveillance and behavior understanding. The system persistence is based on a relational database and the user interface is designed using HTML5, Javascript and CSS. Our tool can easily manage datasets with multiple cameras. It allows annotating a person location in the image, its identity, its body and head gaze, as well as a potential occlusion or group membership. We justify each annotation type with regards to current trends of research in the computer vision community. We further detail how our interface can be used to annotate each of these annotations type. We conclude the paper with an usability evaluation of our system.
Federico Bartoli, Lorenzo Seidenari, Giuseppe Lisanti, Svebor Karaman, Alberto Del Bimbo
ACM Multimedia2
2014 Real-time people counting from depth imagery of crowded environments
abstract
In this paper we describe a system for automatic people counting in crowded environments. The approach we propose is a counting-by-detection method based on depth imagery. It is designed to be deployed as an autonomous appliance for crowd analysis in video surveillance application scenarios. Our system performs foreground/background segmentation on depth image streams in order to coarsely segment persons, then depth information is used to localize head candidates which are then tracked in time on an automatically estimated ground plane. The system runs in real-time, at a frame-rate of about 20 fps. We collected a dataset of RGB-D sequences representing three typical and challenging surveillance scenarios, including crowds, queuing and groups. An extensive comparative evaluation is given between our system and more complex, Latent SVM-based head localization for person counting applications.
Enrico Bondi, Lorenzo Seidenari, Andrew D. Bagdanov, Alberto Del Bimbo
AVSS2
2014 Adaptive Structured Pooling for Action Recognition
Svebor Karaman, Lorenzo Seidenari, Shugao Ma, Alberto Del Bimbo, Stan Sclaroff
BMVC2
2014 Fisher Vectors over Random Density Forests for Object Recognition
abstract
In this paper we describe a Fisher vector encoding of images over Random Density Forests. Random Density Forests (RDFs) are an unsupervised variation of Random Decision Forests for density estimation. In this work we train RDFs by splitting at each node in order to minimize the Gaussian differential entropy of each split. We use this as generative model of image patch features and derive the Fisher vector representation using the RDF as the underlying model. Our approach is computationally efficient, reducing the amount of Gaussian derivatives to compute, and allows more flexibility in the feature density modelling. We evaluate our approach on the PASCAL VOC 2007 dataset showing that our approach, that only uses linear classifiers, improves over bag of visual words and is comparable to the traditional Fisher vector encoding over Gaussian Mixture Models for density estimation.
Claudio Baecchi, Francesco Turchini, Lorenzo Seidenari, Andrew D. Bagdanov, Alberto Del Bimbo
ICPR3
2014 A Cross-media Model for Automatic Image Annotation
abstract
Automatic image annotation is still an important open problem in multimedia and computer vision. The success of media sharing websites has led to the availability of large collections of images tagged with human-provided labels. Many approaches previously proposed in the literature do not accurately capture the intricate dependencies between image content and annotations. We propose a learning procedure based on Kernel Canonical Correlation Analysis which finds a mapping between visual and textual words by projecting them into a latent meaning space. The learned mapping is then used to annotate new images using advanced nearest-neighbor voting methods. We evaluate our approach on three popular datasets, and show clear improvements over several approaches relying on more standard representations.
Lamberto Ballan, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo
ICMR3
2014 Local Pyramidal Descriptors for Image Recognition
abstract
In this paper, we present a novel method to improve the flexibility of descriptor matching for image recognition by using local multiresolution pyramids in feature space. We propose that image patches be represented at multiple levels of descriptor detail and that these levels be defined in terms of local spatial pooling resolution. Preserving multiple levels of detail in local descriptors is a way of hedging one's bets on which levels will most relevant for matching during learning and recognition. We introduce the Pyramid SIFT (P-SIFT) descriptor and show that its use in four state-of-the-art image recognition pipelines improves accuracy and yields state-of-the-art results. Our technique is applicable independently of spatial pyramid matching and we show that spatial pyramids can be combined with local pyramids to obtain further improvement. We achieve state-of-the-art results on Caltech-101 (80.1%) and Caltech-256 (52.6%) when compared to other approaches based on SIFT features over intensity images. Our technique is efficient and is extremely easy to integrate into image recognition pipelines.
Lorenzo Seidenari, Giuseppe Serra 0001, Andrew D. Bagdanov, Alberto Del Bimbo
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Real-time hand status recognition from RGB-D imagery
Andrew D. Bagdanov, Alberto Del Bimbo, Lorenzo Seidenari, Lorenzo Usai
ICPR3
2012 Multi-scale and real-time non-parametric approach for anomaly detection and localization
Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari
Comput. Vis. Image Underst.3
2012 Effective Codebooks for Human Action Representation and Classification in Unconstrained Videos
abstract
Recognition and classification of human actions for annotation of unconstrained video sequences has proven to be challenging because of the variations in the environment, appearance of actors, modalities in which the same action is performed by different persons, speed and duration, and points of view from which the event is observed. This variability reflects in the difficulty of defining effective descriptors and deriving appropriate and effective codebooks for action categorization. In this paper, we propose a novel and effective solution to classify human actions in unconstrained videos. It improves on previous contributions through the definition of a novel local descriptor that uses image gradient and optic flow to respectively model the appearance and motion of human actions at interest point regions. In the formation of the codebook, we employ radius-based clustering with soft assignment in order to create a rich vocabulary that may account for the high variability of human actions. We show that our solution scores very good performance with no need of parameter tuning. We also show that a strong reduction of computation time can be obtained by applying codebook size reduction with Deep Belief Networks with little loss of accuracy.
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari, Giuseppe Serra 0001
IEEE Trans. Multim.4
2011 Adaptive Video Compression for Video Surveillance Applications
abstract
This article describes an approach to adaptive video coding for video surveillance applications. Using a combination of low-level features with low computational cost, we show how it is possible to control the quality of video compression so that semantically meaningful elements of the scene are encoded with higher fidelity, while background elements are allocated fewer bits in the transmitted representation. Our approach is based on adaptive smoothing of individual video frames so that image features highly correlated to semantically interesting objects are preserved. Using only low-level image features on individual frames, this adaptive smoothing can be seamlessly inserted into a video coding pipeline as a pre-processing state. Experiments show that our technique is efficient, outperforms standard H.264 encoding at comparable bit rates, and preserves features critical for downstream detection and recognition.
Andrew D. Bagdanov, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari
ISM4
2011 Event detection and recognition for semantic annotation of video
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari, Giuseppe Serra 0001
Multim. Tools Appl.4
2010 Non-parametric anomaly detection exploiting space-time features
abstract
In this paper a real-time anomaly detection system for video streams is proposed. Spatio-temporal features are exploited to capture scene dynamic statistics together with appearance. Anomaly detection is performed in a non-parametric fashion, evaluating directly local descriptor statistics. A method to update scene statistics, to cope with scene changes that typically happen in real world settings, is also provided. The proposed method is tested on publicly available datasets.
Lorenzo Seidenari, Marco Bertini 0001
ACM Multimedia1
2009 Recognizing human actions by fusing spatio-temporal appearance and motion descriptors
abstract
In this paper we propose a new method for human action categorization by using an effective combination of a new 3D gradient descriptor with an optic flow descriptor, to represent spatio-temporal interest points. These points are used to represent video sequences using a bag of spatio-temporal visual words, following the successful results achieved in object and scene classification. We extensively test our approach on the standard KTH and Weizmann actions datasets, showing its validity and good performance. Experimental results outperform state-of-the-art methods, without requiring fine parameter tuning.
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari, Giuseppe Serra 0001
ICIP4