Alberto Del Bimbo

dblp:b/AlbertoDelBimbo · DBLP profile ↗
← Back
373ranked-venue papers
40as first author
74since 2021 · last 2026
0000-0002-1052-8322ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 268 · 28 first-author · 48 since 2021Artificial intelligence and machine learning · 147 · 18 first-author · 35 since 2021Computer networks · 23 · 3 first-author · 16 since 2021Databases, data management, data science and information retrieval · 18 · 2 first-author · 3 since 2021Security and privacy · 6 · 1 since 2021Systems, architecture and hardware · 3Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Mitigating Negative Flips via Margin Preserving Training
abstract
Minimizing inconsistencies across successive versions of an AI system is as crucial as reducing the overall error. In image classification, such inconsistencies manifest as negative flips, where an updated model misclassifies test samples that were previously classified correctly. This issue becomes increasingly pronounced as the number of training classes grows over time, since adding new categories reduces the margin of each class and may introduce conflicting patterns that undermine their learning process, thereby degrading performance on the original subset. To mitigate negative flips, we propose a novel approach that preserves the margins of the original model while learning an improved one. Our method encourages a larger relative margin between the previously learned and newly introduced classes by introducing an explicit margin-calibration term on the logits. However, overly constraining the logit margin for the new classes can significantly degrade their accuracy compared to a new independently trained model. To address this, we integrate a double-source focal distillation loss with the previous model and a new independently trained model, learning an appropriate decision margin from both old and new data, even under a logit margin calibration. Extensive experiments on image classification benchmarks demonstrate that our approach consistently reduces the negative flip rate with high overall accuracy.
Simone Ricci, Niccolò Biondi, Federico Pernici, Alberto Del Bimbo
AAAI4
2026 Spatio-temporal transformers for action unit classification with event cameras
Luca Cultrera, Federico Becattini, Lorenzo Berlincioni, Claudio Ferrari, Alberto Del Bimbo
Comput. Vis. Image Underst.5
2026 A video dataset for Multi-Emotion and Interpersonal relation analysis
Hajer Guerdelli, Claudio Ferrari, Stefano Berretti, Walid Barhoumi, Alberto Del Bimbo
Comput. Vis. Image Underst.5
2026 Revealing GAN-generated faces through local camera surface frame analysis
abstract
The ability of AI to generate highly realistic, fully synthetic images, particularly of human faces, is rapidly advancing, making it increasingly difficult to distinguish between real and artificially generated content. This growing realism highlights the urgent need for reliable methods to detect subtle inconsistencies introduced during the image generation process. A fundamental distinction between authentic and deepfake content lies in the absence, for the latter, of an acquisition process by a real camera. As a result, the intricate relationships among scene elements, such as lighting, reflectance, and spatial positioning, are not captured from the physical world but are artificially reconstructed. Motivated by this observation, we propose the use of local camera surface frames as a feature to encode such environment-specific attributes. Our experimental results demonstrate that this representation not only achieves high detection accuracy but also exhibits strong and robust generalisation capabilities across different GAN-based generative models.
Andrea Ciamarra, Roberto Caldelli, Alberto Del Bimbo
J. Inf. Secur.3
2026 Prediction of wine quality by monitoring climatic data and weather anomalies
Federico Becattini, Andrea Ferracani, Giuseppe Becchi, Alberto Del Bimbo
Multim. Tools Appl.4
2025 Learning Compatible Representations
Alberto Del Bimbo, Niccolò Biondi, Simone Ricci, Federico Pernici
ICPRAM1
2025 FRED: The Florence RGB-Event Drone Dataset
abstract
Small, fast, and lightweight drones present significant challenges for traditional RGB cameras due to their limitations in capturing fast-moving objects, especially under challenging lighting conditions. Event cameras offer an ideal solution, providing high temporal definition and dynamic range, yet existing benchmarks often lack fine temporal resolution or drone-specific motion patterns, hindering progress in these areas. This paper introduces the Florence RGB-Event Drone dataset (FRED), a novel multimodal dataset specifically designed for drone detection, tracking, and trajectory forecasting, combining RGB video and event streams. FRED features more than 7 hours of densely annotated drone trajectories, using 5 different drone models and including challenging scenarios such as rain and adverse lighting conditions. We provide detailed evaluation protocols and standard metrics for each task, facilitating reproducible benchmarking. The authors hope FRED will advance research in high-speed drone perception and multimodal spatiotemporal understanding.
Gabriele Magrini, Niccolò Marini, Federico Becattini, Lorenzo Berlincioni, Niccolò Biondi, Pietro Pala, Alberto Del Bimbo
ACM Multimedia7
2025 λ-Orthogonality Regularization for Compatible Representation Learning
Simone Ricci, Niccolò Biondi, Federico Pernici, Ioannis Patras, Alberto Del Bimbo
NeurIPS5
2025 Navigating social contexts: A transformer approach to relationship recognition
Lorenzo Berlincioni, Luca Cultrera, Marco Bertini 0001, Alberto Del Bimbo
Comput. Vis. Image Underst.4
2025 3D Pose Nowcasting: Forecast the future to improve the present
abstract
Technologies to enable safe and effective collaboration and coexistence between humans and robots have gained significant importance in the last few years. A critical component useful for realizing this collaborative paradigm is the understanding of human and robot 3D poses using non-invasive systems. Therefore, in this paper, we propose a novel vision-based system leveraging depth data to accurately establish the 3D locations of skeleton joints. Specifically, we introduce the concept of Pose Nowcasting, denoting the capability of the proposed system to enhance its current pose estimation accuracy by jointly learning to forecast future poses. The experimental evaluation is conducted on two different datasets, providing accurate and real-time performance and confirming the validity of the proposed method on both the robotic and human scenarios. • We introduce the novel task of 3D Pose Nowcasting. • Our Pose Nowcasting system is based on both 3D Pose Estimation and Forecasting. • We show that knowledge about pose forecasting improves the accuracy of pose estimation. • We apply the proposed system both to human and robots. • Result on different dataset show state-of-the-art performance and robustness.
Alessandro Simoni, Francesco Marchetti, Guido Borghi, Federico Becattini, Lorenzo Seidenari, Roberto Vezzani, Alberto Del Bimbo
Comput. Vis. Image Underst.7
2025 A fine-tuning approach based on spatio-temporal features for few-shot video object detection
abstract
This paper describes a new Fine-Tuning approach for Few-Shot object detection in Videos that exploits spatio-temporal information to boost detection precision. Despite the progress made in the single image domain in recent years, the few-shot video object detection problem remains almost unexplored. A few-shot detector must quickly adapt to a new domain with a limited number of annotations per category. Therefore, it is not possible to include videos in the training set, hindering the spatio-temporal learning process. We propose augmenting each training image with synthetic frames to train the spatio-temporal module of our method. This module employs attention mechanisms to mine relationships between proposals across frames, effectively leveraging spatio-temporal information. A spatio-temporal double head then localizes objects in the current frame while classifying them using both context from nearby frames and information from the current frame. Finally, the predicted scores are fed into a long-term object-linking method that generates object tubes across the video. By optimizing the classification score based on these tubes, our approach ensures spatio-temporal consistency. Classification is the primary challenge in few-shot object detection. Our results show that spatio-temporal information helps to mitigate this issue, paving the way for future research in this direction. FTFSVid achieves 41.9 AP50 on the Few-Shot Video Object Detection (FSVOD-500) and 42.9 AP50 on the Few-Shot YouTube Video (FSYTV-40) dataset, surpassing our spatial baseline by 4.3 and 2.5 points. Additionally, FTFSVid outperforms previous few-shot video object detectors by 3.2 points on FSVOD-500 and 14.5 points on FSYTV-40, setting a new state-of-the-art.
Daniel Cores, Lorenzo Seidenari, Alberto Del Bimbo, Víctor M. Brea 0001, Manuel Mucientes
Eng. Appl. Artif. Intell.3
2025 iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval
abstract
Given a query consisting of a reference image and a relative caption, Composed Image Retrieval (CIR) aims to retrieve target images visually similar to the reference one while incorporating the changes specified in the relative caption. The reliance of supervised methods on labor-intensive manually labeled datasets hinders their broad applicability to CIR. In this work, we introduce a new task, Zero-Shot CIR (ZS-CIR), that addresses CIR without the need for a labeled training dataset. We propose an approach, named iSEARLE (improved zero-Shot composEd imAge Retrieval with textuaL invErsion), that involves mapping the visual information of the reference image into a pseudo-word token in the CLIP token embedding space and combining it with the relative caption. To foster research on ZS-CIR, we present an open-domain benchmarking dataset named CIRCO (Composed Image Retrieval on Common Objects in context), the first CIR dataset where each query is labeled with multiple ground truths and a semantic categorization. The experimental results illustrate that iSEARLE obtains state-of-the-art performance on three different CIR datasets - FashionIQ, CIRR, and the proposed CIRCO - and two additional evaluation settings, namely domain conversion and object composition.
Lorenzo Agnolucci, Alberto Baldrati, Alberto Del Bimbo, Marco Bertini 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Neuromorphic face analysis: A survey
abstract
Neuromorphic sensors, also known as event cameras, are a class of imaging devices mimicking the function of biological visual systems. Unlike traditional frame-based cameras, which capture fixed images at discrete intervals, neuromorphic sensors continuously generate events that represent changes in light intensity or motion in the visual field with high temporal resolution and low latency. These properties have proven to be interesting in modeling human faces, both from an effectiveness and a privacy-preserving point of view. Neuromorphic face analysis however is still a raw and unstructured field of research, with several attempts at addressing different tasks with no clear standard or benchmark. This survey paper presents a comprehensive overview of capabilities, challenges and emerging applications in the domain of neuromorphic face analysis, to outline promising directions and open issues. After discussing the fundamental working principles of neuromorphic vision and presenting an in-depth overview of the related research, we explore the current state of available data, data representations, emerging challenges, and limitations that require further investigation. This paper aims to highlight the recent progress in this evolving field to provide researchers an all-encompassing analysis of the state of the art along with its problems and shortcomings. • Neuromorphic sensors are becoming more common in human-face related vision tasks. • As a new niche field it lacks overview papers helping researchers orient themselves. • We compare the relevant literature to classical methods and highlights pros and cons. • Several tasks and recently published open dataset are presented and compared. • Future research directions are provided after analyzing the relevant shortcomings.
Federico Becattini, Lorenzo Berlincioni, Luca Cultrera, Alberto Del Bimbo
Pattern Recognit. Lett.4
2025 Spike-TBR: A noise resilient neuromorphic event representation
abstract
Event cameras offer significant advantages over traditional frame-based sensors, including higher temporal resolution, lower latency and dynamic range. However, efficiently converting event streams into formats compatible with standard computer vision pipelines remains a challenging problem, particularly in the presence of noise. In this paper, we propose Spike-TBR, a novel event-based encoding strategy based on Temporal Binary Representation (TBR), addressing its vulnerability to noise by integrating spiking neurons. Spike-TBR combines the frame-based advantages of TBR with the noise-filtering capabilities of spiking neural networks, creating a more robust representation of event streams. We evaluate four variants of Spike-TBR, each using different spiking neurons, across multiple datasets, demonstrating superior performance in noise-affected scenarios while improving the results on clean data. Our method bridges the gap between spike-based and frame-based processing, offering a simple noise-resilient solution for event-driven vision applications. • We present Spike-TBR an enhanced version of Temporal Binary Representation. • By using spiking neurons we make the frame-based representation resilient to noise. • We obtain state-of-the-art results on four different neuromorphic datasets.
Gabriele Magrini, Federico Becattini, Luca Cultrera, Lorenzo Berlincioni, Pietro Pala, Alberto Del Bimbo
Pattern Recognit. Lett.6
2025 Parents and Children: Distinguishing Multimodal Deepfakes from Natural Images
abstract
Recent advancements in diffusion models have enabled the generation of realistic deepfakes from textual prompts in natural language. While these models have numerous benefits across various sectors, they have also raised concerns about the potential misuse of fake images and cast new pressures on fake image detection. In this work, we pioneer a systematic study on deepfake detection generated by state-of-the-art diffusion models. Firstly, we conduct a comprehensive analysis of the performance of contrastive and classification-based visual features, respectively, extracted from CLIP-based models and ResNet or Vision Transformer (ViT)-based architectures trained on image classification datasets. Our results demonstrate that fake images share common low-level cues, which render them easily recognizable. Further, we devise a multimodal setting wherein fake images are synthesized by different textual captions, which are used as seeds for a generator. Under this setting, we quantify the performance of fake detection strategies and introduce a contrastive-based disentangling method that lets us analyze the role of the semantics of textual descriptions and low-level perceptual cues. Finally, we release a new dataset, called COCOFake, containing about 1.2 million images generated from the original COCO image–caption pairs using two recent text-to-image diffusion models, namely Stable Diffusion v1.4 and v2.0.
Roberto Amoroso, Davide Morelli, Marcella Cornia, Lorenzo Baraldi 0001, Alberto Del Bimbo, Rita Cucchiara
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Interactive Garment Recommendation with User in the Loop
abstract
Recommending fashion items often leverages rich user profiles and makes targeted suggestions based on past history and previous purchases. In this paper, we work under the assumption that no prior knowledge is given about a user. We propose to build a user profile on the fly by integrating user reactions as we recommend complementary items to compose an outfit. We present a reinforcement learning agent capable of suggesting appropriate garments and ingesting user feedback so to improve its recommendations and maximize user satisfaction. To train such a model, we resort to a proxy model to be able to simulate having user feedback in the training loop. We experiment on the IQON3000 fashion dataset and we find that a reinforcement learning-based agent becomes capable of improving its recommendations by taking into account personal preferences. Furthermore, such task demonstrated to be hard for non-reinforcement models, that cannot exploit exploration during training.
Federico Becattini, Xiaolin Chen 0001, Andrea Puccia, Haokun Wen, Xuemeng Song, Liqiang Nie, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.7
2025 Introduction to the Special Issue on Realistic Synthetic Data: Generation, Learning, Evaluation
abstract
This is the foreword of our special issue volume on Realistic Synthetic Data: Generation, Learning, Evaluation organized with the ACM Transactions on Multimedia Computing, Communications, and Applications. It presents the target of the special issue that relates to synthetic data for various modalities, e.g., signals, images, volumes, audio, etc., controllable generation for learning from synthetic data, transfer learning and generalization of models, causality in data generation, addressing bias, limitations and trustworthiness in data generation, evaluation measures/protocols and benchmarks to assess quality of synthetic content, open synthetic datasets and software tools, and ethical aspects of synthetic data. The call for papers received a record number of 40 submissions out of which 15 were finally accepted for publication. This introduction provides an overview of the topics of each of the articles.
Bogdan Ionescu, Ioannis Patras, Henning Müller, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Stationary Representations: Optimally Approximating Compatibility and Implications for Improved Model Replacements
abstract
Learning compatible representations enables the interchangeable use of semantic features as models are updated over time. This is particularly relevant in search and retrieval systems where it is crucial to avoid reprocessing of the gallery images with the updated model. While recent research has shown promising empirical evidence, there is still a lack of comprehensive theoretical understanding about learning compatible representations. In this paper, we demonstrate that the stationary representations learned by the d-Simplex fixed classifier optimally approximate compatibility representation according to the two inequality constraints of its formal definition. This not only establishes a solid foundation for future works in this line of research but also presents implications that can be exploited in practical learning scenarios. An exemplary application is the nowstandard practice of downloading and fine-tuning new pretrained models. Specifically, we show the strengths and critical issues of stationary representations in the case in which a model undergoing sequential fine-tuning is asynchronously replaced by downloading a better-performing model pretrained elsewhere. Such a representation enables seamless delivery of retrieval service (i.e., no reprocessing of gallery images) and offers improved performance without operational disruptions during model replacement. Code available at: https://github.com/miccunifi/iamcl2r.
Niccolò Biondi, Federico Pernici, Simone Ricci, Alberto Del Bimbo
CVPR4
2024 IMEmo: An Interpersonal Relation Multi-Emotion Dataset
abstract
While engaged in a face-to-face conversation, being capable of understanding the attitude, emotion, and intention of another person allows one guiding his/her behavior establishing a comfortable communication either verbal and non-verbal (i.e., body and face language). This paper introduces the “The IMEmo Interpersonal Multi-Emotion video dataset”, a new in-the-wild dataset of face-to-face interaction, built from movies of romance and drama categories. We manually collected over 100 clips from different movies in different languages. The dataset consists of 79.3 minutes scenes, with a duration of each clip ranging between 0.20 and 2.13 minutes. Each clip contains two people communicating with each other both verbally and with expression and body pose and gestures. Currently, it includes age, gender, emotions, social relationships, actions and valence/arousal annotations for both individuals. Emotion recognition results using a baseline CNN approach are also reported to provide an estimation of the difficulty of the data also in comparison to existing benchmarks.
Hajer Guerdelli, Claudio Ferrari, Stefano Berretti, Alberto Del Bimbo
FG4
2024 Learning Backward Compatible Representations
abstract
In today's multimedia-rich environment, the rapid growth of data poses significant challenges for developing efficient multi-modal retrieval systems essential for retrieving text, images, audio, and video. As data expands, newer, scalable, and high-performance retrieval systems are increasingly necessary. Embedding-based deep neural networks (DNNs) have become key solutions, transforming high-dimensional data into lower-dimensional embeddings for easy comparison and retrieval. However, updating DNNs changes the internal feature representations, necessitating the extraction of new feature vectors for all gallery data, which is costly, especially with gallery sets comprising billions of data. Learning backward-compatible representations addresses this by allowing new representation to be matched with old gallery data without recalculating features. This tutorial aims to equip participants with the knowledge and tools to apply backward-compatible representations, enhancing multimedia retrieval systems' efficiency and scalability. Participants will learn the importance of compatible representations, basic methods and techniques, and explore challenging open questions that are becoming increasingly relevant to multimedia and cross-modal retrieval.
Niccolò Biondi, Simone Ricci, Federico Pernici, Alberto Del Bimbo
ACM Multimedia4
2024 ARNIQA: Learning Distortion Manifold for Image Quality Assessment
abstract
No-Reference Image Quality Assessment (NR-IQA) aims to develop methods to measure image quality in alignment with human perception without the need for a high-quality reference image. In this work, we propose a self-supervised approach named ARNIQA (leArning distoRtion maNifold for Image Quality Assessment) for modeling the image distortion manifold to obtain quality representations in an intrinsic manner. First, we introduce an image degradation model that randomly composes ordered sequences of consecutively applied distortions. In this way, we can synthetically degrade images with a large variety of degradation patterns. Second, we propose to train our model by maximizing the similarity between the representations of patches of different images distorted equally, despite varying content. Therefore, images degraded in the same manner correspond to neighboring positions within the distortion manifold. Finally, we map the image representations to the quality scores with a simple linear regressor, thus without fine-tuning the encoder weights. The experiments show that our approach achieves state-of-the-art performance on several datasets. In addition, ARNIQA demonstrates improved data efficiency, generalization capabilities, and robustness compared to competing methods. The code and the model are publicly available at https://github.com/miccunifi/ARNIQA.
Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini 0001, Alberto Del Bimbo
WACV4
2024 Reference-based Restoration of Digitized Analog Videotapes
abstract
Analog magnetic tapes have been the main video data storage device for several decades. Videos stored on analog videotapes exhibit unique degradation patterns caused by tape aging and reader device malfunctioning that are different from those observed in film and digital video restoration tasks. In this work, we present a reference-based approach for the resToration of digitized Analog videotaPEs (TAPE). We leverage CLIP for zero-shot artifact detection to identify the cleanest frames of each video through textual prompts describing different artifacts. Then, we select the clean frames most similar to the input ones and employ them as references. We design a transformer-based Swin-UNet network that exploits both neighboring and reference frames via our Multi-Reference Spatial Feature Fusion (MRSFF) blocks. MRSFF blocks rely on cross-attention and attention pooling to take advantage of the most useful parts of each reference frame. To address the absence of ground truth in real-world videos, we create a synthetic dataset of videos exhibiting artifacts that closely resemble those commonly found in analog videotapes. Both quantitative and qualitative experiments show the effectiveness of our approach compared to other state-of-the-art methods. The code, the model, and the synthetic dataset are publicly available at https://github.com/miccunifi/TAPE.
Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini 0001, Alberto Del Bimbo
WACV4
2024 Gaze analysis: A survey on its applications
abstract
The examination of ocular movements has a wide range of applications due to the current developments in sensors that are now able to collect this biometric. This type of investigation is known as “gaze analysis”. The gaze has successfully examined a subject's physical and mental status in the past. As a result, over the last few decades, a large and diverse amount of literature on this subject has been generated and presented. The aim of this study is to collect and debate current gaze analysis methods based on their application field. Due to the context-specific needs for performance and efficiency, the eye movements under research are frequently evaluated from completely distinct perspectives. As a result, a collection of data, methods, and discussions ranging from the medical community to virtual and augmented reality, as well as human computer interface and remote learning, has been produced. In addition to providing a peek of novel observation on the issue of gaze analysis, the gaps between and within areas are also discussed to provide points for researchers to pursue.
Carmen Bisogni, Michele Nappi, Genny Tortora, Alberto Del Bimbo
Image Vis. Comput.4
2024 SMEMO: Social Memory for Trajectory Forecasting
abstract
Effective modeling of human interactions is of utmost importance when forecasting behaviors such as future trajectories. Each individual, with its motion, influences surrounding agents since everyone obeys to social non-written rules such as collision avoidance or group following. In this paper we model such interactions, which constantly evolve through time, by looking at the problem from an algorithmic point of view, i.e., as a data manipulation task. We present a neural network based on an end-to-end trainable working memory, which acts as an external storage where information about each agent can be continuously written, updated and recalled. We show that our method is capable of learning explainable cause-effect relationships between motions of different agents, obtaining state-of-the-art results on multiple trajectory forecasting datasets.
Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 FLODCAST: Flow and depth forecasting via multimodal recurrent architectures
abstract
Forecasting motion and spatial positions of objects is of fundamental importance, especially in safety-critical settings such as autonomous driving. In this work, we address the issue by forecasting two different modalities that carry complementary information, namely optical flow and depth. To this end we propose FLODCAST a flow and depth forecasting model that leverages a multitask recurrent architecture, trained to jointly forecast both modalities at once. We stress the importance of training using flows and depth maps together, demonstrating that both tasks improve when the model is informed of the other modality. We train the proposed model to also perform predictions for several timesteps in the future. This provides better supervision and leads to more precise predictions, retaining the capability of the model to yield outputs autoregressively for any future time horizon. We test our model on the challenging Cityscapes dataset, obtaining state of the art results for both flow and depth forecasting. Thanks to the high quality of the generated flows, we also report benefits on the downstream task of segmentation forecasting, injecting our predictions in a flow-based mask-warping framework.
Andrea Ciamarra, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
Pattern Recognit.4
2024 Generating Multiple 4D Expression Transitions by Learning Face Landmark Trajectories
abstract
In this work, we address the problem of 4D facial expressions generation. This is usually addressed by animating a neutral 3D face to reach an expression peak, and then get back to the neutral state. In the real world though, people show more complex expressions, and switch from one expression to another. We thus propose a new model that generates transitions between different expressions, and synthesizes long and composed 4D expressions. This involves three sub-problems: (i) modeling the temporal dynamics of expressions, (ii) learning transitions between them, and (iii) deforming a generic mesh. We propose to encode the temporal evolution of expressions using the motion of a set of 3D landmarks, that we learn to generate by training a manifold-valued GAN (Motion3DGAN). To allow the generation of composed expressions, this model accepts two labels encoding the starting and the ending expressions. The final sequence of meshes is generated by a Sparse2Dense mesh Decoder (S2D-Dec) that maps the landmark displacements to a dense, per-vertex displacement of a known mesh topology. By explicitly working with motion trajectories, the model is totally independent from the identity. Extensive experiments on five public datasets show that our proposed approach brings significant improvements with respect to previous solutions, while retaining good generalization to unseen data.
Naima Otberdout, Claudio Ferrari, Mohamed Daoudi, Stefano Berretti, Alberto Del Bimbo
IEEE Trans. Affect. Comput.5
2024 Perceptual Quality Improvement in Videoconferencing Using Keyframes-Based GAN
abstract
In the latest years, videoconferencing has taken a fundamental role in interpersonal relations, both for personal and business purposes. Lossy video compression algorithms are the enabling technology for videoconferencing, as they reduce the bandwidth required for real-time video streaming. However, lossy video compression decreases the perceived visual quality. Thus, many techniques for reducing compression artifacts and improving video visual quality have been proposed in recent years. In this work, we propose a novel GAN-based method for compression artifacts reduction in videoconferencing. Given that, in this context, the speaker is typically in front of the camera and remains the same for the entire duration of the transmission, we can maintain a set of reference keyframes of the person from the higher-quality I-frames that are transmitted within the video stream and exploit them to guide the visual quality improvement; a novel aspect of this approach is the update policy that maintains and updates a compact and effective set of reference keyframes. First, we extract multi-scale features from the compressed and reference frames. Then, our architecture combines these features in a progressive manner according to facial landmarks. This allows the restoration of the high-frequency details lost after the video compression. Experiments show that the proposed approach improves visual quality and generates photo-realistic results even with high compression rates. Code and pre-trained networks are publicly available at https://github.com/LorenzoAgnolucci/Keyframes-GAN.
Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini 0001, Alberto Del Bimbo
IEEE Trans. Multim.4
2024 Composed Image Retrieval using Contrastive Learning and Task-oriented CLIP-based Features
abstract
Given a query composed of a reference image and a relative caption, the Composed Image Retrieval goal is to retrieve images visually similar to the reference one that integrates the modifications expressed by the caption. Given that recent research has demonstrated the efficacy of large-scale vision and language pre-trained (VLP) models in various tasks, we rely on features from the OpenAI CLIP model to tackle the considered task. We initially perform a task-oriented fine-tuning of both CLIP encoders using the element-wise sum of visual and textual features. Then, in the second stage, we train a Combiner network that learns to combine the image-text features integrating the bimodal information and providing combined features used to perform the retrieval. We use contrastive learning in both stages of training. Starting from the bare CLIP features as a baseline, experimental results show that the task-oriented fine-tuning and the carefully crafted Combiner network are highly effective and outperform more complex state-of-the-art approaches on FashionIQ and CIRR, two popular and challenging datasets for composed image retrieval. Code and pre-trained models are available at https://github.com/ABaldrati/CLIP4Cir .
Alberto Baldrati, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Deep Variational Learning for 360° Adaptive Streaming
abstract
Prediction of head movements in immersive media is key to designing efficient streaming systems able to focus the bandwidth budget on visible areas of the content. However, most of the numerous proposals made to predict user head motion in 360° images and videos do not explicitly consider a prominent characteristic of the head motion data: its intrinsic uncertainty. In this article, we present an approach to generate multiple plausible futures of head motion in 360° videos, given a common past trajectory. To our knowledge, this is the first work that considers the problem of multiple head motion prediction for 360° video streaming. We introduce our discrete variational multiple sequence (DVMS) learning framework, which builds on deep latent variable models. We design a training procedure to obtain a flexible, lightweight stochastic prediction model compatible with sequence-to-sequence neural architectures. Experimental results on four different datasets show that DVMS outperforms competitors adapted from the self-driving domain by up to 41% on prediction horizons up to 5 s, at lower computational and memory costs. To understand how the learned features account for the motion uncertainty, we analyze the structure of the learned latent space and connect it with the physical properties of the trajectories. We also introduce a method to estimate the likelihood of each generated trajectory, enabling the integration of DVMS in a streaming system. We hence deploy an extensive evaluation of the interest of our DVMS proposal for a streaming system. To do so, we first introduce a new Python-based 360° streaming simulator that we make available to the community. On real-world user, video, and networking data, we show that predicting multiple trajectories yields higher fairness between the traces, the gains for 20–30% of the users reaching up to 10% in visual quality for the best number K of trajectories to generate.
Quentin Guimard, Lucile Sassatelli, Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Introduction to Special Issue on "Recent Trends in Multimedia Forensics"
abstract
Multimedia forensics is a subject area which is the need of the hour in this modern era of media-manipulation and generation of fake images/videos assisted with artificial intelligence (AI) models. With the ubiquitous expansion of internet enabled devices, there is a humungous amount of data available to the perusal of forensic experts. This data comprises of audio, video, images, text or a mix of those. Hence multimedia forensics, which involves a set of scientific techniques to collect, scrutinize and analyze this digital content, becomes highly imperative. The increasing threat of compelling media manipulations through machine learning-based technologies is making the situation more alarming. The most common instances are generative adversarial networks (GANs) (to generate artificial yet realistic images/videos) and DeepFake algorithms (to swap faces and expressions in videos). Furthermore, the ease of getting these manipulations done has lowered the skill required from the attacker’s end, which has intensified the problem manifold. This special issue captures a few recent outstanding works beyond trivial research results in order to push the border of the state-of-the-art and record the developments on this subject of research.
Ritesh Vyas, Michele Nappi, Alberto Del Bimbo, Sambit Bakshi
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Downsampling GAN for Small Object Data Augmentation
Daniel Cores, Víctor M. Brea 0001, Manuel Mucientes, Lorenzo Seidenari, Alberto Del Bimbo
CAIP (1)5
2023 The Florence 4D Facial Expression Dataset
abstract
Human facial expressions change dynamically, so their recognition / analysis should be conducted by accounting for the temporal evolution of face deformations either in 2D or 3D. While abundant 2D video data do exist, this is not the case in 3D, where few 3D dynamic (4D) datasets were released for public use. The negative consequence of this scarcity of data is amplified by current deep learning based-methods for facial expression analysis that require large quantities of variegate samples to be effectively trained. With the aim of smoothing such limitations, in this paper we propose a large dataset, named Florence 4D, composed of dynamic sequences of 3D face models, where a combination of synthetic and real identities exhibit an unprecedented variety of 4D facial expressions, with variations that include the classical neutral-apex transition, but generalize to expression-to-expression. All these characteristics are not exposed by any of the existing 4D datasets and they cannot even be obtained by combining more than one dataset. We strongly believe that making such a data corpora publicly available to the community will allow designing and experimenting new applications that were not possible to investigate till now. To show at some extent the difficulty of our data in terms of different identities and varying expressions, we also report a baseline experimentation on the proposed dataset that can be used as baseline.
Filippo Principi, Stefano Berretti, Claudio Ferrari, Naima Otberdout, Mohamed Daoudi, Alberto Del Bimbo
FG6
2023 Zero-Shot Composed Image Retrieval with Textual Inversion
abstract
Composed Image Retrieval (CIR) aims to retrieve a target image based on a query composed of a reference image and a relative caption that describes the difference between the two images. The high effort and cost required for labeling datasets for CIR hamper the widespread usage of existing methods, as they rely on supervised learning. In this work, we propose a new task, Zero-Shot CIR (ZS-CIR), that aims to address CIR without requiring a labeled training dataset. Our approach, named zero-Shot composEd imAge Retrieval with textuaL invErsion (SEARLE), maps the visual features of the reference image into a pseudo-word token in CLIP token embedding space and integrates it with the relative caption. To support research on ZS-CIR, we introduce an open-domain benchmarking dataset named Composed Image Retrieval on Common Objects in context (CIRCO), which is the first dataset for CIR containing multiple ground truths for each query. The experiments show that SEARLE exhibits better performance than the baselines on the two main datasets for CIR tasks, FashionIQ and CIRR, and on the proposed CIRCO. The dataset, the code and the model are publicly available at https://github.com/miccunifi/SEARLE.
Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini 0001, Alberto Del Bimbo
ICCV4
2023 Zero-Shot Image Retrieval with Human Feedback
abstract
Composed image retrieval extends traditional content-based image retrieval (CBIR) combining a query image with additional descriptive text to express user intent and specify supplementary requests related to the visual attributes of the query image. This approach holds significant potential for e-commerce applications, such as interactive multimodal searches and chatbots. In our demo, we present an interactive composed image retrieval system based on the SEARLE approach, which tackles this task in a zero-shot manner efficiently and effectively. The demo allows users to perform image retrieval iteratively refining the results using textual feedback.
Lorenzo Agnolucci, Alberto Baldrati, Marco Bertini 0001, Alberto Del Bimbo
ACM Multimedia4
2023 Fashion recommendation based on style and social events
Federico Becattini, Lavinia De Divitiis, Claudio Baecchi, Alberto Del Bimbo
Multim. Tools Appl.4
2023 CoReS: Compatible Representations via Stationarity
abstract
Compatible features enable the direct comparison of old and new learned features allowing to use them interchangeably over time. In visual search systems, this eliminates the need to extract new features from the gallery-set when the representation model is upgraded with novel data. This has a big value in real applications as re-indexing the gallery-set can be computationally expensive when the gallery-set is large, or even infeasible due to privacy or other concerns of the application. In this paper, we propose CoReS, a new training procedure to learn representations that are compatible with those previously learned, grounding on the stationarity of the features as provided by fixed classifiers based on polytopes. With this solution, classes are maximally separated in the representation space and maintain their spatial configuration stationary as new classes are added, so that there is no need to learn any mappings between representations nor to impose pairwise training with the previously learned model. We demonstrate that our training procedure largely outperforms the current state of the art and is particularly effective in the case of multiple upgrades of the training-set, which is the typical case in real applications.
Niccolò Biondi, Federico Pernici, Matteo Bruni, Alberto Del Bimbo
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Multiple Trajectory Prediction of Moving Agents With Memory Augmented Networks
abstract
Pedestrians and drivers are expected to safely navigate complex urban environments along with several non cooperating agents. Autonomous vehicles will soon replicate this capability. Each agent acquires a representation of the world from an egocentric perspective and must make decisions ensuring safety for itself and others. This requires to predict motion patterns of observed agents for a far enough future. In this paper we propose MANTRA, a model that exploits memory augmented networks to effectively predict multiple trajectories of other agents, observed from an egocentric perspective. Our model stores observations in memory and uses trained controllers to write meaningful pattern encodings and read trajectories that are most likely to occur in future. We show that our method is able to natively perform multi-modal trajectory prediction obtaining state-of-the art results on four datasets. Moreover, thanks to the non-parametric nature of the memory module, we show how once trained our system can continuously improve by ingesting novel patterns.
Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 A full data augmentation pipeline for small object detection based on generative adversarial networks
abstract
Object detection accuracy on small objects, i.e., objects under 32 × 32 pixels, lags behind that of large ones. To address this issue, innovative architectures have been designed and new datasets have been released. Still, the number of small objects in many datasets does not suffice for training. The advent of the generative adversarial networks (GANs) opens up a new data augmentation possibility for training architectures without the costly task of annotating huge datasets for small objects. In this paper, we propose a full pipeline for data augmentation for small object detection which combines a GAN-based object generator with techniques of object segmentation, image inpainting, and image blending to achieve high-quality synthetic data. The main component of our pipeline is DS-GAN, a novel GAN-based architecture that generates realistic small objects from larger ones. Experimental results show that our overall data augmentation method improves the performance of state-of-the-art models up to 11.9% [email protected] on UAVDT and by 4.7% [email protected] on iSAID, both for the small objects subset and for a scenario where the number of training instances is limited.
Brais Bosquet, Daniel Cores, Lorenzo Seidenari, Víctor M. Brea 0001, Manuel Mucientes, Alberto Del Bimbo
Pattern Recognit.6
2023 The Florence multi-resolution 3D facial expression dataset
abstract
In the literature, several 3D face datasets have been collected, aiming at advancing the field of 3D face analysis from different perspectives. Data collection generally follows specific research needs, and the existing 3D face datasets all have different characteristics that are tailored for investigating different tasks, encompassing face recognition, facial expressions and emotions analysis, 3D face reconstruction. However, the majority of these datasets are either collected with high-resolution scanners, or consumer level devices, such as the Kinect, the latter being motivated by the burdensome and costly process of collecting high-quality scans. Differently from 2D imagery, the difference in resolution in 3D data represents a non negligible problem that is under-investigated, and still prevents the successful development of methods that can work in real scenarios. In this paper, we propose a new 3D face dataset, named “Florence Multi-Resolution 3D Facial Expression” (Florence 3DMRE), which aims at bridging the gap between high- and low-resolution 3D face datasets. Its peculiarity consists in (1) including high-resolution (HR) models obtained with a HR scanner, and paired samples collected with a Kinect sensor, (2) LR and HR scans are synchronized and capture extreme and asymmetric facial deformations as used in facial rehabilitation exercises. In total, our dataset consists of 14 subjects, each performing 19 complex and asymmetric expressions. For each of them, we collected a high-resolution scan, and an RGB-D sequence. Finally, to highlight the value of the dataset and the challenges it introduces, we use the collected data to perform baseline experiments for cross-resolution 3D face recognition and reconstruction. The dataset is released for research purposes only, and complies to GDPR for data treatment. The dataset can be found at this link.
Claudio Ferrari, Stefano Berretti, Pietro Pala, Alberto Del Bimbo
Pattern Recognit. Lett.4
2023 Attribute disentanglement with gradient reversal for interactive fashion retrieval
abstract
Interactive fashion search is gaining more and more interest thanks to the rapid diffusion of online retailers. It allows users to browse fashion items and perform attribute manipulations, modifying parts or details of given garments. To successfully model and analyze garments at such a fine-grained level, it is necessary to obtain attribute-wise representations, separating information relative to different characteristics. In this work we propose an attribute disentanglement method based on attribute classifiers and the usage of gradient reversal layers. This combination allows us to learn attribute-specific features, removing unwanted details from each representation. We test the effectiveness of our learned features in a fashion attribute manipulation task, obtaining state of the art results. Furthermore, to favor training stability we present a novel loss balancing approach, preventing reversed losses to diverge during the optimization process.
Giovanna Scaramuzzino, Federico Becattini, Alberto Del Bimbo
Pattern Recognit. Lett.3
2023 VISCOUNTH: A Large-scale Multilingual Visual Question Answering Dataset for Cultural Heritage
abstract
Visual question answering has recently been settled as a fundamental multi-modal reasoning task of artificial intelligence that allows users to get information about visual content by asking questions in natural language. In the cultural heritage domain, this task can contribute to assisting visitors in museums and cultural sites, thus increasing engagement. However, the development of visual question answering models for cultural heritage is prevented by the lack of suitable large-scale datasets. To meet this demand, we built a large-scale heterogeneous and multilingual (Italian and English) dataset for cultural heritage that comprises approximately 500K Italian cultural assets and 6.5M question-answer pairs. We propose a novel formulation of the task that requires reasoning over both the visual content and an associated natural language description, and present baselines for this task. Results show that the current state of the art is reasonably effective but still far from satisfactory; therefore, further research in this area is recommended. Nonetheless, we also present a holistic baseline to address visual and contextual questions and foster future research on the topic.
Federico Becattini, Pietro Bongini, Luana Bulla, Alberto Del Bimbo, Ludovica Marinucci, Misael Mongiovì, Valentina Presutti
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Disentangling Features for Fashion Recommendation
abstract
Online stores have become fundamental for the fashion industry, revolving around recommendation systems to suggest appropriate items to customers. Such recommendations often suffer from a lack of diversity and propose items that are similar to previous purchases of a user. Recently, a novel kind of approach based on Memory Augmented Neural Networks (MANNs) has been proposed, aimed at recommending a variety of garments to create an outfit by complementing a given fashion item. In this article we address the task of compatible garment recommendation developing a MANN architecture by taking into account the co-occurrence of clothing attributes, such as shape and color, to compose an outfit. To this end we obtain disentangled representations of fashion items and store them in external memory modules, used to guide recommendations at inference time. We show that our disentangled representations are able to achieve significantly better performance compared to the state of the art and also provide interpretable latent spaces, giving a qualitative explanation of the recommendations.
Lavinia De Divitiis, Federico Becattini, Claudio Baecchi, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.4
2023 (Compress and Restore)N: A Robust Defense Against Adversarial Attacks on Image Classification
abstract
Modern image classification approaches often rely on deep neural networks, which have shown pronounced weakness to adversarial examples: images corrupted with specifically designed yet imperceptible noise that causes the network to misclassify. In this article, we propose a conceptually simple yet robust solution to tackle adversarial attacks on image classification. Our defense works by first applying a JPEG compression with a random quality factor; compression artifacts are subsequently removed by means of a generative model Artifact Restoration GAN. The process can be iterated ensuring the image is not degraded and hence the classification not compromised. We train different AR-GANs for different compression factors, so that we can change its parameters dynamically at each iteration depending on the current compression, making the gradient approximation difficult. We experiment with our defense against three white-box and two black-box attacks, with a particular focus on the state-of-the-art BPDA attack. Our method does not require any adversarial training, and is independent of both the classifier and the attack. Experiments demonstrate that dynamically changing the AR-GAN parameters is of fundamental importance to obtain significant robustness.
Claudio Ferrari, Federico Becattini, Leonardo Galteri, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.4
2023 On Modality Bias Recognition and Reduction
abstract
Making each modality in multi-modal data contribute is of vital importance to learning a versatile multi-modal model. Existing methods, however, are often dominated by one or few of modalities during model training, resulting in sub-optimal performance. In this article, we refer to this problem as modality bias and attempt to study it in the context of multi-modal classification systematically and comprehensively. After stepping into several empirical analyses, we recognize that one modality affects the model prediction more just because this modality has a spurious correlation with instance labels. To primarily facilitate the evaluation on the modality bias problem, we construct two datasets, respectively, for the colored digit recognition and video action recognition tasks in line with the Out-of-Distribution (OoD) protocol. Collaborating with the benchmarks in the visual question answering task, we empirically justify the performance degradation of the existing methods on these OoD datasets, which serves as evidence to justify the modality bias learning. In addition, to overcome this problem, we propose a plug-and-play loss function method, whereby the feature space for each label is adaptively learned according to the training set statistics. Thereafter, we apply this method on 10 baselines in total to test its effectiveness. From the results on four datasets regarding the above three tasks, our method yields remarkable performance improvements compared with the baselines, demonstrating its superiority on reducing the modality bias problem.
Liqiang Nie, Harry Cheng 0002, Zhiyong Cheng 0001, Mohan Kankanhalli, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.6
2023 Learning Streamed Attention Network from Descriptor Images for Cross-Resolution 3D Face Recognition
abstract
In this article, we propose a hybrid framework for cross-resolution 3D face recognition which utilizes a Streamed Attention Network (SAN) that combines handcrafted features with Convolutional Neural Networks (CNNs). It consists of two main stages: first, we process the depth images to extract low-level surface descriptors and derive the corresponding Descriptor Images (DIs), represented as four-channel images. To build the DIs, we propose a variation of the 3D Local Binary Pattern (3DLBP) operator that encodes depth differences using a sigmoid function. Then, we design a CNN that learns from these DIs. The peculiarity of our solution consists in processing each channel of the input image separately, and fusing the contribution of each channel by means of both self- and cross-attention mechanisms. This strategy showed two main advantages over the direct application of Deep-CNN to depth images of the face; on the one hand, the DIs can reduce the diversity between high- and low-resolution data by encoding surface properties that are robust to resolution differences. On the other, it allows a better exploitation of the richer information provided by low-level features, resulting in improved recognition. We evaluated the proposed architecture in a challenging cross-dataset, cross-resolution scenario. To this aim, we first train the network on scanner-resolution 3D data. Next, we utilize the pre-trained network as feature extractor on low-resolution data, where the output of the last fully connected layer is used as face descriptor. Other than standard benchmarks, we also perform experiments on a newly collected dataset of paired high- and low-resolution 3D faces. We use the high-resolution data as gallery, while low-resolution faces are used as probe, allowing us to assess the real gap existing between these two types of data. Extensive experiments on low-resolution 3D face benchmarks show promising results with respect to state-of-the-art methods.
João Baptista Cardia Neto, Claudio Ferrari, Aparecido Nilceu Marana, Stefano Berretti, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Meta-learning Advisor Networks for Long-tail and Noisy Labels in Social Image Classification
abstract
Deep neural networks (DNNs)for social image classification are prone to performance reduction and overfitting when trained on datasets plagued by noisy or imbalanced labels. Weight loss methods tend to ignore the influence of noisy or frequent category examples during the training, resulting in a reduction of final accuracy and, in the presence of extreme noise, even a failure of the learning process. A new advisor network is introduced to address both imbalance and noise problems, and is able to pilot learning of a main network by adjusting the visual features and the gradient with a meta-learning strategy. In a curriculum learning fashion, the impact of redundant data is reduced while recognizable noisy label images are downplayed or redirected.Meta Feature Re-Weighting (MFRW)andMeta Equalization Softmax (MES)methods are introduced to let the main network focus only on the information in an image deemed relevant by the advisor network and to adjust the training gradient to reduce the adverse effects of frequent or noisy categories. The proposed method is first tested on synthetic versions of CIFAR10 and CIFAR100, and then on the more realistic ImageNet-LT, Places-LT, and Clothing1M datasets, reporting state-of-the-art results.
Simone Ricci, Tiberio Uricchio, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Effective conditioned and composed image retrieval combining CLIP-based features
abstract
Conditioned and composed image retrieval extend CBIR systems by combining a query image with an additional text that expresses the intent of the user, describing additional requests w.r.t. the visual content of the query image. This type of search is interesting for e-commerce applications, e.g. to develop interactive multimodal searches and chat-bots. In this demo, we present an interactive system based on a combiner network, trained using contrastive learning, that combines visual and textual features obtained from the OpenAI CLIP network to address conditioned CBIR. The system can be used to improve e-shop search engines. For example, considering the fashion domain it lets users search for dresses, shirts and toptees using a candidate start image and expressing some visual differences w.r.t. its visual con-tent, e.g. asking to change color, pattern or shape. The pro-posed network obtains state-of-the-art performance on the FashionIQ dataset and on the more recent CIRR dataset, showing its applicability to the fashion domain for conditioned retrieval, and to more generic content considering the more general task of composed image retrieval.
Alberto Baldrati, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo
CVPR4
2022 Sparse to Dense Dynamic 3D Facial Expression Generation
abstract
In this paper, we propose a solution to the task of generating dynamic 3D facial expressions from a neutral 3D face and an expression label. This involves solving two sub-problems: (i) modeling the temporal dynamics of expressions, and (ii) deforming the neutral mesh to obtain the expressive counterpart. We represent the temporal evolution of expressions using the motion of a sparse set of 3D landmarks that we learn to generate by training a manifold-valued GAN (Motion3DGAN). To better encode the expression-induced deformation and disentangle it from the identity information, the generated motion is represented as per-frame displacement from a neutral configuration. To generate the expressive meshes, we train a Sparse2Dense mesh Decoder (S2D-Dec) that maps the landmark displacements to a dense, per-vertex displacement. This allows us to learn how the motion of a sparse set of landmarks influences the deformation of the overall face surface, independently from the identity. Experimental results on the CoMA and D3DFACS datasets show that our solution brings significant improvements with respect to previous solutions in terms of both dynamic expression generation and mesh reconstruction, while retaining good generalization to unseen data. Code and models are available at https://github.com/CRISTAL-3DSAM/Sparse2Dense.
Naima Otberdout, Claudio Ferrari, Mohamed Daoudi, Stefano Berretti, Alberto Del Bimbo
CVPR5
2022 Online Deep Clustering with Video Track Consistency
abstract
Several unsupervised and self-supervised approaches have been developed in recent years to learn visual features from large-scale unlabeled datasets. Their main drawback however is that these methods are hardly able to recognize visual features of the same object if it is simply rotated or the perspective of the camera changes. To overcome this limitation and at the same time exploit a useful source of supervision, we take into account video object tracks. Following the intuition that two patches in a track should have similar visual representations in a learned feature space, we adopt an unsupervised clustering-based approach and constrain such representations to be labeled as the same category since they likely belong to the same object or object part. Experimental results on two downstream tasks on different datasets demonstrate the effectiveness of our Online Deep Clustering with Video Track Consistency (ODCT) approach compared to prior work, which did not leverage temporal information. In addition we show that exploiting an unsupervised class-agnostic, yet noisy, track generator yields to better accuracy compared to relying on costly and precise track annotations.
Alessandra Alfani, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
ICPR4
2022 What makes you, you? Analyzing Recognition by Swapping Face Parts
abstract
Deep learning advanced face recognition to an unprecedented accuracy. However, understanding how local parts of the face affect the overall recognition performance is still mostly unclear. Among others, face swap has been experimented to this end, but just for the entire face. In this paper, we propose to swap facial parts as a way to disentangle the recognition relevance of different face parts, like eyes, nose and mouth. In our method, swapping parts from a source face to a target one is performed by fitting a 3D prior, which establishes dense pixels correspondence between parts, while also handling pose differences. Seamless cloning is then used to obtain smooth transitions between the mapped source regions and the shape and skin tone of the target face. We devised an experimental protocol that allowed us to draw some preliminary conclusions when the swapped images are classified by deep networks, indicating a prominence of the eyes and eyebrows region. Code available at https://github.com/clferrari/FacePartsSwap
Claudio Ferrari, Matteo Serpentoni, Stefano Berretti, Alberto Del Bimbo
ICPR4
2022 Restoration of Analog Videos Using Swin-UNet
abstract
In this paper we present a system to restore analog videos of historical archives. These videos often contain severe visual degradation due to the deterioration of their tape supports that require costly and slow manual interventions to recover the original content. The proposed system uses a multi-frame approach and is able to deal also with severe tape mistracking, which results in completely scrambled frames. Tests on real-world videos from a major historical video archive show the effectiveness of our demo system.
Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini 0001, Alberto Del Bimbo
ACM Multimedia4
2022 An AI Powered Re-Identification System for Real-time Contextual Multimedia Applications
abstract
In this demo we present a person re-identification system, based on cameras installed in an environment and featuring AI, that can be used in the design and development of human-centered multimedia applications intended to provide improved situational awareness and context-sensitive user-interfaces. Possible applications are related but not limited to user profiling and personalisation systems, multimedia recommendation, social context understanding and gamification, security, with objectives spanning from environment monitoring, cultural heritage fruition enhancement, retail trade promotion and assistive technologies provision in industry. In the context of the demonstration we are going to set up a system with several workstations equipped with cameras and contextual to paintings reproductions that simulate a museum exhibition in which users are tracked and re-identified at different locations. The data collected is used to enrich the user experience through a reactive voice interface that considers the user's visit, the artworks that most attracted the visitor's attention and the social context.
Giuseppe Becchi, Andrea Ferracani, Filippo Principi, Alberto Del Bimbo
ACM Multimedia4
2022 Engaging Museum Visitors with Gamification of Body and Facial Expressions
abstract
In this demo we present two applications designed for the cultural heritage domain that exploit gamification techniques in order to improve fruition and learning of museum artworks. The two applications encourage users to replicate the poses and facial expressions of characters from paintings or statues, to help museum visitors make connections with works of art. Both applications challenge the user to fulfill a task in a funny way and provide the user with a visual report of the his/her experience that can be shared on social media, improving the engagement of the museums, and providing information on the artworks replicated in the challenge.
Maria Giovanna Donadio, Filippo Principi, Andrea Ferracani, Marco Bertini 0001, Alberto Del Bimbo
ACM Multimedia5
2022 Search-oriented Micro-video Captioning
abstract
Pioneer efforts have been dedicated to the content-oriented video captioning that generates relevant sentences to describe the visual contents of a given video from the producer perspective. By contrast, this work targets at the search-oriented one that summarizes the given video via generating query-like sentences from the consumer angle. Beyond relevance, diversity is vital in characterizing consumers' seeking intention from different aspects. Towards this end, we devise a large-scale multimodal pre-training network regularized by five tasks to strengthen the downstream video representation, which is well-trained over our collected 11M micro-videos. Thereafter, we present a flow-based diverse captioning model to generate different captions from consumers' search demand. This model is optimized via a reconstruction loss and a KL divergence between the prior and the posterior. We justify our model over our constructed golden dataset comprising 690k pairs and experimental results demonstrate its superiority.
Liqiang Nie, Leigang Qu, Dai Meng, Min Zhang 0005, Qi Tian 0001, Alberto Del Bimbo
ACM Multimedia6
2022 Deep variational learning for multiple trajectory prediction of 360° head movements
abstract
Prediction of head movements in immersive media is key to design efficient streaming systems able to focus the bandwidth budget on visible areas of the content. Numerous proposals have therefore been made in the recent years to predict 360° images and videos. However, the performance of these models is limited by a main characteristic of the head motion data: its intrinsic uncertainty. In this article, we present an approach to generate multiple plausible futures of head motion in 360° videos, given a common past trajectory. Our method provides likelihood estimates of every predicted trajectory, enabling direct integration in streaming optimization. To the best of our knowledge, this is the first work that considers the problem of multiple head motion prediction for 360° video streaming. We first quantify this uncertainty from the data. We then introduce our discrete variational multiple sequence (DVMS) learning framework, which builds on deep latent variable models. We design a training procedure to obtain a flexible and lightweight stochastic prediction model compatible with sequence-to-sequence recurrent neural architectures. Experimental results on 3 different datasets show that our method DVMS outperforms competitors adapted from the self-driving domain by up to 37% on prediction horizons up to 5 sec., at lower computational and memory costs. Finally, we design a method to estimate the respective likelihoods of the multiple predicted trajectories, by exploiting the stationarity of the distribution of the prediction error over the latent space. Experimental results on 3 datasets show the quality of these estimates, and how they depend on the video category.
Quentin Guimard, Lucile Sassatelli, Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
MMSys6
2022 Effective triplet mining improves training of multi-scale pooled CNN for image retrieval
Federico Vaccaro, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo
Mach. Vis. Appl.4
2022 A Sparse and Locally Coherent Morphable Face Model for Dense Semantic Correspondence Across Heterogeneous 3D Faces
abstract
The 3D Morphable Model (3DMM) is a powerful statistical tool for representing 3D face shapes. To build a 3DMM, a training set of face scans in full point-to-point correspondence is required, and its modeling capabilities directly depend on the variability contained in the training data. Thus, to increase the descriptive power of the 3DMM, establishing a dense correspondence across heterogeneous scans with sufficient diversity in terms of identities, ethnicities, or expressions becomes essential. In this manuscript, we present a fully automatic approach that leverages a 3DMM to transfer its dense semantic annotation across raw 3D faces, establishing a dense correspondence between them. We propose a novel formulation to learn a set of sparse deformation components with local support on the face that, together with an original non-rigid deformation algorithm, allow the 3DMM to precisely fit unseen faces and transfer its semantic annotation. We extensively experimented our approach, showing it can effectively generalize to highly diverse samples and accurately establish a dense correspondence even in presence of complex facial expressions. The accuracy of the dense registration is demonstrated by building a heterogeneous, large-scale 3DMM from more than 9,000 fully registered scans obtained by joining three large datasets together.
Claudio Ferrari, Stefano Berretti, Pietro Pala, Alberto Del Bimbo
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Events in crowded places: A smart service management
Federico Becattini, Andrea Ferracani, Giuseppe Becchi, Alberto Del Bimbo
Pattern Recognit. Lett.4
2022 Automatic Estimation of Self-Reported Pain by Trajectory Analysis in the Manifold of Fixed Rank Positive Semi-Definite Matrices
abstract
We propose an automatic method to estimate self-reported pain based on facial landmarks extracted from videos. For each video sequence, we decompose the face into four different regions and the pain intensity is measured by modeling the dynamics of facial movement using the landmarks of these regions. A formulation based on Gram matrices is used for representing the trajectory of landmarks on the Riemannian manifold of symmetric positive semi-definite matrices of fixed rank. A curve fitting algorithm is used to smooth the trajectories and temporal alignment is performed to compute the similarity between the trajectories on the manifold. A Support Vector Regression classifier is then trained to encode extracted trajectories into pain intensity levels consistent with self-reported pain intensity measurement. Finally, a late fusion of the estimation for each region is performed to obtain the final predicted pain level. The proposed approach is evaluated on two publicly available datasets, the UNBCMcMaster Shoulder Pain Archive and the Biovid Heat Pain dataset. We compared our method to the state-of-the-art on both datasets using different testing protocols, showing the competitiveness of the proposed approach.
Benjamin Szczapa, Mohamed Daoudi, Stefano Berretti, Pietro Pala, Alberto Del Bimbo, Zakia Hammal
IEEE Trans. Affect. Comput.5
2022 Understanding Human Reactions Looking at Facial Microexpressions With an Event Camera
abstract
With the establishment ofIndustry 4.0, machines are now required to interact with workers. By observing biometrics they can assess if humans are authorized, or mentally and physically fit to work. Understanding body language, makes human–machine interaction more natural, secure, and effective. Nonetheless, traditional cameras have limitations; low frame rate and dynamic range hinder a comprehensive human understanding. This poses a challenge, since faces undergo frequent instantaneous microexpressions. In addition, this is privacy-sensitive information that must be protected. We propose to model expressions with event cameras, bio-inspired vision sensors that have found application within the Industry 4.0 scope. They capture motion at millisecond rates and work under challenging conditions like low illumination and highly dynamic scenes. Such cameras are also privacy-preserving, making them extremely interesting for industry. We show that using event cameras, we can understand human reactions by only observing facial expressions. Comparison with red-green-blue (RGB)-based modeling demonstrates improved effectiveness and robustness.
Federico Becattini, Federico Palai, Alberto Del Bimbo
IEEE Trans. Ind. Informatics3
2022 Regular Polytope Networks
abstract
Neural networks are widely used as a model for classification in a large variety of tasks. Typically, a learnable transformation (i.e., the classifier) is placed at the end of such models returning a value for each class used for classification. This transformation plays an important role in determining how the generated features change during the learning process. In this work, we argue that this transformation not only can be fixed (i.e., set as nontrainable) with no loss of accuracy and with a reduction in memory usage, but it can also be used to learn stationary and maximally separated embeddings. We show that the stationarity of the embedding and its maximal separated representation can be theoretically justified by setting the weights of the fixed classifier to values taken from the coordinate vertices of the three regular polytopes available in [Formula: see text], namely, the d -Simplex, the d -Cube, and the d -Orthoplex. These regular polytopes have the maximal amount of symmetry that can be exploited to generate stationary features angularly centered around their corresponding fixed weights. Our approach improves and broadens the concept of a fixed classifier, recently proposed by Hoffer et al., to a larger class of fixed classifier models. Experimental results confirm the theoretical analysis, the generalization capability, the faster convergence, and the improved performance of the proposed method. Code will be publicly available.
Federico Pernici, Matteo Bruni, Claudio Baecchi, Alberto Del Bimbo
IEEE Trans. Neural Networks Learn. Syst.4
2022 CL2R: Compatible Lifelong Learning Representations
abstract
In this article, we propose a method to partially mimic natural intelligence for the problem of lifelong learning representations that are compatible. We take the perspective of a learning agent that is interested in recognizing object instances in an open dynamic universe in a way in which any update to its internal feature representation does not render the features in the gallery unusable for visual search. We refer to this learning problem as Compatible Lifelong Learning Representations (CL 2 R), as it considers compatible representation learning within the lifelong learning paradigm. We identify stationarity as the property that the feature representation is required to hold to achieve compatibility and propose a novel training procedure that encourages local and global stationarity on the learned representation. Due to stationarity, the statistical properties of the learned features do not change over time, making them interoperable with previously learned features. Extensive experiments on standard benchmark datasets show that our CL 2 R training procedure outperforms alternative baselines and state-of-the-art methods. We also provide novel metrics to specifically evaluate compatible representation learning under catastrophic forgetting in various sequential learning tasks. Code is available at https://github.com/NiccoBiondi/CompatibleLifelongRepresentation .
Niccolò Biondi, Federico Pernici, Matteo Bruni, Daniele Mugnai, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.5
2022 LANBIQUE: LANguage-based Blind Image QUality Evaluation
abstract
Image quality assessment is often performed with deep networks that are fine-tuned to regress a human provided quality score of a given image. Usually, this approach may lack generalization capabilities and, while being highly precise on similar image distribution, it may yield lower correlation on unseen distortions. In particular, they show poor performances, whereas images corrupted by noise, blur, or compression have been restored by generative models. As a matter of fact, evaluation of these generative models is often performed providing anecdotal results to the reader. In the case of image enhancement and restoration, reference images are usually available. Nevertheless, using signal based metrics often leads to counterintuitive results: Highly natural crisp images may obtain worse scores than blurry ones. However, blind reference image assessment may rank images reconstructed with GANs higher than the original undistorted images. To avoid time-consuming human-based image assessment, semantic computer vision tasks may be exploited instead. In this article, we advocate the use of language generation tasks to evaluate the quality of restored images. We refer to our assessment approach as LANguage-based Blind Image QUality Evaluation (LANBIQUE). We show experimentally that image captioning, used as a downstream task, may serve as a method to score image quality, independently of the distortion process that affects the data. Captioning scores are better aligned with human rankings with respect to classic signal based or No-reference image quality metrics. We show insights on how the corruption, by artefacts, of local image structure may steer image captions in the wrong direction.
Leonardo Galteri, Lorenzo Seidenari, Pietro Bongini, Marco Bertini 0001, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.5
2022 Fine-Grained Adversarial Semi-Supervised Learning
abstract
In this article, we exploit Semi-Supervised Learning ( SSL ) to increase the amount of training data to improve the performance of Fine-Grained Visual Categorization ( FGVC ). This problem has not been investigated in the past in spite of prohibitive annotation costs that FGVC requires. Our approach leverages unlabeled data with an adversarial optimization strategy in which the internal features representation is obtained with a second-order pooling model. This combination allows one to back-propagate the information of the parts, represented by second-order pooling, onto unlabeled data in an adversarial training setting. We demonstrate the effectiveness of the combined use by conducting experiments on six state-of-the-art fine-grained datasets, which include Aircrafts, Stanford Cars, CUB-200-2011, Oxford Flowers, Stanford Dogs, and the recent Semi-Supervised iNaturalist-Aves. Experimental results clearly show that our proposed method has better performance than the only previous approach that examined this problem; it also obtained higher classification accuracy with respect to the supervised learning methods with which we compared.
Daniele Mugnai, Federico Pernici, Francesco Turchini, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Measuring 3D face deformations from RGB images of expression rehabilitation exercises
abstract
The accurate (quantitative) analysis of face deformations in 3D is a problem of increasing interest for the many applications it may have. In particular, defining a 3D model of the face that can deform to a 2D target image, while capturing local and asymmetric deformations is still a challenge in the existing literature. Computing a measure of such local deformations may represent a relevant index for monitoring rehabilitation exercises that are used in Parkinson’s and Alzheimer’s disease or in recovering from a stroke. In this study, we present a complete framework that allows the construction of a 3D Morphable Shape Model (3DMM) of the face and its fitting to a target RGB image. The model has the specific characteristic of being based on localized components of deformation; the fitting transformation is performed from 3D to 2D and is guided by the correspondence between landmarks detected in the target image and landmarks manually annotated on the average 3DMM. The fitting has also the peculiarity of being performed in two steps, disentangling face deformations that are due to the identity of the target subject from those induced by facial actions. In the experimental validation of the method, we used the MICC-3D dataset that includes 11 subjects each acquired in one neutral pose plus 18 facial actions that deform the face in localized and asymmetric ways. For each acquisition, we fit the 3DMM to an RGB frame with an apex facial action and to the neutral frame, and computed the extent of the deformation. Results indicated that the proposed approach can accurately capture the face deformation even for localized and asymmetric ones. The proposed framework proved the idea of measuring the deformations of a reconstructed 3D face model to monitor the facial actions performed in response to a set of target ones. Interestingly, these results were obtained just using RGB targets without the need for 3D scans captured with costly devices. This opens the way to the use of the proposed tool for remote medical monitoring of rehabilitation.
Claudio Ferrari, Stefano Berretti, Pietro Pala, Alberto Del Bimbo
Virtual Real. Intell. Hardw.4
2021 Style-Based Outfit Recommendation
abstract
In this paper we propose a garment recommendation system that leverages emotive color information to give recommendations that adhere to a desired style. We leverage previous work by Shigenobu Kobayashi on how specific color combinations, that pertain to certain pre-defined styles, are able to convey specific emotions in human beings. Leveraging this information, we extend the classic general garment recommendation to a style-driven one, where the user can adapt the suggestions to a specific style that may be more appropriate for a specific social event. Here, first we train a generalized style classifier based on Kobayashi's color triplets, then we lever-age a recent memory network-based garment recommendation system to perform suggestions of bottom garments (e.g. skirts, trousers, etc.) given a user-defined top (e.g. a shirt, T-shirt, etc.). Suggestions are then processed to maintain only the ones that, according to our classifier, are coherent with the user defined style. Experiments show that our system is able to generalise on Kobayashi's color styles and that the recommendation system is able to propose garments that are in line with the user desire while also introducing diversity in the proposed garments.
Lavinia De Divitiis, Federico Becattini, Claudio Baecchi, Alberto Del Bimbo
CBMI4
2021 AdaVQA: Overcoming Language Priors with Adapted Margin Cosine Loss
abstract
A number of studies point out that current Visual Question Answering (VQA) models are severely affected by the language prior problem, which refers to blindly making predictions based on the language shortcut. Some efforts have been devoted to overcoming this issue with delicate models. However, there is no research to address it from the view of the answer feature space learning, despite the fact that existing VQA methods all cast VQA as a classification task. Inspired by this, in this work, we attempt to tackle the language prior problem from the viewpoint of the feature space learning. An adapted margin cosine loss is designed to discriminate the frequent and the sparse answer feature space under each question type properly. In this way, the limited patterns within the language modality can be largely reduced to eliminate the language priors. We apply this loss function to several baseline models and evaluate its effectiveness on two VQA-CP benchmarks. Experimental results demonstrate that our proposed adapted margin cosine loss can enhance the baseline models with an absolute performance gain of 15\% on average, strongly verifying the potential of tackling the language prior problem in VQA from the angle of the answer feature space learning.
Liqiang Nie, Zhiyong Cheng 0001, Ji Zhang 0011, Alberto Del Bimbo
IJCAI6
2021 Partially Fake it Till you Make It: Mixing Real and Fake Thermal Images for Improved Object Detection
abstract
In this paper we propose a novel data augmentation approach for visual content domains that have scarce training datasets, compositing synthetic 3D objects within real scenes. We show the performance of the proposed system in the context of object detection in thermal videos, a domain where i) training datasets are very limited compared to visible spectrum datasets and ii) creating full realistic synthetic scenes is extremely cumbersome and expensive due to the difficulty in modeling the thermal properties of the materials of the scene. We compare different augmentation strategies, including state of the art approaches obtained through RL techniques, the injection of simulated data and the employment of a generative model, and study how to best combine our proposed augmentation with these other techniques. Experimental results demonstrate the effectiveness of our approach, and our single-modality detector achieves state-of-the-art results on the FLIR ADAS dataset.
Francesco Bongini, Lorenzo Berlincioni, Marco Bertini 0001, Alberto Del Bimbo
ACM Multimedia4
2021 Fast Video Visual Quality and Resolution Improvement using SR-UNet
abstract
In this paper, we address the problem of real-time video quality enhancement, considering both frame super-resolution and compression artifact-removal. The first operation increases the sampling resolution of video frames, the second removes visual artifacts such as blurriness, noise, aliasing, or blockiness introduced by lossy compression techniques, such as JPEG encoding for single-images, or H.264/H.265 for video data. We propose to use SR-UNet, a novel network architecture based on UNet, that has been specialized for fast visual quality improvement (i.e. capable of operating in less than 40ms, to be able to operate on videos at 25FPS). We show how this network can be used in a streaming context where the content is generated live, e.g. in video calls, and how it can be optimized when video to be streamed are prepared in advance. The network can be used as a final post processing, to optimize the visual appearance of a frame before showing it to the end-user in a video player. Thus, it can be applied without any change to existing video coding and transmission pipelines. Experiments carried on standard video datasets, also considering the H.265 compression, show that the proposed approach is able to either improve visual quality metrics given a fixed bandwidth budget, or video distortion given a fixed quality goal.
Federico Vaccaro, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo
ACM Multimedia4
2021 Conditioned Image Retrieval for Fashion using Contrastive Learning and CLIP-based Features
abstract
Building on the recent advances in multimodal zero-shot representation learning, in this paper we explore the use of features obtained from the recent CLIP model to perform conditioned image retrieval. Starting from a reference image and an additive textual description of what the user wants with respect to the reference image, we learn a Combiner network that is able to understand the image content, integrate the textual description and provide combined feature used to perform the conditioned image retrieval. Starting from the bare CLIP features and a simple baseline, we show that a carefully crafted Combiner network, based on such multimodal features, is extremely effective and outperforms more complex state of the art approaches on the popular FashionIQ dataset.
Alberto Baldrati, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo
MMAsia4
2021 PLM-IPE: A Pixel-Landmark Mutual Enhanced Framework for Implicit Preference Estimation
abstract
In this paper, we are interested in understanding how customers perceive fashion recommendations, in particular when observing a proposed combination of garments to compose an outfit. Automatically understanding how a suggested item is perceived, without any kind of active engagement, is in fact an essential block to achieve interactive applications. We propose a pixel-landmark mutual enhanced framework for implicit preference estimation, named PLM-IPE, which is capable of inferring the user’s implicit preferences exploiting visual cues, without any active or conscious engagement. PLM-IPE consists of three key modules: pixel-based estimator, landmark-based estimator and mutual learning based optimization. The former two modules work on capturing the implicit reaction of the user from the pixel level and landmark level, respectively. The last module serves to transfer knowledge between the two parallel estimators. Towards evaluation, we collected a real-world dataset, named SentiGarment, which contains 3,345 facial reaction videos paired with suggested outfits and human labeled reaction scores. Extensive experiments show the superiority of our model over state-of-the-art approaches.
Federico Becattini, Xuemeng Song, Claudio Baecchi, Shi-Ting Fang, Claudio Ferrari, Liqiang Nie, Alberto Del Bimbo
MMAsia7
2021 Language Based Image Quality Assessment
abstract
Evaluation of generative models, in the visual domain, is often performed providing anecdotal results to the reader. In the case of image enhancement, reference images are usually available. Nonetheless, using signal based metrics often leads to counterintuitive results: highly natural crisp images may obtain worse scores than blurry ones. On the other hand, blind reference image assessment may rank images reconstructed with GANs higher than the original undistorted images. To avoid time consuming human based image assessment, semantic computer vision tasks may be exploited instead [9, 25, 33]. In this paper we advocate the use of language generation tasks to evaluate the quality of restored images. We show experimentally that image captioning, used as a downstream task, may serve as a method to score image quality. Captioning scores are better aligned with human rankings with respect to signal based metrics or no-reference image quality metrics. We show insights on how the corruption, by artifacts, of local image structure may steer image captions in the wrong direction.
Lorenzo Seidenari, Leonardo Galteri, Pietro Bongini, Marco Bertini 0001, Alberto Del Bimbo
MMAsia5
2021 Optical Flow based CNN for detection of unlearnt deepfake manipulations
abstract
A new phenomenon named Deepfakes constitutes a serious threat in video manipulation. AI-based technologies have provided easy-to-use methods to create extremely realistic videos. On the side of multimedia forensics, being able to individuate this kind of fake contents becomes ever more crucial. In this work, a new forensic technique able to detect fake and original video sequences is proposed; it is based on the use of CNNs trained to distinguish possible motion dissimilarities in the temporal structure of a video sequence by exploiting optical flow fields. The results obtained highlight comparable performances with the state-of-the-art methods which, in general, only resort to single video frames. Furthermore, the proposed optical flow based detection scheme also provides a superior robustness in the more realistic cross-forgery operative scenario and can even be combined with frame-based approaches to improve their global effectiveness.
Roberto Caldelli, Leonardo Galteri, Irene Amerini, Alberto Del Bimbo
Pattern Recognit. Lett.4
2021 Am I Done? Predicting Action Progress in Videos
abstract
In this article, we deal with the problem of predicting action progress in videos. We argue that this is an extremely important task, since it can be valuable for a wide range of interaction applications. To this end, we introduce a novel approach, named ProgressNet, capable of predicting when an action takes place in a video, where it is located within the frames, and how far it has progressed during its execution. To provide a general definition of action progress, we ground our work in the linguistics literature, borrowing terms and concepts to understand which actions can be the subject of progress estimation. As a result, we define a categorization of actions and their phases. Motivated by the recent success obtained from the interaction of Convolutional and Recurrent Neural Networks, our model is based on a combination of the Faster R-CNN framework, to make framewise predictions, and LSTM networks, to estimate action progress through time. After introducing two evaluation protocols for the task at hand, we demonstrate the capability of our model to effectively predict action progress on the UCF-101 and J-HMDB datasets.
Federico Becattini, Tiberio Uricchio, Lorenzo Seidenari, Lamberto Ballan, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.5
2020 MANTRA: Memory Augmented Networks for Multiple Trajectory Prediction
abstract
Autonomous vehicles are expected to drive in complex scenarios with several independent non cooperating agents. Path planning for safely navigating in such environments can not just rely on perceiving present location and motion of other agents. It requires instead to predict such variables in a far enough future. In this paper we address the problem of multimodal trajectory prediction exploiting a Memory Augmented Neural Network. Our method learns past and future trajectory embeddings using recurrent neural networks and exploits an associative external memory to store and retrieve such embeddings. Trajectory prediction is then performed by decoding in-memory future encodings conditioned with the observed past. We incorporate scene knowledge in the decoding state by learning a CNN on top of semantic scene maps. Memory growth is limited by learning a writing controller based on the predictive capability of existing embeddings. We show that our method is able to natively perform multi-modal trajectory prediction obtaining state-of-the art results on three datasets. Moreover, thanks to the non-parametric nature of the memory module, we show how once trained our system can continuously improve by ingesting novel patterns.
Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
CVPR4
2020 Task-Conditioned Domain Adaptation for Pedestrian Detection in Thermal Imagery
My Kieu, Andrew D. Bagdanov, Marco Bertini 0001, Alberto Del Bimbo
ECCV (22)4
2020 Modelling the Statistics of Cyclic Activities by Trajectory Analysis on the Manifold of Positive-Semi-Definite Matrices
abstract
In this paper, a model is presented to extract statistical summaries to characterize the repetition of a cyclic body action, for instance a gym exercise, for the purpose of checking the compliance of the observed action to a template one and highlighting the parts of the action that are not correctly executed (if any). The proposed system relies on a Riemannian metric to compute the distance between two poses in such a way that the geometry of the manifold where the pose descriptors lie is preserved; a model to detect the begin and end of each cycle; a model to temporally align the poses of different cycles so as to accurately estimate the cross-sectional mean and variance of poses across different cycles. The proposed model is demonstrated using gym videos taken from the Internet.
Ettore Maria Celozzi, Luca Ciabini, Luca Cultrera, Pietro Pala, Stefano Berretti, Mohamed Daoudi, Alberto Del Bimbo
FG7
2020 Multiple Future Prediction Leveraging Synthetic Trajectories
abstract
Trajectory prediction is an important task, especially in autonomous driving. The ability to forecast the position of other moving agents can yield to an effective planning, ensuring safety for the autonomous vehicle as well for the observed entities. In this work we propose a data driven approach based on Markov Chains to generate synthetic trajectories, which are useful for training a multiple future trajectory predictor. The advantages are twofold: on the one hand synthetic samples can be used to augment existing datasets and train more effective predictors; on the other hand, it allows to generate samples with multiple ground truths, corresponding to diverse equally likely outcomes of the observed trajectory. We define a trajectory prediction model and a loss that explicitly address the multimodality of the problem and we show that combining synthetic and real data leads to prediction improvements, obtaining state of the art results.
Lorenzo Berlincioni, Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
ICPR4
2020 Probability Guided Maxout
abstract
In this paper, we propose an original CNN training strategy that brings together ideas from both dropout-like regularization methods and solutions that learn discriminative features. We propose a dropping criterion that, differently from dropout and its variants, is deterministic rather than random. It grounds on the empirical evidence that feature descriptors with larger L2-norm and highly-active nodes are strongly correlated to confident class predictions. Thus, our criterion guides towards dropping a percentage of the most active nodes of the descriptors, proportionally to the estimated class probability. We simultaneously train a per-sample scaling factor to balance the expected output across training and inference. This further allows us to keep high the descriptor's L2-norm, which we show enforces confident predictions. The combination of these two strategies resulted in our “Probability Guided Maxout” solution that acts as a training regularizer. We prove the above behaviors by reporting extensive image classification results on the CIFAR10, CIFAR100, and Caltech256 datasets. Code is available at https://github.com/clferrari/probability-guided-maxout.
Claudio Ferrari, Stefano Berretti, Alberto Del Bimbo
ICPR3
2020 Inner Eye Canthus Localization for Human Body Temperature Screening
abstract
In this paper, we propose an automatic approach for localizing the inner eye canthus in thermal face images. We first coarsely detect 5 facial keypoints corresponding to the center of the eyes, the nosetip and the ears. Then we compute a sparse 2D-3D points correspondence using a 3D Morphable Face Model (3DMM). This correspondence is used to project the entire 3D face onto the image, and subsequently locate the inner eye canthus. Detecting this location allows to obtain the most precise body temperature measurement for a person using a thermal camera. We evaluated the approach on a thermal face dataset provided with manually annotated landmarks. However, such manual annotations are normally conceived to identify facial parts such as eyes, nose and mouth, and are not specifically tailored for localizing the eye canthus region. As additional contribution, we enrich the original dataset by using the annotated landmarks to deform and project the 3DMM onto the images. Then, by manually selecting a small region corresponding to the eye canthus, we enrich the dataset with additional annotations. By using the manual landmarks, we ensure the correctness of the 3DMM projection, which can be used as ground-truth for future evaluations. Moreover, we supply the dataset with the 3D head poses and per-point visibility masks for detecting self-occlusions. The data is publicly available at https://www.micc.unifi.it/resources/datasets/thermal-face/.
Claudio Ferrari, Lorenzo Berlincioni, Marco Bertini 0001, Alberto Del Bimbo
ICPR4
2020 Temporal Binary Representation for Event-Based Action Recognition
abstract
In this paper we present an event aggregation strategy to convert the output of an event camera into frames processable by traditional Computer Vision algorithms. The proposed method first generates sequences of intermediate binary representations, which are then losslessly transformed into a compact format by simply applying a binary-to-decimal conversion. This strategy allows us to encode temporal information directly into pixel values, which are then interpreted by deep learning models. We apply our strategy, called Temporal Binary Representation, to the task of Gesture Recognition, obtaining state of the art results on the popular DVS128 Gesture Dataset. To underline the effectiveness of the proposed method compared to existing ones, we also collect an extension of the dataset under more challenging conditions on which to perform experiments.
Simone Undri Innocenti, Federico Becattini, Federico Pernici, Alberto Del Bimbo
ICPR4
2020 Robust pedestrian detection in thermal imagery using synthesized images
abstract
In this paper we propose a method for improving pedestrian detection in the thermal domain using two stages: first, a generative data augmentation approach is used, then a domain adaptation method using generated data adapts an RGB pedestrian detector. Our model, based on the Least-Squares Generative Adversarial Network, is trained to synthesize realistic thermal versions of input RGB images which are then used to augment the limited amount of labeled thermal pedestrian images available for training. We apply our generative data augmentation strategy in order to adapt a pretrained YOLOv3 pedestrian detector to detection in the thermal-only domain. Experimental results demonstrate the effectiveness of our approach: using less than 50% of available real thermal training data, and relying on synthesized data generated by our model in the domain adaptation phase, our detector achieves state-of-the-art results on the KAIST Multispectral Pedestrian Detection Benchmark; even if more real thermal data is available adding GAN generated images to the training data results in improved performance, thus showing that these images act as an effective form of data augmentation. To the best of our knowledge, our detector achieves the best single-modality detection results on KAIST with respect to the state-of-the-art.
My Kieu, Lorenzo Berlincioni, Leonardo Galteri, Marco Bertini 0001, Andrew D. Bagdanov, Alberto Del Bimbo
ICPR6
2020 A NoGAN approach for image and video restoration and compression artifact removal
abstract
Lossy image and video compression algorithms introduce several different types of visual artifacts that reduce the visual quality of the compressed media, and the higher the compression rate the higher is the strength of these artifacts. In this work, we describe an approach for visual quality improvement of compressed images and videos to be performed at presentation time, as to obtain the benefits of fast data transfer and reduced data storage, while enjoying a visual quality that could be obtained only reducing the compression rate. To obtain this result we propose to use a deep neural network trained using the NoGAN approach, adapting the popular DeOldify architecture used for colorization. We show how the proposed method can be applied both to image and video compression artifact removal and restoration.
Filippo Mameli, Marco Bertini 0001, Leonardo Galteri, Alberto Del Bimbo
ICPR4
2020 Class-incremental Learning with Pre-allocated Fixed Classifiers
abstract
In class-incremental learning, a learning agent faces a stream of data with the goal of learning new classes while not forgetting previous ones. Neural networks are known to suffer under this setting, as they forget previously acquired knowledge. To address this problem, effective methods exploit past data stored in an episodic memory while expanding the final classifier nodes to accommodate the new classes. In this work, we substitute the expanding classifier with a novel fixed classifier in which a number of pre-allocated output nodes are subject to the classification loss right from the beginning of the learning phase. Contrarily to the standard expanding classifier, this allows: (a) the output nodes of future unseen classes to firstly see negative samples since the beginning of learning together with the positive samples that incrementally arrive; (b) to learn features that do not change their geometric configuration as novel classes are incorporated in the learning model. Experiments with public datasets show that the proposed approach is as effective as the expanding classifier while exhibiting novel intriguing properties of the internal feature representation that are otherwise not-existent. Our ablation study on pre-allocating a large number of classes further validates the approach.
Federico Pernici, Matteo Bruni, Claudio Baecchi, Francesco Turchini, Alberto Del Bimbo
ICPR5
2020 Automatic Estimation of Self-Reported Pain by Interpretable Representations of Motion Dynamics
abstract
We propose an automatic method for pain intensity measurement from video. For each video, pain intensity was measured using the dynamics of facial movement using 66 facial points. Gram matrices formulation was used for facial points trajectory representations on the Riemannian manifold of symmetric positive semi-definite matrices of fixed rank. Curve fitting and temporal alignment were then used to smooth the extracted trajectories. A Support Vector Regression model was then trained to encode the extracted trajectories into ten pain intensity levels consistent with the Visual Analogue Scale for pain intensity measurement. The proposed approach was evaluated using the UNBC McMaster Shoulder Pain Archive and was compared to the state-of-the-art on the same data. Using both 5-fold cross-validation and leave-one-subject-out cross-validation, our results are competitive with respect to state-of-the-art methods.
Benjamin Szczapa, Mohamed Daoudi, Stefano Berretti, Pietro Pala, Alberto Del Bimbo, Zakia Hammal
ICPR5
2020 Learning Group Activities from Skeletons without Individual Action Labels
abstract
To understand human behavior we must not just recognize individual actions but model possibly complex group activity and interactions. Hierarchical models obtain the best results in group activity recognition but require fine grained individual action annotations at the actor level. In this paper we show that using only skeletal data we can train a state-of-the art end-to-end system using only group activity labels at the sequence level. Our experiments show that models trained without individual action supervision perform poorly. On the other hand we show that pseudo-labels can be computed from any pre-trained feature extractor with comparable final performance. Finally our carefully designed lean pose only architecture shows highly competitive results versus more complex multimodal approaches even in the self-supervised variant.
Fabio Zappardino, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo
ICPR4
2020 Image Retrieval using Multi-scale CNN Features Pooling
abstract
In this paper, we address the problem of image retrieval by learning images representation based on the activations of a Convolutional Neural Network. We present an end-to-end trainable network architecture that exploits a novel multi-scale local pooling based on NetVLAD and a triplet mining procedure based on samples difficulty to obtain an effective image representation. Extensive experiments show that our approach is able to reach state-of-the-art results on three standard datasets.
Federico Vaccaro, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo
ICMR4
2020 Automatic Interest Recognition from Posture and Behaviour
abstract
In the last years, the clothing industry has attracted a lot of interest from researchers. Increasing research efforts have been devoted into giving the buyer a way to improve the shopping experience by suggesting meaningful items to purchase. These efforts result in works aiming at suggesting good matches for clothes, but seem to lack one important aspect: understanding the user's interest. In fact, to suggest something it is first necessary to collect the user's personal interests, or something about his or her previous purchases. Without this information, no personalized suggestion can be made. User interest understanding allows to recognize if a user is showing interest in a product he or she is looking at, acquiring precious information that can be later leveraged. Usually user interest is associated to facial expressions, but these are known to be easily falsifiable. Moreover, when privacy is a concern, faces are often impossible to exploit. To address all these aspects, we propose an automatic system that aims to recognize the user's interest towards a garment by just looking at body posture and behaviour. To train and evaluate our system we create a body pose interest dataset, named BodyInterest, which consists of 30 users looking at garments for a total of approximately 6 hours of videos. Extensive evaluations show the effectiveness of our proposed method.
Wolmer Bigi, Claudio Baecchi, Alberto Del Bimbo
ACM Multimedia3
2020 Increasing Video Perceptual Quality with GANs and Semantic Coding
abstract
We have seen a rise in video based user communication in the last year, unfortunately fueled by the spread of COVID-19 disease. Efficient low-latency delay of transmission of video is a challenging problem which must also deal with the segmented nature of network infrastructure not always allowing a high throughput. Lossy video compression is a basic requirement to enable such technology widely. While this may compromise the quality of the streamed video there are recent deep learning based solutions to restore quality of a lossy compressed video.
Leonardo Galteri, Marco Bertini 0001, Lorenzo Seidenari, Tiberio Uricchio, Alberto Del Bimbo
ACM Multimedia5
2020 Image and Video Restoration and Compression Artefact Removal Using a NoGAN Approach
abstract
Lossy image and video compression algorithms introduce several types of visual artefacts that reduce the visual quality of the compressed media. In this work, we report results obtained using the NoGAN training approach and adapting the popular DeOldify architecture used for colorization, for image and video compression artefact removal and restoration.
Filippo Mameli, Marco Bertini 0001, Leonardo Galteri, Alberto Del Bimbo
ACM Multimedia4
2020 Self-supervised on-line cumulative learning from video streams
Federico Pernici, Matteo Bruni, Alberto Del Bimbo
Comput. Vis. Image Underst.3
2020 MIFTel: a multimodal interactive framework based on temporal logic rules
Danilo Avola, Luigi Cinque, Alberto Del Bimbo, Marco Raoul Marini
Multim. Tools Appl.3
2020 Learning Visual Elements of Images for Discovery of Brand Posts
abstract
Online Social Network Sites have become a primary platform for brands and organizations to engage their audience by sharing image and video posts on their timelines. Different from traditional advertising, these posts are not restricted to the products or logo but include visual elements that express more in general the values and attributes of the brand, called brand associations. Since marketers are increasingly spending time in discovering and re-posting user generated posts that reflect the brand attributes, there is an increasing demand for such discovery systems. The goal of these systems is to assist brand experts in filtering through online collections of new user media to discover actionable posts, which match the brand value and have the potential to engage the consumers. Driven by this real-life application, we define and formulate a new task of content discovery for brands and propose a framework that learns to rank posts for brands from their historical timeline. We design a Personalized Content Discovery (PCD) framework to address the three challenges of high inter-brand similarity, sparsity of brand--post interactions, and diversification of timeline. To learn fine-grained brand representation and to generate explanations for the ranking, we automatically learn visual elements of posts from the timeline of brands and from a set of brand attributes in the domain of marketing. To test our framework we use two large-scale Instagram datasets that contain a total of more than 1.5 million image and video posts from the historical timeline of hundreds of brands from multiple verticals such as food and fashion. Extensive experiments indicate that our model can effectively learn fine-grained brand representations and outperform the closest state-of-the-art solutions.
Francesco Gelli, Tiberio Uricchio, Xiangnan He 0001, Alberto Del Bimbo, Tat-Seng Chua
ACM Trans. Multim. Comput. Commun. Appl.4
2019 Towards Real-Time Image Enhancement GANs
Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo
CAIP (1)4
2019 Incremental Learning of People Identities
Federico Bartoli, Federico Pernici, Matteo Bruni, Alberto Del Bimbo
CIARP4
2019 Discovering Identity Specific Activation Patterns in Deep Descriptors for Template Based Face Recognition
abstract
The majority of recent face recognition systems are based on Deep Convolutional Neural Networks (DCNNs). These networks are trained on massive amounts of face images so as to learn a compact representation (deep descriptor) aimed at capturing the identity information. Recognition is then performed by computing some similarity (or distance) measure between descriptors. However, in practice, descriptors encode also other intra-class variabilities such as pose and expressions. This well-known problem is usually addressed by designing specific loss-functions or metric learning modules such that the learned descriptors maximize the inter-class (identity) distances and minimize the intra-class differences in the feature space. We tackle this problem from a different perspective by observing that descriptors associated with images of the same subject, on average, share similar patterns in the highest activation units. We demonstrate this assumption by showing that improved accuracy can be obtained in a template-based recognition scenario by retaining the descriptor bins with the average highest activation, and dropping all the others to zero. These activation patterns are also employed to build identity-representative binary masks that are effectively used in place of the descriptors to match templates. We investigate this strategy by performing experiments on the IJB-A dataset, and show that it can significantly boost the recognition accuracy.
Claudio Ferrari, Stefano Berretti, Alberto Del Bimbo
FG3
2019 Coarse to Fine 3D Face Reconstruction from Single Image
abstract
In this demo we propose a coarse to fine reconstruction pipeline, which takes a single RGB image as input and outputs a detailed 3D model of the face. The pipeline is composed by two main blocks, the coarse reconstruction block, which is based on a 3D Morphable Model, and the refinement block, which instead grounds on a Generative Adversarial Network (GAN).
Leonardo Galteri, Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo
FG5
2019 3D Face Reconstruction from RGB-D Data by Morphable Model to Point Cloud Dense Fitting
abstract
3D cameras for face capturing are quite common today thanks to their ease of use and affordable cost. The depth information they provide is mainly used to enhance face pose estimation and tracking, and face-background segmentation, while applications that require finer face details are usually not possible due to the low-resolution data acquired by such devices. In this paper, we propose a framework that allows us to derive high-quality 3D models of the face starting from corresponding low-resolution depth sequences acquired with a depth camera. To this end, we start by defining a solution that exploits temporal redundancy in a short-sequence of adjacent depth frames to remove most of the acquisition noise and produce an aggregated point cloud output with intermediate level details. Then, using a 3DMM specifically designed to support local and expression-related deformations of the face, we propose a two-steps 3DMM fitting solution: initially the model is deformed under the effect of landmarks correspondences; subsequently, it is iteratively refined using points closeness updating guided by a mean-square optimization. Preliminary results show that the proposed solution is able to derive 3D models of the face with high visual quality; quantitative results also evidence the superiority of our approach with respect to methods that use one step fitting based on landmarks.
Claudio Ferrari, Stefano Berretti, Pietro Pala, Alberto Del Bimbo
ICPRAM4
2019 NeuronUnityIntegration2.0. A Unity Based Application for Motion Capture and Gesture Recognition
abstract
NeuronUnityIntgration2.0 (demo video is avilable at http://tiny.cc/u1lz6y) is a plugin for Unity which provides gesture recognition functionalities through the Perception Neuron motion capture suit. The system offers a recording mode, which guides the user through the collection of a dataset of gestures, and a recognition mode, capable of detecting the recorded actions in real time. Gestures are recognized by training Support Vector Machines directly within our plugin. We demonstrate the effectiveness of our application through an experimental evaluation on a newly collected dataset. Furthermore, external applications can exploit NeuronUnityIntgration2.0's recognition capabilities thanks to a set of exposed API.
Federico Becattini, Andrea Ferracani, Filippo Principi, Marioemanuele Ghianni, Alberto Del Bimbo
ACM Multimedia5
2019 PANEL: Challenges for Multimedia/Multimodal Research in the Next Decade
abstract
The multimedia and multi-modal community is witnessing an explosive transformation in the recent years with major societal impact. With the unprecedented deployment of multimedia devices and systems, multimedia research is critical to our abilities and prospects in advancing state-of-the-art technologies and solving real-world challenges facing the society and the nation. To respond to these challenges and further advance the frontiers of the field of multimedia, this panel will discuss the challenges and visions that may guide future research in the next ten years.
Shih-Fu Chang, Louis-Philippe Morency, Alex Hauptmann 0001, Alberto Del Bimbo, Cathal Gurrin, Hayley Hung, Heng Ji 0001, Alan F. Smeaton
ACM Multimedia4
2019 Fast Video Quality Enhancement using GANs
abstract
Video compression algorithms result in a reduction of image quality, because of their lossy approach to reduce the required bandwidth. This affects commercial streaming services such as Netflix, or Amazon Prime Video, but affects also video conferencing and video surveillance systems. In all these cases it is possible to improve the video quality, both for human view and for automatic video analysis, without changing the compression pipeline, through a post-processing that eliminates the visual artifacts created by the compression algorithms. Generative Adversarial Networks have obtained extremely high quality results in image enhancement tasks; however, to obtain such results large generators are usually employed, resulting in high computational costs and processing time. In this work we present an architecture that can be used to reduce the computational cost and that has been implemented on mobile devices. A possible application is to improve video conferencing, or live streaming. In these cases there is no original uncompressed video stream available. Therefore, we report results using no-reference video quality metric showing high naturalness and quality even for efficient networks.
Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo
ACM Multimedia5
2019 Learning Subjective Attributes of Images from Auxiliary Sources
abstract
Recent years have seen unprecedented research on using artificial intelligence to understand the subjective attributes of images and videos. These attributes are not objective properties of the content but are highly dependent on the perception of the viewers. Subjective attributes are extremely valuable in many applications where images are tailored to the needs of a large group, which consists of many individuals with inherently different ideas and preferences. For instance, marketing experts choose images to establish specific associations in the consumers' minds, while psychologists look for pictures with adequate emotions for therapy. Unfortunately, most of the existing frameworks either focus on objective attributes or rely on large scale datasets of annotated images, making them costly and unable to clearly measure multiple interpretations of a single input. Meanwhile, we can see that users or organizations often interact with images in a multitude of real-life applications, such as the sharing of photographs by brands on social media or the re-posting of image microblogs by users. We argue that these aggregated interactions can serve as auxiliary information to infer image interpretations. To this end, we propose a probabilistic learning framework capable of transferring such subjective information to the image-level labels based on a known aggregated distribution. We use our framework to rank images by subjective attributes from the domain knowledge of social media marketing and personality psychology. Extensive studies and visualizations show that using auxiliary information is a viable line of research for the multimedia community to perform subjective attributes prediction.
Francesco Gelli, Tiberio Uricchio, Xiangnan He 0001, Alberto Del Bimbo, Tat-Seng Chua
ACM Multimedia4
2019 DeepPhysio: Monitored Physiotherapeutic Exercise in the Comfort of your Own Home
abstract
This paper describes an action classification pipeline for detecting and evaluating correct execution of actions in video recorded by smartphone cameras; the use case is that of simplifying monitoring of how physiotherapeutic exercises are performed by patients in the comfort of their own home, reducing the need of physical presence of therapists. Our approach is based on applying DensePose to every frame of acquired video and subsequent sequence analysis by an LSTM network. We validate our proposed recognition approach on a subset of the NTU RGB+D dataset in order to determine the best classification pipeline for this application. We also describe a mobile, cross-platform application called DeepPhysio that is designed to allow at physiotherapy patients to obtain immediate feedback about the correctness of the physical exercises. Preliminary usability analysis shows that this type of application can be effective at monitoring physiotherapy exercises.
Gianmarco Sanesi, Andrew D. Bagdanov, Marco Bertini 0001, Alberto Del Bimbo
ACM Multimedia4
2019 Enhanced skeleton and face 3D data for person re-identification from depth cameras
Pietro Pala, Lorenzo Seidenari, Stefano Berretti, Alberto Del Bimbo
Comput. Graph.4
2019 Deep 3D morphable model refinement via progressive growing of conditional Generative Adversarial Networks
Leonardo Galteri, Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo
Comput. Vis. Image Underst.5
2019 From person to group re-identification via unsupervised transfer of sparse features
Giuseppe Lisanti, Niki Martinel, Christian Micheloni, Alberto Del Bimbo, Gian Luca Foresti
Image Vis. Comput.4
2019 Real-time demographic profiling from face imagery with Fisher vectors
Lorenzo Seidenari, Alessandro Rozza, Alberto Del Bimbo
Mach. Vis. Appl.3
2019 Scene-dependent proposals for efficient person detection
Federico Bartoli, Giuseppe Lisanti, Svebor Karaman, Alberto Del Bimbo
Pattern Recognit.4
2019 Webly-supervised zero-shot learning for artwork instance recognition
Riccardo Del Chiaro, Andrew D. Bagdanov, Alberto Del Bimbo
Pattern Recognit. Lett.3
2019 Award winning papers from the 23rd International Conference on Pattern Recognition (ICPR)
Larry Davis 0001, Alberto Del Bimbo, Brian C. Lovell
Pattern Recognit. Lett.2
2019 Deep Universal Generative Adversarial Compression Artifact Removal
abstract
Image compression is a need that arises in many circumstances. Unfortunately, whenever a lossy compression algorithm is used, artifacts will manifest. Image artifacts, caused by compression tend to eliminate higher frequency details and, in certain cases, may add noise or small image structures. There are two main drawbacks of this phenomenon. First, images appear much less pleasant to the human eye. Second, computer vision algorithms, such as object detectors, may be hindered and their performance reduced. Removing such artifacts means recovering the original image from a perturbed version of it. This means that one ideally should invert the compression process through a complicated nonlinear image transformation. We propose an image transformation approach based on a feedforward fully convolutional residual network model. We show that this model can be optimized either traditionally, directly optimizing an image similarity loss (SSIM), or using a generative adversarial approach (GAN). Our GAN is able to produce images with more photorealistic details than SSIM-based networks. We describe a novel training procedure based on subpatches and devise a novel testing protocol to evaluate restored images quantitatively. We show that our approach can be used as a preprocessing step for different computer vision tasks in case images are degraded by compression to a point that state-of-the art algorithms fail. In this case, our GAN-based approach obtains better performance than MSE or SSIM trained networks. Different from previously proposed approaches, we are able to remove artifacts generated at any QF by inferring the image quality directly from data.
Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo
IEEE Trans. Multim.4
2018 Memory Based Online Learning of Deep Representations From Video Streams
abstract
We present a novel online unsupervised method for face identity learning from video streams. The method exploits deep face descriptors together with a memory based learning mechanism that takes advantage of the temporal coherence of visual data. Specifically, we introduce a discriminative descriptor matching solution based on Reverse Nearest Neighbour and a forgetting strategy that detect redundant descriptors and discard them appropriately while time progresses. It is shown that the proposed learning procedure is asymptotically stable and can be effectively used in relevant applications like multiple face identification and tracking from unconstrained video streams. Experimental results show that the proposed method achieves comparable results in the task of multiple face tracking and better performance in face identification with offline approaches exploiting future information. Code will be publicly available.
Federico Pernici, Federico Bartoli, Matteo Bruni, Alberto Del Bimbo
CVPR4
2018 Context-Aware Trajectory Prediction
abstract
Human motion and behaviour in crowded spaces is influenced by several factors, such as the dynamics of other moving agents in the scene, as well as the static elements that might be perceived as points of attraction or obstacles. In this work, we present a new model for human trajectory prediction which is able to take advantage of both human-human and human-space interactions. The future trajectory of humans, are generated by observing their past positions and interactions with the surroundings. To this end, we propose a “context-aware” recurrent neural network LSTM model, which can learn and predict human motion in crowded spaces such as a sidewalk, a museum or a shopping mall. We evaluate our model on a public pedestrian datasets, and we contribute a new challenging dataset that collects videos of humans that navigate in a (real) crowded space such as a big museum. Results show that our approach can predict human trajectories better when compared to previous state-of-the-art forecasting models.
Federico Bartoli, Giuseppe Lisanti, Lamberto Ballan, Alberto Del Bimbo
ICPR4
2018 Extended YouTube Faces: a Dataset for Heterogeneous Open-Set Face Identification
abstract
In this paper, we propose an extension of the famous YouTube Faces (YTF) dataset. In the YTF dataset, the goal was to state whether two videos contained the same subject or not (video-based face verification). We enrich YTF with still images and an identification protocol. In the classic face identification, given a probe image (or video), the correct identity has to be retrieved among the gallery ones; the main peculiarity of such protocol is that each probe identity has a correspondent in the gallery (closed-set). To resemble a realistic and practical scenario, we devised a protocol in which probe identities are not guaranteed to be in the gallery (open-set). Compared to a closed-set identification, the latter is definitely more challenging in as much as the system needs firstly to reject impostors (i.e., probe identities missing from the gallery), and subsequently, if the probe is accepted as genuine, retrieve the correct identity. In our case, the probe set is composed of full-length videos from the original dataset, while the gallery is composed of templates, i.e., sets of still images. To collect the images, an automatic application was developed. The main motivations behind this work can be found in both the lack of open-set identification protocols defined in the literature and the undeniable complexity of such. We also argued that extending an existing and widely used dataset could make its distribution easier and that data heterogeneity would make the problem even more challenging and realistic. We named the dataset Extended YTF (E-YTF). Finally, we report baseline recognition results using two well known DCNN architectures.
Claudio Ferrari, Stefano Berretti, Alberto Del Bimbo
ICPR3
2018 Video Compression for Object Detection Algorithms
abstract
Video compression algorithms have been designed aiming at pleasing human viewers, and are driven by video quality metrics that are designed to account for the capabilities of the human visual system. However, thanks to the advances in computer vision systems more and more videos are going to be watched by algorithms, e.g. implementing video surveillance systems or performing automatic video tagging. This paper describes an adaptive video coding approach for computer vision-based systems. We show how to control the quality of video compression so that automatic object detectors can still process the resulting video, improving their detection performance, by preserving the elements of the scene that are more likely to contain meaningful content. Our approach is based on computation of saliency maps exploiting a fast objectness measure. The computational efficiency of this approach makes it usable in a real-time video coding pipeline. Experiments show that our technique outperforms standard H.265 in speed and coding efficiency, and can be applied to different types of video domains, from surveillance to web videos.
Leonardo Galteri, Marco Bertini 0001, Lorenzo Seidenari, Alberto Del Bimbo
ICPR4
2018 Beyond the Product: Discovering Image Posts for Brands in Social Media
abstract
Brands and organizations are using social networks such as Instagram to share image or video posts regularly, in order to engage and maximize their presence to the users. Differently from the traditional advertising paradigm, these posts feature not only specific products, but also the value and philosophy of the brand, known as brand associations in marketing literature. In fact, marketers are spending considerable resources to generate their content in-house, and increasingly often, to discover and repost the content generated by users. However, to choose the right posts for a brand in social media remains an open problem. Driven by this real-life application, we define the new task of content discovery for brands, which aims to discover posts that match the marketing value and brand associations of a target brand. We identify two main challenges in this new task: high inter-brand similarity and brand-post sparsity; and propose a tailored content-based learning-to-rank system to discover content for a target brand. Specifically, our method learns fine-grained brand representation via explicit modeling of brand associations, which can be interpreted as visual words shared among brands. We collected a new large-scale Instagram dataset, consisting of more than 1.1 million image and video posts from the history of 927 brands of fourteen verticals such as food and fashion. Extensive experiments indicate that our model can effectively learn fine-grained brand representations and outperform the closest state-of-the-art solutions.
Francesco Gelli, Tiberio Uricchio, Xiangnan He 0001, Alberto Del Bimbo, Tat-Seng Chua
ACM Multimedia4
2018 A multi-camera image processing and visualization system for train safety assessment
Giuseppe Lisanti, Svebor Karaman, Daniele Pezzatini, Alberto Del Bimbo
Multim. Tools Appl.4
2018 Investigating Nuisances in DCNN-Based Face Recognition
abstract
Face recognition "in the wild" has been revolutionized by the deployment of deep learning based approaches. In fact, it has been extensively demonstrated that Deep Convolutional Neural Networks (DCNNs) are powerful enough to overcome most of the limits that affected face recognition algorithms based on hand-crafted features. These include variations in illumination, pose, expression and occlusion, to mention some. The DCNNs discriminative power comes from the fact that low- and high-level representations are learned directly from the raw image data. As a consequence, we expect the performance of a DCNN to be influenced by the characteristics of the image/video data that are fed to the network, and their preprocessing. In this work, we present a thorough analysis of several aspects that impact on the use of DCNN for face recognition. The evaluation has been carried out from two main perspectives: the network architecture and the similarity measures used to compare deeply learned features; the data (source and quality) and their preprocessing (bounding box and alignment). Results obtained on the IJB-A, MegaFace, UMDFaces and YouTube Faces datasets indicate viable hints for designing, training and testing DCNNs. Taking into account the outcomes of the experimental evaluation, we show how competitive performance with respect to the state-of-the-art can be reached even with standard DCNN architectures and pipeline.
Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo
IEEE Trans. Image Process.4
2017 Deep Generative Adversarial Compression Artifact Removal
abstract
Compression artifacts arise in images whenever a lossy compression algorithm is applied. These artifacts eliminate details present in the original image, or add noise and small structures; because of these effects they make images less pleasant for the human eye, and may also lead to decreased performance of computer vision algorithms such as object detectors. To eliminate such artifacts, when decompressing an image, it is required to recover the original image from a disturbed version. To this end, we present a feed-forward fully convolutional residual network model trained using a generative adversarial framework. To provide a baseline, we show that our model can be also trained optimizing the Structural Similarity (SSIM), which is a better loss with respect to the simpler Mean Squared Error (MSE). Our GAN is able to produce images with more photorealistic details than MSE or SSIM based networks. Moreover we show that our approach can be used as a pre-processing step for object detection in case images are degraded by compression to a point that state-of-the art detectors fail. In this task, our GAN method obtains better performance than MSE or SSIM trained networks.
Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo
ICCV4
2017 Group Re-identification via Unsupervised Transfer of Sparse Features Encoding
abstract
Person re-identification is best known as the problem of associating a single person that is observed from one or more disjoint cameras. The existing literature has mainly addressed such an issue, neglecting the fact that people usually move in groups, like in crowded scenarios. We believe that the additional information carried by neighboring individuals provides a relevant visual context that can be exploited to obtain a more robust match of single persons within the group. Despite this, re-identifying groups of people compound the common single person re-identification problems by introducing changes in the relative position of persons within the group and severe self-occlusions. In this paper, we propose a solution for group re-identification that grounds on transferring knowledge from single person reidentification to group re-identification by exploiting sparse dictionary learning. First, a dictionary of sparse atoms is learned using patches extracted from single person images. Then, the learned dictionary is exploited to obtain a sparsity-driven residual group representation, which is finally matched to perform the re-identification. Extensive experiments on the i-LIDS groups and two newly collected datasets show that the proposed solution outperforms stateof-the-art approaches.
Giuseppe Lisanti, Niki Martinel, Alberto Del Bimbo, Gian Luca Foresti
ICCV3
2017 Deep Sentiment Features of Context and Faces for Affective Video Analysis
abstract
Given the huge quantity of hours of video available on video sharing platforms such as YouTube, Vimeo, etc. development of automatic tools that help users find videos that fit their interests has attracted the attention of both scientific and industrial communities. So far the majority of the works have addressed semantic analysis, to identify objects, scenes and events depicted in videos, but more recently affective analysis of videos has started to gain more attention. In this work we investigate the use of sentiment driven features to classify the induced sentiment of a video, i.e. the sentiment reaction of the user. Instead of using standard computer vision features such as CNN features or SIFT features trained to recognize objects and scenes, we exploit sentiment related features such as the ones provided by Deep-SentiBank, and features extracted from models that exploit deep networks trained on face expressions. We experiment on two recently introduced datasets: LIRIS-ACCEDE and MEDIAEVAL-2015, that provide sentiment annotations of a large set of short videos. We show that our approach not only outperforms the current state-of-the-art in terms of valence and arousal classification accuracy, but it also uses a smaller number of features, requiring thus less video processing.
Claudio Baecchi, Tiberio Uricchio, Marco Bertini 0001, Alberto Del Bimbo
ICMR4
2017 PACE: Prediction-based Annotation for Crowded Environments
abstract
We present a new tool we have developed to ease the annotation of crowded environments, typical of visual surveillance datasets. Our tool is developed using HTML5 and Javascript and has two back-ends. A PHP based back-end implement the persistence using a relational database and manage the dynamic creation of pages and the authentication procedure. A python based REST server implement all the computer vision facilities to assist annotators. Our tool allows collaborative annotation of person identity, group membership, location, gaze and occluded parts. PACE supports multiple cameras and if calibration is provided the geometry is used to improve computer vision based assistance. We detail the whole interface comprising an administrative view that ease the setup of the system.
Federico Bartoli, Giuseppe Lisanti, Lorenzo Seidenari, Alberto Del Bimbo
ICMR4
2017 Making a Cultural Visit with a Smart Mate
abstract
Digital and mobile technologies have become increasingly popular to support and improve the quality of experience during cultural visits. The portability of the device, the daily adaptation of most people to its usage, the easy access to information and the opportunity of interactive augmented reality have been key factors of this popularity. We believe that computer vision may help to improve such quality of experience, by making the mobile device smarter and capable of inferring the visitor interests directly from his/her behavior, so triggering the delivery of the appropriate information at the right time without any specific user actions. At MICC University of Florence, we have developed two prototypes of smart audio guides, respectively for indoor and outdoor cultural visits, that exploit the availability of multi-core CPUs and GPUs on mobile devices and computer vision to feed information according to the interests of the visitor, in a non intrusive and natural way. In the first one [Seidenari et al. 2017], the YOLO network [Redmon et al. 2016] is used to distinguish between artworks and people in the camera view. If an artwork is detected, it predicts a specific artwork label. The artwork's description is hence given in audio in the visitor's language. In the second one, the GPS coordinates are used to search Google Places and obtain the interest points closeby. To determine what landmark the visitor is actually looking at, the actual view of the camera is matched against the Google Street Map database using SIFT features. Matched views are classified as either artwork or background and for artworks, descriptions are obtained from Wikipedia. Both prototypes were conceived as a smart mate for visits in museums and outdoor sites or cities of art, respectively. In both prototypes, voice activity detection provides hints about what is happening in the surrounding context of the visitor and triggers the audio description only when the visitor is not talking with the accompanying persons. They were developed on NVIDIA Jetson TK1 and deployed on a NVIDIA Shield K1 Tablet, run in real time and were tested in real contexts in a musum and the city of Florence.
Alberto Del Bimbo
ICMR1
2017 Outdoor Object Recognition for Smart Audio Guides
abstract
We present a smart audio guide that adapts itself to the environment the user is navigating into. The system builds automatically a point of interest database exploiting Wikipedia and Google APIs as source. We rely on a computer vision system, to overcome the likely sensor limitations, and determine with high accuracy if the user is facing a certain landmark or if he is not facing any. Thanks to this the guide presents audio description at the most appropriate moment without any user intervention, using text-to-speech augmenting the experience.
Claudio Baecchi, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo
ACM Multimedia4
2017 Natural Experiences in Museums through Virtual Reality and Voice Commands
abstract
In this demo we present a system for immersive experiences in museums using Voice Commands (VCs) and Virtual Reality (VR). The system has been specifically designed for use by people with motor disabilities. Natural interaction is provided through Automatic Speech Recognition (ASR) and allows to experience VR environments wearing an Head Mounted Display (HMD), i.e. the Oculus Rift. Insights gathered during the implementation and results from an initial usability evaluation are reported.
Andrea Ferracani, Marco Faustino, Gabriele Xavier Giannini, Lea Landucci, Alberto Del Bimbo
ACM Multimedia5
2017 Understanding and localizing activities from correspondences of clustered trajectories
Francesco Turchini, Lorenzo Seidenari, Alberto Del Bimbo
Comput. Vis. Image Underst.3
2017 Indexing quantized ensembles of exemplar-SVMs with rejecting taxonomies
Federico Becattini, Lorenzo Seidenari, Alberto Del Bimbo
Multim. Tools Appl.3
2017 Motion segment decomposition of RGB-D sequences for human behavior understanding
Maxime Devanne, Stefano Berretti, Pietro Pala, Hazem Wannous, Mohamed Daoudi, Alberto Del Bimbo
Pattern Recognit.6
2017 Automatic image annotation via label transfer in the semantic space
Tiberio Uricchio, Lamberto Ballan, Lorenzo Seidenari, Alberto Del Bimbo
Pattern Recognit.4
2017 Spatio-Temporal Closed-Loop Object Detection
abstract
Object detection is one of the most important tasks of computer vision. It is usually performed by evaluating a subset of the possible locations of an image, that are more likely to contain the object of interest. Exhaustive approaches have now been superseded by object proposal methods. The interplay of detectors and proposal algorithms has not been fully analyzed and exploited up to now, although this is a very relevant problem for object detection in video sequences. We propose to connect, in a closed-loop, detectors and object proposal generator functions exploiting the ordered and continuous nature of video sequences. Different from tracking we only require a previous frame to improve both proposal and detection: no prediction based on local motion is performed, thus avoiding tracking errors. We obtain three to four points of improvement in mAP and a detection time that is lower than Faster Regions with CNN features (R-CNN), which is the fastest Convolutional Neural Network (CNN) based generic object detector known at the moment.
Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo
IEEE Trans. Image Process.4
2017 Compact Hash Codes for Efficient Visual Descriptors Retrieval in Large Scale Databases
abstract
In this paper, we present an efficient method for visual descriptors retrieval based on compact hash codes computed using a multiple k-means assignment. The method has been applied to the problem of approximate nearest neighbor (ANN) search of local and global visual content descriptors, and it has been tested on different datasets: three large scale standard datasets of engineered features of up to one billion descriptors (BIGANN) and, supported by recent progress in convolutional neural networks (CNNs), on CIFAR-10, MNIST, INRIA Holidays, Oxford 5K, and Paris 6K datasets; also, the recent DEEP1B dataset, composed by one billion CNN-based features, has been used. Experimental results show that, despite its simplicity, the proposed method obtains a very high performance that makes it superior to more complex state-of-the-art methods.
Simone Ercoli, Marco Bertini 0001, Alberto Del Bimbo
IEEE Trans. Multim.3
2017 A Dictionary Learning-Based 3D Morphable Shape Model
abstract
Face analysis from 2D images and videos is a central task in many multimedia applications. Methods developed to this end perform either face recognition or facial expression recognition, and in both cases results are negatively influenced by variations in pose, illumination, and resolution of the face. Such variations have a lower impact on 3D face data, which has given the way to the idea of using a 3D morphable model as an intermediate tool to enhance face analysis on 2D data. In this paper, we propose a new approach for constructing a 3D morphable shape model (called DL-3DMM) and show our solution can reach the accuracy of deformation required in applications where fine details of the face are concerned. For constructing the model, we start from a set of 3D face scans with large variability in terms of ethnicity and expressions. Across these training scans, we compute a point-topoint dense alignment, which is accurate also in the presence of topological variations of the face. The DL-3DMM is constructed by learning a dictionary of basis components on the aligned scans. The model is then fitted to 2D target faces using an efficient regularized ridge-regression guided by 2D/3D facial landmark correspondences in order to generate pose-normalized face images. Comparison between the DL-3DMM and the standard PCA-based 3DMM demonstrates that in general a lower reconstruction error can be obtained with our solution. Application to action unit detection and emotion recognition from 2D images and videos shows competitive results with state of the art methods on two benchmark datasets.
Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo
IEEE Trans. Multim.4
2017 Deep Artwork Detection and Retrieval for Automatic Context-Aware Audio Guides
abstract
In this article, we address the problem of creating a smart audio guide that adapts to the actions and interests of museum visitors. As an autonomous agent, our guide perceives the context and is able to interact with users in an appropriate fashion. To do so, it understands what the visitor is looking at, if the visitor is moving inside the museum hall, or if he or she is talking with a friend. The guide performs automatic recognition of artworks, and it provides configurable interface features to improve the user experience and the fruition of multimedia materials through semi-automatic interaction. Our smart audio guide is backed by a computer vision system capable of working in real time on a mobile device, coupled with audio and motion sensors. We propose the use of a compact Convolutional Neural Network (CNN) that performs object classification and localization. Using the same CNN features computed for these tasks, we perform also robust artwork recognition. To improve the recognition accuracy, we perform additional video processing using shape-based filtering, artwork tracking, and temporal filtering. The system has been deployed on an NVIDIA Jetson TK1 and a NVIDIA Shield Tablet K1 and tested in a real-world environment (Bargello Museum of Florence).
Lorenzo Seidenari, Claudio Baecchi, Tiberio Uricchio, Andrea Ferracani, Marco Bertini 0001, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.6
2016 User interest profiling using tracking-free coarse gaze estimation
abstract
Understanding where people attention focuses is a challenging and extremely valuable task that can be solved using computer vision technologies. In this paper we address this problem on surveillance-like scenarios, where head and body imagery are usually low resolution. We propose a method to profile the attention of people moving in a known space. We exploit coarse gaze estimation and a novel model based on optical flow to improve attention prediction without the need of a tracker. Removing the tracker dependency makes the method applicable also on highly crowded scenarios. The proposed method is able to obtain comparable performance with respect to state of the art solutions in terms of Mean Average Angular Error (MAAE) on the TownCentre dataset. We also test our approach on the publicly available MuseumVisitors dataset showing an improvement both in terms of MAAE and in terms of accuracy in the estimation of visitors' profile.
Federico Bartoli, Giuseppe Lisanti, Lorenzo Seidenari, Alberto Del Bimbo
ICPR4
2016 Learning shape variations of motion trajectories for gait analysis
abstract
The analysis of human gait is more and more investigated due to its large panel of potential applications in various domains, like rehabilitation, deficiency diagnosis, surveillance and movement optimization. In addition, the release of depth sensors offers new opportunities to achieve gait analysis in a non-intrusive context. In this paper, we propose a gait analysis method from depth sequences by analyzing separately each step so as to be robust to gait duration and incomplete cycles. We analyze the shape of the motion trajectory as signature of the gait and consider shape variations within a Riemannian manifold to learn step models. During classification, the derivation of each performed step is evaluated in an online manner to qualitatively analyze the gait. Experiments are carried out in the context of abnormal gait detection and person re-identification trough gait recognition. Results demonstrated the potential of the method in both scenarios.
Maxime Devanne, Hazem Wannous, Mohamed Daoudi, Stefano Berretti, Alberto Del Bimbo, Pietro Pala
ICPR5
2016 Effective 3D based frontalization for unconstrained face recognition
abstract
In this paper, we propose a new and effective frontalization algorithm for frontal rendering of unconstrained face images, and experiment it for face recognition. Initially, a 3DMM is fit to the image, and an interpolating function maps each pixel inside the face region on the image to the 3D model's. Thus, we can render a frontal view without introducing artifacts in the final image thanks to the exact correspondence between each pixel and the 3D coordinate of the model. The 3D model is then back projected onto the frontalized image allowing us to localize image patches where to extract the feature descriptors, and thus enhancing the alignment between the same descriptor over different images. Our solution outperforms other frontalization techniques in terms of face verification. Results comparable to state-of-the-art on two challenging benchmark datasets are also reported, supporting our claim of effectiveness of the proposed face image representation.
Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo
ICPR4
2016 Bloom Filters and Compact Hash Codes for Efficient and Distributed Image Retrieval
abstract
This paper presents a novel method for efficient image retrieval, based on a simple and effective hashing of CNN features and the use of an indexing structure based on Bloom filters. These filters are used as gatekeepers for the database of image features, allowing to avoid to perform a query if the query features are not stored in the database and speeding up the query process, without affecting retrieval performance. Thanks to the limited memory requirements the system is suitable for mobile applications and distributed databases, associating each filter to a distributed portion of the database (database shard), addressing large scale archives and allowing query parallelization. Experimental validation has been performed on three standard image retrieval datasets, outperforming state-of-the-art hashing methods in terms of precision, while the proposed indexing method obtains a 2x speedup.
Andrea Salvi, Simone Ercoli, Marco Bertini 0001, Alberto Del Bimbo
ISM4
2016 Item-Based Video Recommendation: An Hybrid Approach considering Human Factors
abstract
In this paper we propose a method for video recommendation in Social Networks based on crowdsourced and automatic video annotations of salient frames. We show how two human factors, users' self-expression in user profiles and perception of visual saliency in videos, can be exploited in order to stimulate annotations and to obtain an efficient representation of video content features. Results are assessed through experiments conducted on a prototype of social network for video sharing. Several baseline approaches are evaluated and we show how the proposed method improves over them.
Andrea Ferracani, Daniele Pezzatini, Marco Bertini 0001, Alberto Del Bimbo
ICMR4
2016 Web Video Popularity Prediction using Sentiment and Content Visual Features
abstract
Hundreds of hours of videos are uploaded every minute on YouTube and other video sharing sites: some will be viewed by millions of people and other will go unnoticed by all but the uploader. In this paper we propose to use visual sentiment and content features to predict the popularity of web videos. The proposed approach outperforms current state-of-the-art methods on two publicly available datasets.
Giulia Fontanini, Marco Bertini 0001, Alberto Del Bimbo
ICMR3
2016 Do Textual Descriptions Help Action Recognition?
abstract
We present a novel method to improve action recognition by leveraging a set of captioned videos. By learning linear projections to map videos and text onto a common space, our approach shows that improved results on unseen videos can be obtained. We also propose a novel structure preserving loss that further ameliorates the quality of the projections. We tested our method on the challenging, realistic, Hollywood2 action recognition dataset where a considerable gain in performance is obtained. We show that the gain is proportional to the number of training samples used to learn the projections.
Matteo Bruni, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo
ACM Multimedia4
2016 Real-time Wearable Computer Vision System for Improved Museum Experience
abstract
The goal of this work is to implement a real-time computer vision system that can run on wearable devices to perform object classification and artwork recognition, to improve the experience of a museum visit through understanding the interests of users. Object classification helps to understand the context of the visit, e.g. differentiating when a visitor is talking with people, or just wandering through the museum, or if he is looking at an exhibit that interests him. Artwork recognition allows to provide automatically information of the observed item or to create a user profile based on what and how long a user has observed artworks.
Giovanni Taverriti, Stefano Lombini, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo
ACM Multimedia5
2016 A multimodal feature learning approach for sentiment analysis of social network multimedia
Claudio Baecchi, Tiberio Uricchio, Marco Bertini 0001, Alberto Del Bimbo
Multim. Tools Appl.4
2016 Guest Editorial: Learning Multimedia for Real World Applications
Bing-Kun Bao, Congyan Lang, Tao Mei 0001, Alberto Del Bimbo
Multim. Tools Appl.4
2016 Personalized multimedia content delivery on an interactive table by passive observation of museum visitors
Svebor Karaman, Andrew D. Bagdanov, Lea Landucci, Gianpaolo D'Amico, Andrea Ferracani, Daniele Pezzatini, Alberto Del Bimbo
Multim. Tools Appl.7
2016 Continuous localization and mapping of a pan-tilt-zoom camera for wide area tracking
Giuseppe Lisanti, Iacopo Masi, Federico Pernici, Alberto Del Bimbo
Mach. Vis. Appl.4
2016 Reconstructing High-Resolution Face Models From Kinect Depth Sequences
abstract
Performing face recognition across 3D scans with different resolution is now attracting an increasing interest thanks to the introduction of a new generation of depth cameras, capable of acquiring color/depth images over time. In fact, these devices acquire and provide depth data with much lower resolution compared with the 3D high-resolution scanners typically used for face recognition applications. If data are acquired without user cooperation, the problem is even more challenging, and the gap of resolution between probe and gallery scans can yield to a severe loss in terms of recognition accuracy. Based on these premises, we propose a method to build a higher resolution 3D face model from 3D data acquired by a low-resolution scanner. This face model is built using data acquired when a person passes in front of the scanner, without assuming any particular cooperation. The 3D data are registered and filtered by combining a model of the expected distribution of the acquisition error with a variant of the lowess method to remove outliers and build the final face model. The proposed approach is evaluated in terms of accuracy of face reconstruction and face recognition.
Enrico Bondi, Pietro Pala, Stefano Berretti, Alberto Del Bimbo
IEEE Trans. Inf. Forensics Secur.4
2016 Boosting 3D LBP-Based Face Recognition by Fusing Shape and Texture Descriptors on the Mesh
abstract
In this paper, we present a novel approach for fusing shape and texture local binary patterns (LBPs) on a mesh for 3D face recognition. Using a recently proposed framework, we compute LBP directly on the face mesh surface, then we construct a grid of the regions on the facial surface that can accommodate global and partial descriptions. Compared with its depth-image counterpart, our approach is distinguished by the following features: 1) inherits the intrinsic advantages of mesh surface (e.g., preservation of the full geometry); 2) does not require normalization; and 3) can accommodate partial matching. In addition, it allows early level fusion of texture and shape modalities. Through experiments conducted on the BU-3DFE and Bosphorus databases, we assess different variants of our approach with regard to facial expressions and missing data, also in comparison to the state-of-the-art solutions.
Naoufel Werghi, Claudio Tortorici, Stefano Berretti, Alberto Del Bimbo
IEEE Trans. Inf. Forensics Secur.4
2016 From the Past Editor-In-Chief
abstract
editorial Free Access Share on Editorials from the Past and Current Editors-in-Chief Author: Ralf Steinmetz Technical University of Darmstadt Technical University of DarmstadtView Profile , Editor: Alberto del Bimbo Università di Firenze Università di FirenzeView Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 12Issue 3June 2016 Article No.: 37epp 1–3https://doi.org/10.1145/2903774Published:15 June 2016Publication History 0citation185DownloadsMetricsTotal Citations0Total Downloads185Last 12 Months12Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.1
2015 Dictionary Learning Based 3D Morphable Model Construction for Face Recognition with Varying Expression and Pose
abstract
In this paper, we propose a new approach for constructing a 3D morph able model (3DMM) and experiment its application to face recognition. Differently from existing solutions, the proposed 3DMM is constructed from a training set that includes a large spectrum of variability in terms of ethnicity and facial expressions. By exploiting annotated landmarks available in the training data, we are able of establishing dense correspondence across training scans also in the presence of strong facial expressions. The 3DMM is then constructed by learning a dictionary of basis components, instead of using the traditional approach based on PCA decomposition. Finally, we cast the proposed dictionary learning DL-3DMM to a rigid/non-rigid deformation framework, which includes pose estimation and regularized ridge-regression fitting to 2D images. Comparative results between the DL-3DMM and its PCA counterpart are reported, together with face recognition results for images with large pose and expression variations.
Claudio Ferrari, Giuseppe Lisanti, Stefano Berretti, Alberto Del Bimbo
3DV4
2015 MORF: Multi-Objective Random Forests for face characteristic estimation
abstract
In this paper we describe a technique for joint estimation of head pose and multiple soft biometrics from faces (Age, Gender and Ethnicity). Our proposed Multi-Objective Random Forests (MORF) framework is a unified model for the joint estimation of multiple characteristics that automatically adapts the measure of information gain used for evaluating the quality of weak learners. Since facial characteristics are related in the feature space, estimating all of them jointly can be beneficial as trees can learn to condition the estimation of some characteristics on others. We reformulate the splitting criterion of random trees in our multi-objective formulation and evaluate it on publicly available face characteristic estimation imagery. These preliminary experiments show promising results.
Dario Di Fina, Svebor Karaman, Andrew D. Bagdanov, Alberto Del Bimbo
AVSS4
2015 Representing 3D texture on mesh manifolds for retrieval and recognition applications
abstract
In this paper, we present and experiment a novel approach for representing texture of 3D mesh manifolds using local binary patterns (LBP). Using a recently proposed framework [37], we compute LBP directly on the mesh surface, either using geometric or photometric appearance. Compared to its depth-image counterpart, our approach is distinguished by the following features: a) inherits the intrinsic advantages of mesh surface (e.g., preservation of the full geometry); b) does not require normalization; c) can accommodate partial matching. In addition, it allows early-level fusion of the geometry and photometric texture modalities. Through experiments conducted on two application scenarios, namely, 3D texture retrieval and 3D face recognition, we assess the effectiveness of the proposed solution with respect to state of the art approaches.
Naoufel Werghi, Claudio Tortorici, Stefano Berretti, Alberto Del Bimbo
CVPR4
2015 WATTS: a Web Annotation Tool for Surveillance Scenarios
abstract
In this paper, we present a web based annotation tool we developed allowing creating collaboratively a detailed ground truth for datasets related to visual surveillance and behavior understanding. The system persistence is based on a relational database and the user interface is designed using HTML5, Javascript and CSS. Our tool can easily manage datasets with multiple cameras. It allows annotating a person location in the image, its identity, its body and head gaze, as well as a potential occlusion or group membership. We justify each annotation type with regards to current trends of research in the computer vision community. We further detail how our interface can be used to annotate each of these annotations type. We conclude the paper with an usability evaluation of our system.
Federico Bartoli, Lorenzo Seidenari, Giuseppe Lisanti, Svebor Karaman, Alberto Del Bimbo
ACM Multimedia5
2015 Movie's Affect Communication Using Multisensory Modalities
abstract
The goal of the system presented in this demo is to make possible for the visually and hearing impaired audience to live empathetic viewing experiences using their home theatre. In this work we suggest the incorporation of new emotion communication modalities into the standard television, to provide the targeted audience with sensations that they do not have the opportunity to enjoy because of their disability.
Joël Dumoulin, Diana Affi, Elena Mugellini, Omar Abou Khaled, Marco Bertini 0001, Alberto Del Bimbo
ACM Multimedia6
2015 smArt: Open and Interactive Indoor Cultural Data
abstract
In this demo we present smArt, a low-cost framework to quickly set up indoor exhibits featuring a smart navigation system for museums. The framework is web-based and allows the design on a digital map of a sensorized museum environment and the dynamic and assisted definition of the multimedia materials and sensors associated to the artworks. The knowledge-base uses semantic technologies and it is exploited by museum visitors to get directions and to have multimedia insights in a natural way. Indoor localisation and routing is provided taking advantage of active and passive sensors advertisements and user interactions. In this way we overcome the Global Positioning System (GPS) unavailability issue in indoor environments.
Andrea Ferracani, Daniele Pezzatini, Alberto Del Bimbo, Riccardo Del Chiaro, Franco Yang, Maurizio Sanesi
ACM Multimedia3
2015 PITAGORA: Recommending Users and Local Experts in an Airport Social Network
abstract
In this demo we present PITAGORA\footnote{Demo video available at http://bit.ly/1GgtUrN}: a mobile web contextual social network designed for the check-in area of an airport. The app provides recommendation of potential friends, local experts and targeted services. Recommendation is hybrid and combines social media analysis and collaborative filtering techniques. Users' recommendation has been evaluated through a user study with good results.
Andrea Ferracani, Daniele Pezzatini, Andrea Benericetti, Marco Guiducci, Alberto Del Bimbo
ACM Multimedia5
2015 A System for Video Recommendation using Visual Saliency, Crowdsourced and Automatic Annotations
abstract
In this paper we present a system for content-based video recommendation that exploits visual saliency to better represent video features and content\footnote{Demo video available at http://bit.ly/1FYloeQ}. Visual saliency is used to select relevant frames to be presented in a web-based interface to tag and annotate video frames in a social network; it is also employed to summarize video content to create a more effective video representation used in the recommender system. The system exploits automatic annotations from CNN-based classifiers on salient frames and user generated annotations. We evaluate several baseline approaches and show how the proposed method improves over them.
Andrea Ferracani, Daniele Pezzatini, Marco Bertini 0001, Saverio Meucci, Alberto Del Bimbo
ACM Multimedia5
2015 Image Popularity Prediction in Social Media Using Sentiment and Context Features
abstract
Images in social networks share different destinies: some are going to become popular while others are going to be completely unnoticed. In this paper we propose to use visual sentiment features together with three novel context features to predict a concise popularity score of social images. Experiments on large scale datasets show the benefits of proposed features on the performance of image popularity prediction. Exploiting state-of-the-art sentiment features, we report a qualitative analysis of which sentiments seem to be related to good or poor popularity. To the best of our knowledge, this is the first work understanding specific visual sentiments that positively or negatively influence the eventual popularity of images.
Francesco Gelli, Tiberio Uricchio, Marco Bertini 0001, Alberto Del Bimbo, Shih-Fu Chang
ACM Multimedia4
2015 Image Tag Assignment, Refinement and Retrieval
abstract
This tutorial focuses on challenges and solutions for content-based image annotation and retrieval in the context of online image sharing and tagging. We present a unified review on three closely linked problems, i.e., tag assignment, tag refinement, and tag-based image retrieval. We introduce a taxonomy to structure the growing literature, understand the ingredients of the main works, clarify their connections and difference, and recognize their merits and limitations. Moreover, we present an open-source testbed, with training sets of varying sizes and three test datasets, to evaluate methods of varied learning complexity. A selected set of eleven representative works have been implemented and evaluated. During the tutorial we provide a practice session for hands on experience with the methods, software and datasets. For repeatable experiments all data and code are online at http://www.micc.unifi.it/tagsurvey
Xirong Li 0001, Tiberio Uricchio, Lamberto Ballan, Marco Bertini 0001, Cees Snoek, Alberto Del Bimbo
ACM Multimedia6
2015 A data-driven approach for tag refinement and localization in web videos
Lamberto Ballan, Marco Bertini 0001, Giuseppe Serra 0001, Alberto Del Bimbo
Comput. Vis. Image Underst.4
2015 Non-myopic information theoretic sensor management of a single pan-tilt-zoom camera for multiple object detection and tracking
Pietro Salvagnini, Federico Pernici, Marco Cristani, Giuseppe Lisanti, Alberto Del Bimbo, Vittorio Murino
Comput. Vis. Image Underst.5
2015 Local binary patterns on triangular meshes: Concept and applications
Naoufel Werghi, Claudio Tortorici, Stefano Berretti, Alberto Del Bimbo
Comput. Vis. Image Underst.4
2015 Data-driven approaches for social image and video tagging
Lamberto Ballan, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo
Multim. Tools Appl.4
2015 Person Re-Identification by Iterative Re-Weighted Sparse Ranking
abstract
In this paper we introduce a method for person re-identification based on discriminative, sparse basis expansions of targets in terms of a labeled gallery of known individuals. We propose an iterative extension to sparse discriminative classifiers capable of ranking many candidate targets. The approach makes use of soft- and hard- re-weighting to redistribute energy among the most relevant contributing elements and to ensure that the best candidates are ranked at each iteration. Our approach also leverages a novel visual descriptor which we show to be discriminative while remaining robust to pose and illumination variations. An extensive comparative evaluation is given demonstrating that our approach achieves state-of-the-art performance on single- and multi-shot person re-identification scenarios on the VIPeR, i-LIDS, ETHZ, and CAVIAR4REID datasets. The combination of our descriptor and iterative sparse basis expansion improves state-of-the-art rank-1 performance by six percentage points on VIPeR and by 20 on CAVIAR4REID compared to other methods with a single gallery image per person. With multiple gallery and probe images per person our approach improves by 17 percentage points the state-of-the-art on i-LIDS and by 72 on CAVIAR4REID at rank-1. The approach is also quite efficient, capable of single-shot person re-identification over galleries containing hundreds of individuals at about 30 re-identifications per second.
Giuseppe Lisanti, Iacopo Masi, Andrew D. Bagdanov, Alberto Del Bimbo
IEEE Trans. Pattern Anal. Mach. Intell.4
2015 3-D Human Action Recognition by Shape Analysis of Motion Trajectories on Riemannian Manifold
abstract
Recognizing human actions in 3-D video sequences is an important open problem that is currently at the heart of many research domains including surveillance, natural interfaces and rehabilitation. However, the design and development of models for action recognition that are both accurate and efficient is a challenging task due to the variability of the human pose, clothing and appearance. In this paper, we propose a new framework to extract a compact representation of a human action captured through a depth sensor, and enable accurate action recognition. The proposed solution develops on fitting a human skeleton model to acquired data so as to represent the 3-D coordinates of the joints and their change over time as a trajectory in a suitable action space. Thanks to such a 3-D joint-based framework, the proposed solution is capable to capture both the shape and the dynamics of the human body, simultaneously. The action recognition problem is then formulated as the problem of computing the similarity between the shape of trajectories in a Riemannian manifold. Classification using k-nearest neighbors is finally performed on this manifold taking advantage of Riemannian geometry in the open curve shape space. Experiments are carried out on four representative benchmarks to demonstrate the potential of the proposed solution in terms of accuracy/latency for a low-latency action recognition. Comparative results with state-of-the-art methods are reported.
Maxime Devanne, Hazem Wannous, Stefano Berretti, Pietro Pala, Mohamed Daoudi, Alberto Del Bimbo
IEEE Trans. Cybern.6
2015 The Mesh-LBP: A Framework for Extracting Local Binary Patterns From Discrete Manifolds
abstract
In this paper, we present a novel and original framework, which we dubbed mesh-local binary pattern (LBP), for computing local binary-like-patterns on a triangular-mesh manifold. This framework can be adapted to all the LBP variants employed in 2D image analysis. As such, it allows extending the related techniques to mesh surfaces. After describing the foundations, the construction and the main features of the mesh-LBP, we derive its possible variants and show how they can extend most of the 2D-LBP variants to the mesh manifold. In the experiments, we give evidence of the presence of the uniformity aspect in the mesh-LBP, similar to the one noticed in the 2D-LBP. We also report repeatability experiments that confirm, in particular, the rotation-invariance of mesh-LBP descriptors. Furthermore, we analyze the potential of mesh-LBP for the task of 3D texture classification of triangular-mesh surfaces collected from public data sets. Comparison with state-of-the-art surface descriptors, as well as with 2D-LBP counterparts applied on depth images, also evidences the effectiveness of the proposed framework. Finally, we illustrate the robustness of the mesh-LBP with respect to the class of mesh irregularity typical to 3D surface-digitizer scans.
Naoufel Werghi, Stefano Berretti, Alberto Del Bimbo
IEEE Trans. Image Process.3
2014 Real-time people counting from depth imagery of crowded environments
abstract
In this paper we describe a system for automatic people counting in crowded environments. The approach we propose is a counting-by-detection method based on depth imagery. It is designed to be deployed as an autonomous appliance for crowd analysis in video surveillance application scenarios. Our system performs foreground/background segmentation on depth image streams in order to coarsely segment persons, then depth information is used to localize head candidates which are then tracked in time on an automatically estimated ground plane. The system runs in real-time, at a frame-rate of about 20 fps. We collected a dataset of RGB-D sequences representing three typical and challenging surveillance scenarios, including crowds, queuing and groups. An extensive comparative evaluation is given between our system and more complex, Latent SVM-based head localization for person counting applications.
Enrico Bondi, Lorenzo Seidenari, Andrew D. Bagdanov, Alberto Del Bimbo
AVSS4
2014 Adaptive Structured Pooling for Action Recognition
Svebor Karaman, Lorenzo Seidenari, Shugao Ma, Alberto Del Bimbo, Stan Sclaroff
BMVC4
2014 Fisher Vectors over Random Density Forests for Object Recognition
abstract
In this paper we describe a Fisher vector encoding of images over Random Density Forests. Random Density Forests (RDFs) are an unsupervised variation of Random Decision Forests for density estimation. In this work we train RDFs by splitting at each node in order to minimize the Gaussian differential entropy of each split. We use this as generative model of image patch features and derive the Fisher vector representation using the RDF as the underlying model. Our approach is computationally efficient, reducing the amount of Gaussian derivatives to compute, and allows more flexibility in the feature density modelling. We evaluate our approach on the PASCAL VOC 2007 dataset showing that our approach, that only uses linear classifiers, improves over bag of visual words and is comparable to the traditional Fisher vector encoding over Gaussian Mixture Models for density estimation.
Claudio Baecchi, Francesco Turchini, Lorenzo Seidenari, Andrew D. Bagdanov, Alberto Del Bimbo
ICPR5
2014 Unsupervised Scene Adaptation for Faster Multi-scale Pedestrian Detection
abstract
In this paper we describe an approach to automatically improving the efficiency of soft cascade-based person detectors. Our technique addresses the two fundamental bottlenecks in cascade detectors: the number of weak classifiers that need to be evaluated in each cascade, and the total number of detection windows to be evaluated. By simply observing a soft cascade operating on a scene, we learn scale specific linear approximations of cascade traces that allows us to eliminate a large fraction of the classifier evaluation. Independently, this time by observing regions of support in the soft cascade on a training set, we learn a coarse geometric model of the scene that allows our detector to propose candidate detection windows and significantly reduce the number of windows run through the cascade. Our approaches are unsupervised and require no additional labeled person images for learning. Our linear cascade approximation results in about 28% savings in detection, while our geometric model gives a saving of over 95%, without appreciable loss of accuracy.
Federico Bartoli, Giuseppe Lisanti, Svebor Karaman, Andrew D. Bagdanov, Alberto Del Bimbo
ICPR5
2014 Pose Independent Face Recognition by Localizing Local Binary Patterns via Deformation Components
abstract
In this paper we address the problem of pose independent face recognition with a gallery set containing one frontal face image per enrolled subject while the probe set is composed by just a face image undergoing pose variations. The approach uses a set of aligned 3D models to learn deformation components using a 3D Morph able Model (3DMM). This further allows fitting a 3DMM efficiently on an image using a Ridge regression solution, regularized on the face space estimated via PCA. Then the approach describes each profile face by computing Local Binary Pattern (LBP) histograms localized on each deformed vertex, projected on a rendered frontal view. In the experimental result we evaluate the proposed method on the CMU Multi-PIE to assess face recognition algorithm across pose. We show how our process leads to higher performance than regular baselines reporting high recognition rate considering a range of facial poses in the probe set, up to ±45°. Finally we remark that our approach can handle continuous pose variations and it is comparable with recent state-of-the-art approaches.
Iacopo Masi, Claudio Ferrari, Alberto Del Bimbo, Gérard G. Medioni
ICPR3
2014 Computing Local Binary Patterns on Discrete Manifolds
abstract
In this paper, we present a novel and original framework for computing Local Binary Pattern (LBP)-like patterns on a triangular mesh manifold. This framework, that we called mesh-LBP, can be adapted to all the LBP variants employed in 2D image analysis. As such, it allows extending the related techniques to mesh surfaces. First, we describe the foundations, the construction and the features of the mesh-LBP. In the experiments, we first show evidence of the presence of the "uniformity" aspect in the mesh-LBP patterns. Then, we report about the application of mesh-LBP to the problem of 3D texture-classification in comparison to standard 3D surface descriptors and show the mesh-LBP robustness to mesh irregularities.
Naoufel Werghi, Stefano Berretti, Alberto Del Bimbo
ICPR3
2014 A Cross-media Model for Automatic Image Annotation
abstract
Automatic image annotation is still an important open problem in multimedia and computer vision. The success of media sharing websites has led to the availability of large collections of images tagged with human-provided labels. Many approaches previously proposed in the literature do not accurately capture the intricate dependencies between image content and annotations. We propose a learning procedure based on Kernel Canonical Correlation Analysis which finds a mapping between visual and textual words by projecting them into a latent meaning space. The learned mapping is then used to annotate new images using advanced nearest-neighbor voting methods. We evaluate our approach on three popular datasets, and show clear improvements over several approaches relying on more standard representations.
Lamberto Ballan, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo
ICMR4
2014 Loki+Lire: a framework to create web-based multimedia search engines
abstract
In this paper we present Loki+Lire, a framework for the creation of web-based interfaces for search, annotation and presentation of multimedia data. The framework provides tools to ingest, transcode, present, annotate and index different types of media such as images, videos, audio files and textual documents. The front-end is compliant with the latest HTML5 standards, while the back-end allows system administrators to create processing pipelines that can be adapted for different tasks and purposes.
Giuseppe Becchi, Marco Bertini 0001, Lorenzo Cioni, Alberto Del Bimbo, Andrea Ferracani, Daniele Pezzatini, Mathias Lux
ACM Multimedia4
2014 Information theoretic sensor management for multi-target tracking with a single pan-tilt-zoom camera
abstract
Automatic multiple target tracking with pan-tilt-zoom (PTZ) cameras is a hard task, with few approaches in the literature, most of them proposing simplistic scenarios. In this paper, we present a PTZ camera management framework which lies on information theoretic principles: at each time step, the next camera pose (pan, tilt, focal length) is chosen, according to a policy which ensures maximum information gain. The formulation takes into account occlusions, physical extension of targets, realistic pedestrian detectors and the mechanical constraints of the camera. Convincing comparative results on synthetic data, realistic simulations and the implementation on a real video surveillance camera validate the effectiveness of the proposed method.
Pietro Salvagnini, Federico Pernici, Marco Cristani, Giuseppe Lisanti, Iacopo Masi, Alberto Del Bimbo, Vittorio Murino
WACV6
2014 Special Issue on Large-Scale Computer Vision: Geometry, Inference, and Learning
Roberto Cipolla, Carlo Colombo, Alberto Del Bimbo
Int. J. Comput. Vis.3
2014 Preface: Internet multimedia computing and service
Shuqiang Jiang, Changsheng Xu, Yong Rui, Alberto Del Bimbo, Hongxun Yao
Multim. Tools Appl.4
2014 Object Tracking by Oversampling Local Features
abstract
In this paper, we present the ALIEN tracking method that exploits oversampling of local invariant representations to build a robust object/context discriminative classifier. To this end, we use multiple instances of scale invariant local features weakly aligned along the object template. This allows taking into account the 3D shape deviations from planarity and their interactions with shadows, occlusions, and sensor quantization for which no invariant representations can be defined. A non-parametric learning algorithm based on the transitive matching property discriminates the object from the context and prevents improper object template updating during occlusion. We show that our learning rule has asymptotic stability under mild conditions and confirms the drift-free capability of the method in long-term tracking. A real-time implementation of the ALIEN tracker has been evaluated in comparison with the state-of-the-art tracking systems on an extensive set of publicly available video sequences that represent most of the critical conditions occurring in real tracking environments. We have reported superior or equal performance in most of the cases and verified tracking with no drift in very long video sequences.
Federico Pernici, Alberto Del Bimbo
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Local Pyramidal Descriptors for Image Recognition
abstract
In this paper, we present a novel method to improve the flexibility of descriptor matching for image recognition by using local multiresolution pyramids in feature space. We propose that image patches be represented at multiple levels of descriptor detail and that these levels be defined in terms of local spatial pooling resolution. Preserving multiple levels of detail in local descriptors is a way of hedging one's bets on which levels will most relevant for matching during learning and recognition. We introduce the Pyramid SIFT (P-SIFT) descriptor and show that its use in four state-of-the-art image recognition pipelines improves accuracy and yields state-of-the-art results. Our technique is applicable independently of spatial pyramid matching and we show that spatial pyramids can be combined with local pyramids to obtain further improvement. We achieve state-of-the-art results on Caltech-101 (80.1%) and Caltech-256 (52.6%) when compared to other approaches based on SIFT features over intensity images. Our technique is efficient and is extremely easy to integrate into image recognition pipelines.
Lorenzo Seidenari, Giuseppe Serra 0001, Andrew D. Bagdanov, Alberto Del Bimbo
IEEE Trans. Pattern Anal. Mach. Intell.4
2014 Leveraging local neighborhood topology for large scale person re-identification
Svebor Karaman, Giuseppe Lisanti, Andrew D. Bagdanov, Alberto Del Bimbo
Pattern Recognit.4
2014 Special Issue on ICPR 2012 Awarded Papers
Kim Boyer, Horst Bunke, Alberto Del Bimbo, Katsushi Ikeuchi, Ken-ichi Maeda
Pattern Recognit. Lett.3
2014 Face Recognition by Super-Resolved 3D Models From Consumer Depth Cameras
abstract
Face recognition based on the analysis of 3D scans has been an active research subject over the last few years. However, the impact of the resolution of 3D scans on the recognition process has not been addressed explicitly, yet being an element of primal importance to enable the use of the new generation of consumer depth cameras for biometric purposes. In fact, these devices perform depth/color acquisition over time at standard frame-rate, but with a low resolution compared to the 3D scanners typically used for acquiring 3D faces in recognition applications. Motivated by these considerations, in this paper, we define a super-resolution approach for 3D faces by which a sequence of low-resolution 3D face scans is processed to extract a higher resolution 3D face model. The proposed solution relies on the scaled iterative closest point procedure to align the low-resolution scans with each other, and estimates the value of the high-resolution 3D model through a 2D box-spline functions approximation. To evaluate the approach, we built-and made it publicly available-the Florence Superface dataset that collects high-resolution and low-resolution data for about 50 different persons. Qualitative and quantitative results are reported to demonstrate the accuracy of the proposed solution, also in comparison with alternative techniques.
Stefano Berretti, Pietro Pala, Alberto Del Bimbo
IEEE Trans. Inf. Forensics Secur.3
2014 Guest Editorial Special Section on Socio-Mobile Media Analysis and Retrieval
abstract
The four papers in this special section cover several hot and emerging topics in socio-mobile media analysis and retrieval.
Alberto Del Bimbo, K. Selçuk Candan, Yu-Gang Jiang 0001, Jiebo Luo 0001, Tao Mei 0001, Nicu Sebe, Heng Tao Shen, Cees Snoek
IEEE Trans. Multim.1
2014 Selecting stable keypoints and local descriptors for person identification using 3D face scans
Stefano Berretti, Naoufel Werghi, Alberto Del Bimbo, Pietro Pala
Vis. Comput.3
2013 Local descriptors matching for 3D face recognition
abstract
An original solution to 3D face recognition, which supports face matching also in the case of probes with varying expressions and missing parts is proposed in this work. Distinguishing traits of the face are captured by first extracting 3D keypoints of the face scan, then measuring how the face surface changes in the neighborhood of the keypoints using a local descriptor. To this end, an adaptation of the meshDOG detector to the case of 3D faces is proposed, together with a multi-ring geometric histogram descriptor. Face similarity is then evaluated by comparing local keypoint descriptors across inlier pairs of matching keypoints between probe and gallery scans. Experiments have been performed on the Bosphorus database, showing competitive results with respect to existing solutions for 3D face biometrics.
Naoufel Werghi, Stefano Berretti, Alberto Del Bimbo, Pietro Pala
ICIP3
2013 An evaluation of nearest-neighbor methods for tag refinement
abstract
The success of media sharing and social networks has led to the availability of extremely large quantities of images that are tagged by users. The need of methods to manage efficiently and effectively the combination of media and metadata poses significant challenges. In particular, automatic image annotation of social images has become an important research topic for the multimedia community. In this paper we propose and thoroughly evaluate the use of nearest-neighbor methods for tag refinement. Extensive and rigorous evaluation using two standard large-scale datasets shows that the performance of these methods is comparable with that of more complex and computationally intensive approaches and that, differently from these latter approaches, nearest-neighbor methods can be applied to `web-scale' data.
Tiberio Uricchio, Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo
ICME4
2013 A novel framework for collaborative video recommendation, interest discovery and friendship suggestion based on semantic profiling
abstract
Two important challenges for social networks are the creation of targeted and personalized content for their users, selecting the most interesting material from the huge amount of user-generated content, and keeping user engagement , e.g. through creation and curation of users' profiles. In this demo we show a system for video commenting, sharing and interest discovery that combines recommendation algorithms, clustering techniques, tools for video tagging and evaluation of semantic resources relatedness. Combining these tools and techniques it becomes possible to provide personalized multimedia services and to improve and propagate interests and inter-personal connections through the network.
Marco Bertini 0001, Alberto Del Bimbo, Andrea Ferracani, Francesco Gelli, Daniele Maddaluno, Daniele Pezzatini
ACM Multimedia2
2013 euTV: a system for media monitoring and publishing
abstract
In this paper, we describe the euTV system, which provides a flexible approach to collect, manage, annotate and publish collections of images, videos and textual documents. The system is based on a Service Oriented Architecture that allows to combine and orchestrate a large set of web services for automatic and manual annotation, retrieval, browsing, ingestion and authoring of multimedia sources. euTV tools have been used to create several publicly available vertical applications, addressing different use cases. Positive results of user evaluations have shown that the system can be effectively used to create different types of applications.
Marco Bertini 0001, Alberto Del Bimbo, George Ioannidis, Emile Bijk, Isabel Trancoso, Hugo Meinedo
ACM Multimedia2
2013 Flarty: recommending art routes using check-ins latent topics
abstract
In this demo we present Flarty, a mobile location-based social network for the dynamic construction and recommendation of art routes in the city of Florence, Italy, via item based similarity algorithms, places topic extraction and user interest modeling. To achieve this goal Flarty derives knowledge from users check-ins and combines clustering techniques and recommendation algorithms, as well as features such as geo-location, to define groups of similar artworks or POIs (Points Of Interest) and to compute the most efficient routes likely to meet user's interests. Model analysis takes into account ratings, topics extracted from textual features associated with the POIs, and users preferences computed exploiting collaborative filtering techniques on their past behavior.
Alberto Del Bimbo, Andrea Ferracani, Daniele Pezzatini
ACM Multimedia1
2013 Matching 3D face scans using interest points and local histogram descriptors
abstract
In this work, we propose and experiment an original solution to 3D face recognition that supports face matching also in the case of probe scans with missing parts. In the proposed approach, distinguishing traits of the face are captured by first extracting 3D keypoints of the scan and then measuring how the face surface changes in the keypoints neighborhood using local shape descriptors. In particular: 3D keypoints detection relies on the adaptation to the case of 3D faces of the meshDOG algorithm that has been demonstrated to be effective for 3D keypoints extraction from generic objects; as 3D local descriptors we used the HOG descriptor and also proposed two alternative solutions that develop, respectively, on the histogram of orientations and the geometric histogram descriptors. Face similarity is evaluated by comparing local shape descriptors across inlier pairs of matching keypoints between probe and gallery scans. The face recognition accuracy of the approach has been first experimented on the difficult probes included in the new 2D/3D Florence face dataset that has been recently collected and released at the University of Firenze, and on the Binghamton University 3D facial expression dataset. Then, a comprehensive comparative evaluation has been performed on the Bosphorus, Gavab and UND/FRGC v2.0 databases, where competitive results with respect to existing solutions for 3D face biometrics have been obtained.
Stefano Berretti, Naoufel Werghi, Alberto Del Bimbo, Pietro Pala
Comput. Graph.3
2013 Interactive multi-user video retrieval systems
Marco Bertini 0001, Alberto Del Bimbo, Andrea Ferracani, Lea Landucci, Daniele Pezzatini
Multim. Tools Appl.2
2013 Copy-move forgery detection and localization by means of robust clustering with J-Linkage
Irene Amerini, Lamberto Ballan, Roberto Caldelli, Alberto Del Bimbo, Luca Del Tongo, Giuseppe Serra 0001
Signal Process. Image Commun.4
2013 Sparse Matching of Salient Facial Curves for Recognition of 3-D Faces With Missing Parts
abstract
In this work, we propose and experiment a 3-D face recognition approach capable of performing accurate face matching also in the case where just parts of probe scans are available. This is obtained through an original face representation and matching solution that first extracts keypoints of the 3-D depth image of the face and then measures how the face depth changes along facial curves connecting pairs of keypoints. Face similarity is evaluated by sparse comparison of facial curves defined across inlier pairs of matching keypoints between probe and gallery scans. In doing so, a statistical model is also proposed to associate facial curves of the gallery scans with a saliency measure so that curves that model characterizing traits of some subjects are distinguished from curves that are frequently observed in the face of many different subjects. Following recent related work, the recognition accuracy of the approach is experimented using two datasets, both comprising scans with missing parts: the Face Recognition Grand Challenge v2.0 dataset combined with the University of Notre Dame probes; the Gavab dataset.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
IEEE Trans. Inf. Forensics Secur.2
2013 Context-Dependent Logo Matching and Recognition
abstract
We contribute, through this paper, to the design of a novel variational framework able to match and recognize multiple instances of multiple reference logos in image archives. Reference logos and test images are seen as constellations of local features (interest points, regions, etc.) and matched by minimizing an energy function mixing: 1) a fidelity term that measures the quality of feature matching, 2) a neighborhood criterion that captures feature co-occurrence/geometry, and 3) a regularization term that controls the smoothness of the matching solution. We also introduce a detection/recognition procedure and study its theoretical consistency. Finally, we show the validity of our method through extensive experiments on the challenging MICC-Logos dataset. Our method overtakes, by 20%, baseline as well as state-of-the-art matching/recognition procedures.
Hichem Sahbi, Lamberto Ballan, Giuseppe Serra 0001, Alberto Del Bimbo
IEEE Trans. Image Process.4
2013 Automatic facial expression recognition in real-time from dynamic sequences of 3D face scans
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
Vis. Comput.2
2012 Multi-pose face detection for accurate face logging
Andrew D. Bagdanov, Alberto Del Bimbo, Giuseppe Lisanti, Iacopo Masi
ICPR2
2012 Real-time hand status recognition from RGB-D imagery
Andrew D. Bagdanov, Alberto Del Bimbo, Lorenzo Seidenari, Lorenzo Usai
ICPR2
2012 Combining generative and discriminative models for classifying social images from 101 object categories
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Andrea M. Serain, Giuseppe Serra 0001, Benito F. Zaccone
ICPR3
2012 Social and automatic annotation of videos for semantic profiling and content discovery
abstract
This demo presents a system based on social relationships, social knowledge and automatic video and textual content analysis for the discovery of videos in social networks. The system, developed as a web application, allows users to annotate, manually and automatically, and comment video frames and scenes enriching their content with tags, references to Facebook users and pages and Wikipedia resources. These annotations are used to semantically model the profile of each user extracting and expanding his interests and folksonomy, as well as resources of interest in his social graph. The automatically generated profile page is used to suggest to users new resources, Facebook friends and videos whose content is related to their interests and allows profile curation. A screencast showing an example of these functionalities is publicly available at: http://vimeo.com/miccunifi/facetube
Marco Bertini 0001, Alberto Del Bimbo, Andrea Ferracani, Daniele Pezzatini
ACM Multimedia2
2012 Indoor and outdoor profiling of users in multimedia installations
abstract
We present a work-in-progress interactive exhibit for the mu- seum of Onna (L'Aquila, Italy). The Onna Onlus and the citizens of Onna have the ambitious project to rebuild the town affected by the earthquake of April 2009 and to create a museum in memory of Onna. In addition to the tradi- tional fruition tools for museums, we have been asked for an interactive system capable to communicate the past events and the efforts invested. We are working on a multi-modal system composed of an indoor environment in which visitors can interact with a natural interface system and an outdoor module based on a cross-platform mobile application. A pro- filing method is also exploited in order to extract a profile of interest of each visitor and then use it to suggest in-depth personalized and geo-located information about the disaster and the local history via multimedia contents.
Gianpaolo D'Amico, Alberto Del Bimbo, Andrea Ferracani, Lea Landucci, Daniele Pezzatini
ACM Multimedia2
2012 Multi-scale and real-time non-parametric approach for anomaly detection and localization
Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari
Comput. Vis. Image Underst.2
2012 LIT: transcription, annotation, search and visualization tools for the Lexicon of the Italian Television
Thomas M. Alisi, Alberto Del Bimbo, Andrea Ferracani, Tiberio Uricchio, Ervin Hoxha, Besmir Bregasi
Multim. Tools Appl.2
2012 Intelligent multimedia interactivity
Ling Shao 0001, Qi Tian 0001, Alberto Del Bimbo, Changsheng Xu
Pattern Recognit. Lett.3
2012 Distinguishing Facial Features for Ethnicity-Based 3D Face Recognition
abstract
Among different approaches for 3D face recognition, solutions based on local facial characteristics are very promising, mainly because they can manage facial expression variations by assigning different weights to different parts of the face. However, so far, a few works have investigated the individual relevance that local features play in 3D face recognition with very simple solutions applied in the practice. In this article, a local approach to 3D face recognition is combined with a feature selection model to study the relative relevance of different regions of the face for the purpose of discriminating between different subjects. The proposed solution is experimented using facial scans of the Face Recognition Grand Challenge dataset. Results of the experimentation are two-fold: they quantitatively demonstrate the assumption that different regions of the face have different relevance for face discrimination and also show that the relevance of facial regions changes for different ethnic groups.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
ACM Trans. Intell. Syst. Technol.2
2012 Introduction to the Special Section on Intelligent Multimedia Systems and Technology Part II
abstract
No abstract available.
Xian-Sheng Hua 0001, Qi Tian 0001, Alberto Del Bimbo, Ramesh Jain 0001
ACM Trans. Intell. Syst. Technol.3
2012 Effective Codebooks for Human Action Representation and Classification in Unconstrained Videos
abstract
Recognition and classification of human actions for annotation of unconstrained video sequences has proven to be challenging because of the variations in the environment, appearance of actors, modalities in which the same action is performed by different persons, speed and duration, and points of view from which the event is observed. This variability reflects in the difficulty of defining effective descriptors and deriving appropriate and effective codebooks for action categorization. In this paper, we propose a novel and effective solution to classify human actions in unconstrained videos. It improves on previous contributions through the definition of a novel local descriptor that uses image gradient and optic flow to respectively model the appearance and motion of human actions at interest point regions. In the formation of the codebook, we employ radius-based clustering with soft assignment in order to create a rich vocabulary that may account for the high variability of human actions. We show that our solution scores very good performance with no need of parameter tuning. We also show that a strong reduction of computation time can be obtained by applying codebook size reduction with Deep Belief Networks with little loss of accuracy.
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari, Giuseppe Serra 0001
IEEE Trans. Multim.3
2011 Continuous recovery for real time pan tilt zoom localization and mapping
abstract
We propose a method for real time recovering from tracking failure in monocular localization and mapping with a Pan Tilt Zoom camera (PTZ). The method automatically detects and seamlessly recovers from tracking failure while preserving map integrity. By extending recent advances in the PTZ localization and mapping, the system can quickly and continuously resume tracking failures by determining the best way to task two different localization modalities. The tradeoff involved when choosing between the two modalities is captured by maximizing the information expected to be extracted from the scene map. This is especially helpful in four main viewing condition: blurred frames, weak textured scene, not up to date map and occlusions due to sensor quantization or moving objects. Extensive tests show that the resulting system is able to recover from several different failures while zooming-in weak textured scene, all in real time.
Alberto Del Bimbo, Giuseppe Lisanti, Iacopo Masi, Federico Pernici
AVSS1
2011 A web system for ontology-based multimedia annotation, browsing and search
abstract
In this paper we present a complete system for semantic and syn tactic annotation, browsing and search of multimedia data, that is based on a service oriented architecture, with web-based interfaces developed following the Rich Internet Application paradigm. The system has been designed to be: i) flexible and extendable, allowing users to select only the services they need or to add their own tools to the multimedia processing pipelines; ii) distributed, with services that can be executed in a cloud computing infrastructure and accessed through web applications; Hi) user-friendly, with interfaces that have a uniform interface on every platform and that have an interaction level similar to that of desktop applications. Extensive user trials in real-world setup, performed by archive and broadcaster professionals, have shown the efficacy and usability of the proposed solution.
Marco Bertini 0001, Giuseppe Becchi, Alberto Del Bimbo, Andrea Ferracani, Daniele Pezzatini
ICME3
2011 Adaptive Video Compression for Video Surveillance Applications
abstract
This article describes an approach to adaptive video coding for video surveillance applications. Using a combination of low-level features with low computational cost, we show how it is possible to control the quality of video compression so that semantically meaningful elements of the scene are encoded with higher fidelity, while background elements are allocated fewer bits in the transmitted representation. Our approach is based on adaptive smoothing of individual video frames so that image features highly correlated to semantically interesting objects are preserved. Using only low-level image features on individual frames, this adaptive smoothing can be seamlessly inserted into a video coding pipeline as a pre-processing state. Experiments show that our technique is efficient, outperforms standard H.264 encoding at comparable bit rates, and preserves features critical for downstream detection and recognition.
Andrew D. Bagdanov, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari
ISM3
2011 RFID-based Solutions for User Profiling in Interactive Exhibits
abstract
In this paper we present a work-in-progress interactive exhibit for the museum of Onna, a little town near to L'Aquila (Italy), almost completely destroyed by the earthquake of April 2009. The installation will be developed as an environment in which visitors of the museum can interact with a natural interaction system and then discover the history of the disaster via rich multimedia contents. Visitors are detected through the adoption of an RFID-based technology, which allows to store their interaction history and build an interest profile used to enrich the experience. Different scenarios have been implemented and tested in order to evaluate the effectiveness of the proposed solution.
Gianpaolo D'Amico, Alberto Del Bimbo, Andrea Ferracani, Lea Landucci, Daniele Pezzatini, Luca Santi
ISM2
2011 A flexible environment for multimedia management and publishing
abstract
In this paper, we describe the IM3I system, which provides a flexible approach to managing and publishing collections of images and videos. The system is based on web services that allow automatic and manual annotation, retrieval, browsing and authoring of multimedia. Results of user evaluations, performed by professional archivists and archive managers on a real-world system deployment have confirmed that the system is easy to be used and delivers a complete set of functionalities.
Marco Bertini 0001, Alberto Del Bimbo, George Ioannidis, Alexandru Stan, Emile Bijk
ICMR2
2011 Enriching and localizing semantic tags in internet videos
abstract
Tagging of multimedia content is becoming more and more widespread as web 2.0 sites, like Flickr and Facebook for images, YouTube and Vimeo for videos, have popularized tagging functionalities among their users. These user-generated tags are used to retrieve multimedia content, and to ease browsing and exploration of media collections, e.g.~using tag clouds. However, not all media are equally tagged by users: using the current browsers is easy to tag a single photo, and even tagging a part of a photo, like a face, has become common in sites like Flickr and Facebook; on the other hand tagging a video sequence is more complicated and time consuming, so that users just tag the overall content of a video. In this paper we present a system for automatic video annotation that increases the number of tags originally provided by users, and localizes them temporally, associating tags to shots. This approach exploits collective knowledge embedded in tags and Wikipedia, and visual similarity of keyframes and images uploaded to social sites like YouTube and Flickr.
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Giuseppe Serra 0001
ACM Multimedia3
2011 Joint ACM workshop on human gesture and behavior understanding: (J-HGBU'11)
abstract
The ability to understand social signals of a person we are communicating with is the core of social intelligence. Social Intelligence is a facet of human intelligence that has been argued to be indispensable and perhaps the most important for success in life. At the same time, human-centric multimedia applications for humans and about humans are becoming increasingly important. 3D modeled human-objects, like bodies, heads and faces are exploited for animation, security, and human computer interaction, while three dimensional motion of arms, legs and local body features is used for more complete human gesture, activity and behavior analysis. The Joint Human Gesture and Behavior Understanding (J-HGBU) workshop event consists of two parts focusing on these complementary challenges: the Workshop on Multimedia Access to 3D Human Objects (MA3HO'11) and the Workshop on Social Signal Processing (SSPW'11).
Maja Pantic, Alex Pentland, Alessandro Vinciarelli, Rita Cucchiara, Mohamed Daoudi, Alberto Del Bimbo
ACM Multimedia6
2011 Particle filter-based visual tracking with a first order dynamic model and uncertainty adaptation
Alberto Del Bimbo, Fabrizio Dini
Comput. Vis. Image Underst.1
2011 Shape reconstruction and texture sampling by active rectification and virtual view synthesis
Carlo Colombo, Dario Comanducci, Alberto Del Bimbo
Comput. Vis. Image Underst.3
2011 Event detection and recognition for semantic annotation of video
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari, Giuseppe Serra 0001
Multim. Tools Appl.3
2011 Survey papers in multimedia - guest editorial
Ramesh Jain 0001, Alberto Del Bimbo, Tat-Seng Chua, Borko Furht
Multim. Tools Appl.2
2011 Hot research topics - guest editorial
Ramesh Jain 0001, Alberto Del Bimbo, Tat-Seng Chua, Borko Furht
Multim. Tools Appl.2
2011 A SIFT-Based Forensic Method for Copy-Move Attack Detection and Transformation Recovery
abstract
One of the principal problems in image forensics is determining if a particular image is authentic or not. This can be a crucial task when images are used as basic evidence to influence judgment like, for example, in a court of law. To carry out such forensic analysis, various technological instruments have been developed in the literature. In this paper, the problem of detecting if an image has been forged is investigated; in particular, attention has been paid to the case in which an area of an image is copied and then pasted onto another zone to create a duplication or to cancel something that was awkward. Generally, to adapt the image patch to the new context a geometric transformation is needed. To detect such modifications, a novel methodology based on scale invariant features transform (SIFT) is proposed. Such a method allows us to both understand if a copy-move attack has occurred and, furthermore, to recover the geometric transformation used to perform cloning. Extensive experimental results are presented to confirm that the technique is able to precisely individuate the altered area and, in addition, to estimate the geometric transformation parameters with high reliability. The method also deals with multiple cloning.
Irene Amerini, Lamberto Ballan, Roberto Caldelli, Alberto Del Bimbo, Giuseppe Serra 0001
IEEE Trans. Inf. Forensics Secur.4
2011 Introduction to the special issue on intelligent multimedia systems and technology
abstract
No abstract available.
Xian-Sheng Hua 0001, Qi Tian 0001, Alberto Del Bimbo, Ramesh Jain 0001
ACM Trans. Intell. Syst. Technol.3
2011 3D facial expression recognition using SIFT descriptors of automatically detected keypoints
Stefano Berretti, Boulbaba Ben Amor, Mohamed Daoudi, Alberto Del Bimbo
Vis. Comput.4
2010 Geometric tampering estimation by means of a SIFT-based forensic analysis
abstract
In many application scenarios digital images play a basic role and often it is important to assess if their content is realistic or has been manipulated to mislead watcher's opinion. Image forensics tools provide answers to similar questions. This paper, in particular, focuses on the problem of detecting if a feigned image has been created by cloning an area of the image onto another zone to make a duplication or to cancel something awkward. The proposed method is based on SIFT features and allows both to understand which are the image points involved in the counterfeit attack and, furthermore, to recover the parameters of the geometric transformation. Experimental results are provided to witness the powerfulness of the proposed technique.
Irene Amerini, Lamberto Ballan, Roberto Caldelli, Alberto Del Bimbo, Giuseppe Serra 0001
ICASSP4
2010 A Set of Selected SIFT Features for 3D Facial Expression Recognition
abstract
In this paper, the problem of person-independent facial expression recognition is addressed on 3D shapes. To this end, an original approach is proposed that computes SIFT descriptors on a set of facial landmarks of depth images, and then selects the subset of most relevant features. Using SVM classification of the selected features, an average recognition rate of 77.5% on the BU-3DFE database has been obtained. Comparative evaluation on a common experimental setup, shows that our solution is able to obtain state of the art results.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala, Boulbaba Ben Amor, Mohamed Daoudi
ICPR2
2010 Sensor Fusion for Cooperative Head Localization
abstract
In modern video surveillance systems, pan-tilt-zoom (PTZ) cameras certainly have the potential to allow the coverage of wide areas with a much smaller number of sensors, compared to the common approach of fixed camera networks. This paper describes a general framework that aims at exploiting the capabilities of modern PTZ cameras in order to acquire high resolution images of body parts, such as the head, from the observation of pedestrians moving in a wide outdoor area. The framework allows to organize the sensors in a network with arbitrary topology, and to establish pairwise master-slave relationship between them. In this way a slave camera can be steered to acquire imagery of a target keeping into account both target and zooming uncertainties. Experiments show good performance in localizing target's head, independently from the zooming factor of the slave camera.
Alberto Del Bimbo, Fabrizio Dini, Giuseppe Lisanti, Federico Pernici
ICPR1
2010 Person Detection Using Temporal and Geometric Context with a Pan Tilt Zoom Camera
abstract
In this paper we present a system that integrates automatic camera geometry estimation and object detection from a Pan Tilt Zoom camera. We estimate camera pose with respect to a world scene plane in real-time and perform human detection exploiting the relative space-time context. Using camera self-localization, 2D object detections are clustered in a 3D world coordinate frame. Target scale inference is further exploited to reduce the number of false alarms and to increase also the detection rate in the final non-maximum suppression stage. Our integrated system applied on real-world data shows superior performance with respect to the standard detector used.
Alberto Del Bimbo, Giuseppe Lisanti, Iacopo Masi, Federico Pernici
ICPR1
2010 Exploiting distinctive visual landmark maps in pan-tilt-zoom camera networks
Alberto Del Bimbo, Fabrizio Dini, Giuseppe Lisanti, Federico Pernici
Comput. Vis. Image Underst.1
2010 Video event classification using string kernels
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Giuseppe Serra 0001
Multim. Tools Appl.3
2010 Semantic annotation of soccer videos by visual instance clustering and spatial/temporal reasoning in ontologies
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Giuseppe Serra 0001
Multim. Tools Appl.3
2010 3D Face Recognition Using Isogeodesic Stripes
abstract
In this paper, we present a novel approach to 3D face matching that shows high effectiveness in distinguishing facial differences between distinct individuals from differences induced by nonneutral expressions within the same individual. The approach takes into account geometrical information of the 3D face and encodes the relevant information into a compact representation in the form of a graph. Nodes of the graph represent equal width isogeodesic facial stripes. Arcs between pairs of nodes are labeled with descriptors, referred to as 3D Weighted Walkthroughs (3DWWs), that capture the mutual relative spatial displacement between all the pairs of points of the corresponding stripes. Face partitioning into isogeodesic stripes and 3DWWs together provide an approximate representation of local morphology of faces that exhibits smooth variations for changes induced by facial expressions. The graph-based representation permits very efficient matching for face recognition and is also suited to being employed for face identification in very large data sets with the support of appropriate index structures. The method obtained the best ranking at the SHREC 2008 contest for 3D face recognition. We present an extensive comparative evaluation of the performance with the FRGC v2.0 data set and the SHREC08 data set.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Matching Trajectories between Video Sequences by Exploiting a Sparse Projective Invariant Representation
abstract
Identifying correspondences between trajectory segments observed from nonsynchronized cameras is important for reconstruction of the complete trajectory of moving targets in a large scene. Such a reconstruction can be obtained from motion data by comparing the trajectory segments and estimating both the spatial and temporal alignments. Exhaustive testing of all possible correspondences of trajectories over a temporal window is only viable in the cases with a limited number of moving targets and large view overlaps. Therefore, alternative solutions are required for situations with several trajectories that are only partially visible in each view. In this paper, we propose a new method that is based on view-invariant representation of trajectories, which is used to produce a sparse set of salient points for trajectory segments observed in each view. Only the neighborhoods at these salient points in the view--invariant representation are then used to estimate the spatial and temporal alignment of trajectory pairs in different views. It is demonstrated that, for planar scenes, the method is able to recover with good precision and efficiency both spatial and temporal alignments, even given relatively small overlap between views and arbitrary (unknown) temporal shifts of the cameras. The method also provides the same capabilities in the case of trajectories that are only locally planar, but exhibit some nonplanarity at a global level.
Walter Nunziati, Stan Sclaroff, Alberto Del Bimbo
IEEE Trans. Pattern Anal. Mach. Intell.3
2009 Recognizing human actions by fusing spatio-temporal appearance and motion descriptors
abstract
In this paper we propose a new method for human action categorization by using an effective combination of a new 3D gradient descriptor with an optic flow descriptor, to represent spatio-temporal interest points. These points are used to represent video sequences using a bag of spatio-temporal visual words, following the successful results achieved in object and scene classification. We extensively test our approach on the standard KTH and Weizmann actions datasets, showing its validity and good performance. Experimental results outperform state-of-the-art methods, without requiring fine parameter tuning.
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari, Giuseppe Serra 0001
ICIP3
2009 Deep networks for audio event classification in soccer videos
abstract
In this work is presented a novel approach for the classification of audio concepts in broadcast soccer videos using deep belief network (DBN), a probabilistic neural network with several hidden layers. Comparison with support vector machine (SVM) classifiers has been carried on, showing that our preliminary results are promisingly comparable to the state-of-the-art.
Lamberto Ballan, Alessio Bazzica, Marco Bertini 0001, Alberto Del Bimbo, Giuseppe Serra 0001
ICME4
2009 David: Discriminant analysis for verification of monuments in image data
abstract
In the past few years, several research works have addressed the problems posed by vision-assisted navigation systems. Basically, these systems allow a tourists navigating in an urban environment to take the photograph of a scene and submit it to the navigation system that will recognize what monument is represented in the photograph. However, solutions proposed so far focus on the recognition of approximately planar structures (e.g. building facades) and are inadequate to support the recognition task if the object of interest has a generic 3D structure such as a statue. In this paper we present a model to support the recognition of generic monuments that is based on the information-theoretic notion of mutual information to quantify the saliency of invariant local descriptors. Results are demonstrated on a real-database also in comparison with the baseline method that performs matching without any measure of saliency.
Alberto Del Bimbo, Walter Nunziati, Pietro Pala
ICME1
2009 Arneb: a rich internet application for ground truth annotation of videos
abstract
In this technical demonstration we show the current version of Arneb, a web-based system for manual annotation of videos, developed within the EU VidiVideo project. This tool has been developed with the aim of creating ground truth annotations, that can be used for training and evaluating automatic video annotation systems. Annotations can be exported to MPEG-7 and OWL ontologies. The system has been developed according to the Rich Internet Application paradigm, allowing collaborative web-based annotation.
Thomas M. Alisi, Marco Bertini 0001, Gianpaolo D'Amico, Alberto Del Bimbo, Andrea Ferracani, Federico Pernici, Giuseppe Serra 0001
ACM Multimedia4
2009 Sirio: an ontology-based web search engine for videos
abstract
In this technical demonstration we show a web video search engine based on ontologies, the Sirio system, that has been developed within the EU VidiVideo project. The goal of the system is to provide a search engine for videos for both technical and non-technical users. In fact, the system has different interfaces that permit different query modalities: free-text, natural language, graphical composition of concepts using boolean and temporal relations and query by visual example. In addition, the ontology structure is exploited to encode semantic relations between concepts permitting, for example, to expand queries to synonyms and concept specializations.
Thomas M. Alisi, Marco Bertini 0001, Gianpaolo D'Amico, Alberto Del Bimbo, Andrea Ferracani, Federico Pernici, Giuseppe Serra 0001
ACM Multimedia4
2009 3D Mesh decomposition using Reeb graphs
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
Image Vis. Comput.2
2009 Introduction to the special section for the best papers of ACM multimedia 2008
abstract
introduction Share on Introduction to the special section for the best papers of ACM multimedia 2008 Authors: K. Selçuk Candan Arizona State University, USA Arizona State University, USAView Profile , Alberto Del Bimbo Università degli Studi di Firenze, Italy Università degli Studi di Firenze, ItalyView Profile , Carsten Griwodz Simula Research Laboratory, Norway Simula Research Laboratory, NorwayView Profile , Alejandro Jaimes Telefonica Research, Spain Telefonica Research, SpainView Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 5Issue 3August 2009 Article No.: 18pp 1–3https://doi.org/10.1145/1556134.1556135Published:14 August 2009Publication History 0citation329DownloadsMetricsTotal Citations0Total Downloads329Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
K. Selçuk Candan, Alberto Del Bimbo, Carsten Griwodz, Alejandro Jaimes
ACM Trans. Multim. Comput. Commun. Appl.2
2008 Uncalibrated Framework for On-line Camera Cooperation to Acquire Human Head Imagery in Wide Areas
abstract
This paper considers the problem of estimating on-line the time-variant transformation relating a person's feet position in the image of a first, fixed camera, to his head position in the image of a second, pan-tilt-zoom camera. The transformation allows to acquire high-resolution images by steering the PTZ camera at targets detected in a fixed camera view. Assuming a planar scene and modeling humans as vertical segments, we present the development of an uncalibrated framework which does not require any 3D known location to be specified, and it allows to take into account both zooming camera and target uncertainties. Results show good performances in slave camera target head localization, degrading when the high zoom factor causes a lack of feature points in the slave camera.
Alberto Del Bimbo, Fabrizio Dini, Andrea Grifoni, Federico Pernici
AVSS1
2008 Automatic trademark detection and recognition in sport videos
abstract
In this paper we describe a system for automatic detection and recognition of trademarks in sports videos. We propose a compact representation of trademarks based on SIFT feature points and a matching algorithm to robustly detect and retrieve trademarks in a variety of different sports video types. Trademark localization is performed through robust clustering of matched feature points in the video frame. A supervised machine learning approach is used to automatically adapt the similarity threshold used to assess the trademark matches. Experimental results are provided, along with an analysis of the precision and recall. Results show that our proposed technique is efficient and effectively detects and classifies trademarks.
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Arjun Jain
ICME3
2008 Face recognition by SVMS classification of 2D and 3D Radial Geodesics
abstract
An original approach to represent 2D and 3D faces using radial geodesic distances (RGDs) is proposed in this work. In 3D, the RGD of a generic point of the face surface is computed as the length of the geodesic connecting the point with a reference point along a radial direction. In 2D, the RGD of a pixel with respect to a reference pixel accounts for the difference of gray level intensities of the two pixels and the Euclidean distance between them. Support Vector Machines (SVMs) are used to perform face recognition using 2D- and 3D-RGDs. Due to the high dimensionality of face representations based on RGDs, embedding into lower-dimensional spaces is applied before SVMs classification. Experimental results are reported for 3D-3D and 2D-3D face recognition using the proposed approach.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala, Francisco Silva-Mata
ICME2
2008 Connecting artists and scientists in multimedia research
abstract
Historically, the ACM Multimedia Conference is split into a "technical" program and an "arts" program. These programs sometimes seem completely separate from one another, victims of a "semantic gap" between disciplines. The goal of this panel is to create a space in which scientists learn from artists, and arts from science. We need to discover new connections between modalities of research. In order to create the most exciting and powerful future forms of interactive multimedia systems, the ones that will create the most beneficial broader impact on humanity, we need to foster new collaborations between artists and scientists. This panel seeks to bridge the great divide of language and communities that has fragmented us, creating a new space for developing connections between the arts and sciences of multimedia research, as embodied through the artists and scientists of ACM Multimedia. The goal is to make this conference a premier site for catalyzing emergent connections.
Andruid Kerne, Ron Wakkary, Frank Nack, Amanda Steggell, Alejandro Jaimes, K. Selçuk Candan, Alberto Del Bimbo, Pamela Jennings, Aleksandra Dulic
ACM Multimedia7
2008 SHREC'08 entry: 3D face recognition using integral shape information
abstract
In this work, we shortly describe an original 3D face recognition approach and its performance as resulted from the 3D Shape Retrieval Contest of 3D Face Scans organized by SHREC 2008 with the support of the Network of Excellence AIM@SHAPE. In particular, the evaluation shows that the proposed approach attains the highest performance on the SHREC 2008 data set.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
Shape Modeling International2
2008 Natural interaction on tabletops
Stefano Baraldi, Alberto Del Bimbo, Lea Landucci
Multim. Tools Appl.2
2008 Special issue on natural interaction
Alberto Del Bimbo
Multim. Tools Appl.1
2008 A real-time full body tracking and humanoid animation system
Carlo Colombo, Alberto Del Bimbo, Alessandro Valli
Parallel Comput.2
2008 Editorial: Introduction to the Special Issue on Multimedia Data Mining
abstract
The twelve papers in this special issue focus on multimedia data mining. The special issue evolved from a successful workshop organized in conjunction with the 2006 ACM KDD conference, but the special issue was open to the whole community.
Zhongfei Zhang, Florent Masseglia, Ramesh Jain 0001, Alberto Del Bimbo
IEEE Trans. Multim.4
2007 Improving the robustness of particle filter-based visual trackers using online parameter adaptation
abstract
In particle filter-based visual trackers, dynamic velocity components are typically incorporated into the state update equations. In these cases, there is a risk that the uncertainty in the model update stage can become amplified in unexpected and undesirable ways, leading to erroneous behavior of the tracker. Moreover, the use of a weak appearance model can make the estimates provided by the particle filter inaccurate. To deal with this problem, we propose a continuously adaptive approach to estimating uncertainty in the particle filter, one that balances the uncertainty in its static and dynamic elements. We provide quantitative performance evaluation of the resulting particle filter tracker on a set of ten video sequences. Results are reported in terms of a metric that can be used to objectively evaluate the performance of visual trackers. This metric is used to compare our modified particle filter tracker and the continuously adaptive mean shift tracker. Results show that the performance of the particle filter is significantly improved through adaptive parameter estimation, particularly in cases of occlusion and erratic, nonlinear target motion.
Andrew D. Bagdanov, Alberto Del Bimbo, Fabrizio Dini, Walter Nunziati
AVSS2
2007 Accurate self-calibration of two cameras by observations of a moving person on a ground plane
abstract
A calibration algorithm of two cameras using observations of a moving person is presented. Similar methods have been proposed for self-calibration with a single camera, but internal parameter estimation is only limited to the focal length. Recently it has been demonstrated that principal point supposed in the center of the image causes inaccuracy of all estimated parameters. Our method exploits two cameras, using image points of head and foot locations of a moving person, to determine for both cameras the focal length and the principal point. Moreover with the increasing number of cameras there is a demand of procedures to determine their relative placements. In this paper we also describe a method to find the relative position and orientation of two cameras: the rotation matrix and the translation vector which describe the rigid motion between the coordinate frames fixed in two cameras. Results in synthetic and real scenes are presented to evaluate the performance of the proposed method.
Tsuhan Chen, Alberto Del Bimbo, Federico Pernici, Giuseppe Serra 0001
AVSS2
2007 Compact representation and probabilistic classification of human actions in videos
abstract
This paper addresses the problem of classifying human actions in a video sequence. A representation eigenspace approach based on the PCA algorithm is used to train the classifier according to an incremental learning scheme based on a “one action, one eigenspace” approach. Before dimensionality reduction, a high dimensional description of each frame of the video sequence is constructed, based on foreground blob analysis. Classification is performed by matching incrementally the reduced representation of the test image sequence against each of the learned ones, and accumulating matching scores according to a probabilistic framework, until a decision is obtained. Experimental results with real video sequences are presented and discussed.
Carlo Colombo, Dario Comanducci, Alberto Del Bimbo
AVSS3
2007 Geodesic Distances for 3D-3D and 2D-3D Face Recognition
abstract
In this paper, we propose an original framework for representing 2D and 3D face information using geodesic distances. This aims to define a representation enabling the direct comparison between 2D face images of an individual against its 3D face model. This representation is extracted by measuring geodesic distances in 2D and 3D. In 3D, the geodesic distance between two points on a surface is computed as the length of the shortest path connecting the two points. In 2D, the geodesic distance between two pixels is computed based on the differences of gray level intensities along the segment connecting the two pixels. Experimental results are shown to demonstrate the viability of the proposed solution.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala, Francisco Silva-Mata
ICME2
2007 Introducing tangerine: a tangible interactive natural environment
abstract
In this paper we describe TANGerINE, a tangible tabletop environment in which users can interact with digital contents manipulating tangible smart objects. Such objects provide continuous data about their status through the embedded wireless sensors, while an overhead computer vision module tracks their position and orientation. Merging sensing data, the system is able to detect a richer language of gestures and manipulations both on the tabletop and in its surroundings, enabling for a more expressive interaction language across different contexts.
Stefano Baraldi, Alberto Del Bimbo, Lea Landucci, Nicola Torpei, Omar Cafini, Elisabetta Farella, Augusto Pieracci, Luca Benini
ACM Multimedia2
2007 Content-Based Retrieval of 3-D Objects Using Spin Image Signatures
abstract
Retrieval by content of 3-D models is becoming more and more important due to the advancements in 3-D hardware and software technologies for acquisition, authoring and display of 3-D objects, their ever-increasing availability at affordable costs, and the establishment of open standards for 3-D data interchange. In this paper, we present a new method, referred to as Spin Image Signatures, that develops on the original spin images approach, with adaptations to support effective retrieval by content. According to the method proposed, a set of spin images is derived for each model, to obtain a view-independent description of its 3-D shape and a signature is evaluated for each spin image in the set. Clustering is hence performed on the set of Spin Image Signatures to obtain a compact representation. Experimental results are presented, showing the effectiveness of the Spin Image Signatures method for retrieval, also in comparison with other methods, and its sensitivity to model deformations.
Jürgen Assfalg, Marco Bertini 0001, Alberto Del Bimbo, Pietro Pala
IEEE Trans. Multim.3
2007 Robust tracking and remapping of eye appearance with passive computer vision
abstract
A single-camera iris-tracking and remapping approach based on passive computer vision is presented. Tracking is aimed at obtaining accurate and robust measurements of the iris/pupil position. To this purpose, a robust method for ellipse fitting is used, employing search constraints so as to achieve better performance with respect to the standard RANSAC algorithm. Tracking also embeds an iris localization algorithm (working as a bootstrap multiple-hypotheses generation step), and a blink detector that can detect voluntary eye blinks in human-computer interaction applications. On-screen remapping incorporates a head-tracking method capable of compensating for small user-head movements. The approach operates in real time under different light conditions and in the presence of distractors. An extensive set of experiments is presented and discussed. In particular, an evaluation method for the choice of layout of both hardware components and calibration points is described. Experiments also investigate the importance of providing a visual feedback to the user, and the benefits gained from performing head compensation, especially during image-to-screen map calibration.
Carlo Colombo, Dario Comanducci, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.3
2006 Learning Foveal Sensing Strategies in Unconstrained Surveillance Environments
abstract
In this paper we report on techniques for automatically learning foveal sensing strategies for an active pan-tiltzoom camera. The approach uses reinforcement learning to discover foveal actions maximizing the performance of visual detectors, that are in turn assumed to be highly correlated with the task at hand. In our case, the main goal is to recognize people, hence a frontal face detection module is employed. The system uses reinforcement learning to learn if, when and how to foveate on a subject, based on its previous experience in terms or successful actions in similar situations. An action is successful if it leads to a correct face detection in the high resolution images obtained when the subject is zoomed in. In contrast with existing methods, the proposed approach obviates the need for camera calibration and camera performance modeling. Also, the method does not rely on active tracking of targets. Experimental results show how the system is capable of learning foveation strategies without requiring extensive a priori information or environmental models. Results also illustrate how the system effectively learns a strategy that allows the camera to foveate only in situations where successful detection is highly likely.
Andrew D. Bagdanov, Alberto Del Bimbo, Walter Nunziati, Federico Pernici
AVSS2
2006 Camera Calibration with Two Arbitrary Coaxial Circles
Carlo Colombo, Dario Comanducci, Alberto Del Bimbo
ECCV (1)3
2006 Using Knowledge Representation Languages for Video Annotation and Retrieval
Marco Bertini 0001, Gianpaolo D'Amico, Alberto Del Bimbo, Carlo Torniai
FQAS3
2006 3D Face Identification Based on Arrangement of Salient Wrinkles
abstract
In this paper, we propose an original framework for three dimensional face representation and matching for identification purposes. Basic traits of a face are encoded by extracting curves of salient ridges and ravines from the surface of a dense mesh. A compact graph representation is then extracted from these curves through an original modeling technique capable to quantitatively measure spatial relationships between curves in a three dimensional space. In this way, face recognition is obtained by matching 3D graph representations of faces. Experimental results on a 3D face database show that the proposed solution attains high recognition accuracy and is quite robust to facial expression and pose changes
Gianni Antini, Stefano Berretti, Alberto Del Bimbo, Pietro Pala
ICME3
2006 Bringing the Wiki Collaboration Model to the Tabletop World
abstract
We present an interactive workspace that integrates wiki collaboration in knowledge-building activities with face-to-face scenarios like brainstorming or problem solving sessions. A wiki serves as the repository for knowledge elements, which are presented to co-located users in form of a concept map projected on a table. A computer vision module tracks multiple hands and fingers on the table surface and map elements can be manipulated using a simple gesture language. Also smart devices, like PDAs or tablets, can be connected to the system and participate in the interaction with their local input. The concept map and the wiki are synchronized in real-time, providing notifications to both co-located and distributed users and allowing a community shared awareness that enhances and enriches the knowledge building experience
Stefano Baraldi, Alberto Del Bimbo, Alessandro Valli
ICME2
2006 Matching Faces with Textual Cues in Soccer Videos
abstract
In soccer videos, most significant actions are usually followed by close-up shots of players that take part in the action itself. Automatically annotating the identity of the players present in these shots would be considerably valuable for indexing and retrieval applications. Due to high variations in pose and illumination across shots however, current face recognition methods are not suitable for this task. We show how the inherent multiple media structure of soccer videos can be exploited to understand the players' identity without relying on direct face recognition. The proposed method is based on a combination of interest point detector to "read" textual cues that allow to label a player with its name, such as the number depicted on its jersey, or the superimposed text caption showing its name. Players not identified by this process are then assigned to one of the labeled faces by means of a face similarity measure, again based on the appearance of local salient patches. We present results obtained from soccer videos taken from various recent games between national teams
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati
ICME2
2006 A Desktop 3D Scanner Exploiting Rotation and Visual Rectification of Laser Profiles
abstract
We describe a low cost system for metric 3D scanning from uncalibrated images based on rotational kinematic constraints. The system is composed by a turntable, an offthe- shelf camera and a laser stripe illuminator. System operation is based on the construction of the virtual image of a surface of revolution (SOR), from which two imaged SOR cross-sections are obtained in an automatic way, and internal camera calibration is performed by exploiting the same object being scanned. Shape acquisition is finally obtained by laser profile rectification and collation. Experiments with real data are shown, providing an insight into both camera calibration and shape reconstruction performance. System accuracy appears to be adequate for desktop applications.
Carlo Colombo, Dario Comanducci, Alberto Del Bimbo
ICVS3
2006 Automatic detection of player's identity in soccer videos using faces and text cues
abstract
In soccer videos, most significant actions are usually followed by close--up shots of players that take part in the action itself. Automatically annotating the identity of the players present in these shots would be considerably valuable for indexing and retrieval applications. Due to high variations in pose and illumination across shots however, current face recognition methods are not suitable for this task. We show how the inherent multiple media structure of soccer videos can be exploited to understand the players' identity without relying on direct face recognition. The proposed method is based on a combination of interest point detector to "read" textual cues that allow to label a player with its name, such as the number depicted on its jersey, or the superimposed text caption showing its name. Players not identified by this process are then assigned to one of the labeled faces by means of a face similarity measure, again based on the appearance of local salient patches. We present results obtained from soccer videos taken from various recent games between national teams.
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati
ACM Multimedia2
2006 Automatic annotation and semantic retrieval of video sequences using multimedia ontologies
abstract
Effective usage of multimedia digital libraries has to deal with the problem of building efficient content annotation and retrieval tools. MOM (Multimedia Ontology Manager) is a complete system that allows the creation of multimedia ontologies, supports automatic annotation and creation of extended text (and audio) commentaries of video sequences, and permits complex queries by reasoning on the ontology.
Marco Bertini 0001, Alberto Del Bimbo, Carlo Torniai
ACM Multimedia2
2006 MOM: multimedia ontology manager. A framework for automatic annotation and semantic retrieval of video sequences
abstract
Effective usage of multimedia digital libraries has to deal with the problem of building efficient content annotation and retrieval tools. MOM (Multimedia Ontology Manager) is a complete system that allows the creation of multimedia ontologies, supports automatic annotation and creation of extended text (and audio) commentaries of video sequences, and permits complex queries by reasoning on the ontology.
Marco Bertini 0001, Alberto Del Bimbo, Carlo Torniai, Rita Cucchiara, Costantino Grana
ACM Multimedia2
2006 PEANO: pictorial enriched annotation of video
abstract
In this DEMO, we present a tool set for video digital library management that allows i) structural annotation of edited videos in MPEG-7 by automatically extracting shots and clips; ii) automatic semantic annotation based on perceptual similarity against a taxonomy enriched with pictorial concepts iii) video clip access and hierarchical summarization with stand-alone and web interface iv) access to clips from mobile platform in GPRS-UMTS video-streaming. The tools can be applied in different domain-specific Video Digital Libraries. The main novelty is the possibility to enrich the annotation with pictorial concepts that are added to a textual taxonomy in order to make the automatic annotation process more fast and often effective. The resulting multimedia ontology is described in the MPEG-7 framework. The PEANO (Perceptual Annotation of Video) tool has been tested over video art , sport (Soccer, Olimpic Games 2006, Formula 1) and news clips.
Costantino Grana, Roberto Vezzani, Daniele Bulgarelli, Giovanni Gualdi, Rita Cucchiara, Marco Bertini 0001, Carlo Torniai, Alberto Del Bimbo
ACM Multimedia8
2006 Content-based retrieval of 3D models through curvature maps: a CBR approach exploiting media conversion
Jürgen Assfalg, Alberto Del Bimbo, Pietro Pala
Multim. Tools Appl.2
2006 Towards on-line saccade planning for high-resolution image sensing
Alberto Del Bimbo, Federico Pernici
Pattern Recognit. Lett.1
2006 Semantic adaptation of sport videos with user-centred performance analysis
abstract
In semantic video adaptation measures of performance must consider the impact of the errors in the automatic annotation over the adaptation in relationship with the preferences and expectations of the user. In this paper, we define two new performance measures Viewing Quality Loss and Bit-rate Cost Increase,that are obtained from classical peak signal-to-noise ration (PSNR) and bitrate, and relate the results of semantic adaptation to the errors in the annotation of events and objects and the user's preferences and expectations. We present and discuss results obtained with a system that performs automatic annotation of soccer sport video highlights and applies different coding strategies to different parts of the video according to their relative importance for the end user. With reference to this framework, we analyze how highlights' statistics and the errors of the annotation engine influence the performance of semantic adaptation and reflect into the quality of the video displayed at the user's client and the increase of transmission costs.
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001
IEEE Trans. Multim.3
2006 Content-based retrieval of 3D models
abstract
In the past few years, there has been an increasing availability of technologies for the acquisition of digital 3D models of real objects and the consequent use of these models in a variety of applications, in medicine, engineering, and cultural heritage. In this framework, content-based retrieval of 3D objects is becoming an important subject of research, and finding adequate descriptors to capture global or local characteristics of the shape has become one of the main investigation goals. In this article, we present a comparative analysis of a few different solutions for description and retrieval by similarity of 3D models that are representative of the principal classes of approaches proposed. We have developed an experimental analysis by comparing these methods according to their robustness to deformations, the ability to capture an object's structural complexity, and the resolution at which models are considered.
Alberto Del Bimbo, Pietro Pala
ACM Trans. Multim. Comput. Commun. Appl.1
2005 View registration using interesting segments of planar trajectories
abstract
We introduce a method for recovering the spatial and temporal alignment between two or more views of objects moving over a ground plane. Existing approaches either assume that the streams are globally synchronized, so that only solving the spatial alignment is needed, or that the temporal misalignment is small enough so that exhaustive search can be performed. In contrast, our approach can recover both the spatial and temporal alignment, regardless of their magnitude. We compute for each trajectory a number of interesting segments, and we use their description to form putative matches between trajectories. Each pair of corresponding interesting segments induces a temporal alignment, and defines an interval of common support across two views of an object that is used to recover the spatial alignment. Interesting segments and their descriptors are defined using algebraic projective invariants measured along the trajectories. Similarity between interesting segments is computed taking into account the statistics of such invariants. Candidate alignment parameters are verified by checking the consistency, in terms of the symmetric transfer error, of all the putative pairs of corresponding interesting segments. Experiments are conducted with two different sets of data, one with two views of an outdoor scene featuring moving people and cars, and one with four views of a laboratory sequence featuring moving radio-controlled cars.
Walter Nunziati, Jonathan Alon, Stan Sclaroff, Alberto Del Bimbo
AVSS4
2005 Domain Knowledge Extension with Pictorially Enriched Ontologies
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Carlo Torniai
CAIP3
2005 Computer Vision Based System for Interactive Cooperation of Multiple Users
Alberto Del Bimbo, Lea Landucci, Alessandro Valli
CAIP1
2005 Automatic Annotation of Sport Video Content
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati
CIARP2
2005 Retrieval of 3D objects using curvature correlograms
abstract
Along with images and videos, 3D models have raised a certain interest for a number of reasons, including advancements in 3D hardware and software technologies, their ever decreasing prices and increasing availability, affordable 3D authoring tools, and the establishment of open standards for 3D data interchange. The resulting proliferation of 3D models demands for tools supporting their effective and efficient management, including archival and retrieval. In order to support effective retrieval by content of 3D objects and enable retrieval by object parts, information about local object structure should be combined with spatial information on object surface. In this paper, as a solution to this requirement, we present a method relying on curvature correlograms to perform description and retrieval by content of 3D objects. Experimental results are presented both to show results of sample queries by content and to compare-in terms of precision/recall figures-the proposed solution to alternative techniques.
Gianni Antini, Stefano Berretti, Alberto Del Bimbo, Pietro Pala
ICME3
2005 3D Mesh Partitioning for Retrieval by Parts Applications
abstract
A solution for part segmentation of 3D objects is proposed in this paper. The approach is targeted to identify salient visual parts of a mesh by determining its main protrusions and discarding, at the same time, parts originated by un-relevant local properties. This is obtained by first breaking the 3D mesh into seed regions according to the sum of geodesic distances between vertices, then by using topological and curvature information to refine the number of regions and their boundaries. In so doing, effective segmentation is regarded as a prerequisite to enable retrieval of 3D objects based on similarity of parts. Experimental results show the applicability of the proposed solution to complex shapes and its effectiveness in the identification of object parts.
Gianni Antini, Stefano Berretti, Alberto Del Bimbo, Pietro Pala
ICME3
2005 Video Annotation with Pictorially Enriched Ontologies
abstract
Video annotation is typically performed by classifying video elements according to some pre-defined ontology of the video content domain. Ontologies are defined by establishing relationships between linguistic terms, that specify domain concepts at different abstraction levels. However, although linguistic terms are appropriate to distinguish event and object categories, they are inadequate when they must describe specific patterns of events or video entities. Instead, in these cases, pattern specifications are better expressed through visual prototypes that capture the essence of the event or entity. Pictorially enriched ontologies, that include visual concepts together with linguistic keywords, are therefore needed to support video annotation up to the level of detail of pattern specification. This paper presents pictorially enriched ontologies and provide a solution for their implementation in the soccer video domain. The pictorially enriched ontology is used both to directly assign multimedia objects to concepts, providing a more meaningful definition than the linguistics terms, and to extend the initial knowledge of the domain, adding subclasses of highlights or new highlight classes that were not defined in the linguistic ontology. Automatic annotation of soccer clips up to the pattern specification level using a pictorially enriched ontology is discussed.
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Carlo Torniai
ICME3
2005 Segmentation of 3D Objects Using Pulse-Coupled Oscillator Networks
abstract
Along with image and video libraries, archives of 3D models have recently gained increasing attention. Accordingly, there is an increasing demand for solutions enabling retrieval of 3D models based on global properties as well as properties of object parts. In particular, retrieval based on object parts relies on segmentation of 3D objects into their constituent parts. This is a challenging task, as the identification of object parts should conform to human perceptual judgement. Therefore, definition of models and solutions that enable decomposition of 3D objects into perceptually relevant parts is a fundamental step to enable effective retrieval based on object parts. However, a few approaches have been proposed to support segmentation of 3D meshes into perceptually relevant parts. In this paper we propose a model based on pulse-coupled oscillator networks. Preliminary experiments are reported to demonstrate the validity and potential of the proposed solution.
Eva Ceccarelli, Alberto Del Bimbo, Pietro Pala
ICME2
2005 Automatic video annotation using ontologies extended with visual information
abstract
Classifying video elements according to some pre-defined ontology of the video content domain is a typical way to perform video annotation. Ontologies are defined by establishing relationships between linguistic terms that specify domain concepts at different abstraction levels. However, although linguistic terms are appropriate to distinguish event and object categories, they are inadequate when they must describe specific patterns of events or video entities. Instead, in these cases, pattern specifications can be better expressed through visual prototypes that capture the essence of the event or entity. Therefore pictorially enriched ontologies, that include both visual and linguistic concepts, can be useful to support video annotation up to the level of detail of pattern specification.This paper presents pictorially enriched ontologies and discusses a solution for their implementation for the soccer video domain. An unsupervised clustering method is proposed in order to create the enriched ontologies by defining visual prototypes representing specific patterns of highlights and adding them as visual concepts to the ontology.An algorithm that uses pictorially enriched ontologies to perform automatic soccer video annotation is proposed and results for typical highlights are presented. Annotation is performed associating occurrences of events, or entities, to higher level concepts by checking their proximity to visual concepts that are hierarchically linked to higher level semantics.
Marco Bertini 0001, Alberto Del Bimbo, Carlo Torniai
ACM Multimedia2
2005 Common Visual Cues for Sports Highlights Modeling
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati
Multim. Tools Appl.2
2005 An Integrated Framework for Semantic Annotation and Adaptation
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001
Multim. Tools Appl.3
2005 Guest Editorial: Special Issue on Video Segmentation for Semantic Annotation and Transcoding
Rita Cucchiara, Alberto Del Bimbo
Multim. Tools Appl.2
2005 Guest Editors' Introduction to the Special Section on Syntactic and Structural Pattern Recognition
abstract
This paper presents the guest editors' introduction to the special section which was planned in honor of the memory of the late Professor King-Sun Fu. Dr. King-Sun Fu is widely recognized for his paramount contributions in the field of pattern recognition, especially in the area of syntactic and structural pattern recognition. The paper discusses the problems of interest regarding syntactic and structural pattern recognition and in related areas. The concepts of syntactic and structural pattern recognition were briefly discussed and the contributed papers in the section were introduced.
Mitra Basu, Horst Bunke, Alberto Del Bimbo
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 Metric 3D Reconstruction and Texture Acquisition of Surfaces of Revolution from a Single Uncalibrated View
abstract
Image analysis and computer vision can be effectively employed to recover the three-dimensional structure of imaged objects, together with their surface properties. In this paper, we address the problem of metric reconstruction and texture acquisition from a single uncalibrated view of a surface of revolution (SOR). Geometric constraints induced in the image by the symmetry properties of the SOR structure are exploited to perform self-calibration of a natural camera, 3D metric reconstruction, and texture acquisition. By exploiting the analogy with the geometry of single axis motion, we demonstrate that the imaged apparent contour and the visible segments of two imaged cross sections in a single SOR view provide enough information for these tasks. Original contributions of the paper are: single view self-calibration and reconstruction based on planar rectification, previously developed for planar surfaces, has been extended to deal also with the SOR class of curved surfaces; self-calibration is obtained by estimating both camera focal length (one parameter) and principal point (two parameters) from three independent linear constraints for the SOR fixed entities; the invariant-based description of the SOR scaling function has been extended from affine to perspective projection. The solution proposed exploits both the geometric and topological properties of the transformation that relates the apparent contour to the SOR scaling function. Therefore, with this method, a metric localization of the SOR occluded parts can be made, so as to cope with them correctly. For the reconstruction of textured SORs, texture acquisition is performed without requiring the estimation of external camera calibration parameters, but only using internal camera parameters obtained from self-calibration.
Carlo Colombo, Alberto Del Bimbo, Federico Pernici
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Image Mosaicing from Uncalibrated Views of a Surface of Revolution
abstract
We present a novel approach to obtain a mosaic image for the surface texture content of a surface of revolution (SOR) from a collection of uncalibrated views. The SOR scene constraint is used to calibrate each view and align the corresponding pictorial content into a global representation. Metric surface properties are extracted from each view by exploiting special properties of the imaged SOR geometry expressed in terms of homologies. Image alignment is achieved by projecting imaged surface elements onto a reference plane, and then registering them according to a translational motion model. This work extends previous research on calibrated scenes of right circular cylinders to the more general case of uncalibrated SOR scenes. Experimental results with images taken from the web demonstrate the effectiveness and the general applicability of the approach. 1 Introduction and Related Work Image mosaicing consists in merging collections of images having a partially overlapping content. The process can be decomposed into three main steps. First, the transformations
Carlo Colombo, Alberto Del Bimbo, Federico Pernici
BMVC2
2004 Content Based Retrieval of 3D Data
Alberto Del Bimbo, Pietro Pala
CIARP1
2004 3D content-based retrieval with spin images
abstract
Along with images and videos, 3D models have recently gained increasing attention for a number of reasons: advancements in 3D hardware and software technologies; their ever decreasing prices and increasing availability; affordable 3D authoring tools; the establishment of open standards for 3D data interchange. The ever increasing availability of 3D models demands tools to support their effective and efficient management. Among these tools, those enabling content-based retrieval play a key role. We present a novel approach to 3D content-based retrieval that is based on spin images. Spin images are used to derive a view-independent description of both database and query objects: a set of spin images is first created for each object; then, a descriptor is evaluated for each spin image in the set; clustering is performed on the set of image-based descriptors of each object to achieve a compact representation of the object, thus allowing for efficient indexing and matching. Experimental results are presented for a test database of about 300 models. These results indicate that spin images can be successfully exploited for content-based retrieval of 3D objects
Jürgen Assfalg, Gianpaolo D'Amico, Alberto Del Bimbo, Pietro Pala
ICME3
2004 Shape representation by spatial partitioning for content based retrieval applications
abstract
Shape representations for classification or retrieval purposes, have been extensively investigated, but only a few methods have tried to represent shapes as extended entities without reducing them to their boundary profiles or to synthetic geometric descriptors. An original solution for shape representation is proposed; it relies on a modelling technique originally developed to express directional spatial relationships between extended spatial entities. The representation selects a discrete set of points with respect to which a relationship matrix is computed, accounting for the spatial distribution of shape pixels. This is accomplished at different levels of resolution by a tree based structural representation. Properties of the representation and a measure of shape similarity are discussed. The efficiency and effectiveness of the proposed solution have also been assessed in the context of content based retrieval applications, through an experimental evaluation using a shape collection.
Stefano Berretti, Gianpaolo D'Amico, Alberto Del Bimbo
ICME3
2004 Common visual cues for sports highlights detection
abstract
Automatic annotation of semantic events allows effective retrieval of video content. We present automatic annotation of sports highlights for some of the principal sports types. They are obtained by detecting and tracking a limited number of visual cues common to each sport. Highlights are represented as atomic entities at the semantic level. They have a limited temporal extension and can be modeled as the spatio-temporal concatenation of specific events. Visual cues encode position and speed information coming from the camera and from the objects/athletes that are present in the scene, and are estimated automatically from the video stream. Algorithms for model checking and for visual cue estimation are discussed. as well as applications of the representation to different sports domains
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati
ICME2
2004 Content-based video adaptation with user's preferences
abstract
We present an integrated system that has been designed to support automatic semantic extraction of highlights in sports video and automatic video adaptation according to user's preferences. To analyze the user's satisfaction, we propose a new performance measure that explicitly takes into account the user's preferences and considers the number and type of errors produced by the annotation engine and the way in which these errors affect the compressed video quality and bandwidth allocation. We provide experimental results with application to soccer and swimming.
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001
ICME3
2004 Automatic annotation of video streams
abstract
Broadcasters are demonstrating interest in systems that ease the process of annotating huge amount of live and archived video materials. Exploitation of such assets is considered a key method for the improvement of production quality and sport videos (one of the most marketable assets). In Europe, soccer is one of the most relevant sport types. This paper deals with detection and recognition of soccer highlights, using an approach based on temporal logic models.
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati
MMSP2
2004 Merging Results for Distributed Content Based Image Retrieval
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
Multim. Tools Appl.2
2004 Highlights modeling and detection in sports videos
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati
Pattern Anal. Appl.2
2003 Annotation and Retrieval of Structured Video Documents
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati
ECIR2
2003 Automatic extraction and annotation of soccer video highlights
abstract
Broadcasters are demonstrating interest in systems that ease the process of annotation the huge amount of live and archived video materials. Exploitation of such assets is considered a key method for the improvement of production quality, and sport videos are one of the most marketable assets. In particular, in Europe, soccer is one of the most relevant sport types. This paper deals with detection and recognition of soccer highlights, using an approach based on temporal logic models.
Jürgen Assfalg, Marco Bertini 0001, Carlo Colombo, Alberto Del Bimbo, Walter Nunziati
ICIP (2)4
2003 Curvature maps for 3D CBR
abstract
Along with images and videos, 3D models have recently gained increasing attention for a number of reasons: advancements in 3D hardware and software technologies, their ever decreasing prices and increasing availability, affordable 3D authoring tools, and the establishment of open standards for 3D data interchange. In this paper we address the problem of content-based retrieval of 3D models. The solution we present here relies on the description of each model by means a curvature map: after an initial pre-processing of the model, differential properties of points on the surface of the 3D object are evaluated; the model surface is then warped into an ellipsoid, and is mapped onto a 2D image retaining curvature information of the original model. Matching is performed by comparing the 2D map of the query against the 2D maps of the database models. This method has been implemented in a prototype system supporting retrieval by content of 3D objects through a Web interface.
Jürgen Assfalg, Alberto Del Bimbo, Pietro Pala
ICME2
2003 Merging results of distributed image libraries
abstract
Exploitation of information repositories available on the Internet requires users to separately query each repository and manually gather retrieved results. Such a solution could be simplified by using a centralized server that acts as a gateway between the user and repositories: the centralized server forwards the user query to federated repositories and fuses retrieved documents for presentation to the user. To perform these tasks efficiently, the centralized server should perform two main functions: resource selection and data fusion. The former is required to forward the user query only to the repositories that are candidate to contain relevant documents. The latter is used to gather all retrieved documents and conveniently arrange them for presentation to the user. In the case of image repositories, data fusion is particularly challenging owing to the difficulty to normalize document scores returned by different repositories. In this paper a novel solution is presented for fusion of results returned by different image repositories. Experimental results are presented that show the potential of the proposed approach.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
ICME2
2003 Object and event detection for semantic annotation and transcoding
abstract
Video annotation provides a suitable way to describe, organize, and index stored videos. On the other hand, transcoding aims at adapting content to the user/client capabilities and requirements. Both cues are now mandatory, given the tremendous demand of multimedia access from remote clients, in particular nowadays that new terminals with limited resources (PDAs, HCCs, Smart phones) have access to the network. In this paper we propose a unified framework to define event-based and object-based semantic extraction from video to provide both semantic video annotation for video stored and semantic on-line transcoding from live cameras. Two case studies (highlights' extraction from soccer videos for the annotation and people behavior detection in domotic application for transcoding) and corresponding experimental results are reported.
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001
ICME3
2003 Semantic annotation for live and posterity logging of video documents
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati
VCIP2
2003 Semantic annotation of soccer videos: automatic highlights identification
Jürgen Assfalg, Marco Bertini 0001, Carlo Colombo, Alberto Del Bimbo, Walter Nunziati
Comput. Vis. Image Underst.4
2003 Weighted walkthroughs between extended entities for retrieval by spatial arrangement
abstract
In the access to image databases, queries based on the appearing visual features of searched data reduce the gap between the user and the engineering representation. To support this access modality, image content can be modeled in terms of different types of features such as shape, texture, color, and spatial arrangement. An original framework is presented which supports quantitative nonsymbolic representation and comparison of the mutual positioning of extended nonrectangular spatial entities. Properties of the model are expounded to develop an efficient computation technique and to motivate and assess a metric of similarity for quantitative comparison of spatial relationships. Representation and comparison of binary relationships between entities is then embedded into a graph-theoretical framework supporting representation and comparison of the spatial arrangements of a picture. Two prototype applications are described.
Stefano Berretti, Alberto Del Bimbo, Enrico Vicario
IEEE Trans. Multim.2
2003 Visual capture and understanding of hand pointing actions in a 3-D environment
abstract
We present a nonintrusive system based on computer vision for human-computer interaction in three-dimensional (3-D) environments controlled by hand pointing gestures. Users are allowed to walk around in a room and manipulate information displayed on its walls by using their own hands as pointing devices. Once captured and tracked in real-time using stereo vision, hand pointing gestures are remapped onto the current point of interest, thus reproducing in an advanced interaction scenario the "drag and click" behavior of traditional mice. The system, called PointAt (patent pending), enjoys a careful modeling of both user and optical subsystem, and visual algorithms for self-calibration and adaptation to both user peculiarities and environmental changes. The concluding sections provide an insight into system characteristics, performance, and relevance for real applications.
Carlo Colombo, Alberto Del Bimbo, Alessandro Valli
IEEE Trans. Syst. Man Cybern. Part B2
2002 Soccer highlights detection and recognition using HMMs
abstract
In this paper we report on our experience in the detection and recognition of soccer highlights in videos using hidden Markov models. A first approach relies on camera motion only, whereas a second one also includes information regarding the location of players on the playing field. While the former approach requires less information, the latter has proven to be more precise. Our experimental evaluation yields interesting results.
Jürgen Assfalg, Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati, Pietro Pala
ICME (1)3
2002 Using indexing structures for resource descriptors extraction from distributed image repositories
abstract
Content based retrieval from distributed libraries raises new and challenging issues with respect to retrieval from a single repository. In particular, an effective management of distributed libraries develops upon three main processes: resource description (extraction of descriptors that qualify the content of a given archive), resource selection (given a user query, analyze resource descriptions and select the resources that contain relevant documents) and results merging (organize and present items returned by individual libraries). So far, these issues have been mainly addressed for text archives. We present a solution to resource descriptors extraction, developing on the use of techniques for multidimensional data indexing. In particular, we implement and compare the extraction of resource descriptors computed through two different indexing approaches; namely m-tree indexing and fuzzy clustering. Comparative results are presented for a test database of about 1000 images.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
ICME (2)2
2002 Shape reconstruction from a single photograph for 3D object retrieval and visualization
abstract
We describe a geometric approach for reconstructing 3D textured graphical models of surfaces of revolution (SOR) from a single uncalibrated view. Metric reconstruction of 3D shape is complemented with the extraction of flattened 2D texture, so as to support visual retrieval from 2D/3D cues and to generate realistic 3D visualization models. The approach developed is quite simple, yet accurate and robust; its applications range from the preservation, analysis and classification of cultural heritage, to advanced graphics and multimedia.
Carlo Colombo, Alberto Del Bimbo, Federico Pernici
ICME (1)2
2002 Semantic Annotation and Indexing of News and Sports Videos
Jürgen Assfalg, Marco Bertini 0001, Carlo Colombo, Alberto Del Bimbo, Walter Nunziati
SOFSEM4
2002 Spatial arrangement of color in retrieval by visual similarity
Stefano Berretti, Alberto Del Bimbo, Enrico Vicario
Pattern Recognit.2
2002 Indexing for reuse of TV news shots
Marco Bertini 0001, Alberto Del Bimbo, Pietro Pala
Pattern Recognit.2
2002 Introduction to the special issue on multimedia database
Sankar Basu, Alberto Del Bimbo, Ahmed H. Tewfik, HongJiang Zhang
IEEE Trans. Multim.2
2002 Three-Dimensional Interfaces for Querying by Example in Content-Based Image Retrieval
abstract
Image databases are widely exploited in a number of different contexts, ranging from history of art, through medicine, to education. Existing querying paradigms are based either on the usage of textual strings, for high-level semantic queries or on 2D visual examples for the expression of perceptual queries. Semantic queries require manual annotation of the database images. Instead, perceptual queries only require that image analysis is performed on the database images in order to extract salient perceptual features that are matched with those of the example. However, usage of 2D examples is generally inadequate as effective authoring of query images, attaining a realistic reproduction of complex scenes, needs manual editing and sketching ability. Investigation of new querying paradigms is therefore an important-yet still marginally investigated-factor for the success of content-based image retrieval. In this paper, a novel querying paradigm is presented which is based on usage of 3D interfaces exploiting navigation and editing of 3D virtual environments. Query images are obtained by taking a snapshot of the framed environment and by using the snapshot as an example to retrieve similar database images. A comparative analysis is carried out between the usage of 3D and 2D interfaces and their related query paradigms. This analysis develops on a user test on retrieval efficiency and effectiveness, as well as on an evaluation of users' satisfaction.
Jürgen Assfalg, Alberto Del Bimbo, Pietro Pala
IEEE Trans. Vis. Comput. Graph.2
2001 Advanced man-machine interface for cultural heritage
abstract
This paper describes a system for advanced man-machine interaction based on computer vision technology allowing users to manipulate information displayed on large wall panels by using their own hands as pointing devices. The system gets its input from a pair of color video cameras (placed so as to have the user in view); a personal computer performs image analysis and updates interaction parameters. Such parameters, which reflect current user status, are then transformed into graphic interface commands and output to the screen through a beamer. Graphic interface operation is the same as with standard computer mice; the main points of innovation concern naturality of interaction, low intrusiveness and a priori training, and low equipment cost. An experimental version of the system is currently being employed to provide museum visitors with advanced interactive services.
Gabriele Baggiani, Carlo Colombo, Alberto Del Bimbo
ICIP (1)3
2001 Description and retrieval of 3D cellular structures
abstract
Recent advances in management of multimedia digital libraries enable effective retrieval of information in the form of audio, image and video. Many archives of 3D objects already exist and are expected to grow both in relevance and size. However, retrieval of information in the form of 3D objects has received limited attention. We address the problem of effective description and retrieval of 3D data representing intracellular structures. These structures are represented as image stacks, where an image stack is constituted by a set of 2D images representing sections of a cellular body at different heights. In the proposed approach 2D visual feature descriptors and hidden Markov models are combined to obtain a representation model which is able to distinguish such intracellular structures as Golgi, nucleus, endoplasmic reticulum and lysosomes. Preliminary results are presented to show the effectiveness of the proposed representation model.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
ICIP (1)2
2001 Content Based Retrieval Of 3d Cellular Structures
abstract
Recent advances in management of multimedia digital libraries enable effective retrieval of information in the form of audio, image and video. However, retrieval of information in the form of 3D objects has received limited attention so far. Yet many archives of 3D objects already exist and are expected to grow both in relevance and size. In this paper, we address the problem of effective description and retrieval of 3D data representing intracellular structures. These structures are represented in the form of image stacks, being an image stack a set of 2D images representing planar sections of a cellular body at different heights. In the proposed method, 2D visual feature descriptors and Hidden Markov Models are combined to obtain a representation model which is able to distinguish such intracellular structures as Golgi, nucleus, endoplasmic reticulum and lysosomes. Preliminary results are presented to show the effectiveness of the proposed representation model.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
ICME2
2001 Automatic Caption Localization in Videos Using Salient Points
abstract
Broadcasters are demonstrating interest in building digital archives of their assets for reuse of archive materials for TV programs, on-line availability, and archiving. This requires tools for video indexing and retrieval by content exploiting high-level video information such as that contained in super-imposed text captions. In this paper we present a method to automatically detect and localize captions in digital video using temporal and spatial local properties of salient points in video frames. Results of experiments on both high-resolutionDV sequences and standard VHS videos are presented and discussed. 1.
Marco Bertini 0001, Carlo Colombo, Alberto Del Bimbo
ICME3
2001 Spatial Arrangement Of Color Flows For Video Retrieval
Alberto Del Bimbo, Enrico Vicario, Pietro Pala
ICME1
2001 Classification Of Rawmaterial Sports Videos For Broadcasting Using Color And Edge Features
abstract
The authors discuss the method to classify raw material sports videos for broadcasting. Because the raw material sports videos sometimes do not get edited, one cannot use the knowledge on edited videos. The authors use the color and edge features and evaluate whether one can classify the sports videos with those features. Also introduced is the "player" and "audience" class - apart from each sport class - to improve the classification results.
Masayuki Mukunoki, Marco Bertini 0001, Jürgen Assfalg, Alberto Del Bimbo
ICME4
2001 Retrieval of Commercials by Semantic Content: The Semiotic Perspective
Carlo Colombo, Alberto Del Bimbo, Pietro Pala
Multim. Tools Appl.2
2001 Modelling Spatial Relationships between Colour Clusters
Stefano Berretti, Alberto Del Bimbo, Enrico Vicario
Pattern Anal. Appl.2
2001 Efficient Matching and Indexing of Graph Models in Content-Based Retrieval
abstract
In retrieval from image databases, evaluation of similarity, based both on the appearance of spatial entities and on their mutual relationships, depends on content representation based on attributed relational graphs. This kind of modeling entails complex matching and indexing, which presently prevents its usage within comprehensive applications. In this paper, we provide a graph-theoretical formulation for the problem of retrieval based on the joint similarity of individual entities and of their mutual relationships and we expound its implications on indexing and matching. In particular, we propose the usage of metric indexing to organize large archives of graph models, and we propose an original look-ahead method which represents an efficient solution for the (sub)graph error correcting isomorphism problem needed to compute object distances. Analytic comparison and experimental results show that the proposed lookahead improves the state-of-the-art in state-space search methods and that the combined use of the proposed matching and indexing scheme permits for the management of the complexity of a typical application of retrieval by spatial arrangement.
Stefano Berretti, Alberto Del Bimbo, Enrico Vicario
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Content-based indexing and retrieval of TV news
Marco Bertini 0001, Alberto Del Bimbo, Pietro Pala
Pattern Recognit. Lett.2
2000 Semantics-Based Retrieval by Content
abstract
For the practical use of visual information retrieval systems, access to visual data must be provided so as to bridge the semantic gap between the system and users. Although the problem seems to face only the level of representation, indeed the user interface and feature organization play equally essential roles in obtaining effective retrieval systems. Semiotics, which is concerned with the description of conditions for the production of sense and of the way in which it is received by humans appears as the formal background to casting the development of retrieval by semantic information.
Alberto Del Bimbo
ICIP1
2000 Image Retrieval by Positive and Negative Examples
abstract
Systems for content based image retrieval typically support access to database images through the query-by-example paradigm. This includes query-by-image and query-by-sketch. Since query-by-sketch can be difficult in some cases-lack of sketching abilities, difficulty to detect distinguishing image features-generally querying is performed through the query-by-image paradigm. A limiting factor of this paradigm is that a single sample image rarely includes all and only the characterizing elements the user is looking for. In this paper a system is presented that supports query-by-image using multiple image examples. Examples can be positive and negative and can be edited in order to disregard irrelevant image features.
Jürgen Assfalg, Alberto Del Bimbo, Pietro Pala
ICPR2
2000 The Computational Aspect of Retrieval by Spatial Arrangement
abstract
Image retrieval by spatial arrangement underlies a matching problem for the interpretation of entities specified in the user query on the entities appearing in the image of the database, and for the joint comparison of their features and spatial relationships. In this paper, we provide a graph-theoretical formulation of the problem and discuss its implication on indexing and matching. We first identify an indexing scheme which may fit the characteristics of the problem, and then expound and evaluate an efficient graph matching technique which makes this indexing approach viable.
Stefano Berretti, Alberto Del Bimbo, Enrico Vicario
ICPR2
2000 Issues and Directions in Visual Information Retrieval
Alberto Del Bimbo
ICPR1
2000 Video Retrieval Based on Dynamics of Color Flows
Alberto Del Bimbo, Pietro Pala, L. Tanganelli
ICPR1
2000 Guest Editor's Introduction
Alberto Del Bimbo
Multim. Tools Appl.1
2000 Retrieval by Shape Similarity with Perceptual Distance and Effective Indexing
abstract
An important problem in accessing and retrieving visual information is to provide efficient similarity matching in large databases. Though much work is being done on the investigation of suitable perceptual models and the automatic extraction of features, little attention is given to the combination of useful representations and similarity models with efficient index structures. In this paper we propose retrieval by shape similarity using local descriptors and effective indexing. Shapes are partitioned into tokens in correspondence with their protrusions, and each token is modeled according to a set of perceptually salient attributes. Shape indexing is obtained by arranging shape tokens into a suitably modified M-tree index structure. Two distinct distance functions model respectively, token and shape perceptual similarity. Examples from a prototype system and computational experiences are reported for both retrieval accuracy and indexing efficiency. Shape retrieval has been tested under shape scaling, orientation changes, and partial shape occlusions. A comparative analysis of different indexing structures, for shape retrieval is presented.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
IEEE Trans. Multim.2
1999 Efficient shape retrieval by parts
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
CAIP2
1999 Generalized Bounds for Time to Collision from First-Order Image Motion
abstract
This paper addresses the problem of estimating time to collision from focal motion field measurements in the case of unconstrained relative rigid motion and surface orientation. It is first observed that, as long as time to collision is regarded as a scaled depth, the above problem does not admit a solution unless a narrow camera field of view is assumed. By a careful generalization of the time to collision concept, it is then expounded how to compute novel solutions which hold however wide the field of view. The formulation, which reduces to known literature approaches in the narrow field of view case, extends the applicability range of time to collision based techniques in areas such as mobile robotics and visual surveillance. The experimental validation of the main theoretical results includes a comparison of narrow- and wide-field of view time to collision approaches using both dense and sparse motion estimates.
Carlo Colombo, Alberto Del Bimbo
ICCV2
1999 GUEST EDITORS' INTRODUCTION: Content-Based Access of Image and Video Libraries
Alberto Del Bimbo, Vittorio Castelli, Shih-Fu Chang, Chung-Sheng Li
Comput. Vis. Image Underst.1
1999 Shape indexing by multi-scale representation
Alberto Del Bimbo, Pietro Pala
Image Vis. Comput.1
1999 Image Retrieval by Color Semantics
Jacopo M. Corridoni, Alberto Del Bimbo, Pietro Pala
Multim. Syst.2
1999 Color-induced image representation and retrieval
Carlo Colombo, Alberto Del Bimbo
Pattern Recognit.2
1999 Real-time head tracking from the deformation of eye contours using a piecewise affine camera
Carlo Colombo, Alberto Del Bimbo
Pattern Recognit. Lett.2
1998 Retrieval of Commercials by Video Semantics
abstract
Videos convey information through several planes of communication, encompassing what is represented in the images how the images are linked together and how the subject is imaged. This feature is stressed in commercials where colors, editing effects, rhythms, and object motion are exploited to influence human purchasing habits. In this paper, based on researches in the marketing field, a link is formalized between low level features of a commercial video and feelings that the video would inspire in the observer. This link is used to define high level indices capturing the main semantics of the video. These indices are embedded in a video retrieval system to support access to a database of video based on their semantics.
Carlo Colombo, Alberto Del Bimbo, Pietro Pala
CVPR2
1998 Query by dialog: an interactive approach to pictorial querying
Alberto Del Bimbo, Maria De Marsico, Stefano Levialdi, Giuliano Peritore
Image Vis. Comput.1
1998 Image Retrieval by Color Semantics with Incomplete Knowledge
abstract
Retrieval by content from image databases faces the distance between low-level syntactic features that can be automatically detected by conventional image processing tools and high level semantics which captures user's filtering intentions. A system is presented which bridges this gap by resorting to a theory formulated by Johannes Itten in 1960, and widely accepted in the community of fine arts, to support objective interpretation of color arrangements over paintings. The system relies upon a schema distinguishing archiving, querying, and retrieval stages. In the archiving stage, images are associated with a description capturing the spatial arrangement of regions with homogeneous chromatic attributes, as detected by the use of an automatic image processing tool. Imprecise descriptions are supported through the adoption of a hierarchical index providing a multi-resolution representation of image contents. In the querying stage, a visual iconic language allows the expression of sentences about chromatic contents in accordance with a high-level semantic model of colors combinations. By permitting flexible expression of abstract, non-literal, properties, the model supports intentional vagueness and incompleteness in the specification of searching queries. In the retrieval stage, a similarity score is introduced, which accounts for the degree with which a query assertion applies to a given image. The measure of similarity drives the traversal of the hierarchical index up to find the minimum level of description precision, permitting a definite decision about the satisfaction of the query on each stored image. © 1998 John Wiley & Sons, Inc.
Jacopo M. Corridoni, Alberto Del Bimbo, Enrico Vicario
J. Am. Soc. Inf. Sci.2
1998 Visual Querying By Color Perceptive Regions
Alberto Del Bimbo, Mauro Mugnaini, Pietro Pala, F. Turco
Pattern Recognit.1
1998 Structured representation and automatic indexing of movie information content
Jacopo M. Corridoni, Alberto Del Bimbo
Pattern Recognit.2
1997 Sensations and Psychological Effects in Color Image Database
abstract
The development of a system supporting querying of image databases by color content tackles a major design choice about properties of colors which are referenced within user queries. On the one hand, low-level properties directly reflect numerical features and concepts tied to the machine representation of color information. On the other hand, high-level properties address concepts such as the perceptual quality of colors and the sensations that they convey. Color-induced sensations include warmth, accordance or contrast, harmony, excitement, depression, anguish etc. In particular, paintings are an example where the message is contained more in the high-level color qualities and spatial arrangements than in the physical properties of colors. Starting from this observation, Johannes Itten (1961) introduced a formalism to analyze the use of color in art and the effects that this induces on the user's psyche. In this paper, we present a system which translates the Itten theory into a formal language that allows us to express the semantics associated with the combination of chromatic properties of color images.
Stefano Berretti, Alberto Del Bimbo, Pietro Pala
ICIP (1)2
1997 Visual Image Retrieval by Elastic Matching of User Sketches
abstract
Effective image retrieval by content from database requires that visual image properties are used instead of textual labels to properly index and recover pictorial data. Retrieval by shape similarity, given a user-sketched template is particularly challenging, owing to the difficulty to derive a similarity measure that closely conforms to the common perception of similarity by humans. In this paper, we present a technique which is based on elastic matching of sketched templates over the shapes in the images to evaluate similarity ranks. The degree of matching achieved and the elastic deformation energy spent by the sketch to achieve such a match are used to derive a measure of similarity between the sketch and the images in the database and to rank images to be displayed. The elastic matching is integrated with arrangements to provide scale invariance and take into account spatial relationships between objects in multi-object queries. Examples from a prototype system are expounded with considerations about the effectiveness of the approach and comparative performance analysis.
Alberto Del Bimbo, Pietro Pala
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 Effective image retrieval using deformable templates
abstract
Image retrieval by contents from database is a major research subject in advanced multimedia systems. Effective image retrieval by contents requires that visual image properties are used instead of textual labels to properly index and recover pictorial data. Retrieval by shape similarity given a user-sketched template is particularly challenging, owing to the difficulty to derive a similarity measure that closely conforms to the common perception of similarity by humans. In this paper we present a technique which takes into account global shape properties and is based on elastic matching of sketched templates over the shapes in the images to evaluate similarity ranks. The degree of matching achieved and the elastic deformation energy spent by the sketch to achieve such a match are used to derive a measure of similarity between the sketch and the images in the database and to rank images to be displayed.
Alberto Del Bimbo, Pietro Pala
ICPR1
1996 Image indexing using shape-based visual features
abstract
Efficient image retrieval by contents from database requires that selective access methods are provided to prune out uninteresting items during the search process. Image indexing based on visual features is particularly challenging, owing to the difficulty to derive a representation of the shapes that closely models the visual appearance perceived by humans. In this paper we present a novel approach for indexing planar and closed curves, on the basis on their visual appearance. A hierarchical model of the curve is derived from its multi-scale analysis, and it is used to provide a description of the curve which is able to distinguish between its structural parts and its details. To cope with the inherent uncertainty of shape appearance, fuzzy sets are used to represent the visual attributes of the shapes. In this way the shapes which share similar structural parts can be gathered and embedded into an index structure which retains the imprecision of shape description by embedding fuzziness in the index itself.
Alberto Del Bimbo, Pietro Pala
ICPR1
1996 Structured digital video indexing
abstract
This paper addresses the problems of film segmentation into its syntactic elements (the shots) and of their aggregation into semantic sets under constrained conditions: the shot/reverse-shot scenes (SRS). It then addresses the problem of describing the dynamic content of each shot through the analysis of camera motion. New techniques for edit effect detection and classification are herein proposed. A new algorithm is thereby introduced for automatic inferring the belonging of a set of shots to the same SRS scene. Finally, an algorithm is presented, which segments a sequence into regions having different motion types and extracts the camera motion. Extended performance analysis of the algorithms has been made on about 20 hours movies.
Jacopo M. Corridoni, Alberto Del Bimbo
ICPR2
1996 A vision-based 3-D mouse
Paolo Nesi, Alberto Del Bimbo
Int. J. Hum. Comput. Stud.2
1996 3D object classification using multi-object Kohonen networks
Jacopo M. Corridoni, Alberto Del Bimbo, Leonardo Landi
Pattern Recognit.2
1996 Optical flow computation using extended constraints
abstract
Several approaches for optical flow estimation use partial differential equations to model changes in image brightness throughout time. A commonly used equation is the so-called optical flow constraint (OFC), which assumes that the image brightness is stationary with respect to time. More recently, a different constraint referred to as the extended optical flow constraint (EOFC) has been introduced, which also contains the divergence of the flow field of image brightness. There is no agreement in the literature about which of these constraints provides the best estimation of the velocity field. Two new solutions for optical flow computation are proposed, which are based on an approximation of the constraint equations. The two techniques have been used with both EOFC and OFC constraint equations. Results achieved by using these solutions have been compared with several well-known computational methods for optical flow estimation in different motion conditions. Estimation errors have also been measured and compared for different types of motion.
Alberto Del Bimbo, Paolo Nesi, Jorge L. C. Sanz
IEEE Trans. Image Process.1
1995 Combining Head Tracking and Pupil Monitoring in Vision-Based Human-Computer Interaction
Carlo Colombo, Alberto Del Bimbo, Silvio De Magistris
CAIP2
1995 Film Editing Reconstruction and Semantic Analysis
Jacopo M. Corridoni, Alberto Del Bimbo
CAIP2
1995 A Robust Algorithm for Optical Flow Estimation
Paolo Nesi, Alberto Del Bimbo, Doron Ben-Tzvi
Comput. Vis. Image Underst.2
1995 Special section on image technology in Italy
Alberto Del Bimbo
Mach. Vis. Appl.1
1995 Recurrent neural networks can be trained to be maximum a posteriori probability classifiers
Simone Santini, Alberto Del Bimbo
Neural Networks2
1995 Properties of block feedback neural networks
Simone Santini, Alberto Del Bimbo
Neural Networks2
1995 Block-structured recurrent neural networks
Simone Santini, Alberto Del Bimbo, Ramesh Jain 0001
Neural Networks2
1995 Optical flow by nonlinear relaxation
Carlo Colombo, Alberto Del Bimbo, Simone Santini
Pattern Recognit.2
1995 Analysis of optical flow constraints
abstract
Different constraint equations have been proposed in the literature for the derivation of optical flow. Despite of the large number of papers dealing with computational techniques to estimate optical flow, only a few authors have investigated conditions under which these constraints exactly model the velocity field, that is, the perspective projection on the image plane of the true 3-D velocity. These conditions are analyzed under different hypotheses, and the departures of the constraint equations in modeling the velocity field are derived for different motion conditions. Experiments are also presented giving measures of these departures and of the induced errors in the estimation of the velocity field.
Alberto Del Bimbo, Paolo Nesi, Jorge L. C. Sanz
IEEE Trans. Image Process.1
1995 Symbolic Description and Visual Querying of Image Sequences Using Spatio-Temporal Logic
abstract
The emergence of advanced multimedia applications is emphasizing the relevance of retrieval by contents within databases of images and image sequences. Matching the inherent visuality of the information stored in such databases, visual specification by example provides an effective and natural way to express content-oriented queries. To support this querying approach, the system must be able to interpret example scenes reproducing the contents of images and sequences to be retrieved, and to match them against the actual contents of the database. In the accomplishment of this task, to avoid a direct access to raw image data, the system must be provided with an appropriate description language supporting the representation of the contents of pictorial data. An original language for the symbolic representation of the contents of image sequences is presented. This language, referred to as spatio-temporal logic, comprises a framework for the qualitative representation of the contents of image sequences, which allows for treatment and operation of content structures at a higher level than pixels or image features. Organization and operation principles of a prototype system exploiting spatio-temporal logic to support querying by example through visual iconic interaction are expounded.>
Alberto Del Bimbo, Enrico Vicario, Daniele Zingoni
IEEE Trans. Knowl. Data Eng.1
1995 Specification by-Example of Virtual Agents Behavior
abstract
The development of virtual agents running within graphic environments which emulate real-life contexts may largely benefit from the use of visual specification by-example. To support this specification, the development system must be able to interpret the examples and cast their underlying rules into an internal representation language. This language must find a suitable trade-off among a number of contrasting requirements regarding expressiveness, automatic executability, and suitability to the automatic representation of rules deriving from the analysis of examples. A language is presented which attains this trade-off by combining together an operational and a declarative fragment to separately represent the autonomous execution of each individual agent and its interaction with the environment, respectively. While the declarative part permits to capture interaction rules emerging from specification examples, the operational part supports the automatic execution in the operation of the virtual environment. A system is presented which embeds this language within a visual shell to support a behavioral training in which the animation rules of virtual agents are defined through visual examples.
Alberto Del Bimbo, Enrico Vicario
IEEE Trans. Vis. Comput. Graph.1
1994 OCR from poor quality images by deformation of elastic templates
abstract
We present a method for elastic pattern matching that we apply to segmentation-free OCR. We use a number of digit models (templates) that are allowed to undergo elastic deformation while trying to adapt to the digit image they are exposed to. Each template must satisfy two competing requirements: (1) it must match the image as much as possible and (2) it must keep its elastic deformation energy as little as possible. These two requirements can be expressed in the form of a variational problem, whose solution gives the optimal deformation for a template.
Alberto Del Bimbo, Simone Santini, Jorge L. C. Sanz
ICPR (2)1
1994 Optical flow through relaxation in the velocity space
Carlo Colombo, Alberto Del Bimbo, Simone Santini
Pattern Recognit. Lett.2
1994 Transport measurements over an Ethernet LAN
abstract
The performance evaluation of transport protocols is a relevant step for designing efficient distributed applications. This analysis can be performed either with an unloaded medium or in the presence of background load. In this latter case, the effects of the interaction with other users are made evident, but remarkable difficulties exist to obtain a reproducible network load and to achieve accuracy in measurements. In this paper, performance measurements in the presence of background load are presented for TCP and UDP in the Internet family. And SPP and IDP in the Xerox Network System family. Experiments were carried out over an Ethernet LAN with MS-DOS AT-class stations.>
Alberto Del Bimbo, Enrico Vicario
IEEE Trans. Commun.1
1994 Performance Analysis of Two Different Algorithms for Ethernet-FDDI Interconnection
abstract
Fiber Distributed Data Interface (FDDI) local area networks (LAN's) are used either as high-speed links between computers and peripherals, or as backbones for lower-speed LAN's, such as Ethernet and Token Ring. The availability of such a high-speed channel will lead to the implementation of high-performance distributed environments spread over a wider area than that allowed by commonly used LAN's. The performance of such distributed environments will strongly depend on that of the interconnecting devices. In this paper, two different algorithms for packet filtering are discussed, referring to bridges interconnecting Ethernet LAN's to FDDI backbones. Algorithm performance is compared with respect to 1) the traffic increase produced on a local Ethernet, and 2) the maximum allowed traffic on remote Ethernets.>
Giacomo Bucci, Alberto Del Bimbo, Simone Santini
IEEE Trans. Parallel Distributed Syst.2
1993 Optical flow from constraint lines parametrization
Doron Ben-Tzvi, Alberto Del Bimbo, Paolo Nesi
Pattern Recognit.2
1993 A programming environment for imaging applications
Roberto Cecchini, Alberto Del Bimbo
Pattern Recognit. Lett.2
1993 Determination of road directions using feedback neural nets
Alberto Del Bimbo, Leonardo Landi, Simone Santini
Signal Process.1
1993 A Three-Dimensional Iconic Environment for Image Database Querying
abstract
Retrieval by contents of images from pictorial databases can be effectively performed through visual icon-based systems. In these systems, the representation of pictures with 2D strings, which are derived from symbolic projections, provides an efficient and natural way to construct iconic indexes for pictures and is also an ideal representation for the visual query. With this approach, retrieval is reduced to matching two symbolic strings. However, using 2D-string representations, spatial relationships between the objects represented in the image might not be exactly specified. Ambiguities arise for the retrieval of images of 3D scenes. In order to allow the unambiguous description of object spatial relationships, in this paper, following the symbolic projections approach, images are referred to by considering spatial relationships in the 3D imaged scene. A representation language is introduced that expresses positional and directional relationships between objects in three dimensions, still preserving object spatial extensions after projections. Iconic retrieval from pictorial databases with 3D interfaces is discussed and motivated. A system for querying by example with 3D icons, which supports this language, is also presented.>
Alberto Del Bimbo, Maurizio Campanai, Paolo Nesi
IEEE Trans. Software Eng.1
1992 Dynamic neural estimation for autonomous vehicles driving
abstract
Mobile robots and vehicles may be driven by dynamical neural networks which utilize image data of real-world scenes collected through a TV camera for learning and performance. An innovative system for road direction detection is proposed which is comprised of three specialized blocks performing edge extraction, image-segments detection, and road direction estimation. The road direction estimation block is implemented as a feedback neural network.>
Alberto Del Bimbo, Leonardo Landi, Simone Santini
ICPR (2)1
1992 A multilayer massively parallel architecture for optical flow computation
abstract
A two-layers architecture for optical flow computation is presented, which uses neural nets in the lower layer and a special relaxation system in the upper layer. The high degree of parallelism of the architecture makes it particularly suitable for real-time applications.>
Carlo Colombo, Alberto Del Bimbo, Simone Santini
ICPR (4)2
1992 On the perspective projection of 3D planar-faced junctions
abstract
A method to determine the spatial orientation of a 3D planar-faced object from a single perspective view is presented. In this method the junctions, that is the concurrences of three object edges in a vertex, are considered as key features. A geometric constraint, related to the perspective projection of the object junctions, is shown and exploited to make the process faster and more efficient. In this way, the knowledge of only two parameters is sufficient to verify the object orientation in 3D space. A system of equations is proposed which define the orientations space for a junction as a subset of R/sup 2/, instead of R/sup 3/ as in the previous literature.>
R. Consales, Alberto Del Bimbo, Paolo Nesi
ICPR (1)2
1992 Behavioral object recognition from multiple image frames
Alberto Del Bimbo, Paolo Nesi
Signal Process.1
1991 Interprocess Communication Dependency on Network Load
abstract
Results of an analysis of the communication performance as perceived by the application layer in the presence of background load are presented. The analysis was carried out on two distinct protocol suites, TCP/IP and XNS, on a single Ethernet LAN with personal computer workstations. Service times for a reliable transfer and an unreliable transmission as perceived by the application layer were measured. An appropriate model of the reliable transfer process was used to derive service time analytical estimates. Owing to the close relation between measured and estimated values, the model was used to identify the major sources of delay and to evaluate the effects of possible improvements. The most important causes affecting the transport layer performance as background load increases are identified, and different strategies are examined in order to reduce their impact. The improvement margin offered by the implementation of a lightweight transport protocol specifically tailored to the requirements of a bulk data transfer over an Ethernet is addressed.>
Alessandro Braccini, Alberto Del Bimbo, Enrico Vicario
IEEE Trans. Software Eng.2
1990 Integrating object oriented programming paradigm concepts in designing a vision and pattern recognition system architecture
abstract
A fully object-oriented approach to a general-purpose object-recognition system is presented. This, together with a blackboard architecture, provides a flexible and easily maintainable system. The knowledge information is distributed into two object-oriented databases: one, hierarchically structures into three layers according to visual abstractions, contains the object models (the knowledge source of the recognition process); the other, which reflects in the most natural way the problem domain knowledge, helps in the pruning of the decision trees and allows for the analysis of the behavior of the objects.>
V. Cappelini, Alberto Del Bimbo, Paolo Nesi
ICPR (2)2
1990 CSL: A class specification language for object-oriented design
Giacomo Bucci, Roberto Cecchini, Alberto Del Bimbo
Microprocessing and Microprogramming3
1989 An object oriented approach for object recognition and classification
abstract
Scene analysis with processing of multiple frames is considered for automatic recognition and tracking of moving objects. The importance of taking into account information about the objects in the design of flexible systems is emphasized. This goal is reached through an object-oriented approach to problem-solving. It is shown that abstract data type and inheritance relationship definitions over real-world entities can be efficiently used for object classification purposes. This has been verified in a traffic monitoring application. An efficient algorithmic solution that satisfies the real-time constraints of the problem is proposed.>
Vito Cappellini, Alberto Del Bimbo, Alessandro Mecocci
ICASSP2
1984 Two special digital processing systems for recognition of moving objects
abstract
Some special digital operators and recognition techniques are presented for digital image processing. Two complete digital systems are described. The first fast and relatively simple system uses non-linear smoothing and adaptive thresholding as a pre-processing and inertial invariant components to recognize moving objects. The second more complex system uses non-linear smoothing, edge detection and noise spike reduction as a pre-processing and FFT of the boundary distances from the centroid to recognize and track moving objects. Experiments of recognition and tracking of mechanical objects are reported.
P. Borghesi, Vito Cappellini, Alberto Del Bimbo, Alessandro Mecocci
ICASSP3
1984 Object decomposition and subpart identification - classification algorithms
Vito Cappellini, Alberto Del Bimbo, Alessandro Mecocci
Image Vis. Comput.2