Rita Cucchiara

dblp:c/RitaCucchiara · DBLP profile ↗
← Back
291ranked-venue papers
26as first author
104since 2021 · last 2026
0000-0002-2239-283XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 198 · 13 first-author · 65 since 2021Artificial intelligence and machine learning · 165 · 11 first-author · 77 since 2021Databases, data management, data science and information retrieval · 11 · 4 since 2021Computer networks · 10 · 1 first-author · 7 since 2021Systems, architecture and hardware · 7 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
Silvia Cappelletti, Tobia Poppi, Samuele Poppi, Diego Garcia-Olano, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
ICPR (3)8
2026 HyperMIL: Hypergraph-Based Channel Reasoning for Multiple Instance Learning on Multivariate Time Series
Livia Del Gaudio, Vittorio Cuculo, Rita Cucchiara
ICPR (3)3
2026 RaTA-Tool: Retrieval-Based Tool Selection with Multimodal Large Language Models
Gabriele Mattioli, Evelyn Turri, Sara Sarto, Lorenzo Baraldi 0002, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
ICPR (7)7
2026 Sketch2Stitch: GANs for Abstract Sketch-Based Dress Synthesis
abstract
In the realm of creative expression, not everyone possesses the gift of effortlessly translating their imaginative visions into flawless sketches. More often than not, the outcome resembles an abstract, perhaps even slightly distorted representation. The art of producing impeccable sketches is not only challenging but also a time-consuming process. Our work is the first of this kind in transforming abstract, sometimes deformed garment sketches into photorealistic catalog images, to empower the everyday individual to become their own fashion designer. We create Sketch2Stitch, a dataset featuring over 65,000 abstract sketch images generated from garments of Dress Code [40] and VITON HD [6], two benchmark datasets in the virtual try-on task. Sketch2Stitch is the first dataset in the literature to provide abstract sketches in the fashion domain. We propose a StyleGAN-based generative framework that bridges freehand sketching with photorealistic garment synthesis. We demonstrate that our framework allows users to sketch rough outlines and optionally provide color hints, producing realistic designs in seconds. Experimental results demonstrate, both quantitatively and qualitatively, that the proposed framework achieves superior performance against various baselines and existing methods on both subsets of our dataset. Our work highlights a pathway toward AI-assisted fashion design tools, democratizing garment ideation for students, independent designers, and casual creators.
Faizan Farooq Khan, Eslam Abdelrahman, Davide Morelli, Marcella Cornia, Rita Cucchiara, Mohamed Elhoseiny 0001
WACV5
2026 Autoregressive Styled Text Image Generation, but Make it Reliable
abstract
Generating faithful and readable styled text images (especially for Styled Handwritten Text generation-HTG) is an open problem with several possible applications across graphic design, document understanding, and image editing. A lot of research effort in this task is dedicated to developing strategies that reproduce the stylistic characteristics of a given writer, with promising results in terms of style fidelity and generalization achieved by the recently proposed Autoregressive Transformer paradigm for HTG. However, this method requires additional inputs, lacks a proper stop mechanism, and might end up in repetition loops, generating visual artifacts. In this work, we rethink the autoregressive formulation by framing HTG as a multimodal prompt-conditioned generation task, and tackle the content controllability issues by introducing special textual input tokens for better alignment with the visual ones. Moreover, we devise a Classifier-Free-Guidance-based strategy for our autoregressive model. Through extensive experimental validation, we demonstrate that our approach, dubbed Eruku, compared to previous solutions requires fewer inputs, generalizes better to unseen styles, and follows more faithfully the textual prompt, improving content adherence.
Carmine Zaccagnino, Fabio Quattrini, Vittorio Pippi, Silvia Cascianelli, Alessio Tonioni, Rita Cucchiara
WACV6
2026 Hallucination Early Detection in Diffusion Models
Federico Betti 0001, Lorenzo Baraldi 0002, Lorenzo Baraldi 0001, Rita Cucchiara, Nicu Sebe
Int. J. Comput. Vis.4
2026 A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models
abstract
In federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Vision (CV) and Natural Language Processing (NLP) tasks. It has been revealed that the gradient of Position Embeddings (PEs) in Transformer contains sufficient information, which can be used to reconstruct the input data. To mitigate this issue, we introduce a Masked Jigsaw Puzzle (MJP) framework. MJP starts with random token shuffling to break the token order, and then a learnable unknown (unk) position embedding is used to mask out the PEs of the shuffled tokens. In this manner, the local spatial information which is encoded in the position embeddings is disrupted, and the models are forced to learn feature representations that are less reliant on the local spatial information. Notably, with the careful use of MJP, we can not only improve models' robustness against gradient attacks, but also boost their performance in both vision and text application scenarios, such as classification for images (e.g., ImageNet-1 K) and sentiment analysis for text (e.g., Yelp and Amazon). Experimental results suggest that MJP is a unified framework for different Transformer-based models in both vision and language tasks.
Weixin Ye, Wei Wang 0108, Yue Song 0002, Bin Ren 0005, Wei Bi, Rita Cucchiara, Nicu Sebe
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing
abstract
Fashion illustration is a crucial medium for designers to convey their creative vision and transform design concepts into tangible representations that showcase the interplay between clothing and the human body. In the context of fashion design, computer vision techniques have the potential to enhance and streamline the design process. Departing from prior research primarily focused on virtual try-on, this article tackles the task of multimodal-conditioned fashion image editing. Our approach aims to generate human-centric fashion images guided by multimodal prompts, including text, human body poses, garment sketches, and fabric textures. To address this problem, we propose extending latent diffusion models to incorporate these multiple modalities and modifying the structure of the denoising network, taking multimodal prompts as input. To condition the proposed architecture on fabric textures, we employ textual inversion techniques and let diverse cross-attention layers of the denoising network attend to textual and texture information, thus incorporating different granularity conditioning details. Given the lack of datasets for the task, we extend two existing fashion datasets, Dress Code and VITON-HD, with multimodal annotations. Experimental evaluations demonstrate the effectiveness of our proposed approach in terms of realism and coherence concerning the provided multimodal inputs.
Alberto Baldrati, Davide Morelli, Marcella Cornia, Marco Bertini 0001, Rita Cucchiara
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval
abstract
Cross-modal retrieval is gaining increasing efficacy and interest from the research community, thanks to large-scale training, novel architectural and learning designs, and its application in LLMs and multimodal LLMs. In this paper, we move a step forward and design an approach that allows for multimodal queries – composed of both an image and a text – and can search within collections of multi-modal documents, where images and text are interleaved. Our model, ReT, employs multi-level representations extracted from different layers of both visual and textual backbones, both at the query and document side. To allow for multi-level and cross-modal understanding and feature extraction, ReT employs a novel Transformer-based recurrent cell that integrates both textual and visual features at different layers, and leverages sigmoidal gates inspired by the classical design of LSTMs. Extensive experiments on M2KR and M-BEIR benchmarks show that ReT achieves state-of-the-art performance across diverse settings. Our source code and trained models are publicly available at: https://github.com/aimagelab/ReT.
Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
CVPR5
2025 Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
abstract
Multimodal LLMs (MLLMs) are the natural extension of large language models to handle multimodal inputs, combining text and image data. They have recently garnered attention due to their capability to address complex tasks involving both modalities. However, their effectiveness is limited to the knowledge acquired during training, which restricts their practical utility. In this work, we introduce a novel method to enhance the adaptability of MLLMs by integrating external knowledge sources. Our proposed model, Reflective LLaVA (ReflectiVA), utilizes reflective tokens to dynamically determine the need for external knowledge and predict the relevance of information retrieved from an external database. Tokens are trained following a two-stage two-model training recipe. This ultimately enables the MLLM to manage external knowledge while preserving fluency and performance on tasks where external knowledge is not needed. Through our experiments, we demonstrate the efficacy of ReflectiVA for knowledge-based visual question answering, highlighting its superior performance compared to existing methods. Source code and trained models are publicly available at https://aimagelab.github.io/ReflectiVA.
Federico Cocchi, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
CVPR5
2025 Zero-Shot Styled Text Image Generation, but Make It Autoregressive
abstract
Styled Handwritten Text Generation (HTG) has recently received attention from the computer vision and document analysis communities, which have developed several solutions, either GAN- or diffusion-based, that achieved promising results. Nonetheless, these strategies fail to generalize to novel styles and have technical constraints, particularly in terms of maximum output length and training efficiency. To overcome these limitations, in this work, we propose a novel framework for text image generation, dubbed Emuru. Our approach leverages a powerful text image representation model (a variational autoencoder) combined with an autoregressive Transformer. Our approach enables the generation of styled text images conditioned on textual content and style examples, such as specific fonts or handwriting styles. We train our model solely on a diverse, synthetic dataset of English text rendered in over 100,000 typewritten and calligraphy fonts, which gives it the capability to reproduce unseen styles (both fonts and users’ handwriting) in zero-shot. To the best of our knowledge, Emuru is the first autoregressive model for HTG, and the first designed specifically for generalization to novel styles. Moreover, our model generates images without background artifacts, which are easier to use for downstream applications. Extensive evaluation on both typewritten and handwritten, anylength text image generation scenarios demonstrates the effectiveness of our approach.
Vittorio Pippi, Fabio Quattrini, Silvia Cascianelli, Alessio Tonioni, Rita Cucchiara
CVPR5
2025 Hyperbolic Safety-Aware Vision-Language Models
abstract
Addressing the retrieval of unsafe content from vision-language models such as CLIP is an important step towards real-world integration. Current efforts have relied on unlearning techniques that try to erase the model’s knowledge of unsafe concepts. While effective in reducing unwanted outputs, unlearning limits the model’s capacity to discern between safe and unsafe content. In this work, we introduce a novel approach that shifts from unlearning to an awareness paradigm by leveraging the inherent hierarchical properties of the hyperbolic space. We propose to encode safe and unsafe content as an entailment hierarchy, where both are placed in different regions of hyperbolic space. Our HySAC, Hyperbolic Safety-Aware CLIP, employs entailment loss functions to model the hierarchical and asymmetrical relations between safe and unsafe image-text pairs. This modelling – ineffective in standard vision-language models due to their reliance on Euclidean embeddings – endows the model with awareness of unsafe content, enabling it to serve as both a multimodal unsafe classifier and a flexible content retriever, with the option to dynamically redirect unsafe queries toward safer alternatives or retain the original output. Extensive experiments show that our approach not only enhances safety recognition but also establishes a more adaptable and interpretable framework for content moderation in vision-language models. Our source code is available at: https://github.com/aimagelab/HySAC
Tobia Poppi, Tejaswi Kasarla, Pascal Mettes, Lorenzo Baraldi 0001, Rita Cucchiara
CVPR5
2025 Generating Synthetic Data with Large Language Models for Low-Resource Sentence Retrieval
Davide Caffagni, Federico Cocchi, Anna Mambelli, Fabio Tutrone, Marco Zanella, Marcella Cornia, Rita Cucchiara
TPDL7
2025 Multimodal Emotion Recognition in Conversation via Possible Speaker's Audio and Visual Sequence Selection
abstract
Multimodal Emotion Recognition in Conversation (MERC) is an important element in human-machine interaction. It allows machines to automatically identify and track the emotional status of speakers during a conversation in a multimodal setting. However, the conversations involving various audio and visual cues aligned with textual cues are very complex. Recent works have tried integrating the audio and visual modalities with textual to improve the performance of emotion recognition in conversation. Although many MERC models leverage textual, audio, and visual modalities, those models assume that the speaker’s textual utterance, audio speech, and facial sequences are present. However, a conversation may contain multiple parties, among which only one is the speaker. Previous MERC assumed the availability of all modalities, but in many instances, one or more modalities may be unavailable during multiparty conversations. To tackle these issues, we propose the Possible Speaker Informed Multimodal Emotion Recognition in Conversation framework (PSI). PSI is specifically tasked to extract audio (speech) and visual (face) sequences of a possible speaker in the presence of multiple parties. Further, PSI seamlessly extracts the rich unimodal features and fuses them while addressing the unavailability of specific modalities. PSI demonstrates competitive performance with existing state-of-the-art models through experiments with a benchmark dataset.
Rahul Singh Maharjan, Niyati Rawal, Marta Romeo, Lorenzo Baraldi 0001, Rita Cucchiara, Angelo Cangelosi
ICASSP5
2025 What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
abstract
Instruction-based image editing models offer increased personalization opportunities in generative tasks. However, properly evaluating their results is challenging, and most of the existing metrics lag in terms of alignment with human judgment and explainability. To tackle these issues, we introduce DICE (DIfference Coherence Estimator), a model designed to detect localized differences between the original and the edited image and to assess their relevance to the given modification request. DICE consists of two key components: a difference detector and a coherence estimator, both built on an autoregressive Multimodal Large Language Model (MLLM) and trained using a strategy that leverages self-supervision, distillation from inpainting networks, and full supervision. Through extensive experiments, we evaluate each stage of our pipeline, comparing different MLLMs within the proposed framework. We demonstrate that DICE effectively identifies coherent edits, effectively evaluating images generated by different editing models with a strong correlation with human judgment. We publicly release our source code, models, and data.
Lorenzo Baraldi 0001, Davide Bucciarelli, Federico Betti 0001, Marcella Cornia, Lorenzo Baraldi 0002, Nicu Sebe, Rita Cucchiara
ICCV7
2025 Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation
abstract
Open-Vocabulary Segmentation (OVS) aims at segmenting images from free-form textual concepts without predefined training classes. While existing vision-language models such as CLIP can generate segmentation masks by leveraging coarse spatial information from Vision Transformers, they face challenges in spatial localization due to their global alignment of image and text features. Conversely, self-supervised visual models like DINO excel in fine-grained visual encoding but lack integration with language. To bridge this gap, we present Talk2DINO, a novel hybrid approach that combines the spatial accuracy of DINOv2 with the language understanding of CLIP. Our approach aligns the textual embeddings of CLIP to the patch-level features of DINOv2 through a learned mapping function without the need to fine-tune the underlying backbones. At training time, we exploit the attention maps of DINOv2 to selectively align local visual patches with textual embeddings. We show that the powerful semantic and localization abilities of Talk2DINO can enhance the segmentation process, resulting in more natural and less noisy segmentations, and that our approach can also effectively distinguish foreground objects from the background. Experimental results demonstrate that Talk2DINO achieves state-of-the-art performance across several unsupervised OVS benchmarks. Source code and models are publicly available at: https://lorebianchi98.github.io/Talk2DINO/.
Luca Barsellotti, Lorenzo Bianchi 0001, Nicola Messina, Fabio Carrara, Marcella Cornia, Lorenzo Baraldi 0001, Fabrizio Falchi, Rita Cucchiara
ICCV8
2025 Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
Giuseppe Cartella, Vittorio Cuculo, Alessandro D'Amelio, Marcella Cornia, Giuseppe Boccignone, Rita Cucchiara
ICCV6
2025 MISSRAG: Addressing the Missing Modality Challenge in Multimodal Large Language Models
Vittorio Pipoli, Alessia Saporita, Federico Bolelli, Marcella Cornia, Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara, Elisa Ficarra
ICCV7
2025 Diffusion Transformers for Tabular Data Time Series Generation
abstract
Tabular data generation has recently attracted a growing interest due to its different application scenarios. However, generating time series of tabular data, where each element of the series depends on the others, remains a largely unexplored domain. This gap is probably due to the difficulty of jointly solving different problems, the main of which are the heterogeneity of tabular data (a problem common to non-time-dependent approaches) and the variable length of a time series. In this paper, we propose a Diffusion Transformers (DiTs) based approach for tabular data series generation. Inspired by the recent success of DiTs in image and video generation, we extend this framework to deal with heterogeneous data and variable-length sequences. Using extensive experiments on six datasets, we show that the proposed approach outperforms previous work by a large margin.
Fabrizio Garuti, Enver Sangineto, Simone Luetto, Lorenzo Forni, Rita Cucchiara
ICLR5
2025 Causal Graphical Models for Vision-Language Compositional Understanding
abstract
Recent work has empirically shown that Vision-Language Models (VLMs) struggle to fully understand the compositional properties of the human language, usually modeling an image caption as a “bag of words”. As a result, they perform poorly on compositional tasks, which require a deeper understanding of the different entities of a sentence (subject, verb, etc.) jointly with their mutual relationships in order to be solved. In this paper, we model the dependency relations among textual and visual tokens using a Causal Graphical Model (CGM), built using a dependency parser, and we train a decoder conditioned by the VLM visual encoder. Differently from standard autoregressive or parallel predictions, our decoder’s generative process is partially-ordered following the CGM structure. This structure encourages the decoder to learn only the main causal dependencies in a sentence discarding spurious correlations. Using extensive experiments on five compositional benchmarks, we show that our method significantly outperforms all the state-of-the-art compositional approaches by a large margin, and it also improves over methods trained using much larger datasets. Our model weights and code are publicly available.
Fiorenzo Parascandolo, Nicholas Moratelli, Enver Sangineto, Lorenzo Baraldi 0001, Rita Cucchiara
ICLR5
2025 A Second-Order Perspective on Model Compositionality and Incremental Learning
abstract
The fine-tuning of deep pre-trained models has revealed compositional properties, with multiple specialized modules that can be arbitrarily composed into a single, multi-task model. However, identifying the conditions that promote compositionality remains an open issue, with recent efforts concentrating mainly on linearized networks. We conduct a theoretical study that attempts to demystify compositionality in standard non-linear networks through the second-order Taylor approximation of the loss function. The proposed formulation highlights the importance of staying within the pre-training basin to achieve composable modules. Moreover, it provides the basis for two dual incremental training algorithms: the one from the perspective of multiple models trained individually, while the other aims to optimize the composed model as a whole. We probe their application in incremental classification tasks and highlight some valuable skills. In fact, the pool of incrementally learned modules not only supports the creation of an effective multi-task model but also enables unlearning and specialization in certain tasks. Code available at <https://github.com/aimagelab/mammoth>
Angelo Porrello, Lorenzo Bonicelli, Pietro Buzzega, Monica Millunzi, Simone Calderara, Rita Cucchiara
ICLR6
2025 Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
abstract
The evaluation of machine-generated captions is a complex and evolving challenge. With the advent of Multimodal Large Language Models (MLLMs), image captioning has become a core task, increasing the need for robust and reliable evaluation metrics. This survey provides a comprehensive overview of advancements in image captioning evaluation, analyzing the evolution, strengths, and limitations of existing metrics. We assess these metrics across multiple dimensions, including correlation with human judgment, ranking accuracy, and sensitivity to hallucinations. Additionally, we explore the challenges posed by the longer and more detailed captions generated by MLLMs and examine the adaptability of current metrics to these stylistic variations. Our analysis highlights some limitations of standard evaluation approaches and suggests promising directions for future research in image captioning assessment. For a comprehensive overview of captioning evaluation refer to our project page available at https://github.com/aimagelab/awesome-captioning-evaluation.
Sara Sarto, Marcella Cornia, Rita Cucchiara
IJCAI3
2025 Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
abstract
In recent years, the fashion industry has increasingly adopted AI technologies to enhance customer experience, driven by the proliferation of e-commerce platforms and virtual applications. Among the various tasks, virtual try-on and multimodal fashion image editing–which utilizes diverse input modalities such as text, garment sketches, and body poses–have become a key area of research. Diffusion models have emerged as a leading approach for such generative tasks, offering superior image quality and diversity. However, most existing virtual try-on methods rely on having a specific garment input, which is often impractical in real-world scenarios where users may only provide textual specifications. To address this limitation, in this work we introduce Fashion Retrieval-Augmented Generation (Fashion-RAG), a novel method that enables the customization of fashion items based on user preferences provided in textual form. Our approach retrieves multiple garments that match the input specifications and generates a personalized image by incorporating attributes from the retrieved items. To achieve this, we employ textual inversion techniques, where retrieved garment images are projected into the textual embedding space of the Stable Diffusion text encoder, allowing seamless integration of retrieved elements into the generative process. Experimental results on the Dress Code dataset demonstrate that Fashion-RAG outperforms existing methods both qualitatively and quantitatively, effectively capturing fine-grained visual details from retrieved garments. To the best of our knowledge, this is the first work to introduce a retrieval-augmented generation approach specifically tailored for multimodal fashion image editing.
Fulvio Sanguigni, Davide Morelli, Marcella Cornia, Rita Cucchiara
IJCNN4
2025 BRUM: Robust 3D Vehicle Reconstruction from 360° Sparse Images
abstract
Accurate 3D reconstruction of vehicles is vital for applications such as vehicle inspection, predictive maintenance, and urban planning. Existing methods like Neural Radiance Fields and Gaussian Splatting have shown impressive results but remain limited by their reliance on dense input views, which hinders real-world applicability. This paper addresses the challenge of reconstructing vehicles from sparse-view inputs, leveraging depth maps and a robust pose estimation architecture to synthesize novel views and augment training data. Specifically, we enhance Gaussian Splatting by integrating a selective photometric loss, applied only to high-confidence pixels, and replacing standard Structure-from-Motion pipelines with the DUSt3R architecture to improve camera pose estimation. Furthermore, we present a novel dataset featuring both synthetic and real-world public transportation vehicles, enabling extensive evaluation of our approach. Experimental results demonstrate state-of-the-art performance across multiple benchmarks, showcasing the method's ability to achieve high-quality reconstructions even under constrained input conditions. Code and data are publicly available at https://aimagelab.ing.unimore.it/go/brum.
Davide Di Nucci, Matteo Tomei, Guido Borghi, Luca Ciuffreda, Roberto Vezzani, Rita Cucchiara
IV6
2025 DitHub: A Modular Framework for Incremental Open-Vocabulary Object Detection
abstract
Open-Vocabulary object detectors can generalize to an unrestricted set of categories through simple textual prompting. However, adapting these models to rare classes or reinforcing their abilities on multiple specialized domains remains essential. While recent methods rely on monolithic adaptation strategies with a single set of weights, we embrace modular deep learning. We introduce DitHub, a framework designed to build and maintain a library of efficient adaptation modules. Inspired by Version Control Systems, DitHub manages expert modules as branches that can be fetched and merged as needed. This modular approach allows us to conduct an in-depth exploration of the compositional properties of adaptation modules, marking the first such study in Object Detection. Our method achieves state-of-the-art performance on the ODinW-13 benchmark and ODinW-O, a newly introduced benchmark designed to assess class reappearance.
Chiara Cappellino, Gianluca Mancusi, Matteo Mosconi, Angelo Porrello, Simone Calderara, Rita Cucchiara
NeurIPS6
2025 Perceive. Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries
abstract
Video Question Answering (Video QA) is a challenging video understanding task that requires models to compre-hend entire videos, identify the most relevant information based on contextual cues from a given question, and rea-son accurately to provide answers. Recent advancements in Multimodal Large Language Models (MLLMs) have trans-formed video QA by leveraging their exceptional common-sense reasoning capabilities. This progress is largely driven by the effective alignment between visual data and the language space of MLLMs. However, for video QA, an ad-ditional space-time alignment poses a considerable chal-lenge for extracting question-relevant information across frames. In this work, we investigate diverse temporal modeling techniques to integrate with MLLMs, aiming to achieve question-guided temporal modeling that leverages pre-trained visual and textual alignment in MLLMs. We propose T-Former, a novel temporal modeling method that creates a question-guided temporal bridge between frame-wise visual perception and the reasoning capabilities of LLMs. Our evaluation across multiple video QA bench-marks demonstrates that T-Former competes favorably with existing temporal modeling approaches and aligns with re-cent advancements in video QA.
Roberto Amoroso, Gengyuan Zhang, Rajat Koner, Lorenzo Baraldi 0001, Rita Cucchiara, Volker Tresp
WACV5
2025 TPP-Gaze: Modelling Gaze Dynamics in Space and Time with Neural Temporal Point Processes
abstract
Attention guides our gaze to fixate the proper location of the scene and holds it in that location for the de-served amount of time given current processing demands, before shifting to the next one. As such, gaze deploy-ment crucially is a temporal process. Existing computational models have made significant strides in predicting spatial aspects of observer's visual scanpaths (where to look), while often putting on the background the tempo-ral facet of attention dynamics (when). In this paper we present TPP-Gaze, a novel and principled approach to model scanpath dynamics based on Neural Temporal Point Process (TPP), that Jointly learns the temporal dynamics of fixations position and duration, integrating deep learning methodologies with point process theory. We conduct ex-tensive experiments across five publicly available datasets. Our results show the overall superior performance of the proposed model compared to state-of-the-art approaches. Source code and trained models are publicly available at: https://github.com/phuselab/tppgaze.
Alessandro D'Amelio, Giuseppe Cartella, Vittorio Cuculo, Manuele Lucchi, Marcella Cornia, Rita Cucchiara, Giuseppe Boccignone
WACV6
2025 Semantically Conditioned Prompts for Visual Recognition Under Missing Modality Scenarios
abstract
This paper tackles the domain of multimodal prompting for visual recognition, specifically when dealing with missing modalities through multimodal Transformers. It presents two main contributions: (i) we introduce a novel prompt learning module which is designed to produce sample-specific prompts and (ii) we show that modalityagnostic prompts can effectively adjust to diverse missing modality scenarios. Our model, termed SCP, exploits the semantic representation of available modalities to query a learnable memory bank, which allows the generation of prompts based on the semantics of the input. Notably, SCP distinguishes itself from existing methodologies for its capacity of self-adjusting to both the missing modality scenario and the semantic context of the input, without prior knowledge about the specific missing modality and the number of modalities. Through extensive experiments, we show the effectiveness of the proposed prompt learning framework and demonstrate enhanced performance and robustness across a spectrum of missing modality cases. Our source code is available at https://github.com/vittoriopipoli/SCP_WACV2025.
Vittorio Pipoli, Federico Bolelli, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara, Elisa Ficarra
WACV7
2025 Learning to mask and permute visual tokens for Vision Transformer pre-training
abstract
The use of self-supervised pre-training has emerged as a promising approach to enhance the performance of many different visual tasks. In this context, recent approaches have employed the Masked Image Modeling paradigm, which pre-trains a backbone by reconstructing visual tokens associated with randomly masked image patches. This masking approach, however, introduces noise into the input data during pre-training, leading to discrepancies that can impair performance during the fine-tuning phase. Furthermore, input masking neglects the dependencies between corrupted patches, increasing the inconsistencies observed in downstream fine-tuning tasks. To overcome these issues, we propose a new self-supervised pre-training approach, named Masked and Permuted Vision Transformer (MaPeT), that employs autoregressive and permuted predictions to capture intra-patch dependencies. In addition, MaPeT employs auxiliary positional information to reduce the disparity between the pre-training and fine-tuning phases. In our experiments, we employ a fair setting to ensure reliable and meaningful comparisons and conduct investigations on multiple visual tokenizers, including our proposed k -CLIP which directly employs discretized CLIP features. Our results demonstrate that MaPeT achieves competitive performance on ImageNet, compared to baselines and competitors under the same model setting. We release an implementation of our code and models at https://github.com/aimagelab/MaPeT .
Lorenzo Baraldi 0002, Roberto Amoroso, Marcella Cornia, Lorenzo Baraldi 0001, Andrea Pilzer, Rita Cucchiara
Comput. Vis. Image Underst.6
2025 Monocular per-object distance estimation with Masked Object Modeling
abstract
Per-object distance estimation is critical in surveillance and autonomous driving, where safety is crucial. While existing methods rely on geometric or deep supervised features, only a few attempts have been made to leverage self-supervised learning. In this respect, our paper draws inspiration from Masked Image Modeling (MiM) and extends it to multi-object tasks . While MiM focuses on extracting global image-level representations, it struggles with individual objects within the image. This is detrimental for distance estimation, as objects far away correspond to negligible portions of the image. Conversely, our strategy, termed Masked Object Modeling ( MoM ), enables a novel application of masking techniques. In a few words, we devise an auxiliary objective that reconstructs the portions of the image pertaining to the objects detected in the scene. The training phase is performed in a single unified stage, simultaneously optimizing the masking objective and the downstream loss ( i.e ., distance estimation). We evaluate the effectiveness of MoM on a novel reference architecture (DistFormer) on the standard KITTI, NuScenes, and MOTSynth datasets. Our evaluation reveals that our framework surpasses the SoTA and highlights its robust regularization properties. The MoM strategy enhances both zero-shot and few-shot capabilities, from synthetic to real domain. Finally, it furthers the robustness of the model in the presence of occluded or poorly detected objects. • We introduce a novel self-supervised approach called Masked Object Modeling (MoM). • The paper proposes a novel framework for distance estimation from monocular images. • We show improved transfer capabilities for zero-shot and fine-tuning when using MoM. • We reach state-of-the-art performance on several distance estimation benchmarks.
Aniello Panariello, Gianluca Mancusi, Fedy Haj Ali, Angelo Porrello, Simone Calderara, Rita Cucchiara
Comput. Vis. Image Underst.6
2025 Positive-Augmented Contrastive Learning for Vision-and-Language Evaluation and Training
Sara Sarto, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
Int. J. Comput. Vis.5
2025 Augmenting and mixing Transformers with synthetic data for image captioning
abstract
Image captioning has attracted significant attention within the Computer Vision and Multimedia research domains, resulting in the development of effective methods for generating natural language descriptions of images. Concurrently, the rise of generative models has facilitated the production of highly realistic and high-quality images, particularly through recent advancements in latent diffusion models. In this paper, we propose to leverage the recent advances in Generative AI and create additional training data that can be effectively used to boost the performance of an image captioning model. Specifically, we combine real images with their synthetic counterparts generated by Stable Diffusion using a Mixup data augmentation technique to create novel training examples. Extensive experiments on the COCO dataset demonstrate the effectiveness of our solution in comparison to different baselines and state-of-the-art methods and validate the benefits of using synthetic data to augment the training stage of an image captioning model and improve the quality of the generated captions. Source code and trained models are publicly available at: https://github.com/aimagelab/synthcap_pp .
Davide Caffagni, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
Image Vis. Comput.4
2025 One transformer for all time series: representing and training with time-dependent heterogeneous tabular data
abstract
Abstract There is a recent growing interest in applying Deep Learning techniques to tabular data in order to replicate the success of other Artificial Intelligence areas in this structured domain. Particularly interesting is the case in which tabular data have a time dependence, such as, for instance, financial transactions. However, the heterogeneity of the tabular values, in which categorical elements are mixed with numerical features, makes this adaptation difficult. In this paper we propose UniTTab, a Transformer based architecture whose goal is to uniformly represent heterogeneous time-dependent tabular data, in which both numerical and categorical features are described using continuous embedding vectors. Moreover, differently from common approaches, which use a combination of different loss functions for training with both numerical and categorical targets, UniTTab is uniformly trained with a unique Masked Token pretext task. Finally, UniTTab can also represent time series in which the individual row components have a variable internal structure with a variable number of fields, which is a common situation in many application domains, such as in real world transactional data. Using extensive experiments with five datasets of variable size and complexity, we empirically show that UniTTab consistently and significantly improves the prediction accuracy over several downstream tasks and with respect to both Deep Learning and more standard Machine Learning approaches. Our code and our models are available at: https://github.com/fabriziogaruti/UniTTab .
Simone Luetto, Fabrizio Garuti, Enver Sangineto, Lorenzo Forni, Rita Cucchiara
Mach. Learn.5
2025 VATr++: Choose Your Words Wisely for Handwritten Text Generation
abstract
Styled Handwritten Text Generation (HTG) has received significant attention in recent years, propelled by the success of learning-based solutions employing GANs, Transformers, and, preliminarily, Diffusion Models. Despite this surge in interest, there remains a critical yet understudied aspect - the impact of the input, both visual and textual, on the HTG model training and its subsequent influence on performance. This work extends the VATr (Pippi et al. 2023) Styled-HTG approach by addressing the pre-processing and training issues that it faces, which are common to many HTG models. In particular, we propose generally applicable strategies for input preparation and training regularization that allow the model to achieve better performance and generalization capabilities. Moreover, in this work, we go beyond performance optimization and address a significant hurdle in HTG research - the lack of a standardized evaluation protocol. In particular, we propose a standardization of the evaluation protocol for HTG and conduct a comprehensive benchmarking of existing approaches. By doing so, we aim to establish a foundation for fair and meaningful comparisons between HTG strategies, fostering progress in the field.
Bram Vanherle, Vittorio Pippi, Silvia Cascianelli, Nick Michiels, Frank Van Reeth, Rita Cucchiara
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Continual Facial Features Transfer for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) models based on deep learning mostly rely on a supervised train-once-test-all approach. These approaches assume that a model trained on an in-the-wild facial expression dataset with one type of domain distribution will perform well on a test dataset with a domain distribution shift. However, facial images in real-world can be from different domain distributions from which the model has been trained. However, re-training models on only new domain distributions will severely affect the performance of the previous domain. Re-training on all previous and new data can improve overall performance but is computationally expansive. In this study, we oppose the train-once-test-all approach and propose a buffer-based continual learning approach to enhance the performance of multiple in-the-wild datasets. We propose a model that continually leverages attention to important facial features from the pre-trained model to improve performance in multiple datasets. We validated our model using split-in-the-wild datasets where the dataset is provided to the model in an incremental setting instead of all at once. Furthermore, to evaluate the model performance, we continually used three in-the-wild datasets representing different domains (Domain-FER). Extensive experiments on these datasets reveal that the proposed model achieves better results than other Continual FER models.
Rahul Singh Maharjan, Lorenzo Bonicelli, Marta Romeo, Simone Calderara, Angelo Cangelosi, Rita Cucchiara
IEEE Trans. Affect. Comput.6
2025 Parents and Children: Distinguishing Multimodal Deepfakes from Natural Images
abstract
Recent advancements in diffusion models have enabled the generation of realistic deepfakes from textual prompts in natural language. While these models have numerous benefits across various sectors, they have also raised concerns about the potential misuse of fake images and cast new pressures on fake image detection. In this work, we pioneer a systematic study on deepfake detection generated by state-of-the-art diffusion models. Firstly, we conduct a comprehensive analysis of the performance of contrastive and classification-based visual features, respectively, extracted from CLIP-based models and ResNet or Vision Transformer (ViT)-based architectures trained on image classification datasets. Our results demonstrate that fake images share common low-level cues, which render them easily recognizable. Further, we devise a multimodal setting wherein fake images are synthesized by different textual captions, which are used as seeds for a generator. Under this setting, we quantify the performance of fake detection strategies and introduce a contrastive-based disentangling method that lets us analyze the role of the semantics of textual descriptions and low-level perceptual cues. Finally, we release a new dataset, called COCOFake, containing about 1.2 million images generated from the original COCO image–caption pairs using two recent text-to-image diffusion models, namely Stable Diffusion v1.4 and v2.0.
Roberto Amoroso, Davide Morelli, Marcella Cornia, Lorenzo Baraldi 0001, Alberto Del Bimbo, Rita Cucchiara
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-supervised Learning
Bin Ren 0005, Guofeng Mei, Danda Pani Paudel, Weijie Wang 0002, Yawei Li 0001, Mengyuan Liu 0001, Rita Cucchiara, Luc Van Gool, Nicu Sebe
ACCV (7)7
2024 Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
Nicholas Moratelli, Davide Caffagni, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
BMVC5
2024 Training-Free Open-Vocabulary Segmentation with Offline Diffusion-Augmented Prototype Generation
abstract
Open-vocabulary semantic segmentation aims at segmenting arbitrary categories expressed in textual form. Pre-vious works have trained over large amounts of image-caption pairs to enforce pixel-level multimodal alignments. However, captions provide global information about the semantics of a given image but lack direct localization of individual concepts. Further, training on large-scale datasets inevitably brings significant computational costs. In this paper, we propose FreeDA, a training-free diffusion-augmented method for open-vocabulary semantic segmentation, which leverages the ability of diffusion models to visually localize generated concepts and local-global similarities to match class-agnostic regions with semantic classes. Our approach involves an offline stage in which textual-visual reference embeddings are collected, starting from a large set of captions and leveraging visual and semantic contexts. At test time, these are queried to support the visual matching process, which is carried out by jointly considering class-agnostic regions and global semantic similarities. Extensive analyses demonstrate that FreeDA achieves state-of-the-art performance on five datasets, surpassing previous methods by more than 7.0 average points in terms of mIoU and without requiring any training. Our source code is available at aimagelab.github. io/freeda.
Luca Barsellotti, Roberto Amoroso, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
CVPR5
2024 Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities
Federico Cocchi, Marcella Cornia, Lorenzo Baraldi 0001, Alessandro Nicolosi, Rita Cucchiara
ECCV (63)5
2024 Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models
Samuele Poppi, Tobia Poppi, Federico Cocchi, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
ECCV (53)6
2024 Merging and Splitting Diffusion Paths for Semantically Coherent Panoramas
Fabio Quattrini, Vittorio Pippi, Silvia Cascianelli, Rita Cucchiara
ECCV (79)4
2024 BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues
Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
ECCV (78)4
2024 Binarizing Documents by Leveraging both Space and Frequency
Fabio Quattrini, Vittorio Pippi, Silvia Cascianelli, Rita Cucchiara
ICDAR (3)4
2024 Trajectory Forecasting Through Low-Rank Adaptation of Discrete Latent Codes
Riccardo Benaglia, Angelo Porrello, Pietro Buzzega, Simone Calderara, Rita Cucchiara
ICPR (16)5
2024 Adapt to Scarcity: Few-Shot Deepfake Detection via Low-Rank Adaptation
Silvia Cappelletti, Lorenzo Baraldi 0002, Federico Cocchi, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
ICPR (21)6
2024 Fluent and Accurate Image Captioning with a Self-trained Reward Model
Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
ICPR (18)4
2024 Mask and Compress: Efficient Skeleton-Based Action Recognition in Continual Learning
Matteo Mosconi, Andriy Sorokin, Aniello Panariello, Angelo Porrello, Jacopo Bonato, Marco Cotogni, Luigi Sabetta, Simone Calderara, Rita Cucchiara
ICPR (9)9
2024 Unlearning Vision Transformers Without Retaining Data via Low-Rank Decompositions
Samuele Poppi, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
ICPR (3)5
2024 Mapping High-level Semantic Regions in Indoor Environments without Object Recognition
abstract
Robots require a semantic understanding of their surroundings to operate in an efficient and explainable way in human environments. In the literature, there has been an extensive focus on object labeling and exhaustive scene graph generation; less effort has been focused on the task of purely identifying and mapping large semantic regions. The present work proposes a method for semantic region mapping via embodied navigation in indoor environments, generating a high-level representation of the knowledge of the agent. To enable region identification, the method uses a vision-to-language model to provide scene information for mapping. By projecting egocentric scene understanding into the global frame, the proposed method generates a semantic map as a distribution over possible region labels at each location. This mapping procedure is paired with a trained navigation policy to enable autonomous map generation. The proposed method significantly outperforms a variety of baselines, including an object-based system and a pretrained scene classifier, in experiments in a photorealistic simulator.
Roberto Bigazzi, Lorenzo Baraldi 0001, Shreyas Kousik, Rita Cucchiara, Marco Pavone 0001
ICRA4
2024 Trends, Applications, and Challenges in Human Attention Modelling
Giuseppe Cartella, Marcella Cornia, Vittorio Cuculo, Alessandro D'Amelio, Dario Zanca, Giuseppe Boccignone, Rita Cucchiara
IJCAI7
2024 Personalized Instance-based Navigation Toward User-Specific Objects in Realistic Environments
abstract
In the last years, the research interest in visual navigation towards objects in indoor environments has grown significantly. This growth can be attributed to the recent availability of large navigation datasets in photo-realistic simulated environments, like Gibson and Matterport3D. However, the navigation tasks supported by these datasets are often restricted to the objects present in the environment at acquisition time. Also, they fail to account for the realistic scenario in which the target object is a user-specific instance that can be easily confused with similar objects and may be found in multiple locations within the environment. To address these limitations, we propose a new task denominated Personalized Instance-based Navigation (PIN), in which an embodied agent is tasked with locating and reaching a specific personal object by distinguishing it among multiple instances of the same category. The task is accompanied by PInNED, a dedicated new dataset composed of photo-realistic scenes augmented with additional 3D objects. In each episode, the target object is presented to the agent using two modalities: a set of visual reference images on a neutral background and manually annotated textual descriptions. Through comprehensive evaluations and analyses, we showcase the challenges of the PIN task as well as the performance and shortcomings of currently available methods designed for object-driven navigation, considering modular and end-to-end agents.
Luca Barsellotti, Roberto Bigazzi, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
NeurIPS5
2024 Is Multiple Object Tracking a Matter of Specialization?
abstract
End-to-end transformer-based trackers have achieved remarkable performance on most human-related datasets. However, training these trackers in heterogeneous scenarios poses significant challenges, including negative interference - where the model learns conflicting scene-specific parameters - and limited domain generalization, which often necessitates expensive fine-tuning to adapt the models to new domains. In response to these challenges, we introduce Parameter-efficient Scenario-specific Tracking Architecture (PASTA), a novel framework that combines Parameter-Efficient Fine-Tuning (PEFT) and Modular Deep Learning (MDL). Specifically, we define key scenario attributes (e.g, camera-viewpoint, lighting condition) and train specialized PEFT modules for each attribute. These expert modules are combined in parameter space, enabling systematic generalization to new domains without increasing inference time. Extensive experiments on MOTSynth, along with zero-shot evaluations on MOT17 and PersonPath22 demonstrate that a neural tracker built from carefully selected modules surpasses its monolithic counterpart. We release models and code.
Gianluca Mancusi, Mattia Bernardi, Aniello Panariello, Angelo Porrello, Rita Cucchiara, Simone Calderara
NeurIPS5
2024 Sharing Key Semantics in Transformer Makes Efficient Image Restoration
abstract
Image Restoration (IR), a classic low-level vision task, has witnessed significant advancements through deep models that effectively model global information. Notably, the emergence of Vision Transformers (ViTs) has further propelled these advancements. When computing, the self-attention mechanism, a cornerstone of ViTs, tends to encompass all global cues, even those from semantically unrelated objects or regions. This inclusivity introduces computational inefficiencies, particularly noticeable with high input resolution, as it requires processing irrelevant information, thereby impeding efficiency. Additionally, for IR, it is commonly noted that small segments of a degraded image, particularly those closely aligned semantically, provide particularly relevant information to aid in the restoration process, as they contribute essential contextual cues crucial for accurate reconstruction. To address these challenges, we propose boosting IR's performance by sharing the key semantics via Transformer for IR (i.e., SemanIR) in this paper. Specifically, SemanIR initially constructs a sparse yet comprehensive key-semantic dictionary within each transformer stage by establishing essential semantic connections for every degraded patch. Subsequently, this dictionary is shared across all subsequent transformer blocks within the same stage. This strategy optimizes attention calculation within each block by focusing exclusively on semantically related components stored in the key-semantic dictionary. As a result, attention calculation achieves linear computational complexity within each window. Extensive experiments across 6 IR tasks confirm the proposed SemanIR's state-of-the-art performance, quantitatively and qualitatively showcasing advancements. The visual results, code, and trained models are available at: https://github.com/Amazingren/SemanIR.
Bin Ren 0005, Yawei Li 0001, Jingyun Liang, Mengyuan Liu 0001, Rita Cucchiara, Luc Van Gool, Ming-Hsuan Yang 0001, Nicu Sebe
NeurIPS6
2024 FOSSIL: Free Open-Vocabulary Semantic Segmentation through Synthetic References Retrieval
abstract
Unsupervised Open-Vocabulary Semantic Segmentation aims to segment an image into regions referring to an arbitrary set of concepts described by text, without relying on dense annotations that are available only for a subset of the categories. Previous works rely on inducing pixel-level alignment in a multi-modal space through contrastive training over vast corpora of image-caption pairs. However, representing a semantic category solely through its textual embedding is insufficient to encompass the wide-ranging variability in the visual appearances of the images associated with that category. In this paper, we propose FOSSIL, a pipeline that enables a self-supervised backbone to perform open-vocabulary segmentation relying only on the visual modality. In particular, we decouple the task into two components: (1) we leverage text-conditioned diffusion models to generate a large collection of visual embeddings, starting from a set of captions. These can be retrieved at inference time to obtain a support set of references for the set of textual concepts. Further, (2) we exploit self-supervised dense features to partition the image into semantically coherent regions. We demonstrate that our approach provides strong performance on different semantic segmentation datasets, without requiring any additional training.
Luca Barsellotti, Roberto Amoroso, Lorenzo Baraldi 0001, Rita Cucchiara
WACV4
2024 What's Outside the Intersection? Fine-grained Error Analysis for Semantic Segmentation Beyond IoU
abstract
Semantic segmentation represents a fundamental task in computer vision with various application areas such as autonomous driving, medical imaging, or remote sensing. For evaluating and comparing semantic segmentation models, the mean intersection over union (mIoU) is currently the gold standard. However, while mIoU serves as a valuable benchmark, it does not offer insights into the types of errors incurred by a model. Moreover, different types of errors may have different impacts on downstream applications. To address this issue, we propose an intuitive method for the systematic categorization of errors, thereby enabling a fine-grained analysis of semantic segmentation models. Since we assign each erroneous pixel to precisely one error type, our method seamlessly extends the popular IoU-based evaluation by shedding more light on the false positive and false negative predictions. Our approach is model- and dataset-agnostic, as it does not rely on additional information besides the predicted and ground-truth segmentation masks. In our experiments, we demonstrate that our method accurately assesses model strengths and weaknesses on a quantitative basis, thus reducing the dependence on time-consuming qualitative model inspection. We analyze a variety of state-of-the-art semantic segmentation models, revealing systematic differences across various architectural paradigms. Exploiting the gained insights, we showcase that combining two models with complementary strengths in a straightforward way is sufficient to consistently improve mIoU, even for models setting the current state of the art on ADE20K. We release a toolkit for our evaluation method at https://github.com/mxbh/beyond-iou.
Maximilian Bernhard, Roberto Amoroso, Yannic Kindermann, Lorenzo Baraldi 0001, Rita Cucchiara, Volker Tresp, Matthias Schubert
WACV5
2024 Generating More Pertinent Captions by Leveraging Semantics and Style on Multi-Source Datasets
Marcella Cornia, Lorenzo Baraldi 0001, Giuseppe Fiameni, Rita Cucchiara
Int. J. Comput. Vis.4
2024 Unveiling the Truth: Exploring Human Gaze Patterns in Fake Images
abstract
Creating high-quality and realistic images is now possible thanks to the impressive advancements in image generation. A description in natural language of your desired output is all you need to obtain breathtaking results. However, as the use of generative models grows, so do concerns about the propagation of malicious content and misinformation. Consequently, the research community is actively working on the development of novel fake detection techniques, primarily focusing on low-level features and possible fingerprints left by generative models during the image generation process. In a different vein, in our work, we leverage human semantic knowledge to investigate the possibility of being included in frameworks of fake image detection. To achieve this, we collect a novel dataset of partially manipulated images using diffusion models and conduct an eye-tracking experiment to record the eye movements of different observers while viewing real and fake stimuli. A preliminary statistical analysis is conducted to explore the distinctive patterns in how humans perceive genuine and altered images. Statistical findings reveal that, when perceiving counterfeit samples, humans tend to focus on more confined regions of the image, in contrast to the more dispersed observational pattern observed when viewing genuine images. Our dataset is publicly available at:https://github.com/aimagelab/unveiling-the-truth.
Giuseppe Cartella, Vittorio Cuculo, Marcella Cornia, Rita Cucchiara
IEEE Signal Process. Lett.4
2024 Towards Retrieval-Augmented Architectures for Image Captioning
abstract
The objective of image captioning models is to bridge the gap between the visual and linguistic modalities by generating natural language descriptions that accurately reflect the content of input images. In recent years, researchers have leveraged deep learning-based models and made advances in the extraction of visual features and the design of multimodal connections to tackle this task. This work presents a novel approach toward developing image captioning models that utilize an externalkNN memory to improve the generation process. Specifically, we propose two model variants that incorporate a knowledge retriever component that is based on visual similarities, a differentiable encoder to represent input images, and akNN-augmented language model to predict tokens based on contextual cues and text retrieved from the external memory. We experimentally validate our approach on COCO and nocaps datasets and demonstrate that incorporating an explicit external memory can significantly enhance the quality of captions, especially with a larger retrieval corpus. This work provides valuable insights into retrieval-augmented captioning models and opens up new avenues for improving image captioning at a larger scale.
Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Alessandro Nicolosi, Rita Cucchiara
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Superpixel Positional Encoding to Improve ViT-based Semantic Segmentation Models
Roberto Amoroso, Matteo Tomei, Lorenzo Baraldi 0001, Rita Cucchiara
BMVC4
2023 HWD: A Novel Evaluation Score for Styled Handwritten Text Generation
Vittorio Pippi, Fabio Quattrini, Silvia Cascianelli, Rita Cucchiara
BMVC4
2023 Handwritten Text Generation from Visual Archetypes
abstract
Generating synthetic images of handwritten text in a writer-specific style is a challenging task, especially in the case of unseen styles and new words, and even more when these latter contain characters that are rarely encountered during training. While emulating a writer's style has been recently addressed by generative models, the generalization towards rare characters has been disregarded. In this work, we devise a Transformer-based model for Few-Shot styled handwritten text generation and focus on obtaining a robust and informative representation of both the text and the style. In particular, we propose a novel representation of the textual content as a sequence of dense vectors obtained from images of symbols written as standard GNU Unifont glyphs, which can be considered their visual archetypes. This strategy is more suitable for generating characters that, despite having been seen rarely during training, possibly share visual details with the frequently observed ones. As for the style, we obtain a robust representation of unseen writers' calligraphy by exploiting specific pre-training on a large synthetic dataset. Quantitative and qualitative results demonstrate the effectiveness of our proposal in generating words in unseen styles and with rare characters more faithfully than existing approaches relying on independent one-hot encodings of the characters.
Vittorio Pippi, Silvia Cascianelli, Rita Cucchiara
CVPR3
2023 Masked Jigsaw Puzzle: A Versatile Position Embedding for Vision Transformers
abstract
Position Embeddings (PEs), an arguably indispensable component in Vision Transformers (ViTs), have been shown to improve the performance of ViTs on many vision tasks. However, PEs have a potentially high risk of privacy leakage since the spatial information of the input patches is exposed. This caveat naturally raises a series of interesting questions about the impact of PEs on accuracy, privacy, prediction consistency, etc. To tackle these issues, we propose a Masked Jigsaw Puzzle (MJP) position embedding method. In particular, MJP first shuffles the selected patches via our block-wise random jigsaw puzzle shuffle algorithm, and their corresponding PEs are occluded. Meanwhile, for the non-occluded patches, the PEs remain the original ones but their spatial relation is strengthened via our dense absolute localization regressor. The experimental results reveal that 1) PEs explicitly encode the 2D spatial relationship and lead to severe privacy leakage problems under gradient inversion attack; 2) Training ViTs with the naively shuffled patches can alleviate the problem, but it harms the accuracy; 3) Under a certain shuffle ratio, the proposed MJP not only boosts the performance and robustness on large-scale datasets (i.e., ImageNet-1K and ImageNet-C, -A/O) but also improves the privacy preservation ability under typical gradient attacks by a large margin. The source code and trained models are available at https://github.com/yhlleo/MJP.
Bin Ren 0005, Yue Song 0002, Wei Bi, Rita Cucchiara, Nicu Sebe, Wei Wang 0108
CVPR5
2023 Positive-Augmented Contrastive Learning for Image and Video Captioning Evaluation
abstract
The CLIP model has been recently proven to be very effective for a variety of cross-modal tasks, including the evaluation of captions generated from vision-and-language architectures. In this paper, we propose a new recipe for a contrastive-based evaluation metric for image captioning, namely Positive-Augmented Contrastive learning Score (PAC-S), that in a novel way unifies the learning of a contrastive visual-semantic space with the addition of generated images and text on curated data. Experiments spanning several datasets demonstrate that our new metric achieves the highest correlation with human judgments on both images and videos, outperforming existing referencebased metrics like CIDEr and SPICE and reference-free metrics like CLIP-Score. Finally, we test the system-level correlation of the proposed metric when considering popular image captioning approaches, and assess the impact of employing different cross-modal features. Our source code and trained models are publicly available at: https://github.com/aimagelab/pacscore.
Sara Sarto, Manuele Barraco, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
CVPR5
2023 Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing
abstract
Fashion illustration is used by designers to communicate their vision and to bring the design idea from conceptualization to realization, showing how clothes interact with the human body. In this context, computer vision can thus be used to improve the fashion design process. Differently from previous works that mainly focused on the virtual try-on of garments, we propose the task of multimodal-conditioned fashion image editing, guiding the generation of human-centric fashion images by following multimodal prompts, such as text, human body poses, and garment sketches. We tackle this problem by proposing a new architecture based on latent diffusion models, an approach that has not been used before in the fashion domain. Given the lack of existing datasets suitable for the task, we also extend two existing fashion datasets, namely Dress Code and VITON-HD, with multimodal annotations collected in a semi-automatic manner. Experimental results on these new datasets demonstrate the effectiveness of our proposal, both in terms of realism and coherence with the given multimodal inputs. Source code and collected multimodal annotations are publicly available at: https://github.com/aimagelab/multimodal-garment-designer.
Alberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia, Marco Bertini 0001, Rita Cucchiara
ICCV6
2023 With a Little Help from your own Past: Prototypical Memory Networks for Image Captioning
abstract
Image captioning, like many tasks involving vision and language, currently relies on Transformer-based architectures for extracting the semantics in an image and translating it into linguistically coherent descriptions. Although successful, the attention operator only considers a weighted summation of projections of the current input sample, therefore ignoring the relevant semantic information which can come from the joint observation of other samples. In this paper, we devise a network which can perform attention over activations obtained while processing other training samples, through a prototypical memory model. Our memory models the distribution of past keys and values through the definition of prototype vectors which are both discriminative and compact. Experimentally, we assess the performance of the proposed model on the COCO dataset, in comparison with carefully designed baselines and state-of-the-art approaches, and by investigating the role of each of the proposed components. We demonstrate that our proposal can increase the performance of an encoder-decoder Transformer by 3.7 CIDEr points both when training in cross-entropy only and when fine-tuning with self-critical sequence training. Source code and trained models are available at: https://github.com/aimagelab/PMA-Net.
Manuele Barraco, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
ICCV5
2023 TrackFlow: Multi-Object Tracking with Normalizing Flows
abstract
The field of multi-object tracking has recently seen a renewed interest in the good old schema of tracking-by-detection, as its simplicity and strong priors spare it from the complex design and painful babysitting of tracking-by-attention approaches. In view of this, we aim at extending tracking-by-detection to multi-modal settings, where a comprehensive cost has to be computed from heterogeneous information e.g., 2D motion cues, visual appearance, and pose estimates. More precisely, we follow a case study where a rough estimate of 3D information is also available and must be merged with other traditional metrics (e.g., the IoU). To achieve that, recent approaches resort to either simple rules or complex heuristics to balance the contribution of each cost. However, i) they require careful tuning of tailored hyperparameters on a hold-out set, and ii) they imply these costs to be independent, which does not hold in reality. We address these issues by building upon an elegant probabilistic formulation, which considers the cost of a candidate association as the negative log-likelihood yielded by a deep density estimator, trained to model the conditional joint probability distribution of correct associations. Our experiments, conducted on both simulated and real benchmarks, show that our approach consistently enhances the performance of several tracking-by-detection algorithms.
Gianluca Mancusi, Aniello Panariello, Angelo Porrello, Matteo Fabbri, Simone Calderara, Rita Cucchiara
ICCV6
2023 How to Choose Pretrained Handwriting Recognition Models for Single Writer Fine-Tuning
Vittorio Pippi, Silvia Cascianelli, Christopher Kermorvant, Rita Cucchiara
ICDAR (2)4
2023 Input Perturbation Reduces Exposure Bias in Diffusion Models
abstract
Denoising Diffusion Probabilistic Models have shown an impressive generation quality although their long sampling chain leads to high computational costs. In this paper, we observe that a long sampling chain also leads to an error accumulation phenomenon, which is similar to the exposure bias problem in autoregressive text generation. Specifically, we note that there is a discrepancy between training and testing, since the former is conditioned on the ground truth samples, while the latter is conditioned on the previously generated results. To alleviate this problem, we propose a very simple but effective training regularization, consisting in perturbing the ground truth samples to simulate the inference time prediction errors. We empirically show that, without affecting the recall and precision, the proposed input perturbation leads to a significant improvement in the sample quality while reducing both the training and the inference times. For instance, on CelebA 64x64, we achieve a new state-of-the-art FID score of 1.27, while saving 37.5% of the training time. The code is available at https://github.com/forever208/DDPM-IP
Mang Ning, Enver Sangineto, Angelo Porrello, Simone Calderara, Rita Cucchiara
ICML5
2023 Embodied Agents for Efficient Exploration and Smart Scene Description
abstract
The development of embodied agents that can communicate with humans in natural language has gained increasing interest over the last years, as it facilitates the diffusion of robotic platforms in human-populated environments. As a step towards this objective, in this work, we tackle a setting for visual navigation in which an autonomous agent needs to explore and map an unseen indoor environment while portraying interesting scenes with natural language descriptions. To this end, we propose and evaluate an approach that combines recent advances in visual robotic exploration and image captioning on images generated through agent-environment interaction. Our approach can generate smart scene descriptions that maximize semantic knowledge of the environment and avoid repetitions. Further, such descriptions offer user-understandable insights into the robot's representation of the environment by high-lighting the prominent objects and the correlation between them as encountered during the exploration. To quantitatively assess the performance of the proposed approach, we also devise a specific score that takes into account both exploration and description skills. The experiments carried out on both photorealistic simulated environments and real-world ones demonstrate that our approach can effectively describe the robot's point of view during exploration, improving the human-friendly interpretability of its observations.
Roberto Bigazzi, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara
ICRA5
2023 Let's ViCE! Mimicking Human Cognitive Behavior in Image Generation Evaluation
abstract
Research in Image Generation has recently made significant progress, particularly boosted by the introduction of VisionLanguage models which are able to produce high-quality visual content based on textual inputs. Despite ongoing advancements in terms of generation quality and realism, no methodical frameworks have been defined yet to quantitatively measure the quality of the generated content and the adherence with the prompted requests: so far, only humanbased evaluations have been adopted for quality satisfaction and for comparing different generative methods. We introduce a novel automated method for Visual Concept Evaluation (ViCE), i.e. to assess consistency between a generated/edited image and the corresponding prompt/instructions, with a process inspired by the human cognitive behaviour. ViCE combines the strengths of Large Language Models (LLMs) and Visual Question Answering (VQA) into a unified pipeline, aiming to replicate the human cognitive process in quality assessment. This method outlines visual concepts, formulates image-specific verification questions, utilizes the Q&A system to investigate the image, and scores the combined outcome. Although this brave new hypothesis of mimicking humans in the image evaluation process is in its preliminary assessment stage, results are promising and open the door to a new form of automatic evaluation which could have significant impact as the image generation or the image target editing tasks become more and more sophisticated.
Federico Betti 0001, Jacopo Staiano, Lorenzo Baraldi 0002, Lorenzo Baraldi 0001, Rita Cucchiara, Nicu Sebe
ACM Multimedia5
2023 LaDI-VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-On
abstract
The rapidly evolving fields of e-commerce and metaverse continue to seek innovative approaches to enhance the consumer experience. At the same time, recent advancements in the development of diffusion models have enabled generative networks to create remarkably realistic images. In this context, image-based virtual try-on, which consists in generating a novel image of a target model wearing a given in-shop garment, has yet to capitalize on the potential of these powerful generative solutions. This work introduces LaDI-VTON, the first Latent Diffusion textual Inversion-enhanced model for the Virtual Try-ON task. The proposed architecture relies on a latent diffusion model extended with a novel additional autoencoder module that exploits learnable skip connections to enhance the generation process preserving the model's characteristics. To effectively maintain the texture and details of the in-shop garment, we propose a textual inversion component that can map the visual features of the garment to the CLIP token embedding space and thus generate a set of pseudo-word token embeddings capable of conditioning the generation process. Experimental results on Dress Code and VITON-HD datasets demonstrate that our approach outperforms the competitors by a consistent margin, achieving a significant milestone for the task. Source code and trained models are publicly available at: https://github.com/miccunifi/ladi-vton.
Davide Morelli, Alberto Baldrati, Giuseppe Cartella, Marcella Cornia, Marco Bertini 0001, Rita Cucchiara
ACM Multimedia6
2023 Fully-attentive iterative networks for region-based controllable image and video captioning
abstract
Controllable image captioning has recently gained attention as a way to increase the diversity and the applicability to real-world scenarios of image captioning algorithms. In this task, a captioner is conditioned on an external control signal, which needs to be followed during the generation of the caption. We aim to overcome the limitations of current controllable captioning methods by proposing a fully-attentive and iterative network that can generate grounded and controllable captions from a control signal given as a sequence of visual regions from the image. Our architecture is based on a set of novel attention operators, which take into account the hierarchical nature of the control signal, and is endowed with a decoder which explicitly focuses on each part of the control signal. We demonstrate the effectiveness of the proposed approach by conducting experiments on three datasets, where our model surpasses the performances of previous methods and achieves a new state of the art on both image and video controllable captioning.
Marcella Cornia, Lorenzo Baraldi 0001, Ayellet Tal, Rita Cucchiara
Comput. Vis. Image Underst.4
2023 From Show to Tell: A Survey on Deep Learning-Based Image Captioning
abstract
Connecting Vision and Language plays an essential role in Generative Intelligence. For this reason, large research efforts have been devoted to image captioning, i.e. describing images with syntactically and semantically meaningful sentences. Starting from 2015 the task has generally been addressed with pipelines composed of a visual encoder and a language model for text generation. During these years, both components have evolved considerably through the exploitation of object regions, attributes, the introduction of multi-modal connections, fully-attentive approaches, and BERT-like early-fusion strategies. However, regardless of the impressive results, research in image captioning has not reached a conclusive answer yet. This work aims at providing a comprehensive overview of image captioning approaches, from visual encoding and text generation to training strategies, datasets, and evaluation metrics. In this respect, we quantitatively compare many relevant state-of-the-art approaches to identify the most impactful technical innovations in architectures and training strategies. Moreover, many variants of the problem and its open challenges are discussed. The final goal of this work is to serve as a tool for understanding the existing literature and highlighting the future directions for a research area where Computer Vision and Natural Language Processing can find an optimal synergy.
Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi 0001, Silvia Cascianelli, Giuseppe Fiameni, Rita Cucchiara
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Depth-based 3D human pose refinement: Evaluating the refinet framework
abstract
In recent years, Human Pose Estimation has achieved impressive results on RGB images. The advent of deep learning architectures and large annotated datasets have contributed to these achievements. However, little has been done towards estimating the human pose using depth maps, and especially towards obtaining a precise 3D body joint localization. To fill this gap, this paper presents RefiNet, a depth-based 3D human pose refinement framework. Given a depth map and an initial coarse 2D human pose, RefiNet regresses a fine 3D pose. The framework is composed of three modules, based on different data representations, i.e. 2D depth patches, 3D human skeletons, and point clouds. An extensive experimental evaluation is carried out to investigate the impact of the model hyper-parameters and to compare RefiNet with off-the-shelf 2D methods and literature approaches. Results confirm the effectiveness of the proposed framework and its limited computational requirements.
Andrea D'Eusanio, Alessandro Simoni, Stefano Pini, Guido Borghi, Roberto Vezzani, Rita Cucchiara
Pattern Recognit. Lett.6
2023 Evaluating synthetic pre-Training for handwriting processing tasks
abstract
In this work, we explore massive pre-training on synthetic word images for enhancing the performance on four benchmark downstream handwriting analysis tasks. To this end, we build a large synthetic dataset of word images rendered in several handwriting fonts, which offers a complete supervision signal. We use it to train a simple convolutional neural network (ConvNet) with a fully supervised objective. The vector representations of the images obtained from the pre-trained ConvNet can then be considered as encodings of the handwriting style. We exploit such representations for Writer Retrieval, Writer Identification, Writer Verification, and Writer Classification and demonstrate that our pre-training strategy allows extracting rich representations of the writers’ style that enable the aforementioned tasks with competitive results with respect to task-specific State-of-the-Art approaches.
Vittorio Pippi, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara
Pattern Recognit. Lett.4
2022 ALADIN: Distilling Fine-grained Alignment Scores for Efficient Image-Text Matching and Retrieval
abstract
Image-text matching is gaining a leading role among tasks involving the joint understanding of vision and language. In literature, this task is often used as a pre-training objective to forge architectures able to jointly deal with images and texts. Nonetheless, it has a direct downstream application: cross-modal retrieval, which consists in finding images related to a given query text or vice-versa. Solving this task is of critical importance in cross-modal search engines. Many recent methods proposed effective solutions to the image-text matching problem, mostly using recent large vision-language (VL) Transformer networks. However, these models are often computationally expensive, especially at inference time. This prevents their adoption in large-scale cross-modal retrieval scenarios, where results should be provided to the user almost instantaneously. In this paper, we propose to fill in the gap between effectiveness and efficiency by proposing an ALign And DIstill Network (ALADIN). ALADIN first produces high-effective scores by aligning at fine-grained level images and texts. Then, it learns a shared embedding space – where an efficient kNN search can be performed – by distilling the relevance scores obtained from the fine-grained alignments. We obtained remarkable results on MS-COCO, showing that our method can compete with state-of-the-art VL Transformers while being almost 90 times faster. The code for reproducing our results is available at https://github.com/mesnico/ALADIN.
Nicola Messina, Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi 0001, Fabrizio Falchi, Giuseppe Amato 0001, Rita Cucchiara
CBMI7
2022 Retrieval-Augmented Transformer for Image Captioning
abstract
Image captioning models aim at connecting Vision and Language by providing natural language descriptions of input images. In the past few years, the task has been tackled by learning parametric models and proposing visual feature extraction advancements or by modeling better multi-modal connections. In this paper, we investigate the development of an image captioning approach with a kNN memory, with which knowledge can be retrieved from an external corpus to aid the generation process. Our architecture combines a knowledge retriever based on visual similarities, a differentiable encoder, and a kNN-augmented attention layer to predict tokens based on the past context and on text retrieved from the external memory. Experimental results, conducted on the COCO dataset, demonstrate that employing an explicit external memory can aid the generation process and increase caption quality. Our work opens up new avenues for improving image captioning models at larger scale.
Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
CBMI4
2022 How many Observations are Enough? Knowledge Distillation for Trajectory Forecasting
abstract
Accurate prediction of future human positions is an essential task for modern video-surveillance systems. Current state-of-the-art models usually rely on a “history” of past tracked locations (e.g., 3 to 5 seconds) to predict a plausible sequence of future locations (e.g., up to the next 5 seconds). We feel that this common schema neglects critical traits of realistic applications: as the collection of input trajectories involves machine perception (i.e., detection and tracking), incorrect detection and fragmentation errors may accumulate in crowded scenes, leading to tracking drifts. On this account, the model would be fed with corrupted and noisy input data, thus fatally affecting its prediction performance. In this regard, we focus on delivering accurate predictions when only few input observations are used, thus potentially lowering the risks associated with automatic perception. To this end, we conceive a novel distillation strategy that allows a knowledge transfer from a teacher network to a student one, the latter fed with fewer observations (just two ones). We show that a properly defined teacher super-vision allows a student network to perform comparably to state-of-the-art approaches that demand more observations. Besides, extensive experiments on common trajectory forecasting datasets highlight that our student network better generalizes to unseen scenarios.
Alessio Monti, Angelo Porrello, Simone Calderara, Pasquale Coscia, Lamberto Ballan, Rita Cucchiara
CVPR6
2022 Dress Code: High-Resolution Multi-category Virtual Try-On
Davide Morelli, Matteo Fincato, Marcella Cornia, Federico Landi, Fabio Cesari, Rita Cucchiara
ECCV (8)6
2022 CaMEL: Mean Teacher Learning for Image Captioning
abstract
Describing images in natural language is a fundamental step towards the automatic modeling of connections between the visual and textual modalities. In this paper we present CaMEL, a novel Transformer-based architecture for image captioning. Our proposed approach leverages the interaction of two interconnected language models that learn from each other during the training phase. The interplay between the two language models follows a mean teacher learning paradigm with knowledge distillation. Experimentally, we assess the effectiveness of the proposed solution on the COCO dataset and in conjunction with different visual feature extractors. When comparing with existing proposals, we demonstrate that our model provides state-of-the-art caption quality with a significantly reduced number of parameters. According to the CIDEr metric, we obtain a new state of the art on COCO when training without using external data. The source code and trained models will be made publicly available at: https://github.com/aimagelab/camel.
Manuele Barraco, Matteo Stefanini, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara
ICPR6
2022 The LAM Dataset: A Novel Benchmark for Line-Level Handwritten Text Recognition
abstract
Handwritten Text Recognition (HTR) is an open problem at the intersection of Computer Vision and Natural Language Processing. The main challenges, when dealing with historical manuscripts, are due to the preservation of the paper support, the variability of the handwriting – even of the same author over a wide time-span – and the scarcity of data from ancient, poorly represented languages. With the aim of fostering the research on this topic, in this paper we present the Ludovico Antonio Muratori (LAM) dataset, a large line-level HTR dataset of Italian ancient manuscripts edited by a single author over 60 years. The dataset comes in two configurations: a basic splitting and a date-based splitting which takes into account the age of the author. The first setting is intended to study HTR on ancient documents in Italian, while the second focuses on the ability of HTR systems to recognize text written by the same writer in time periods for which training data are not available. For both configurations, we analyze quantitative and qualitative characteristics, also with respect to other line-level HTR benchmarks, and present the recognition performance of state-of-the-art HTR architectures. The dataset is available for download at https://aimagelab.ing.unimore.it/go/lam.
Silvia Cascianelli, Vittorio Pippi, Martin Maarand, Marcella Cornia, Lorenzo Baraldi 0001, Christopher Kermorvant, Rita Cucchiara
ICPR7
2022 Spot the Difference: A Novel Task for Embodied Agents in Changing Environments
abstract
Embodied AI is a recent research area that aims at creating intelligent agents that can move and operate inside an environment. Existing approaches in this field demand the agents to act in completely new and unexplored scenes. However, this setting is far from realistic use cases that instead require executing multiple tasks in the same environment. Even if the environment changes over time, the agent could still count on its global knowledge about the scene while trying to adapt its internal representation to the current state of the environment. To make a step towards this setting, we propose Spot the Difference: a novel task for Embodied AI where the agent has access to an outdated map of the environment and needs to recover the correct layout in a fixed time budget. To this end, we collect a new dataset of occupancy maps starting from existing datasets of 3D spaces and generating a number of possible layouts for a single environment. This dataset can be employed in the popular Habitat simulator and is fully compliant with existing methods that employ reconstructed occupancy maps during navigation. Furthermore, we propose an exploration policy that can take advantage of previous knowledge of the environment and identify changes in the scene faster and more effectively than existing agents. Experimental results show that the proposed architecture outperforms existing state-of-the-art models for exploration on this new setting.
Federico Landi, Roberto Bigazzi, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara
ICPR6
2022 Maximum Class Separation as Inductive Bias in One Matrix
abstract
Maximizing the separation between classes constitutes a well-known inductive bias in machine learning and a pillar of many traditional algorithms. By default, deep networks are not equipped with this inductive bias and therefore many alternative solutions have been proposed through differential optimization. Current approaches tend to optimize classification and separation jointly: aligning inputs with class vectors and separating class vectors angularly. This paper proposes a simple alternative: encoding maximum separation as an inductive bias in the network by adding one fixed matrix multiplication before computing the softmax activations. The main observation behind our approach is that separation does not require optimization but can be solved in closed-form prior to training and plugged into a network. We outline a recursive approach to obtain the matrix consisting of maximally separable vectors for any number of classes, which can be added with negligible engineering effort and computational overhead. Despite its simple nature, this one matrix multiplication provides real impact. We show that our proposal directly boosts classification, long-tailed recognition, out-of-distribution detection, and open-set recognition, from CIFAR to ImageNet. We find empirically that maximum separation works best as a fixed bias; making the matrix learnable adds nothing to the performance. The closed-form implementation and code to reproduce the experiments are available on github.
Tejaswi Kasarla, Gertjan J. Burghouts, Max van Spengler, Elise van der Pol, Rita Cucchiara, Pascal Mettes
NeurIPS5
2022 Boosting modern and historical handwritten text recognition with deformable convolutions
Silvia Cascianelli, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
Int. J. Document Anal. Recognit.4
2022 Warp and Learn: Novel Views Generation for Vehicles and Other Objects
abstract
In this article we introduce a new self-supervised, semi-parametric approach for synthesizing novel views of a vehicle starting from a single monocular image. Differently from parametric (i.e., entirely learning-based) methods, we show how a-priori geometric knowledge about the object and the 3D world can be successfully integrated into a deep learning based image generation framework. As this geometric component is not learnt, we call our approach semi-parametric. In particular, we exploit man-made object symmetry and piece-wise planarity to integrate rich a-priori visual information into the novel viewpoint synthesis process. An Image Completion Network (ICN) is then trained to generate a realistic image starting from this geometric guidance. This careful blend between parametric and non-parametric components allows us to i) operate in a real-world scenario, ii) preserve high-frequency visual information such as textures, iii) handle truly arbitrary 3D roto-translations of the input, and iv) perform shape transfer to completely different 3D models. Eventually, we show that our approach can be easily complemented with synthetic data and extended to other rigid objects with completely different topology, even in presence of concave structures and holes (e.g., chairs). A comprehensive experimental analysis against state-of-the-art competitors shows the efficacy of our method both from a quantitative and a perceptive point of view. Supplementary material, animated results, code, and data are available at: https://github.com/ndrplz/semiparametric.
Andrea Palazzi, Luca Bergamini, Simone Calderara, Rita Cucchiara
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 A computational approach for progressive architecture shrinkage in action recognition
abstract
Abstract Efficiency plays a key role in video understanding modeling, and developing more efficient spatiotemporal deep networks is a key ingredient for enabling their usage in production scenarios. In this work, we propose a methodology for reducing the computational complexity of a video understanding backbone while limiting the drop in accuracy caused by architectural changes. Our approach, named, Progressive Architecture Shrinkage, applies a sequence of reduction operators to the hyperparameters of a network to reduce its computational footprint. The choice of the sequence of operations is automatically optimized in a coordinate‐descent schema, and the approach transfers knowledge from both the initial network and previous stages of the shrinking process by employing a Knowledge Distillation and an adaptive fine‐tuning strategy. As each iteration of the shrinking algorithm requires to train a large‐scale video understanding network, we perform experiments on MARCONI 100—a supercomputer equipped with an IBM Power9 architecture and Volta NVIDIA GPUs. Experimental evaluations are conducted using two backbones and three different action recognition benchmarks. We show that, through our approach, high accuracy levels can be maintained while reducing the number of multiply–adds operations by four times with respect to the original architectures. Code will be made available.
Matteo Tomei, Lorenzo Baraldi 0001, Giuseppe Fiameni, Simone Bronzin, Rita Cucchiara
Softw. Pract. Exp.5
2022 Wind Turbine Power Curve Monitoring Based on Environmental and Operational Data
abstract
The power produced by a wind turbine depends on environmental conditions, working parameters, and interactions with nearby turbines. However, these aspects are often neglected in the design of data-driven models for wind farms’ performance analysis. In this article, we propose to predict the active power and to provide reliable prediction intervals via ensembles of multivariate polynomial regression models that exploit a higher number of inputs (compared to most approaches in the literature), including operational and thermal variables. We present two main strategies: the former considers the environmental measurements collected at the other wind turbines in the farm as additional modeling information for the turbine under analysis; the latter combines multiple models relative to different operative conditions. We validate our approach on real data from the SCADA system of a wind farm in Italy and obtain a MAE of the order of 1.0% of the rated power of the turbine. Moreover, due to the structure of our approach, we can gain quantitative insights on the covariates most frequently selected depending on the working region of the wind turbines.
Silvia Cascianelli, Davide Astolfi, Francesco Castellani, Rita Cucchiara, Mario Luca Fravolini
IEEE Trans. Ind. Informatics4
2022 Matching Faces and Attributes Between the Artistic and the Real Domain: the PersonArt Approach
abstract
In this article, we present an approach for retrieving similar faces between the artistic and the real domain. The application we refer to is an interactive exhibition inside a museum, in which a visitor can take a photo of himself and search for a lookalike in the collection of paintings. The task requires not only to identify faces but also to extract discriminative features from artistic and photo-realistic images, tackling a significant domain shift. Our method integrates feature extraction networks which account for the aesthetic similarity of two faces and their correspondences in terms of semantic attributes. Also, it addresses the domain shift between realistic images and paintings by translating photo-realistic images into the artistic domain. Noticeably, by exploiting the same technique, our model does not need to rely on annotated data in the artistic domain. Experimental results are conducted on different paired datasets to show the effectiveness of the proposed solution in terms of identity and attribute preservation. The approach is also evaluated on unpaired settings and in combination with an interactive relevance feedback strategy. Finally, we show how the proposed algorithm has been implemented in a real showcase at the Gallerie Estensi museum in Italy, with the participation of more than 1,100 visitors in just three days.
Marcella Cornia, Matteo Tomei, Lorenzo Baraldi 0001, Rita Cucchiara
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Fine-grained Human Analysis under Occlusions and Perspective Constraints in Multimedia Surveillance
abstract
Human detection in the wild is a research topic of paramount importance in computer vision, and it is the starting step for designing intelligent systems oriented to human interaction that work in complete autonomy. To achieve this goal, computer vision and machine learning should aim at superhuman capabilities. In this work, we address the problem of fine-grained human analysis under occlusions and perspective constraints. More specifically, we discuss some issues and some possible solutions to effectively detect people using pose estimation methods and to detect humans under occlusions both in the two-dimensional (2D) image plane and in the 3D space exploiting single monocular cameras. Dealing with occlusion can be done at the joint level or pixel level: We discuss two different solutions, the former based on a supervised neural network architecture for detecting occluded joints and the latter based on a semi-supervised specialized GAN that exploits both appearance and human shape attributes to determine the missing parts of the visible shape. To deal with perspective constraints, we further discuss a neural approach based on a double architecture that learns to create an optimal neural representation, which is useful to reconstruct the 3D position of human keypoints starting with simple RGB images. All these approaches have a critical point in common: the need for large annotated datasets. To have large, fair, consistent, transparent, and ethical datasets, we propose the adoption of synthetic datasets as, for example, JTA and MOTSynth. In this article, we discuss the pros and cons of using synthetic datasets while tackling several human-centered AI issues with respect to European GDPR rules for privacy. We further explore and discuss an application in the field of risk assessment by space occupancy estimation during the COVID-19 pandemic called Inter-Homines.
Rita Cucchiara, Matteo Fabbri
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Transform, Warp, and Dress: A New Transformation-guided Model for Virtual Try-on
abstract
Virtual try-on has recently emerged in computer vision and multimedia communities with the development of architectures that can generate realistic images of a target person wearing a custom garment. This research interest is motivated by the large role played by e-commerce and online shopping in our society. Indeed, the virtual try-on task can offer many opportunities to improve the efficiency of preparing fashion catalogs and to enhance the online user experience. The problem is far to be solved: current architectures do not reach sufficient accuracy with respect to manually generated images and can only be trained on image pairs with a limited variety. Existing virtual try-on datasets have two main limits: they contain only female models, and all the images are available only in low resolution. This not only affects the generalization capabilities of the trained architectures but makes the deployment to real applications impractical. To overcome these issues, we present Dress Code , a new dataset for virtual try-on that contains high-resolution images of a large variety of upper-body clothes and both male and female models. Leveraging this enriched dataset, we propose a new model for virtual try-on capable of generating high-quality and photo-realistic images using a three-stage pipeline. The first two stages perform two different geometric transformations to warp the desired garment and make it fit into the target person’s body pose and shape. Then, we generate the new image of that same person wearing the try-on garment using a generative network. We test the proposed solution on the most widely used dataset for this task as well as on our newly collected dataset and demonstrate its effectiveness when compared to current state-of-the-art methods. Through extensive analyses on our Dress Code dataset, we show the adaptability of our model, which can generate try-on images even with a higher resolution.
Matteo Fincato, Marcella Cornia, Federico Landi, Fabio Cesari, Rita Cucchiara
ACM Trans. Multim. Comput. Commun. Appl.5
2022 Special Section on AI-empowered Multimedia Data Analytics for Smart Healthcare
abstract
No abstract available.
M. Shamim Hossain, Rita Cucchiara, Muhammad Ghulam, Diana P. Tobón, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Multi-Category Mesh Reconstruction From Image Collections
abstract
Recently, learning frameworks have shown the capability of inferring the accurate shape, pose, and texture of an object from a single RGB image. However, current methods are trained on image collections of a single category in order to exploit specific priors, and they often make use of category-specific 3D templates. In this paper, we present an alternative approach that infers the textured mesh of objects combining a series of deformable 3D models and a set of instance-specific deformation, pose, and texture. Differently from previous works, our method is trained with images of multiple object categories using only foreground masks and rough camera poses as supervision. Without specific 3D templates, the framework learns category-level models which are deformed to recover the 3D shape of the depicted object. The instance-specific deformations are predicted independently for each vertex of the learned 3D mesh, enabling the dynamic subdivision of the mesh during the training process. Experiments show that the proposed framework can distinguish between different object categories and learn category-specific shape priors in an unsupervised manner. Predicted shapes are smooth and can leverage from multiple steps of subdivision during the training process, obtaining comparable or state-of-the-art results on two public datasets. Models and code are publicly released1.
Alessandro Simoni, Stefano Pini, Roberto Vezzani, Rita Cucchiara
3DV4
2021 Assessing the Role of Boundary-Level Objectives in Indoor Semantic Segmentation
Roberto Amoroso, Lorenzo Baraldi 0001, Rita Cucchiara
CAIP (1)3
2021 Out of the Box: Embodied Navigation in the Real World
Roberto Bigazzi, Federico Landi, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara
CAIP (1)6
2021 Learning to Read L'Infinito: Handwritten Text Recognition with Synthetic Training Data
Silvia Cascianelli, Marcella Cornia, Lorenzo Baraldi 0001, Maria Ludovica Piazzi, Rosiana Schiuma, Rita Cucchiara
CAIP (2)6
2021 MOTSynth: How Can Synthetic Data Help Pedestrian Detection and Tracking?
abstract
Deep learning-based methods for video pedestrian detection and tracking require large volumes of training data to achieve good performance. However, data acquisition in crowded public environments raises data privacy concerns – we are not allowed to simply record and store data without the explicit consent of all participants. Furthermore, the annotation of such data for computer vision applications usually requires a substantial amount of manual effort, especially in the video domain. Labeling instances of pedestrians in highly crowded scenarios can be challenging even for human annotators and may introduce errors in the training data. In this paper, we study how we can advance different aspects of multi-person tracking using solely synthetic data. To this end, we generate MOTSynth, a large, highly diverse synthetic dataset for object detection and tracking using a rendering game engine. Our experiments show that MOTSynth can be used as a replacement for real data on tasks such as pedestrian detection, re-identification, segmentation, and tracking.
Matteo Fabbri, Guillem Brasó, Gianluca Maugeri, Orcun Cetintas, Riccardo Gasparini, Aljosa Osep, Simone Calderara, Laura Leal-Taixé, Rita Cucchiara
ICCV9
2021 Learning to Select: A Fully Attentive Approach for Novel Object Captioning
abstract
Image captioning models have lately shown impressive results when applied to standard datasets. Switching to real-life scenarios, however, constitutes a challenge due to the larger variety of visual concepts which are not covered in existing training sets. For this reason, novel object captioning (NOC) has recently emerged as a paradigm to test captioning models on objects which are unseen during the training phase. In this paper, we present a novel approach for NOC that learns to select the most relevant objects of an image, regardless of their adherence to the training set, and to constrain the generative process of a language model accordingly. Our architecture is fully-attentive and end-to-end trainable, also when incorporating constraints. We perform experiments on the held-out COCO dataset, where we demonstrate improvements over the state of the art, both in terms of adaptability to novel objects and caption quality.
Marco Cagrandi, Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi 0001, Rita Cucchiara
ICMR5
2021 SHREC 2021: Skeleton-based hand gesture recognition in the wild
Ariel Caputo, Andrea Giachetti 0001, Simone Soso, Deborah Pintani, Andrea D'Eusanio, Stefano Pini, Guido Borghi, Alessandro Simoni, Roberto Vezzani, Rita Cucchiara, Andrea Ranieri, Franca Giannini, Katia Lupinetti, Marina Monti, Mehran Maghoumi, Joseph J. LaViola Jr., Minh-Quan Le, Hai-Dang Nguyen, Minh-Triet Tran
Comput. Graph.10
2021 AC-VRNN: Attentive Conditional-VRNN for multi-future trajectory prediction
abstract
Anticipating human motion in crowded scenarios is essential for developing intelligent transportation systems, social-aware robots and advanced video surveillance applications. A key component of this task is represented by the inherently multi-modal nature of human paths which makes socially acceptable multiple futures when human interactions are involved. To this end, we propose a generative architecture for multi-future trajectory predictions based on Conditional Variational Recurrent Neural Networks (C-VRNNs). Conditioning mainly relies on prior belief maps, representing most likely moving directions and forcing the model to consider past observed dynamics in generating future positions. Human interactions are modelled with a graph-based attention mechanism enabling an online attentive hidden state refinement of the recurrent estimation. To corroborate our model, we perform extensive experiments on publicly-available datasets (e.g., ETH/UCY, Stanford Drone Dataset, STATS SportVU NBA, Intersection Drone Dataset and TrajNet++) and demonstrate its effectiveness in crowded scenes compared to several state-of-the-art methods.
Alessia Bertugli, Simone Calderara, Pasquale Coscia, Lamberto Ballan, Rita Cucchiara
Comput. Vis. Image Underst.5
2021 Multimodal attention networks for low-level vision-and-language navigation
Federico Landi, Lorenzo Baraldi 0001, Marcella Cornia, Massimiliano Corsini, Rita Cucchiara
Comput. Vis. Image Underst.5
2021 Video action detection by learning graph-based spatio-temporal interactions
abstract
Action Detection is a complex task that aims to detect and classify human actions in video clips. Typically, it has been addressed by processing fine-grained features extracted from a video classification backbone. Recently, thanks to the robustness of object and people detectors, a deeper focus has been added on relationship modeling. Following this line, we propose a graph-based framework to learn high-level interactions between people and objects, in both space and time. In our formulation, spatio-temporal relationships are learned through self-attention on a multi-layer graph structure which can connect entities from consecutive clips, thus considering long-range spatial and temporal dependencies. The proposed module is backbone independent by design and does not require end-to-end training. Extensive experiments are conducted on the AVA dataset, where our model demonstrates state-of-the-art results and consistent improvements over baselines built with different backbones. Code is publicly available at https://github.com/aimagelab/STAGE_action_detection.
Matteo Tomei, Lorenzo Baraldi 0001, Simone Calderara, Simone Bronzin, Rita Cucchiara
Comput. Vis. Image Underst.5
2021 Unifying tensor factorization and tensor nuclear norm approaches for low-rank tensor completion
Shiqiang Du, Qingjiang Xiao, Yuqing Shi, Rita Cucchiara, Yide Ma
Neurocomputing4
2021 Working Memory Connections for LSTM
Federico Landi, Lorenzo Baraldi 0001, Marcella Cornia, Rita Cucchiara
Neural Networks4
2020 A Transformer-Based Network for Dynamic Hand Gesture Recognition
abstract
Transformer-based neural networks represent a successful self-attention mechanism that achieves state-of-the-art results in language understanding and sequence modeling. However, their application to visual data and, in particular, to the dynamic hand gesture recognition task has not yet been deeply investigated. In this paper, we propose a transformer-based architecture for the dynamic hand gesture recognition task. We show that the employment of a single active depth sensor, specifically the usage of depth maps and the surface normals estimated from them, achieves state-of-the-art results, overcoming all the methods available in the literature on two automotive datasets, namely NVidia Dynamic Hand Gesture and Briareo. Moreover, we test the method with other data types available with common RGB-D devices, such as infrared and color data. We also assess the performance in terms of inference time and number of parameters, showing that the proposed framework is suitable for an online in-car infotainment system.
Andrea D'Eusanio, Alessandro Simoni, Stefano Pini, Guido Borghi, Roberto Vezzani, Rita Cucchiara
3DV6
2020 Conditional Channel Gated Networks for Task-Aware Continual Learning
abstract
Convolutional Neural Networks experience catastrophic forgetting when optimized on a sequence of learning problems: as they meet the objective of the current training examples, their performance on previous tasks drops drastically. In this work, we introduce a novel framework to tackle this problem with conditional computation. We equip each convolutional layer with task-specific gating modules, selecting which filters to apply on the given input. This way, we achieve two appealing properties. Firstly, the execution patterns of the gates allow to identify and protect important filters, ensuring no loss in the performance of the model for previously learned tasks. Secondly, by using a sparsity objective, we can promote the selection of a limited set of kernels, allowing to retain sufficient model capacity to digest new tasks. Existing solutions require, at test time, awareness of the task to which each example belongs to. This knowledge, however, may not be available in many practical scenarios. Therefore, we additionally introduce a task classifier that predicts the task label of each example, to deal with settings in which a task oracle is not available. We validate our proposal on four continual learning datasets. Results show that our model consistently outperforms existing methods both in the presence and the absence of a task oracle. Notably, on Split SVHN and Imagenet-50 datasets, our model yields up to 23.98% and 17.42% improvement in accuracy w.r.t. competing methods.
Davide Abati, Jakub M. Tomczak, Tijmen Blankevoort, Simone Calderara, Rita Cucchiara, Babak Ehteshami Bejnordi
CVPR5
2020 Meshed-Memory Transformer for Image Captioning
abstract
Transformer-based architectures represent the state of the art in sequence modeling tasks like machine translation and language understanding. Their applicability to multi-modal contexts like image captioning, however, is still largely under-explored. With the aim of filling this gap, we present M2- a Meshed Transformer with Memory for Image Captioning. The architecture improves both the image encoding and the language generation steps: it learns a multi-level representation of the relationships between image regions integrating learned a priori knowledge, and uses a mesh-like connectivity at decoding stage to exploit low- and high-level features. Experimentally, we investigate the performance of the M2Transformer and different fully-attentive models in comparison with recurrent ones. When tested on COCO, our proposal achieves a new state of the art in single-model and ensemble configurations on the "Karpathy" test split and on the online test server. We also assess its performances when describing objects unseen in the training set. Trained models and code for reproducing the experiments are publicly available at: https://github.com/aimagelab/meshed-memory-transformer.
Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi 0001, Rita Cucchiara
CVPR4
2020 Compressed Volumetric Heatmaps for Multi-Person 3D Pose Estimation
abstract
In this paper we present a novel approach for bottom-up multi-person 3D human pose estimation from monocular RGB images. We propose to use high resolution volumetric heatmaps to model joint locations, devising a simple and effective compression method to drastically reduce the size of this representation. At the core of the proposed method lies our Volumetric Heatmap Autoencoder, a fully-convolutional network tasked with the compression of ground-truth heatmaps into a dense intermediate representation. A second model, the Code Predictor, is then trained to predict these codes, which can be decompressed at test time to re-obtain the original representation. Our experimental evaluation shows that our method performs favorably when compared to state of the art on both multi-person and single-person 3D human pose estimation datasets and, thanks to our novel compression strategy, can process full-HD images at the constant runtime of 8 fps regardless of the number of subjects in the scene. Code and models are publicly available.
Matteo Fabbri, Fabio Lanzi, Simone Calderara, Stefano Alletto, Rita Cucchiara
CVPR5
2020 Baracca: a Multimodal Dataset for Anthropometric Measurements in Automotive
abstract
The recent spread of depth sensors has enabled new methods to automatically estimate anthropometric measurements, in place of manual procedures or expensive 3D scanners. Generally, the use of depth data is limited by the lack of depth-based public datasets containing accurate anthropometric annotations. Therefore, in this paper we propose a new dataset, called Baracca, specifically designed for the automotive context, including in-car and outside views. The dataset is multimodal: it has been acquired with synchronized depth, infrared, thermal and RGB cameras in order to deal with the requirements imposed by the automotive context. In addition, we propose several baselines to test the challenges of the presented dataset and provide considerations for future work.
Stefano Pini, Andrea D'Eusanio, Guido Borghi, Roberto Vezzani, Rita Cucchiara
IJCB5
2020 Explore and Explain: Self-supervised Navigation and Recounting
abstract
Embodied AI has been recently gaining attention as it aims to foster the development of autonomous and intelligent agents. In this paper, we devise a novel embodied setting in which an agent needs to explore a previously unknown environment while recounting what it sees during the path. In this context, the agent needs to navigate the environment driven by an exploration goal, select proper moments for description, and output natural language descriptions of relevant objects and scenes. Our model integrates a novel self-supervised exploration module with penalty, and a fully-attentive captioning model for explanation. Also, we investigate different policies for selecting proper moments for explanation, driven by information coming from both the environment and the navigation. Experiments are conducted on photorealistic environments from the Matterport3D dataset and investigate the navigation and explanation capabilities of the agent as well as the role of their interactions.
Roberto Bigazzi, Federico Landi, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara
ICPR6
2020 Watch Your Strokes: Improving Handwritten Text Recognition with Deformable Convolutions
abstract
Handwritten Text Recognition (HTR) in free-layout pages is a valuable yet challenging task which aims to automatically understand handwritten texts. State-of-the-art approaches in this field usually encode input images with Convolutional Neural Networks, whose kernels are typically defined on a fixed grid and focus on all input pixels independently. However, this is in contrast with the sparse nature of handwritten pages, in which only pixels representing the ink of the writing are useful for the recognition task. Furthermore, the standard convolution operator is not explicitly designed to take into account the great variability in shape, scale, and orientation of handwritten characters. To overcome these limitations, we investigate the use of deformable convolutions for handwriting recognition. The kernel of this type of convolution deforms according to the content of the neighborhood, and can therefore be more adaptable to geometric variations and other deformations of the text. Experiments conducted on the IAM and RIMES datasets demonstrate that the use of deformable convolutions is a promising direction for the design of novel architectures for handwritten text recognition.
Iulian Cojocaru, Silvia Cascianelli, Lorenzo Baraldi 0001, Massimiliano Corsini, Rita Cucchiara
ICPR5
2020 RefiNet: 3D Human Pose Refinement with Depth Maps
abstract
Human Pose Estimation is a fundamental task for many applications in the Computer Vision community and it has been widely investigated in the 2D domain, i.e. intensity images. Therefore, most of the available methods for this task are mainly based on 2D Convolutional Neural Networks and huge manually-annotated RGB datasets, achieving stunning results. In this paper, we propose RefiNet, a multi-stage framework that regresses an extremely-precise 3D human pose estimation from a given 2D pose and a depth map. The framework consists of three different modules, each one specialized in a particular refinement and data representation, i.e. depth patches, 3D skeleton and point clouds. Moreover, we present a new dataset, called Baracca, acquired with RGB, depth and thermal cameras and specifically created for the automotive context. Experimental results confirm the quality of the refinement procedure that largely improves the human pose estimations of off-the-shelf 2D methods.
Andrea D'Eusanio, Stefano Pini, Guido Borghi, Roberto Vezzani, Rita Cucchiara
ICPR5
2020 VITON-GT: An Image-based Virtual Try-On Model with Geometric Transformations
abstract
The large spread of online shopping has led computer vision researchers to develop different solutions for the fashion domain to potentially increase the online user experience and improve the efficiency of preparing fashion catalogs. Among them, image-based virtual try-on has recently attracted a lot of attention resulting in several architectures that can generate a new image of a person wearing an input try-on garment in a plausible and realistic way. In this paper, we present VITON-G T, a new model for virtual try-on that generates high-quality and photo-realistic images thanks to multiple geometric transformations. In particular, our model is composed of a two-stage geometric transformation module that performs two different projections on the input garment, and a transformation-guided try-on module that synthesizes the new image. We experimentally validate the proposed solution on the most common dataset for this task, containing mainly t-shirts, and we demonstrate its effectiveness compared to different baselines and previous methods. Additionally, we assess the generalization capabilities of our model on a new set of fashion items composed of upper-body clothes from different categories. To the best of our knowledge, we are the first to test virtual try-on architectures in this challenging experimental setting.
Matteo Fincato, Federico Landi, Marcella Cornia, Fabio Cesari, Rita Cucchiara
ICPR5
2020 Anomaly Detection, Localization and Classification for Railway Inspection
abstract
The ability to detect, localize and classify objects that are anomalies is a challenging task in the computer vision community. In this paper, we tackle these tasks developing a framework to automatically inspect the railway during the night. Specifically, it is able to predict the presence, the image coordinates and the class of obstacles. To deal with the low-light environment, the framework is based on thermal images and consists of three different modules that address the problem of detecting anomalies, predicting their image coordinates and classifying them. Moreover, due to the absolute lack of publicly-released datasets collected in the railway context for anomaly detection, we introduce a new multi-modal dataset, acquired from a rail drone, used to evaluate the proposed framework. Experimental results confirm the accuracy of the framework and its suitability, in terms of computational load, performance, and inference time, to be implemented on a self-powered inspection system.
Riccardo Gasparini, Andrea D'Eusanio, Guido Borghi, Stefano Pini, Giuseppe Scaglione, Simone Calderara, Eugenio Fedeli, Rita Cucchiara
ICPR8
2020 DAG-Net: Double Attentive Graph Neural Network for Trajectory Forecasting
abstract
Understanding human motion behaviour is a critical task for several possible applications like self-driving cars or social robots, and in general for all those settings where an autonomous agent has to navigate inside a human-centric environment. This is non-trivial because human motion is inherently multi-modal: given a history of human motion paths, there are many plausible ways by which people could move in the future. Additionally, people activities are often driven by goals, e.g. reaching particular locations or interacting with the environment. We address the aforementioned aspects by proposing a new recurrent generative model that considers both single agents' future goals and interactions between different agents. The model exploits a double attention-based graph neural network to collect information about the mutual influences among different agents and to integrate it with data about agents' possible future objectives. Our proposal is general enough to be applied to different scenarios: the model achieves state-of-the-art results in both urban environments and also in sports applications.
Alessio Monti, Alessia Bertugli, Simone Calderara, Rita Cucchiara
ICPR4
2020 Future Urban Scenes Generation Through Vehicles Synthesis
abstract
In this work we propose a deep learning pipeline to predict the visual future appearance of an urban scene. Despite recent advances, generating the entire scene in an end-to-end fashion is still far from being achieved. Instead, here we follow a two stages approach, where interpretable information is included in the loop and each actor is modelled independently. We leverage a per-object novel view synthesis paradigm; i.e. generating a synthetic representation of an object undergoing a geometrical roto-translation in the 3D space. Our model can be easily conditioned with constraints (e.g. input trajectories) provided by state-of-the-art tracking methods or by the user itself. This allows us to generate a set of diverse realistic futures starting from the same input in a multi-modal fashion. We visually and quantitatively show the superiority of this approach over traditional end-to-end scene-generation methods on CityFlow, a challenging real world dataset.
Alessandro Simoni, Luca Bergamini, Andrea Palazzi, Simone Calderara, Rita Cucchiara
ICPR5
2020 A Novel Attention-based Aggregation Function to Combine Vision and Language
abstract
The joint understanding of vision and language has been recently gaining a lot of attention in both the Computer Vision and Natural Language Processing communities, with the emergence of tasks such as image captioning, image-text matching, and visual question answering. As both images and text can be encoded as sets or sequences of elements - like regions and words - proper reduction functions are needed to transform a set of encoded elements into a single response, like a classification or similarity score. In this paper, we propose a novel fully-attentive reduction method for vision and language. Specifically, our approach computes a set of scores for each element of each modality employing a novel variant of cross-attention, and performs a learnable and cross-modal reduction, which can be used for both classification and ranking. We test our approach on image-text matching and visual question answering, building fair comparisons with other reduction choices, on both COCO and VQA 2.0 datasets. Experimentally, we demonstrate that our approach leads to a performance increase on both tasks. Further, we conduct ablation studies to validate the role of each component of the approach.
Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
ICPR4
2020 RMS-Net: Regression and Masking for Soccer Event Spotting
abstract
The recently proposed action spotting task consists in finding the exact timestamp in which an event occurs. This task fits particularly well for soccer videos, where events correspond to salient actions strictly defined by soccer rules (a goal occurs when the ball crosses the goal line). In this paper, we devise a lightweight and modular network for action spotting, which can simultaneously predict the event label and its temporal offset using the same underlying features. We enrich our model with two training strategies: the first one for data balancing and uniform sampling, the second for masking ambiguous frames and keeping the most discriminative visual cues. When tested on the SoccerNet dataset and using standard features, our full proposal exceeds the current state of the art by 3 Average-mAP points. Additionally, it reaches a gain of more than 10 Average-mAP points on the test set when fine-tuned in combination with a strong 2D backbone.
Matteo Tomei, Lorenzo Baraldi 0001, Simone Calderara, Simone Bronzin, Rita Cucchiara
ICPR5
2020 SMArT: Training Shallow Memory-aware Transformers for Robotic Explainability
abstract
The ability to generate natural language explanations conditioned on the visual perception is a crucial step towards autonomous agents which can explain themselves and communicate with humans. While the research efforts in image and video captioning are giving promising results, this is often done at the expense of the computational requirements of the approaches, limiting their applicability to real contexts. In this paper, we propose a fully-attentive captioning algorithm which can provide state-of-the-art performances on language generation while restricting its computational demands. Our model is inspired by the Transformer model and employs only two Transformer layers in the encoding and decoding stages. Further, it incorporates a novel memory-aware encoding of image regions. Experiments demonstrate that our approach achieves competitive results in terms of caption quality while featuring reduced computational demands. Further, to evaluate its applicability on autonomous agents, we conduct experiments on simulated scenes taken from the perspective of domestic robots.
Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
ICRA3
2020 A unified cycle-consistent neural model for text and image retrieval
Marcella Cornia, Lorenzo Baraldi 0001, Hamed Rezazadegan Tavakoli, Rita Cucchiara
Multim. Tools Appl.4
2020 Face-from-Depth for Head Pose Estimation on Depth Images
abstract
Depth cameras allow to set up reliable solutions for people monitoring and behavior understanding, especially when unstable or poor illumination conditions make unusable common RGB sensors. Therefore, we propose a complete framework for the estimation of the head and shoulder pose based on depth images only. A head detection and localization module is also included, in order to develop a complete end-to-end system. The core element of the framework is a Convolutional Neural Network, called POSEidon+, that receives as input three types of images and provides the 3D angles of the pose as output. Moreover, a Face-from-Depth component based on a Deterministic Conditional GAN model is able to hallucinate a face from the corresponding depth image. We empirically demonstrate that this positively impacts the system performances. We test the proposed framework on two public datasets, namely Biwi Kinect Head Pose and ICT-3DHP, and on Pandora, a new challenging dataset mainly inspired by the automotive setup. Experimental results show that our method overcomes several recent state-of-art works based on both intensity and depth input data, running in real-time at more than 30 frames per second.
Guido Borghi, Matteo Fabbri, Roberto Vezzani, Simone Calderara, Rita Cucchiara
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 Explaining digital humanities by aligning images and textual descriptions
Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi 0001, Massimiliano Corsini, Rita Cucchiara
Pattern Recognit. Lett.5
2019 Classifying Signals on Irregular Domains via Convolutional Cluster Pooling
abstract
We present a novel and hierarchical approach for supervised classification of signals spanning over a fixed graph, reflecting shared properties of the dataset. To this end, we introduce a Convolutional Cluster Pooling layer exploiting a multi-scale clustering in order to highlight, at different resolutions, locally connected regions on the input graph. Our proposal generalises well-established neural models such as Convolutional Neural Networks (CNNs) on irregular and complex domains, by means of the exploitation of the weight sharing property in a graph-oriented architecture. In this work, such property is based on the centrality of each vertex within its soft-assigned cluster. Extensive experiments on NTU RGB+D, CIFAR-10 and 20NEWS demonstrate the effectiveness of the proposed technique in capturing both local and global patterns in graph-structured data out of different domains.
Angelo Porrello, Davide Abati, Simone Calderara, Rita Cucchiara
AISTATS4
2019 Embodied Vision-and-Language Navigation with Dynamic Convolutional Filters
Federico Landi, Lorenzo Baraldi 0001, Massimiliano Corsini, Rita Cucchiara
BMVC4
2019 Latent Space Autoregression for Novelty Detection
abstract
Novelty detection is commonly referred as the discrimination of observations that do not conform to a learned model of regularity. Despite its importance in different application settings, designing a novelty detector is utterly complex due to the unpredictable nature of novelties and its inaccessibility during the training procedure, factors which expose the unsupervised nature of the problem. In our proposal, we design a general unsupervised framework where we equip a deep autoencoder with a parametric density estimator that learns the probability distribution underlying the latent representations with an autoregressive procedure. We show that a maximum likelihood objective, optimized in conjunction with the reconstruction of normal samples, effectively acts as a regularizer for the task at hand, by minimizing the differential entropy of the distribution spanned by latent vectors. In addition to providing a very general formulation, extensive experiments of our model on publicly available datasets deliver on-par or superior performances if compared to state-of-the-art methods in one-class and in video anomaly detection settings. Differently from our competitors, we remark that our proposal does not make any assumption about the nature of the novelties, making our work easily applicable to disparate contexts.
Davide Abati, Angelo Porrello, Simone Calderara, Rita Cucchiara
CVPR4
2019 Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions
abstract
Current captioning approaches can describe images using black-box architectures whose behavior is hardly controllable and explainable from the exterior. As an image can be described in infinite ways depending on the goal and the context at hand, a higher degree of controllability is needed to apply captioning algorithms in complex scenarios. In this paper, we introduce a novel framework for image captioning which can generate diverse descriptions by allowing both grounding and controllability. Given a control signal in the form of a sequence or set of image regions, we generate the corresponding caption through a recurrent architecture which predicts textual chunks explicitly grounded on regions, following the constraints of the given control. Experiments are conducted on Flickr30k Entities and on COCO Entities, an extended version of COCO in which we add grounding annotations collected in a semi-automatic manner. Results demonstrate that our method achieves state of the art performances on controllable image captioning, in terms of caption quality and diversity. Code and annotations are publicly available at: https://github.com/aimagelab/show-control-and-tell.
Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
CVPR3
2019 Art2Real: Unfolding the Reality of Artworks via Semantically-Aware Image-To-Image Translation
abstract
The applicability of computer vision to real paintings and artworks has been rarely investigated, even though a vast heritage would greatly benefit from techniques which can understand and process data from the artistic domain. This is partially due to the small amount of annotated artistic data, which is not even comparable to that of natural images captured by cameras. In this paper, we propose a semantic-aware architecture which can translate artworks to photo-realistic visualizations, thus reducing the gap between visual features of artistic and realistic data. Our architecture can generate natural images by retrieving and learning details from real photos through a similarity matching strategy which leverages a weakly-supervised semantic understanding of the scene. Experimental results show that the proposed technique leads to increased realism and to a reduction in domain shift, which improves the performance of pre-trained architectures for classification, detection, and segmentation. Code is publicly available at: https://github.com/aimagelab/art2real.
Matteo Tomei, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara
CVPR4
2019 Can adversarial networks hallucinate occluded people with a plausible aspect?
abstract
When you see a person in a crowd, occluded by other persons, you miss visual information that can be used to recognize, re-identify or simply classify him or her. You can imagine its appearance given your experience, nothing more. Similarly, AI solutions can try to hallucinate missing information with specific deep learning architectures, suitably trained with people with and without occlusions. The goal of this work is to generate a complete image of a person, given an occluded version in input, that should be a) without occlusion b) similar at pixel level to a completely visible people shape c) capable to conserve similar visual attributes (e.g. male/female) of the original one. For the purpose, we propose a new approach by integrating the state-of-the-art of neural network architectures, namely U-nets and GANs, as well as discriminative attribute classification nets, with an architecture specifically designed to de-occlude people shapes. The network is trained to optimize a Loss function which could take into account the aforementioned objectives. As well we propose two datasets for testing our solution: the first one, occluded RAP, created automatically by occluding real shapes of the RAP dataset created by Li et al. (2016) (which collects also attributes of the people aspect); the second is a large synthetic dataset, AiC, generated in computer graphics with data extracted from the GTA video game, that contains 3D data of occluded objects by construction. Results are impressive and outperform any other previous proposal. This result could be an initial step to many further researches to recognize people and their behavior in an open crowded world.
Federico Fulgeri, Matteo Fabbri, Stefano Alletto, Simone Calderara, Rita Cucchiara
Comput. Vis. Image Underst.5
2019 M-VAD names: a dataset for video captioning with naming
Stefano Pini, Marcella Cornia, Federico Bolelli, Lorenzo Baraldi 0001, Rita Cucchiara
Multim. Tools Appl.5
2019 Predicting the Driver's Focus of Attention: The DR(eye)VE Project
abstract
In this work we aim to predict the driver's focus of attention. The goal is to estimate what a person would pay attention to while driving, and which part of the scene around the vehicle is more critical for the task. To this end we propose a new computer vision model based on a multi-branch deep architecture that integrates three sources of information: raw video, motion and scene semantics. We also introduce DR(eye)VE, the largest dataset of driving scenes for which eye-tracking annotations are available. This dataset features more than 500,000 registered frames, matching ego-centric views (from glasses worn by drivers) and car-centric views (from roof-mounted camera), further enriched by other sensors measurements. Results highlight that several attention patterns are shared across drivers and can be reproduced to some extent. The indication of which elements in the scene are likely to capture the driver's attention may benefit several applications in the context of human-vehicle interaction and driver attention analysis.
Andrea Palazzi, Davide Abati, Simone Calderara, Francesco Solera, Rita Cucchiara
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 Self-Supervised Optical Flow Estimation by Projective Bootstrap
abstract
Dense optical flow estimation is complex and time consuming, with state-of-the-art methods relying either on large synthetic data sets or on pipelines requiring up to a few minutes per frame pair. In this paper, we address the problem of optical flow estimation in the automotive scenario in a self-supervised manner. We argue that optical flow can be cast as a geometrical warping between two successive video frames and devise a deep architecture to estimate such transformation in two stages. First, a dense pixel-level flow is computed with a projective bootstrap on rigid surfaces. We show how such global transformation can be approximated with a homography and extend spatial transformer layers so that they can be employed to compute the flow field implied by such transformation. Subsequently, we refine the prediction by feeding a second, deeper network that accounts for moving objects. A final reconstruction loss compares the warping of frame Xtwith the subsequent frame Xt+1and guides both estimates. The model has the speed advantages of end-to-end deep architectures while achieving competitive performances, both outperforming recent unsupervised methods and showing good generalization capabilities on new automotive data sets.
Stefano Alletto, Davide Abati, Simone Calderara, Rita Cucchiara, Luca Rigazio
IEEE Trans. Intell. Transp. Syst.4
2018 Learning to Generate Facial Depth Maps
abstract
In this paper, an adversarial architecture for facial depth map estimation from monocular intensity images is presented. By following an image-to-image approach, we combine the advantages of supervised learning and adversarial training, proposing a conditional Generative Adversarial Network that effectively learns to translate intensity face images into the corresponding depth maps. Two public datasets, namely Biwi database and Pandora dataset, are exploited to demonstrate that the proposed model generates high-quality synthetic depth images, both in terms of visual appearance and informative content. Furthermore, we show that the model is capable of predicting distinctive facial details by testing the generated depth maps through a deep model trained on authentic depth maps for the face verification task.
Stefano Pini, Filippo Grazioli, Guido Borghi, Roberto Vezzani, Rita Cucchiara
3DV5
2018 Face Verification from Depth using Privileged Information
Guido Borghi, Stefano Pini, Filippo Grazioli, Roberto Vezzani, Rita Cucchiara
BMVC5
2018 LAMV: Learning to Align and Match Videos With Kernelized Temporal Layers
abstract
This paper considers a learnable approach for comparing and aligning videos. Our architecture builds upon and revisits temporal match kernels within neural networks: we propose a new temporal layer that finds temporal alignments by maximizing the scores between two sequences of vectors, according to a time-sensitive similarity metric parametrized in the Fourier domain. We learn this layer with a temporal proposal strategy, in which we minimize a triplet loss that takes into account both the localization accuracy and the recognition rate. We evaluate our approach on video alignment, copy detection and event retrieval. Our approach outperforms the state on the art on temporal video alignment and video copy detection datasets in comparable setups. It also attains the best reported results for particular event search, while precisely aligning videos.
Lorenzo Baraldi 0001, Matthijs Douze, Rita Cucchiara, Hervé Jégou
CVPR3
2018 Learning to Detect and Track Visible and Occluded Body Joints in a Virtual World
Matteo Fabbri, Fabio Lanzi, Simone Calderara, Andrea Palazzi, Roberto Vezzani, Rita Cucchiara
ECCV (4)6
2018 Hands on the wheel: A Dataset for Driver Hand Detection and Tracking
abstract
The ability to detect, localize and track the hands is crucial in many applications requiring the understanding of the person behavior, attitude and interactions. In particular, this is true for the automotive context, in which hand analysis allows to predict preparatory movements for maneuvers or to investigate the driver's attention level. Moreover, due to the recent diffusion of cameras inside new car cockpits, it is feasible to use hand gestures to develop new Human-Car Interaction systems, more user-friendly and safe. In this paper, we propose a new dataset, called Turms, that consists of infrared images of driver's hands, collected from the back of the steering wheel, an innovative point of view. The Leap Motion device has been selected for the recordings, thanks to its stereo capabilities and the wide view-angle. Besides, we introduce a method to detect the presence and the location of driver's hands on the steering wheel, during driving activity tasks.
Guido Borghi, Elia Frigieri, Roberto Vezzani, Rita Cucchiara
FG4
2018 Fully Convolutional Network for Head Detection with Depth Images
abstract
Head detection and localization are one of the most investigated and demanding tasks of the Computer Vision community. These are also a key element for many disciplines, like Human Computer Interaction, Human Behavior Understanding, Face Analysis and Video Surveillance. In last decades, many efforts have been conducted to develop accurate and reliable head or face detectors on standard RGB images, but only few solutions concern other types of images, such as depth maps. In this paper, we propose a novel method for head detection on depth images, based on a deep learning approach. In particular, the presented system overcomes the classic sliding-window approach, that is often the main computational bottleneck of many object detectors, through a Fully Convolutional Network. Two public datasets, namely Pandora and Watch-n-Patch, are exploited to train and test the proposed network. Experimental results confirm the effectiveness of the method, that is able to exceed all the state-of-art works based on depth images and to run with real time performance.
Diego Ballotta, Guido Borghi, Roberto Vezzani, Rita Cucchiara
ICPR4
2018 Aligning Text and Document Illustrations: Towards Visually Explainable Digital Humanities
abstract
While several approaches to bring vision and language together are emerging, none of them has yet addressed the digital humanities domain, which, nevertheless, is a rich source of visual and textual data. To foster research in this direction, we investigate the learning of visual-semantic embeddings for historical document illustrations, devising both supervised and semi-supervised approaches. We exploit the joint visual-semantic embeddings to automatically align illustrations and textual elements, thus providing an automatic annotation of the visual content of a manuscript. Experiments are performed on the Borso d'Este Holy Bible, one of the most sophisticated illuminated manuscript from the Renaissance, which we manually annotate aligning every illustration with textual commentaries written by experts. Experimental results quantify the domain shift between ordinary visual-semantic datasets and the proposed one, validate the proposed strategies, and devise future works on the same line.
Lorenzo Baraldi 0001, Marcella Cornia, Costantino Grana, Rita Cucchiara
ICPR4
2018 Domain Translation with Conditional GANs: from Depth to RGB Face-to-Face
abstract
Can faces acquired by low-cost depth sensors be useful to catch some characteristic details of the face? Typically the answer is no. However, new deep architectures can generate RGB images from data acquired in a different modality, such as depth data. In this paper, we propose a new Deterministic Conditional GAN, trained on annotated RGB-D face datasets, effective for a face-to-face translation from depth to RGB. Although the network cannot reconstruct the exact somatic features for unknown individual faces, it is capable to reconstruct plausible faces; their appearance is accurate enough to be used in many pattern recognition tasks. In fact, we test the network capability to hallucinate with some Perceptual Probes, as for instance face aspect classification or landmark detection. Depth face can be used in spite of the correspondent RGB images, that often are not available due to difficult luminance conditions. Experimental results are very promising and are as far as better than previously proposed approaches: this domain translation can constitute a new way to exploit depth data in new future applications.
Matteo Fabbri, Guido Borghi, Fabio Lanzi, Roberto Vezzani, Simone Calderara, Rita Cucchiara
ICPR6
2018 Human Behaviour Understanding for Automotive and Surveillance
Rita Cucchiara
ICPRAM1
2018 Predicting Human Eye Fixations via an LSTM-Based Saliency Attentive Model
abstract
Data-driven saliency has recently gained a lot of attention thanks to the use of Convolutional Neural Networks for predicting gaze fixations. In this paper we go beyond standard approaches to saliency prediction, in which gaze maps are computed with a feed-forward network, and present a novel model which can predict accurate saliency maps by incorporating neural attentive mechanisms. The core of our solution is a Convolutional LSTM that focuses on the most salient regions of the input image to iteratively refine the predicted saliency map. Additionally, to tackle the center bias typical of human eye fixations, our model can learn a set of prior maps generated with Gaussian functions. We show, through an extensive evaluation, that the proposed architecture outperforms the current state of the art on public saliency prediction datasets. We further study the contribution of each key component to demonstrate their robustness on different scenarios.
Marcella Cornia, Lorenzo Baraldi 0001, Giuseppe Serra 0001, Rita Cucchiara
IEEE Trans. Image Process.4
2018 Paying More Attention to Saliency: Image Captioning with Saliency and Context Attention
abstract
Image captioning has been recently gaining a lot of attention thanks to the impressive achievements shown by deep captioning architectures, which combine Convolutional Neural Networks to extract image representations and Recurrent Neural Networks to generate the corresponding captions. At the same time, a significant research effort has been dedicated to the development of saliency prediction models, which can predict human eye fixations. Even though saliency information could be useful to condition an image captioning architecture, by providing an indication of what is salient and what is not, research is still struggling to incorporate these two techniques. In this work, we propose an image captioning approach in which a generative recurrent neural network can focus on different parts of the input image during the generation of the caption, by exploiting the conditioning given by a saliency prediction model on which parts of the image are salient and which are contextual. We show, through extensive quantitative and qualitative experiments on large-scale datasets, that our model achieves superior performance with respect to captioning baselines with and without saliency and to different state-of-the-art approaches combining saliency and captioning.
Marcella Cornia, Lorenzo Baraldi 0001, Giuseppe Serra 0001, Rita Cucchiara
ACM Trans. Multim. Comput. Commun. Appl.4
2018 Guest Editorial: Special Section on "Multimedia Understanding via Multimodal Analytics"
abstract
No abstract available.
Yan Yan 0002, Liqiang Nie, Rita Cucchiara
ACM Trans. Multim. Comput. Commun. Appl.3
2017 Generative adversarial models for people attribute recognition in surveillance
abstract
In this paper we propose a deep architecture for detecting people attributes (e.g. gender, race, clothing ...) in surveillance contexts. Our proposal explicitly deal with poor resolution and occlusion issues that often occur in surveillance footages by enhancing the images by means of Deep Convolutional Generative Adversarial Networks (DCGAN). Experiments show that by combining both our Generative Reconstruction and Deep Attribute Classification Network we can effectively extract attributes even when resolution is poor and in presence of strong occlusions up to 80% of the whole person figure.
Matteo Fabbri, Simone Calderara, Rita Cucchiara
AVSS3
2017 Hierarchical Boundary-Aware Neural Encoder for Video Captioning
abstract
The use of Recurrent Neural Networks for video captioning has recently gained a lot of attention, since they can be used both to encode the input video and to generate the corresponding description. In this paper, we present a recurrent video encoding scheme which can discover and leverage the hierarchical structure of the video. Unlike the classical encoder-decoder approach, in which a video is encoded continuously by a recurrent layer, we propose a novel LSTM cell which can identify discontinuity points between frames or segments and modify the temporal connections of the encoding layer accordingly. We evaluate our approach on three large-scale datasets: the Montreal Video Annotation dataset, the MPII Movie Description dataset and the Microsoft Video Description Corpus. Experiments show that our approach can discover appropriate hierarchical representations of input videos and improve the state of the art results on movie description datasets.
Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara
CVPR3
2017 POSEidon: Face-from-Depth for Driver Pose Estimation
abstract
Fast and accurate upper-body and head pose estimation is a key task for automatic monitoring of driver attention, a challenging context characterized by severe illumination changes, occlusions and extreme poses. In this work, we present a new deep learning framework for head localization and pose estimation on depth images. The core of the proposal is a regressive neural network, called POSEidon, which is composed of three independent convolutional nets followed by a fusion layer, specially conceived for understanding the pose by depth. In addition, to recover the intrinsic value of face appearance for understanding head position and orientation, we propose a new Face-from-Depth model for learning image faces from depth. Results in face reconstruction are qualitatively impressive. We test the proposed framework on two public datasets, namely Biwi Kinect Head Pose and ICT-3DHP, and on Pandora, a new challenging dataset mainly inspired by the automotive setup. Results show that our method overcomes all recent state-of-art works, running in real time at more than 30 frames per second.
Guido Borghi, Marco Venturelli, Roberto Vezzani, Rita Cucchiara
CVPR4
2017 Modeling multimodal cues in a deep learning-based framework for emotion recognition in the wild
abstract
In this paper, we propose a multimodal deep learning architecture for emotion recognition in video regarding our participation to the audio-video based sub-challenge of the Emotion Recognition in the Wild 2017 challenge. Our model combines cues from multiple video modalities, including static facial features, motion patterns related to the evolution of the human expression over time, and audio information. Specifically, it is composed of three sub-networks trained separately: the first and second ones extract static visual features and dynamic patterns through 2D and 3D Convolutional Neural Networks (CNN), while the third one consists in a pretrained audio network which is used to extract useful deep acoustic signals from video. In the audio branch, we also apply Long Short Term Memory (LSTM) networks in order to capture the temporal evolution of the audio features. To identify and exploit possible relationships among different modalities, we propose a fusion network that merges cues from the different modalities in one representation. The proposed architecture outperforms the challenge baselines (38.81 % and 40.47 %): we achieve an accuracy of 50.39 % and 49.92 % respectively on the validation and the testing data.
Stefano Pini, Olfa Ben Ahmed, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara, Benoit Huet
ICMI5
2017 Embedded recurrent network for head pose estimation in car
abstract
An accurate and fast driver's head pose estimation is a rich source of information, in particular in the automotive context. Head pose is a key element for driver's behavior investigation, pose analysis, attention monitoring and also a useful component to improve the efficacy of Human-Car Interaction systems. In this paper, a Recurrent Neural Network is exploited to tackle the problem of driver head pose estimation, directly and only working on depth images to be more reliable in presence of varying or insufficient illumination. Experimental results, obtained from two public dataset, namely Biwi Kinect Head Pose and ICT-3DHP Database, prove the efficacy of the proposed method that overcomes state-of-art works. Besides, the entire system is implemented and tested on two embedded boards with real time performance.
Guido Borghi, Riccardo Gasparini, Roberto Vezzani, Rita Cucchiara
Intelligent Vehicles Symposium4
2017 Learning where to attend like a human driver
abstract
Despite the advent of autonomous cars, it's likely - at least in the near future - that human attention will still maintain a central role as a guarantee in terms of legal responsibility during the driving task. In this paper we study the dynamics of the driver's gaze and use it as a proxy to understand related attentional mechanisms. First, we build our analysis upon two questions: where and what the driver is looking at? Second, we model the driver's gaze by training a coarse-to-fine convolutional network on short sequences extracted from the DR(eye)VE dataset. Experimental comparison against different baselines reveal that the driver's gaze can indeed be learnt to some extent, despite (i) being highly subjective and (ii) having only one driver's gaze available for each sequence due to the irreproducibility of the scene. Eventually, we advocate for a new assisted driving paradigm which suggests to the driver, with no intervention, where she should focus her attention.
Andrea Palazzi, Francesco Solera, Simone Calderara, Stefano Alletto, Rita Cucchiara
Intelligent Vehicles Symposium5
2017 Video registration in egocentric vision under day and night illumination changes
Stefano Alletto, Giuseppe Serra 0001, Rita Cucchiara
Comput. Vis. Image Underst.3
2017 Segmentation models diversity for object proposals
Marco Manfredi, Costantino Grana, Rita Cucchiara, Arnold W. M. Smeulders
Comput. Vis. Image Underst.3
2017 Tracking Social Groups Within and Across Cameras
abstract
We propose a method for tracking groups from single and multiple cameras with disjointed fields of view. Our formulation follows the tracking-by-detection paradigm in which groups are the atomic entities and are linked over time to form long and consistent trajectories. To this end, we formulate the problem as a supervised clustering problem in which a structural SVM classifier learns a similarity measure appropriate for group entities. Multicamera group tracking is handled inside the framework by adopting an orthogonal feature encoding that allows the classifier to learn inter- and intra-camera feature weights differently. Experiments were carried out on a novel annotated group tracking data set, the DukeMTMC-Groups data set. Since this is the first data set on the problem, it comes with the proposal of a suitable evaluation measure. Results of adopting learning for the task are encouraging, scoring a +15% improvement in F1measure over a nonlearning-based clustering baseline. To the best of our knowledge, this is the first proposal of its kind dealing with multicamera group tracking.
Francesco Solera, Simone Calderara, Ergys Ristani, Carlo Tomasi, Rita Cucchiara
IEEE Trans. Circuits Syst. Video Technol.5
2017 Guest Editorial Special Issue on Wearable and Ego-Vision Systems for Augmented Experience
abstract
The papers in this special section focus on the deployment of wearable computing technologies and ego-vision systems for augmented reality applications. Rapid progress in the development of low-level component technologies such as wearable sensors, wearable displays, and wearable computers is making our digital lives grow, connect, and play a relevant role in reality. To name a few examples, body-mounted sensors and displays help athletes in training by presenting real-time performance metrics such as speed, distance, and heart rate.Wearable systems allow medical staff in hospitals to consult specialists located anywhere in the world, in real time, providing optimal patient care. And within the context of assistive technologies, a head-mounted camera can be used to identify and convey the presence of objects, people, or text to a visually impaired user.
Giuseppe Serra 0001, Rita Cucchiara, Kris Makoto Kitani, Javier Civera 0001
IEEE Trans. Hum. Mach. Syst.2
2017 Recognizing and Presenting the Storytelling Video Structure With Deep Multimodal Networks
abstract
In this paper, we propose a novel scene detection algorithm which employs semantic, visual, textual, and audio cues. We also show how the hierarchical decomposition of the storytelling video structure can improve retrieval results presentation with semantically and aesthetically effective thumbnails. Our method is built upon two advancements of the state of the art: first is semantic feature extraction which builds video-specific concept detectors; and second is multimodal feature embedding learning that maps the feature vector of a shot to a space in which the Euclidean distance has task specific semantic properties. The proposed method is able to decompose the video in annotated temporal segments which allow us for a query specific thumbnail extraction. Extensive experiments are performed on different data sets to demonstrate the effectiveness of our algorithm. An in-depth discussion on how to deal with the subjectivity of the task is conducted and a strategy to overcome the problem is suggested.
Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara
IEEE Trans. Multim.3
2017 Personalized Egocentric Video Summarization of Cultural Tour on User Preferences Input
abstract
In this paper, we propose a new method for customized summarization of egocentric videos according to specific user preferences, so that different users can extract different summaries from the same stream. Our approach, tailored on a cultural heritage scenario, relies on creating a short synopsis of the original video focused on key shots, in which concepts relevant to user preferences can be visually detected and the chronological flow of the original video is preserved. Moreover, we release a new dataset, composed of egocentric streams taken in uncontrolled scenarios, capturing tourists cultural visits in six art cities, with geolocalization information. Our experimental results show that the proposed approach is able to leverage user's preferences with an accent on storyline chronological flow and on visual smoothness.
Patrizia Varini, Giuseppe Serra 0001, Rita Cucchiara
IEEE Trans. Multim.3
2017 Affective level design for a role-playing videogame evaluated by a brain-computer interface and machine learning methods
Fabrizio Balducci, Costantino Grana, Rita Cucchiara
Vis. Comput.3
2016 Spotting prejudice with nonverbal behaviours
abstract
Despite prejudice cannot be directly observed, nonverbal behaviours provide profound hints on people inclinations. In this paper, we use recent sensing technologies and machine learning techniques to automatically infer the results of psychological questionnaires frequently used to assess implicit prejudice. In particular, we recorded 32 students discussing with both white and black collaborators. Then, we identified a set of features allowing automatic extraction and measured their degree of correlation with psychological scores. Results confirmed that automated analysis of nonverbal behaviour is actually possible thus paving the way for innovative clinical tools and eventually more secure societies.
Andrea Palazzi, Simone Calderara, Nicola Bicocchi, Loris Vezzali, Gian Antonio di Bernardo, Franco Zambonelli, Rita Cucchiara
UbiComp7
2016 Fast gesture recognition with Multiple Stream Discrete HMMs on 3D skeletons
abstract
HMMs are widely used in action and gesture recognition due to their implementation simplicity, low computational requirement, scalability and high parallelism. They have worth performance even with a limited training set. All these characteristics are hard to find together in other even more accurate methods. In this paper, we propose a novel double-stage classification approach, based on Multiple Stream Discrete Hidden Markov Models (MSD-HMM) and 3D skeleton joint data, able to reach high performances maintaining all advantages listed above. The approach allows both to quickly classify pre-segmented gestures (offline classification), and to perform temporal segmentation on streams of gestures (online classification) faster than real time. We test our system on three public datasets, MSRAction3D, UTKinect-Action and MSRDailyAction, and on a new dataset, Kinteract Dataset, explicitly created for Human Computer Interaction (HCI). We obtain state of the art performances on all of them.
Guido Borghi, Roberto Vezzani, Rita Cucchiara
ICPR3
2016 Historical document digitization through layout analysis and deep content classification
abstract
Document layout segmentation and recognition is an important task in the creation of digitized documents collections, especially when dealing with historical documents. This paper presents an hybrid approach to layout segmentation as well as a strategy to classify document regions, which is applied to the process of digitization of an historical encyclopedia. Our layout analysis method merges a classic top-down approach and a bottom-up classification process based on local geometrical features, while regions are classified by means of features extracted from a Convolutional Neural Network merged in a Random Forest classifier. Experiments are conducted on the first volume of the “Enciclopedia Treccani”, a large dataset containing 999 manually annotated pages from the historical Italian encyclopedia.
Andrea Corbelli, Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara
ICPR4
2016 A deep multi-level network for saliency prediction
abstract
This paper presents a novel deep architecture for saliency prediction. Current state of the art models for saliency prediction employ Fully Convolutional networks that perform a non-linear combination of features extracted from the last convolutional layer to predict saliency maps. We propose an architecture which, instead, combines features extracted at different levels of a Convolutional Neural Network (CNN). Our model is composed of three main blocks: a feature extraction CNN, a feature encoding network, that weights low and high level feature maps, and a prior learning network. We compare our solution with state of the art saliency models on two public benchmarks datasets. Results show that our model outperforms under all evaluation metrics on the SALICON dataset, which is currently the largest public dataset for saliency prediction, and achieves competitive results on the MIT300 benchmark. Code is available at https://github.com/marcellacornia/mlnet.
Marcella Cornia, Lorenzo Baraldi 0001, Giuseppe Serra 0001, Rita Cucchiara
ICPR4
2016 Scene-driven Retrieval in Edited Videos using Aesthetic and Semantic Deep Features
abstract
This paper presents a novel retrieval pipeline for video collections, which aims to retrieve the most significant parts of an edited video for a given query, and represent them with thumbnails which are at the same time semantically meaningful and aesthetically remarkable. Videos are first segmented into coherent and story-telling scenes, then a retrieval algorithm based on deep learning is proposed to retrieve the most significant scenes for a textual query. A ranking strategy based on deep features is finally used to tackle the problem of visualizing the best thumbnail. Qualitative and quantitative experiments are conducted on a collection of edited videos to demonstrate the effectiveness of our approach.
Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara
ICMR3
2016 Motion Segmentation using Visual and Bio-mechanical Features
abstract
Nowadays, egocentric wearable devices are continuously increasing their widespread among both the academic community and the general public. For this reason, methods capable of automatically segment the video based on the recorder motion patterns are gaining attention. These devices present the unique opportunity of both high quality video recordings and multimodal sensors readings. Significant efforts have been made in either analyzing the video stream recorded by these devices or the bio-mechanical sensor information. So far, the integration between these two realities has not been fully addressed, and the real capabilities of these devices are not yet exploited. In this paper, we present a solution to segment a video sequence into motion activities by introducing a novel data fusion technique based on the covariance of visual and bio-mechanical features. The experimental results are promising and show that the proposed integration strategy outperforms the results achieved focusing solely on a single source.
Stefano Alletto, Giuseppe Serra 0001, Rita Cucchiara
ACM Multimedia3
2016 A Browsing and Retrieval System for Broadcast Videos using Scene Detection and Automatic Annotation
abstract
This paper presents a novel video access and retrieval system for edited videos. The key element of the proposal is that videos are automatically decomposed into semantically coherent parts (called scenes) to provide a more manageable unit for browsing, tagging and searching. The system features an automatic annotation pipeline, with which videos are tagged by exploiting both the transcript and the video itself. Scenes can also be retrieved with textual queries; the best thumbnail for a query is selected according to both semantics and aesthetics criteria.
Lorenzo Baraldi 0001, Costantino Grana, Alberto Messina, Rita Cucchiara
ACM Multimedia4
2016 An Indoor Location-Aware System for an IoT-Based Smart Museum
abstract
The new technologies characterizing the Internet of Things (IoT) allow realizing real smart environments able to provide advanced services to the users. Recently, these smart environments are also being exploited to renovate the users' interest on the cultural heritage, by guaranteeing real interactive cultural experiences. In this paper, we design and validate an indoor location-aware architecture able to enhance the user experience in a museum. In particular, the proposed system relies on a wearable device that combines image recognition and localization capabilities to automatically provide the users with cultural contents related to the observed artworks. The localization information is obtained by a Bluetooth low energy (BLE) infrastructure installed in the museum. Moreover, the system interacts with the Cloud to store multimedia contents produced by the user and to share environment-generated events on his/her social networks. Finally, several location-aware services, running in the system, control the environment status also according to users' movements. These services interact with physical devices through a multiprotocol middleware. The system has been designed to be easily extensible to other IoT technologies and its effectiveness has been evaluated in the MUST museum, Lecce, Italy.
Stefano Alletto, Rita Cucchiara, Giuseppe Del Fiore, Luca Mainetti, Vincenzo Mighali, Luigi Patrono, Giuseppe Serra 0001
IEEE Internet Things J.2
2016 Layout analysis and content enrichment of digitized books
Costantino Grana, Giuseppe Serra 0001, Marco Manfredi, Dalia Coppi, Rita Cucchiara
Multim. Tools Appl.5
2016 Socially Constrained Structural Learning for Groups Detection in Crowd
abstract
Modern crowd theories agree that collective behavior is the result of the underlying interactions among small groups of individuals. In this work, we propose a novel algorithm for detecting social groups in crowds by means of a Correlation Clustering procedure on people trajectories. The affinity between crowd members is learned through an online formulation of the Structural SVM framework and a set of specifically designed features characterizing both their physical and social identity, inspired by Proxemic theory, Granger causality, DTW and Heat-maps. To adhere to sociological observations, we introduce a loss function ( G -MITRE) able to deal with the complexity of evaluating group detection performances. We show our algorithm achieves state-of-the-art results when relying on both ground truth trajectories and tracklets previously extracted by available detector/tracker systems.
Francesco Solera, Simone Calderara, Rita Cucchiara
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Transductive People Tracking in Unconstrained Surveillance
abstract
Long-term tracking of people in unconstrained scenarios is still an open problem due to the absence of constant elements in the problem setting. The camera, when active, may move and the appearance of both the background and the target may change abruptly, leading to the inadequacy of most standard tracking techniques. We propose to exploit a learning approach that considers the tracking task as a semisupervised learning problem. Given few target samples, the aim is to search for the target occurrences in the video stream, reinterpreting the problem as label propagation on a similarity graph. We propose a solution based on graph transduction that iteratively works frame by frame. In addition, to avoid drifting, we introduce an update strategy based on an evolutionary clustering technique that chooses the visual templates that better describe target appearance, evolving the model during the processing of the video. Since we model people's appearance by means of covariance matrices on color and gradient information, our framework is directly related to structure learning on Riemannian manifolds. Tests on publicly available data sets and comparisons with state-of-the-art techniques allow us to conclude that our solution exhibits interesting performances in terms of tracking precision and recall in most of the considered scenarios.
Dalia Coppi, Simone Calderara, Rita Cucchiara
IEEE Trans. Circuits Syst. Video Technol.3
2015 Towards the evaluation of reproducible robustness in tracking-by-detection
abstract
Conventional experiments on MTT are built upon the belief that fixing the detections to different trackers is sufficient to obtain a fair comparison. In this work we argue how the true behavior of a tracker is exposed when evaluated by varying the input detections rather than by fixing them. We propose a systematic and reproducible protocol and a MATLAB toolbox for generating synthetic data starting from ground truth detections, a proper set of metrics to understand and compare trackers peculiarities and respective visualization solutions.
Francesco Solera, Simone Calderara, Rita Cucchiara
AVSS3
2015 Automatic configuration and calibration of modular sensing floors
abstract
Sensing floors are becoming an emerging solution for many privacy-compliant and large area surveillance systems. Many research and even commercial technologies have been proposed in the last years. Similarly to distributed camera networks, the problem of calibration is crucial, specially when installed in wide areas. This paper addresses the general problem of automatic calibration and configuration of modular and scalable sensing floors. Working on training data only, the system automatically finds the spatial placement of each sensor module and estimates threshold parameters needed for people detection. Tests on several training sequences captured with a commercial sensing floor are provided to validate the method.
Roberto Vezzani, Martino Lombardi, Rita Cucchiara
AVSS3
2015 Shot and Scene Detection via Hierarchical Clustering for Re-using Broadcast Video
Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara
CAIP (1)3
2015 Learning to Divide and Conquer for Online Multi-target Tracking
abstract
Online Multiple Target Tracking (MTT) is often addressed within the tracking-by-detection paradigm. Detections are previously extracted independently in each frame and then objects trajectories are built by maximizing specifically designed coherence functions. Nevertheless, ambiguities arise in presence of occlusions or detection errors. In this paper we claim that the ambiguities in tracking could be solved by a selective use of the features, by working with more reliable features if possible and exploiting a deeper representation of the target only if necessary. To this end, we propose an online divide and conquer tracker for static camera scenes, which partitions the assignment problem in local subproblems and solves them by selectively choosing and combining the best features. The complete framework is cast as a structural learning task that unifies these phases and learns tracker parameters from examples. Experiments on two different datasets highlights a significant improvement of tracking performances (MOTA +10%) over the state of the art.
Francesco Solera, Simone Calderara, Rita Cucchiara
ICCV3
2015 Scene segmentation using temporal clustering for accessing and re-using broadcast video
abstract
Scene detection is a fundamental tool for allowing effective video browsing and re-using. In this paper we present a model that automatically divides videos into coherent scenes, which is based on a novel combination of local image descriptors and temporal clustering techniques. Experiments are performed to demonstrate the effectiveness of our approach, by comparing our algorithm against two recent proposals for automatic scene segmentation. We also propose improved performance measures that aim to reduce the gap between numerical evaluation and expected results.
Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara
ICME3
2015 Personalized Egocentric Video Summarization for Cultural Experience
abstract
Recent egocentric video summarization approaches have dealt with motion analysis and social interaction without considering that user can be interested in preserving only part of the video related to his interests. In this paper we propose a new method for personalized video summarization of cultural experiences with the goal of extracting from the streams only the scenes corresponding to a user's specific topics request, chosen among the shots in which it's possible to deduce that the visitor was focusing on a point of interest. Preliminary experiments show that our approach is promising and allows visitor to better customize the summary of his experience.
Patrizia Varini, Giuseppe Serra 0001, Rita Cucchiara
ICMR3
2015 A Deep Siamese Network for Scene Detection in Broadcast Videos
abstract
We present a model that automatically divides broadcast videos into coherent scenes by learning a distance measure between shots. Experiments are performed to demonstrate the effectiveness of our approach by comparing our algorithm against recent proposals for automatic scene segmentation. We also propose an improved performance measure that aims to reduce the gap between numerical evaluation and expected results, and propose and release a new benchmark dataset.
Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara
ACM Multimedia3
2015 Egocentric Video Summarization of Cultural Tour based on User Preferences
abstract
In this paper, we propose a new method to obtain customized video summarization according to specific user preferences. Our approach is tailored on Cultural Heritage scenario and is designed on identifying candidate shots, selecting from the original streams only the scenes with behavior patterns related to the presence of relevant experiences, and further filtering them in order to obtain a summary matching the requested user's preferences. Our preliminary results show that the proposed approach is able to leverage user's preferences in order to obtain a customized summary, so that different users may extract from the same stream different summaries.
Patrizia Varini, Giuseppe Serra 0001, Rita Cucchiara
ACM Multimedia3
2015 GOLD: Gaussians of Local Descriptors for image representation
Giuseppe Serra 0001, Costantino Grana, Marco Manfredi, Rita Cucchiara
Comput. Vis. Image Underst.4
2015 Mapping Appearance Descriptors on 3D Body Models for People Re-identification
Davide Baltieri, Roberto Vezzani, Rita Cucchiara
Int. J. Comput. Vis.3
2015 Understanding social relationships in egocentric vision
Stefano Alletto, Giuseppe Serra 0001, Simone Calderara, Rita Cucchiara
Pattern Recognit.4
2015 A General-Purpose Sensing Floor Architecture for Human-Environment Interaction
abstract
Smart environments are now designed as natural interfaces to capture and understand human behavior without a need for explicit human-computer interaction. In this article, we present a general-purpose architecture that acquires and understands human behaviors through a sensing floor. The pressure field generated by moving people is captured and analyzed. Specific actions and events are then detected by a low-level processing engine and sent to high-level interfaces providing different functions. The proposed architecture and sensors are modular, general-purpose, cheap, and suitable for both small- and large-area coverage. Some sample entertainment and virtual reality applications that we developed to test the platform are presented.
Roberto Vezzani, Martino Lombardi, Augusto Pieracci, Paolo Santinelli, Rita Cucchiara
ACM Trans. Interact. Intell. Syst.5
2014 Learning superpixel relations for supervised image segmentation
abstract
In this paper we propose to extend the well known graph cut segmentation framework by learning superpixel relations and use them to weight superpixel-to-superpixel edges in a superpixel graph. Adjacent superpixel-pairs are analyzed to build an object boundary model, able to discriminate between superpixel-pairs belonging to the same object or placed on the edge between the foreground object and the background. Several superpixel-pair features are investigated and exploited to build a non-linear SVM to learn object boundary appearance. The adoption of this modified graph cut enhances the performance of a previously proposed segmentation method on two publicly available datasets, reaching state-of-the-art results.
Marco Manfredi, Costantino Grana, Rita Cucchiara
ICIP3
2014 Head Pose Estimation in First-Person Camera Views
abstract
In this paper we present a new method for head pose real-time estimation in ego-vision scenarios that is a key step in the understanding of social interactions. In order to robustly detect head under changing aspect ratio, scale and orientation we use and extend the Hough-Based Tracker which allows to follow simultaneously each subject in the scene. In an ego-vision scenario where a group interacts in a discussion, each subject's head orientation will be more likely to remain focused for a while on the person who has the floor. In order to encode this behavior we include a stateful Hidden Markov Model technique that enforces the predicted pose with the temporal coherence from a video sequence. We extensively test our approach on several indoor and outdoor ego-vision videos with high illumination variations showing its validity and outperforming other recent related state of the art approaches.
Stefano Alletto, Giuseppe Serra 0001, Simone Calderara, Rita Cucchiara
ICPR4
2014 Learning Graph Cut Energy Functions for Image Segmentation
abstract
In this paper we address the task of learning how to segment a particular class of objects, by means of a training set of images and their segmentations. In particular we propose a method to overcome the extremely high training time of a previously proposed solution to this problem, Kernelized Structural Support Vector Machines. We employ a one-class SVM working with joint kernels to robustly learn significant support vectors (representative image-mask pairs) and accordingly weight them to build a suitable energy function for the graph cut framework. We report results obtained on two public datasets and a comparison of training times on different training set sizes.
Marco Manfredi, Costantino Grana, Rita Cucchiara
ICPR3
2014 Kernelized Structural Classification for 3D Dogs Body Parts Detection
abstract
Despite pattern recognition methods for human behavioral analysis has flourished in the last decade, animal behavioral analysis has been almost neglected. Those few approaches are mostly focused on preserving livestock economic value while attention on the welfare of companion animals, like dogs, is now emerging as a social need. In this work, following the analogy with human behavior recognition, we propose a system for recognizing body parts of dogs kept in pens. We decide to adopt both 2D and 3D features in order to obtain a rich description of the dog model. Images are acquired using the Microsoft Kinect to capture the depth map images of the dog. Upon depth maps a Structural Support Vector Machine (SSVM) is employed to identify the body parts using both 3D features and 2D images. The proposal relies on a kernelized discriminative structural classificator specifically tailored for dogs independently from the size and breed. The classification is performed in an online fashion using the LaRank optimization technique to obtaining real time performances. Promising results have emerged during the experimental evaluation carried out at a dog shelter, managed by IZSAM, in Teramo, Italy.
Simone Pistocchi, Simone Calderara, Shanis Barnard, Nicola Ferri, Rita Cucchiara
ICPR5
2014 On detection of novel categories and subcategories of images using incongruence
abstract
Novelty detection is a crucial task in the development of autonomous vision systems. It aims at detecting if samples do not conform with the learnt models. In this paper, we consider the problem of detecting novelty in object recognition problems in which the set of object classes are grouped to form a semantic hierarchy. We follow the idea that, within a semantic hierarchy, novel samples can be defined as samples whose categorization at a specific level contrasts with the categorization at a more general level. This measure indicates if a sample is novel and, in that case, if it is likely to belong to a novel broad category or to a novel sub-category. We present an evaluation of this approach on two hierarchical subsets of the Caltech256 objects dataset and on the SUN scenes dataset, with different classification schemes. We obtain an improvement over Weinshall et al. and show that it is possible to bypass their normalisation heuristic. We demonstrate that this approach achieves good novelty detection rates as far as the conceptual taxonomy is congruent with the visual hierarchy, but tends to fail if this assumption is not satisfied.
Dalia Coppi, Teófilo Emídio de Campos, Fei Yan 0001, Josef Kittler, Rita Cucchiara
ICMR5
2014 Covariance of Covariance Features for Image Classification
abstract
In this paper we propose a novel image descriptor built by computing the covariance of pixel level features on densely sampled patches and encoding them using their covariance. Appropriate projections to the Euclidean space and feature normalizations are employed in order to provide a strong descriptor usable with linear classifiers. In order to remove border effects, we further enhance the Spatial Pyramid representation with bilinear interpolation. Experimental results conducted on two common datasets for object and texture classification show that the performance of our method is comparable with state of the art techniques, but removing any dataset specific dependency in the feature encoding step.
Giuseppe Serra 0001, Costantino Grana, Marco Manfredi, Rita Cucchiara
ICMR4
2014 Miniature illustrations retrieval and innovative interaction for digital illuminated manuscripts
Daniele Borghesani, Costantino Grana, Rita Cucchiara
Multim. Syst.3
2014 3D Hough transform for sphere recognition on point clouds - A systematic study and a new method proposal
Marco Camurri, Roberto Vezzani, Rita Cucchiara
Mach. Vis. Appl.3
2014 A complete system for garment segmentation and color classification
Marco Manfredi, Costantino Grana, Simone Calderara, Rita Cucchiara
Mach. Vis. Appl.4
2014 Visual Tracking: An Experimental Survey
abstract
There is a large variety of trackers, which have been proposed in the literature during the last two decades with some mixed success. Object tracking in realistic scenarios is a difficult problem, therefore, it remains a most active area of research in computer vision. A good tracker should perform well in a large number of videos involving illumination changes, occlusion, clutter, camera motion, low contrast, specularities, and at least six more aspects. However, the performance of proposed trackers have been evaluated typically on less than ten videos, or on the special purpose datasets. In this paper, we aim to evaluate trackers systematically and experimentally on 315 video fragments covering above aspects. We selected a set of nineteen trackers to include a wide variety of algorithms often cited in literature, supplemented with trackers appearing in 2010 and 2011 for which the code was publicly available. We demonstrate that trackers can be evaluated objectively by survival curves, Kaplan Meier statistics, and Grubs testing. We find that in the evaluation practice the F-score is as effective as the object tracking accuracy (OTA) score. The analysis under a large variety of circumstances provides objective insight into the strengths and weaknesses of trackers.
Arnold W. M. Smeulders, Dung Manh Chu, Rita Cucchiara, Simone Calderara, Afshin Dehghan, Mubarak Shah
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 A fast and effective ellipse detector for embedded vision applications
Michele Fornaciari, Andrea Prati 0001, Rita Cucchiara
Pattern Recognit.3
2014 Pattern recognition and crowd analysis
Stefania Bandini, Simone Calderara, Rita Cucchiara
Pattern Recognit. Lett.3
2014 Detection of static groups and crowds gathered in open spaces by texture classification
Marco Manfredi, Roberto Vezzani, Simone Calderara, Rita Cucchiara
Pattern Recognit. Lett.4
2013 Sensing floors for privacy-compliant surveillance of wide areas
abstract
Surveillance systems can really benefit from the integration of multiple and heterogeneous sensors. In this paper we describe an innovative sensing floor. Thanks to its low cost and ease of installation, the floor is suitable for both private and public environments, from narrow zones to wide areas. The floor is made adding a sensing layer below commercial floating tiles. The sensor is scalable, reliable, and completely invisible to the users. The temporal and spatial resolutions of the data are high enough to identify the presence of people, to recognize their behavior and to detect events in a privacy compliant way. Experimental results on a real prototype implementation confirm the potentiality of the framework.
Martino Lombardi, Augusto Pieracci, Paolo Santinelli, Roberto Vezzani, Rita Cucchiara
AVSS5
2013 A people counting system for business analytics
abstract
This paper deals with people counting in stores for business analytics using stereo vision. Among the several problems in this type of applications, two are the most relevant for our purposes: the management of occlusions and the distinction between adult people (potential customers) and other objects (children, trolleys, strollers, animals, etc.). The proposed solution uses a novel approach for object detection (based on background suppression on a so-called “depth bird-eye view” and the clustering on the 3D point cloud by means of mean shift with a cylindrical kernel) followed by an adult people classifier which exploits a fitness measure with respect to a cylindrical human body model. The fitness is computed using Montecarlo sampling to estimate the volume occupation. Experiments are conducted on two real setups (including a store in a normal day of activity) and compared with a previous work. The results demonstrate the accuracy of the proposed solution.
Carlo Pane, Marco Gasparini, Andrea Prati 0001, Giovanni Gualdi, Rita Cucchiara
AVSS5
2013 Structured learning for detection of social groups in crowd
abstract
Group detection in crowds will play a key role in future behavior analysis surveillance systems. In this work we build a new Structural SVM-based learning framework able to solve the group detection task by exploiting annotated video data to deduce a sociologically motivated distance measure founded on Hall's proxemics and Granger's causality. We improve over state-of-the-art results even in the most crowded test scenarios, while keeping the classification time affordable for quasi-real time applications. A new scoring scheme specifically designed for the group detection task is also proposed.
Francesco Solera, Simone Calderara, Rita Cucchiara
AVSS3
2013 Learning articulated body models for people re-identification
abstract
People re-identification is a challenging problem in surveillance and forensics and it aims at associating multiple instances of the same person which have been acquired from different points of view and after a temporal gap. Image-based appearance features are usually adopted but, in addition to their intrinsically low discriminability, they are subject to perspective and view-point issues. We propose to completely change the approach by mapping local descriptors extracted from RGB-D sensors on a 3D body model for creating a view-independent signature. An original bone-wise color descriptor is generated and reduced with PCA to compute the person signature. The virtual bone set used to map appearance features is learned using a recursive splitting approach. Finally, people matching for re-identification is performed using the Relaxed Pairwise Metric Learning, which simultaneously provides feature reduction and weighting. Experiments on a specific dataset created with the Microsoft Kinect sensor and the OpenNi libraries prove the advantages of the proposed technique with respect to state of the art methods based on 2D or non-articulated 3D body models.
Davide Baltieri, Roberto Vezzani, Rita Cucchiara
ACM Multimedia3
2013 Modeling local descriptors with multivariate gaussians for object and scene recognition
abstract
Common techniques represent images by quantizing local descriptors and summarizing their distribution in a histogram. In this paper we propose to employ a parametric description and compare its capabilities to histogram based approaches. We use the multivariate Gaussian distribution, applied over the SIFT descriptors, extracted with dense sampling on a spatial pyramid. Every distribution is converted to a high-dimensional descriptor, by concatenating the mean vector and the projection of the covariance matrix on the Euclidean space tangent to the Riemannian manifold. Experiments on Caltech-101 and ImageCLEF2011 are performed using the Stochastic Gradient Descent solver, which allows to deal with large scale datasets and high dimensional feature spaces.
Giuseppe Serra 0001, Costantino Grana, Marco Manfredi, Rita Cucchiara
ACM Multimedia4
2013 Video surveillance online repository (ViSOR): www.openvisor.org
abstract
This paper describe the ViSOR (Video Surveillance Online Repository) repository, designed with the aim of establishing an open platform for collecting, annotating, retrieving, and sharing surveillance videos, as well as evaluating the performance of automatic surveillance systems. The repository is free and researchers can collaborate sharing their own videos or datasets. Most of the included videos are annotated. Annotations are based on a reference ontology which has been defined integrating hundreds of concepts, some of them coming from the LSCOM and MediaMill ontologies. A new annotation classification schema is also provided, which is aimed at identifying the spatial, temporal and domain detail level used. The web interface allows video browsing, querying by annotated concepts or by keywords, compressed video previewing, media downloading and uploading. Finally, ViSOR includes a performance evaluation desk which can be used to compare different annotations.
Roberto Vezzani, Rita Cucchiara
MMSys2
2013 Beyond Bag of Words for Concept Detection and Search of Cultural Heritage Archives
Costantino Grana, Giuseppe Serra 0001, Marco Manfredi, Rita Cucchiara
SISAP4
2012 People Orientation Recognition by Mixtures of Wrapped Distributions on Random Trees
Davide Baltieri, Roberto Vezzani, Rita Cucchiara
ECCV (5)3
2012 Class-Based Color Bag of Words for Fashion Retrieval
abstract
Color signatures, histograms and bag of colors are basic and effective strategies for describing the color content of images, for retrieving images by their color appearance or providing color annotation. In some domains, colors assume a specific meaning for users and the color-based classification and retrieval should mirror the initial suggestions given by users in the training set. For instance in fashion world, the names given to the dominant color of a garment or a dress reflect the fashion dictact and not an uniform division of the color space. In this paper we propose a general approach to implement color signature as a trained bag of words, defined on the basis of user defined color classes. The novel Class-based Color Bag of Words is a easy computable bag of words of color, constructed following an approach similar to the Median Cut algorithm, but biased by color distribution in the trained classes. Moreover, to dramatically reduce the computational effort we propose 3D integral histograms, a 3D extension of integral images, easily extensible for many histogram-based signature in 3D color space. Several comparisons in large fashion datasets confirm the discriminant power of this signature.
Costantino Grana, Daniele Borghesani, Rita Cucchiara
ICME3
2012 2D images map warping for improved user interaction
Daniele Borghesani, Costantino Grana, Rita Cucchiara
ICPR3
2012 Learning non-target items for interesting clothes segmentation in fashion images
Costantino Grana, Simone Calderara, Daniele Borghesani, Rita Cucchiara
ICPR4
2012 Veiling Luminance estimation on FPGA-based embedded smart camera
abstract
This paper describes the design and development of a Veiling Luminance estimation system based on the use of a CMOS image sensor, fully implemented on FPGA. The system is composed of the CMOS Image sensor, FPGA, DDR SDRAM, USB controller and SPI (Serial Peripheral Interface) Flash. The FPGA is used to build a system-on-chip integrating a soft processor (Xilinx MicroBlaze) and all the hardware blocks needed to handle the external peripherals and memory. The soft processor is used to handle image acquisition and all computational tasks need to compute the Veiling Luminance value. The advantages of this single chip FPGA implementation include the reduction of the hardware requirements, power consumption, and system complexity. The problem of the high dynamic range images have been addressed with multiple acquisitions at different exposure times. Vignetting, radial distortion and angular weighting, as required by veiling luminance definition, are handled by a single integer look-up table (LUT) access. Results are compared with a state of the art certified instrument.
Costantino Grana, Daniele Borghesani, Paolo Santinelli, Rita Cucchiara
Intelligent Vehicles Symposium4
2012 Real-time object detection and localization with SIFT-based clustering
Paolo Piccinini, Andrea Prati 0001, Rita Cucchiara
Image Vis. Comput.3
2012 Multistage Particle Windows for Fast and Accurate Object Detection
abstract
The common paradigm employed for object detection is the sliding window (SW) search. This approach generates grid-distributed patches, at all possible positions and sizes, which are evaluated by a binary classifier: The tradeoff between computational burden and detection accuracy is the real critical point of sliding windows; several methods have been proposed to speed up the search such as adding complementary features. We propose a paradigm that differs from any previous approach since it casts object detection into a statistical-based search using a Monte Carlo sampling for estimating the likelihood density function with Gaussian kernels. The estimation relies on a multistage strategy where the proposal distribution is progressively refined by taking into account the feedback of the classifiers. The method can be easily plugged into a Bayesian-recursive framework to exploit the temporal coherency of the target objects in videos. Several tests on pedestrian and face detection, both on images and videos, with different types of classifiers (cascade of boosted classifiers, soft cascades, and SVM) and features (covariance matrices, Haar-like features, integral channel features, and histogram of oriented gradients) demonstrate that the proposed method provides higher detection rates and accuracy as well as a lower computational burden w.r.t. sliding window detection.
Giovanni Gualdi, Andrea Prati 0001, Rita Cucchiara
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 Feature Space Warping Relevance Feedback with Transductive Learning
Daniele Borghesani, Dalia Coppi, Costantino Grana, Simone Calderara, Rita Cucchiara
ACIVS5
2011 Appearance tracking by transduction in surveillance scenarios
abstract
We propose a formulation of people tracking problem as a Transductive Learning (TL) problem. TL is an effective semi-supervised learning technique by which many classification problems have been recently reinterpreted as learning labels from incomplete datasets. In our proposal the joint exploitation of spectral graph theory and Riemannian manifold learning tools leads to the formulation of a robust approach for appearance based tracking in Video Surveillance scenarios. The key advantage of the presented method is a continuously updated model of the tracked target, used in the TL process, that allows to on-line learn the target visual appearance and consequently to improve the tracker accuracy. Experiments on public datasets show an encouraging advancement over alternative state-of the-art techniques.
Dalia Coppi, Simone Calderara, Rita Cucchiara
AVSS3
2011 A multi-stage pedestrian detection using monolithic classifiers
abstract
Despite the many efforts in finding effective feature sets or accurate classifiers for people detection, few works have addressed ways for reducing the computational burden introduced by the sliding window paradigm. This paper proposes a multi-stage procedure for refining the search for pedestrians using the HOG features and the monolithic SVM classifier. The multi-stage procedure is based on particle-based estimation of pdfs and exploits the margin provided by the classifier to draw more particles on the areas where the classifier's response is higher. This iterative algorithm achieves the same accuracy than sliding window using less particles (and thus being more efficient) and, conversely, is more accurate when configured to work at the same computational load. Experimental results on publicly available datasets demonstrate that this method, previously proposed for boosted classifiers only, can be successfully applied to monolithic classifiers.
Giovanni Gualdi, Andrea Prati 0001, Rita Cucchiara
AVSS3
2011 Relevance feedback strategies for artistic image collections tagging
abstract
This paper provides an analysis on relevance feedback techniques in a multimedia system designed for the interactive exploration and annotation of artistic collections, in particular illuminated manuscripts. The relevance feedback is presented not only as a very effective technique to improve the performance of the system, but also as a clever way to increase the user experience, mixing the interactive surfing through the artistic content with the possibility to gather valuable information from the user, and consequently improving his retrieval satisfaction. We compare a modification of the Mean-Shift Feature Space Warping algorithm, as representative of the standard RF procedures, and a learning-based technique based on transduction, considered in order to overcome some limitation of the previous technique. Experiments are reported regarding the adopted visual features based on covariance matrices.
Costantino Grana, Daniele Borghesani, Rita Cucchiara
ICMR3
2011 Joint ACM workshop on human gesture and behavior understanding: (J-HGBU'11)
abstract
The ability to understand social signals of a person we are communicating with is the core of social intelligence. Social Intelligence is a facet of human intelligence that has been argued to be indispensable and perhaps the most important for success in life. At the same time, human-centric multimedia applications for humans and about humans are becoming increasingly important. 3D modeled human-objects, like bodies, heads and faces are exploited for animation, security, and human computer interaction, while three dimensional motion of arms, legs and local body features is used for more complete human gesture, activity and behavior analysis. The Joint Human Gesture and Behavior Understanding (J-HGBU) workshop event consists of two parts focusing on these complementary challenges: the Workshop on Multimedia Access to 3D Human Objects (MA3HO'11) and the Workshop on Social Signal Processing (SSPW'11).
Maja Pantic, Alex Pentland, Alessandro Vinciarelli, Rita Cucchiara, Mohamed Daoudi, Alberto Del Bimbo
ACM Multimedia4
2011 Detecting anomalies in people's trajectories using spectral graph analysis
Simone Calderara, Uri Heinemann, Andrea Prati 0001, Rita Cucchiara, Naftali Tishby
Comput. Vis. Image Underst.4
2011 Automatic segmentation of digitalized historical manuscripts
Costantino Grana, Daniele Borghesani, Rita Cucchiara
Multim. Tools Appl.3
2011 Vision based smoke detection system using image energy and color information
Simone Calderara, Paolo Piccinini, Rita Cucchiara
Mach. Vis. Appl.3
2011 Probabilistic people tracking with appearance models and occlusion classification: The AD-HOC system
Roberto Vezzani, Costantino Grana, Rita Cucchiara
Pattern Recognit. Lett.3
2011 Mixtures of von Mises Distributions for People Trajectory Shape Analysis
abstract
People trajectory analysis is a recurrent task in many pattern recognition applications, such as surveillance, behavior analysis, video annotation, and many others. In this paper, we propose a new framework for analyzing trajectory shape, invariant to spatial shifts of the people motion in the scene. In order to cope with the noise and the uncertainty of the trajectory samples, we propose to describe the trajectories as a sequence of angles modeled by distributions of circular statistics, i.e., a mixture of von Mises (MovM) distributions. To deal with MovM, we define a new specific expectation-maximization (EM) algorithm for estimating the parameters and derive a closed form of the Bhattacharyya distance between single von Mises pdfs. Trajectories are then modeled with a sequence of symbols, corresponding to the most suitable distribution in the mixture, and compared each other after a global alignment procedure to cope with trajectories of different lengths. The trajectories in the training set are clustered according to their shape similarity in an off-line phase, and testing trajectories are then classified with a specific on-line EM, based on sufficient statistics. The approach is particularly suitable for classifying people trajectories in video surveillance, searching for abnormal (i.e., infrequent) paths. Tests on synthetic and real data are provided with also a complete comparison with other circular statistical and alignment methods.
Simone Calderara, Andrea Prati 0001, Rita Cucchiara
IEEE Trans. Circuits Syst. Video Technol.3
2010 Fast Background Initialization with Recursive Hadamard Transform
abstract
In this paper, we present a new and fast technique for background estimation from cluttered image sequences. Most of the background initialization approaches developed so far collect a number of initial frames and then require a slow estimation step which introduces a delay whenever it is applied. Conversely, the proposed technique redistributes the computational load among all the frames by means of a patch by patch preprocessing, which makes the overall algorithm more suitable for real-time applications. For each patch location a prototype set is created and maintained. The background is then iteratively estimated by choosing from each set the most appropriate candidate patch, which should verify a sort of frequency coherence with its neighbors. To this aim, the Hadamard transform has been adopted which requires less computation time than the commonly used DCT Finally, a refinement step exploits spatial continuity constraints along the patch borders to prevent erroneous patch selections. The approach has been compared with the state of the art on videos from available datasets (ViSOR and CAVIAR), showing a speed up of about 10 times and an improved accuracy.
Davide Baltieri, Roberto Vezzani, Rita Cucchiara
AVSS3
2010 Multi-stage Sampling with Boosting Cascades for Pedestrian Detection in Images and Videos
Giovanni Gualdi, Andrea Prati 0001, Rita Cucchiara
ECCV (6)3
2010 Alignment-Based Similarity of People Trajectories Using Semi-directional Statistics
abstract
This paper presents a method for comparing people trajectories for video surveillance applications, based on semi-directional statistics. In fact, the modelling of a trajectory as a sequence of angles, speeds and time lags, requires the use of a statistical tool capable to jointly consider periodic and linear variables. Our statistical method is compared with two state-of-the-art methods.
Simone Calderara, Andrea Prati 0001, Rita Cucchiara
ICPR3
2010 Decision Trees for Fast Thinning Algorithms
abstract
We propose a new efficient approach for neighborhood exploration, optimized with decision tables and decision trees, suitable for local algorithms in image processing. In this work, it is employed to speed up two widely used thinning techniques. The performance gain is shown over a large freely available dataset of scanned document images.
Costantino Grana, Daniele Borghesani, Rita Cucchiara
ICPR3
2010 Surfing on artistic documents with visually assisted tagging
abstract
This paper describes a complete architecture for the interactive exploration and annotation of artistic collections. In particular the focus is on Renaissance illuminated manuscripts, which typically contain thousands of pictures, used to comment or embellish the manuscript Gothic text. The final aim is to create a human centered multimedia application allowing the non practitioners to enjoy these masterpieces and expert users to share their knowledge. The system is composed by a modern user interface for browsing, surfing and querying, an automatic segmentation module, to ease the initial picture extraction task, and a similarity based retrieval engine, used to provide visually assisted tagging capabilities. A relevance feedback procedure is included to further refine the results. Experiments are reported regarding the adopted visual features based on covariance matrices and the Mean Shift Feature Space Warping relevance feedback. Finally some hints on the user interface for museum installations are discussed.
Daniele Borghesani, Costantino Grana, Rita Cucchiara
ACM Multimedia3
2010 Rerum novarum: interactive exploration of illuminated manuscripts
abstract
This paper describes an interactive application for the exploration and annotation of illuminated manuscripts, which typically contain thousands of pictures, used to comment or embellish the manuscript Gothic text. The system is composed by a modern user interface for browsing, surfing and querying, an automatic segmentation module, to ease the initial picture extraction task, and a similarity based retrieval engine, used to provide visually assisted tagging capabilities. A relevance feedback procedure is included to further refine the results.
Daniele Borghesani, Costantino Grana, Rita Cucchiara
ACM Multimedia3
2010 Video Surveillance Online Repository (ViSOR): an integrated framework
Roberto Vezzani, Rita Cucchiara
Multim. Tools Appl.2
2010 Optimized Block-Based Connected Components Labeling With Decision Trees
abstract
In this paper, we define a new paradigm for eight-connection labeling, which employs a general approach to improve neighborhood exploration and minimizes the number of memory accesses. First, we exploit and extend the decision table formalism introducing OR-decision tables, in which multiple alternative actions are managed. An automatic procedure to synthesize the optimal decision tree from the decision table is used, providing the most effective conditions evaluation order. Second, we propose a new scanning technique that moves on a 2 x 2 pixel grid over the image, which is optimized by the automatically generated decision tree. An extensive comparison with the state of art approaches is proposed, both on synthetic and real datasets. The synthetic dataset is composed of different sizes and densities random images, while the real datasets are an artistic image analysis dataset, a document analysis dataset for text detection and recognition, and finally a standard resolution dataset for picture segmentation tasks. The algorithm provides an impressive speedup over the state of the art algorithms.
Costantino Grana, Daniele Borghesani, Rita Cucchiara
IEEE Trans. Image Process.3
2009 Learning People Trajectories Using Semi-directional Statistics
abstract
This paper proposes a system for people trajectory shape analysis by exploiting a statistical approach which accounts for sequences of both directional (the directions of the trajectory) and linear (the speeds) data. A semi-directional distribution (AWLG - Approximated Wrapped and Linear Gaussian) is used with a mixture to find main directions and speeds. A variational version of the mutual information criterion is proposed to prove the statistical dependency of the data. Then, in order to compare data sequences, we define an inexact method with a Kullback-Leibler-based distance measure and employ a global alignment technique is to handle sequences of different lengths and with local shifts or deformations. A comprehensive analysis of variable dependency and parameter estimation techniques are reported and evaluated on both synthetic and real data sets.
Simone Calderara, Andrea Prati 0001, Rita Cucchiara
AVSS3
2009 Fast block based connected components labeling
abstract
In this paper we present a new optimization technique for the neighborhood computation in connected component labeling focused on images stored in raster scan order. This new technique is based on a 2 × 2 square block analysis of the image, and it exploits the fact that, when using 8-connection, the pixels of a 2 × 2 square are all connected to each other. This implies that they will share the same label at the end of the computation. To prove the effectiveness of our proposal, we show a comprehensive comparison of the most used and advanced connected components labeling techniques presented so far. The tests are conducted on high resolution images obtained from digitized historical manuscripts and a set of transformations is applied in order to show the algorithms behavior at different image resolutions and with a varying number of labels.
Costantino Grana, Daniele Borghesani, Rita Cucchiara
ICIP3
2009 An efficient Bayesian framework for on-line action recognition
abstract
On-line action recognition from a continuous stream of actions is still an open problem with fewer solutions proposed compared to time-segmented action recognition. The most challenging task is to classify the current action while finding its time boundaries at the same time. In this paper we propose an approach capable of performing on-line action segmentation and recognition by means of batteries of HMM taking into account all the possible time boundaries and action classes. A suitable Bayesian normalization is applied to make observation sequences of different length comparable and computational optimizations are introduce to achieve real-time performances. Results on a well known action dataset prove the efficacy of the proposed method.
Roberto Vezzani, Massimo Piccardi, Rita Cucchiara
ICIP3
2009 Multimedia in forensics
abstract
No abstract available.
Marcel Worring, Rita Cucchiara
ACM Multimedia2
2009 A fast multi-model approach for object duplicate extraction
abstract
This paper presents an innovative approach for localizing and segmenting duplicate objects for industrial applications. The working conditions are challenging, with complex heavily-occluded objects, arranged at random in the scene. To account for high flexibility and processing speed, this approach exploits SIFT keypoint extraction and mean shift clustering to efficiently partition the correspondences between the object model and the duplicates onto the different object instances. The re-projection (by means of an Euclidean transform) of some delimiting points onto the current image is used to segment the object shapes. This procedure is compared in terms of accuracy with existing homography-based solutions which make use of RANSAC to eliminate outliers in the homography estimation. Moreover, in order to improve the extraction in the case of reflective or transparent objects, multiple object models are used and fused together. Experimental results on different and challenging kinds of objects are reported.
Paolo Piccinini, Andrea Prati 0001, Rita Cucchiara
WACV3
2008 Action Signature: A Novel Holistic Representation for Action Recognition
abstract
Recognizing different actions with a unique approach can be a difficult task. This paper proposes a novel holistic representation of actions that we called "action signature". This 1D trajectory is obtained by parsing the 2D image containing the orientations of the gradient calculated on the motion feature map called motion-history image. In this way, the trajectory is a sketch representation of how the object motion varies in time. A robust statistical framework based on mixtures of von Mises distributions and dynamic programming for sequence alignment are used to compare and classify actions/trajectories. The experimental results show a rather high accuracy in distinguishing quite complicated actions, such as drinking, jumping, or abandoning an object.
Simone Calderara, Rita Cucchiara, Andrea Prati 0001
AVSS2
2008 Annotation Collection and Online Performance Evaluation for Video Surveillance: The ViSOR Project
abstract
This paper presents the Visor (video surveillance online repository) project designed with the aim of establishing an open platform for collecting, annotating, retrieving, sharing surveillance videos, and of evaluating the performance of automatic surveillance systems. The main idea is to exploit the collaborative paradigm spreading in the web community to join together the ontology based annotation and retrieval concepts and the requirements of the computer vision and video surveillance communities. The ViSOR open repository is based on a reference ontology which integrates many concepts, also coming from LSCOM and MediaMill ontologies.The web interface allows video browse, query by annotated concepts or by keywords, compressed video preview, media download and upload. The repository contains metadata annotations, which can be either manually created as ground truth or automatically generated by video surveillance systems. Their automatic annotations can be compared each other or with the reference ground-truth exploiting an integrated on-line performance evaluator.
Roberto Vezzani, Rita Cucchiara
AVSS2
2008 Using circular statistics for trajectory shape analysis
abstract
The analysis of patterns of movement is a crucial task for several surveillance applications, for instance to classify normal or abnormal people trajectories on the basis of their occurrence. This paper proposes to model the shape of a single trajectory as a sequence of angles described using a mixture of Von Mises (MoVM) distribution. A complete EM (expectation maximization) algorithm is derived for MoVM parameters estimation and an on-line version proposed to meet real time requirement. Maximum-A-Posteriori is used to encode the trajectory as a sequence of symbols corresponding to the MoVM components. Iterative k-medoids clustering groups trajectories in a variable number of similarity classes. The similarity is computed aligning (with dynamic programming) two sequences and considering as symbol-to-symbol distance the Bhattacharyya distance between von Mises distributions. Extensive experiments have been performed on both synthetic and real data.
Andrea Prati 0001, Simone Calderara, Rita Cucchiara
CVPR3
2008 Reliable smoke detection in the domains of image energy and color
abstract
Smoke detection calls for a reliable and fast distinction between background, moving objects and variable shapes that are recognizable as smoke. In our system we propose a stable background suppression module joined with a smoke detection module working on segmented objects. It exploits two features: the energy variation in wavelet model and a color model of the smoke. The decrease of energy ratio in wavelet domain between background and current image is a clue to detect smoke representing the variations of texture level. A mixture of Gaussians models this texture ratio for temporal evolution. The color model is used as reference to measure the deviation of the current pixel color from the model. The two features have been combined using a Bayesian classifier to detect smoke in the scene. Experiments on real data and a comparison between our background model and Gaussian mixture (MOG) model for smoke detection are presented.
Paolo Piccinini, Simone Calderara, Rita Cucchiara
ICIP3
2008 ViSOR: VIdeo Surveillance On-line Repository for annotation retrieval
abstract
The Imagelab Laboratory of the University of Modena and Reggio Emilia has designed a large video repository, aiming at containing annotated video surveillance footages. The Web interface, named ViSOR (video surveillance online repository), allows video browse, query by annotated concepts or by keywords, compressed preview, video download and upload. The repository contains metadata annotation, both manually annotated ground-truth data and automatically obtained outputs of a particular system. In such a manner, the users of the repository are able to perform validation tasks of their own algorithms as well as comparative activities.
Roberto Vezzani, Rita Cucchiara
ICME2
2008 Describing texture directions with Von Mises distributions
abstract
In this work we describe a new approach for texture characterization. Starting from the autocorrelation matrix an elegant description through a mixture of Von Mises distributions is proposed. A compact 6 valued descriptor is produced for each block and served as input to an SVM classifier. Tests are carried out on high resolution illuminated manuscripts images.
Costantino Grana, Daniele Borghesani, Rita Cucchiara
ICPR3
2008 Smoke Detection in Video Surveillance: A MoG Model in the Wavelet Domain
Simone Calderara, Paolo Piccinini, Rita Cucchiara
ICVS3
2008 HECOL: Homography and epipolar-based consistent labeling for outdoor park surveillance
Simone Calderara, Andrea Prati 0001, Rita Cucchiara
Comput. Vis. Image Underst.3
2008 Bayesian-Competitive Consistent Labeling for People Surveillance
abstract
This paper presents a novel and robust approach to consistent labeling for people surveillance in multi-camera systems. A general framework scalable to any number of cameras with overlapped views is devised. An off-line training process automatically computes ground-plane homography and recovers epipolar geometry. When a new object is detected in any one camera, hypotheses for potential matching objects in the other cameras are established. Each of the hypotheses is evaluated using a prior and likelihood value. The prior accounts for the positions of the potential matching objects, while the likelihood is computed by warping the vertical axis of the new object on the field of view of the other cameras and measuring the amount of match. In the likelihood, two contributions (forward and backward) are considered so as to correctly handle the case of groups of people merged into single objects. Eventually, a maximum-a-posteriori approach estimates the best label assignment for the new object. Comparisons with other methods based on homography and extensive outdoor experiments demonstrate that the proposed approach is accurate and robust in coping with segmentation errors and in disambiguating groups.
Simone Calderara, Rita Cucchiara, Andrea Prati 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Video Streaming for Mobile Video Surveillance
abstract
Mobile video surveillance represents a new paradigm that encompasses, on the one side, ubiquitous video acquisition and, on the other side, ubiquitous video processing and viewing, addressing both computer-based and human-based surveillance. To this aim, systems must provide efficient video streaming with low latency and low frame skipping, even over limited bandwidth networks. This work presents MoSES (MObile Streaming for vidEo Surveillance), an effective system for mobile video surveillance for both PC and PDA clients; it relies over H.264/AVC video coding and GPRS/EDGE-GPRS network. Adaptive control algorithms are employed to achieve the best tradeoff between low latency and good video fluidity. MoSES provides a good-quality video streaming that is used as input to computer-based video surveillance applications for people segmentation and tracking. In this paper new and general-purpose methodologies for streaming performance evaluation are also proposed and used to compare MoSES with existing solutions in terms of different parameters (latency, image quality, video fluidity, and frame losses), as well as in terms of performance in people segmentation and tracking.
Giovanni Gualdi, Andrea Prati 0001, Rita Cucchiara
IEEE Trans. Multim.3
2007 Detection of abnormal behaviors using a mixture of Von Mises distributions
abstract
This paper proposes the use of a mixture of Von Mises distributions to detect abnormal behaviors of moving people. The mixture is created from an unsupervised training set by exploiting k-medoids clustering algorithm based on Bhattacharyya distance between distributions. The extracted medoids are used as modes in the multi-modal mixture whose weights are the priors of the specific medoid. Given the mixture model a new trajectory is verified on the model by considering each direction composing it as independent. Experiments over a real scenario composed of multiple, partially-overlapped cameras are reported.
Simone Calderara, Rita Cucchiara, Andrea Prati 0001
AVSS2
2007 An Open Source Architecture for Low-Latency Video Streaming on PDAs
abstract
This paper presents a open-source system for low- latency video streaming on PDAs, specifically addressing mobile video surveillance requirements. The system is based on H.264 and suitably modified to obtain the best trade-off between image quality and video fluidity, working also at very limited bandwidths. Moreover, the used controls allow to keep the number of lost frames very low. A large set of experiments and comparisons have been carried out and the achieved results demonstrate the efficacy and efficiency of our system.
Giovanni Gualdi, Andrea Prati 0001, Rita Cucchiara
ISM3
2007 A multi-camera vision system for fall detection and alarm generation
abstract
Abstract: In‐house video surveillance can represent an excellent support for people with some difficulties (e.g. elderly or disabled people) living alone and with a limited autonomy. New hardware technologies and in particular digital cameras are now affordable and they have recently gained credit as tools for (semi‐)automatically assuring people's safety. In this paper a multi‐camera vision system for detecting and tracking people and recognizing dangerous behaviours and events such as a fall is presented. In such a situation a suitable alarm can be sent, e.g. by means of an SMS. A novel technique of warping people's silhouette is proposed to exchange visual information between partially overlapped cameras whenever a camera handover occurs. Finally, a multi‐client and multi‐threaded transcoding video server delivers live video streams to operators/remote users in order to check the validity of a received alarm. Semantic and event‐based transcoding algorithms are used to optimize the bandwidth usage. A two‐room setup has been created in our laboratory to test the performance of the overall system and some of the results obtained are reported.
Rita Cucchiara, Andrea Prati 0001, Roberto Vezzani
Expert Syst. J. Knowl. Eng.1
2007 Expert environments: machine intelligence methods for ambient intelligence
abstract
The ideas put forward by Donald A. Norman (1999) in his monograph entitled The Invisible Computer can be considered the main source of inspiration of a new research area, called ambient intelligence. Ambient intelligence, commonly abbreviated as AmI, is primarily concerned with human–environment interactions. An environment is seen anthropomorphically, as an intelligent agent able to interact with users, creating for them processes to interpret, inform, communicate and dialogue (Abowd & Mynatt, 2000; Remagnino & Foresti, 2005). The history of ambient intelligence starts in Europe in 2001 with the Fifth European Framework Program. At that time, the IST Program Advisory Group (ISTAG) of the European Commission (Directorate General on Information Society and the Media) introduced the concept of ambient intelligence by publishing the report Scenarios for Ambient Intelligence in 2010 (Ducatel et al., 2001). Since then, ambient intelligence has been recognized in Europe as one of the key concepts related to the information and communication technology society. An updated version of the report was published in 2003 under the title Ambient Intelligence: from Vision to Reality (Ducatel et al., 2003). Ambient intelligence's emphasis is on support to human interactions with the environment, user-friendliness, ubiquitous accessibility etc. and requires competences from many research areas, ranging from computer vision, machine learning, distributed computing and middleware, context awareness systems, sensor networks etc. Enticing illustrative scenarios have been published, in which the user wears technology that communicates with systems and devices present in the environment in order to provide information and receive services. Since its inception, ambient intelligence has inspired the design and implementation of system prototypes for intelligent spaces, developed and tested in controlled environments. The scope of this special issue is to publish innovative ideas on a selection of topics related to ambient intelligence. Cameras and computer vision algorithms, for instance, can be used to unobtrusively acquire awareness of the environment. In the paper entitled ‘A multi-camera vision system for fall detection and alarm generation’, Cucchiara et al. propose a system for using cameras to detect people's falls. The use of cameras is preferred in ambient intelligence to accelerometers or other devices because they are not intrusive, people tend to get used to their presence and people's actions are less affected. In this paper, a system of multiple cameras is used to monitor people's movements in house environments and detect changes in people's posture with the specific goal to identify falls. In the paper entitled ‘Understanding intention of movement from electroencephalograms’, Lakany and Conway analyse people's intentions, using electroencephalogram waves. Their method is non-intrusive and it uses a brain–computer interface. Support vector machines are used to select features and perform classification of intentions, and tests are performed to detect the direction of users' movements. While the previous papers address ambient intelligence from the point of view of sensing technologies and algorithms, the following two papers are mainly focused on knowledge representation for context awareness. The paper entitled ‘Knowledge representation for ambient security’ by Snidaro and Foresti proposes an ontology-based methodology for representing the interaction between the user and the environment with specific reference to security scenarios. Similarly, in the paper entitled ‘Context-aware environments: from specification to implementation’ Reigner et al. are concerned with the problem of implementing a context model for a smart environment. Their paper proposes interesting approaches based on ‘networks of situations’, introducing a comparison of the use of Petri nets and hidden Markov models. Finally, in the paper entitled ‘Collection, storage and application of human knowledge in expert system development’ Balch et al. propose an analysis of the knowledge engineering flow, encompassing knowledge acquisition, representation and inference. This flow analysis is presented and applied to the petroleum industry application domain, and specific software tools that use fuzzy logic are utilized.
Paolo Remagnino, Andrea Prati 0001, Gian Luca Foresti, Rita Cucchiara
Expert Syst. J. Knowl. Eng.4
2007 Linear Transition Detection as a Unified Shot Detection Approach
abstract
In this paper, we propose an automatic system for video shot segmentation, called linear transition detector, unique for both cuts and linear transitions detection. Comparison with publicly available shot detection systems is reported on different sports (Formula 1 racing, basketball, soccer, and cycling) and TRECVID 2005 results are also reported
Costantino Grana, Rita Cucchiara
IEEE Trans. Circuits Syst. Video Technol.2
2006 Group Detection at Camera Handoff for Collecting People Appearance in Multi-camera Systems
abstract
Logging information on moving objects is crucial in video surveillance systems. Distributed multi-camera systems can provide the appearance of objects/people from different viewpoints and at different resolutions, allowing a more complete and precise logging of the information. This is achieved through consistent labeling to correlate collected information of the same person. This paper proposes a novel approach to consistent labeling also capable to fully characterize groups of people and to manage miss segmentations. The ground-plane homography and the epipolar geometry are automatically learned and exploited to warp objects' principal axes between overlapped cameras. A MAP estimator that exploits two contributions (forward and backward) is used to choose the most probable label configuration to be assigned at the handoff of a new object. Extensive experiments demonstrate the accuracy of the proposed method in detecting single and simultaneous handoffs, miss segmentations, and groups.
Simone Calderara, Rita Cucchiara, Andrea Prati 0001
AVSS2
2006 3-D Virtual Environments on Mobile Devices for Remote Surveillance
abstract
In this paper we present a distributed videosurveillance framework. Our end is the remote monitoring of the behavior of people moving in a scene exploiting a virtual reconstruction on low capabilities devices, like PDAs and cell phones. The main novelty of this system is the effective integration of the computer vision and computer graphics modules. The first, using a probabilistic frameworks, can detect the position, the trajectory and the posture of peoples moving in the scene. The second exploits the new possibility of both standard 3D graphics libraries on mobile (namely JSR184 and M3G graphic format) and new PDAs processing capability in order to reconstruct the remote surveillance data in real-time.
Roberto Vezzani, Rita Cucchiara, Alessio Malizia, Luigi Cinque
AVSS2
2006 Low-Latency Live Video Streaming over Low-Capacity Networks
abstract
This paper presents an effective system for streaming over low-capacity networks (such as GPRS and EGPRS) of live videos with low latency. Existing solutions are either too complex or not suitable to our scope. For this reason, we developed a complete, ready-to-use streaming system based on H.264/AVC codec and UDP/IP stack. The system employs adaptive controls to achieve the best tradeoff between low latency and good video fluency, by keeping the UDP buffer occupancy at the decoder side between two given levels. Our experiments demonstrate that this system is able to transmit live videos at CIF format and 10 fps over GPRS/EGPRS with very low latency (1.73 sec on average, basically due to the network delay), good fluency and average quality, measured with PSNR, of 31 dB on GPRS at 23 kbps at 10 fps
Giovanni Gualdi, Rita Cucchiara, Andrea Prati 0001
ISM2
2006 A Semi-Automatic Video Annotation tool with MPEG-7 Content Collections
abstract
In this work, we present a general purpose system for hierarchical structural segmentation and automatic annotation of video clips, by means of standardized low level features. We propose to automatically extract some prototypes for each class with a context based intra-class clustering. Clips are annotated following the MPEG-7 standard directives to provide easier portability. Results of automatic annotation and semiautomatic metadata creation are provided
Roberto Vezzani, Costantino Grana, Daniele Bulgarelli, Rita Cucchiara
ISM4
2006 MOM: multimedia ontology manager. A framework for automatic annotation and semantic retrieval of video sequences
abstract
Effective usage of multimedia digital libraries has to deal with the problem of building efficient content annotation and retrieval tools. MOM (Multimedia Ontology Manager) is a complete system that allows the creation of multimedia ontologies, supports automatic annotation and creation of extended text (and audio) commentaries of video sequences, and permits complex queries by reasoning on the ontology.
Marco Bertini 0001, Alberto Del Bimbo, Carlo Torniai, Rita Cucchiara, Costantino Grana
ACM Multimedia4
2006 PEANO: pictorial enriched annotation of video
abstract
In this DEMO, we present a tool set for video digital library management that allows i) structural annotation of edited videos in MPEG-7 by automatically extracting shots and clips; ii) automatic semantic annotation based on perceptual similarity against a taxonomy enriched with pictorial concepts iii) video clip access and hierarchical summarization with stand-alone and web interface iv) access to clips from mobile platform in GPRS-UMTS video-streaming. The tools can be applied in different domain-specific Video Digital Libraries. The main novelty is the possibility to enrich the annotation with pictorial concepts that are added to a textual taxonomy in order to make the automatic annotation process more fast and often effective. The resulting multimedia ontology is described in the MPEG-7 framework. The PEANO (Perceptual Annotation of Video) tool has been tested over video art , sport (Soccer, Olimpic Games 2006, Formula 1) and news clips.
Costantino Grana, Roberto Vezzani, Daniele Bulgarelli, Giovanni Gualdi, Rita Cucchiara, Marco Bertini 0001, Carlo Torniai, Alberto Del Bimbo
ACM Multimedia5
2006 Special issue on multimedia Surveillance systems: guest editorial
Jake K. Aggarwal, Rita Cucchiara
Multim. Syst.2
2006 A semi-automatic system for segmentation of cardiac M-mode images
Luca Bertelli, Rita Cucchiara, Giovanni Paternostro, Andrea Prati 0001
Pattern Anal. Appl.2
2006 A system for automatic face obscuration for privacy purposes
Rita Cucchiara, Andrea Prati 0001, Roberto Vezzani
Pattern Recognit. Lett.1
2006 Semantic adaptation of sport videos with user-centred performance analysis
abstract
In semantic video adaptation measures of performance must consider the impact of the errors in the automatic annotation over the adaptation in relationship with the preferences and expectations of the user. In this paper, we define two new performance measures Viewing Quality Loss and Bit-rate Cost Increase,that are obtained from classical peak signal-to-noise ration (PSNR) and bitrate, and relate the results of semantic adaptation to the errors in the annotation of events and objects and the user's preferences and expectations. We present and discuss results obtained with a system that performs automatic annotation of soccer sport video highlights and applies different coding strategies to different parts of the video according to their relative importance for the end user. With reference to this framework, we analyze how highlights' statistics and the errors of the annotation engine influence the performance of semantic adaptation and reflect into the quality of the video displayed at the user's client and the increase of transmission costs.
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001
IEEE Trans. Multim.2
2005 Entry edge of field of view for multi-camera tracking in distributed video surveillance
abstract
Efficient solution to people tracking in distributed video surveillance is requested to monitor crowded and large environments. This paper proposes a novel use of the entry edges of field of view (E/sup 2/oFoV) to solve the consistent labeling problem between partially overlapped views. An automatic and reliable procedure allows obtaining the homographic transformation between two overlapped views, without any manual calibration of the cameras. Through the homography, the consistent labeling is established each time a new track is detected in one of the cameras. A camera transition graph (CTG) is defined to speed up the establishment process by reducing the search space. Experimental results prove the effectiveness of the proposed solution also in challenging conditions.
Simone Calderara, Roberto Vezzani, Andrea Prati 0001, Rita Cucchiara
AVSS4
2005 Domain Knowledge Extension with Pictorially Enriched Ontologies
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Carlo Torniai
CAIP2
2005 Posture classification in a multi-camera indoor environment
abstract
Posture classification is a key process for analyzing the people's behaviour. Computer vision techniques can be helpful in automating this process, but cluttered environments and consequent occlusions make this task often difficult. Different views provided by multiple cameras can be exploited to solve occlusions by warping known object appearance into the occluded view. To this aim, this paper describes an approach to posture classification based on projection histograms, reinforced by HMM for assuring temporal coherence of the posture. The single camera posture classification is then exploited in the multi-camera system to solve the cases in which the occlusions make the classification impossible. Experimental results of the classification from both the single camera and the multi-camera system are provided.
Rita Cucchiara, Andrea Prati 0001, Roberto Vezzani
ICIP (1)1
2005 Video Annotation with Pictorially Enriched Ontologies
abstract
Video annotation is typically performed by classifying video elements according to some pre-defined ontology of the video content domain. Ontologies are defined by establishing relationships between linguistic terms, that specify domain concepts at different abstraction levels. However, although linguistic terms are appropriate to distinguish event and object categories, they are inadequate when they must describe specific patterns of events or video entities. Instead, in these cases, pattern specifications are better expressed through visual prototypes that capture the essence of the event or entity. Pictorially enriched ontologies, that include visual concepts together with linguistic keywords, are therefore needed to support video annotation up to the level of detail of pattern specification. This paper presents pictorially enriched ontologies and provide a solution for their implementation in the soccer video domain. The pictorially enriched ontology is used both to directly assign multimedia objects to concepts, providing a more meaningful definition than the linguistics terms, and to extend the initial knowledge of the domain, adding subclasses of highlights or new highlight classes that were not defined in the linguistic ontology. Automatic annotation of soccer clips up to the pattern specification level using a pictorially enriched ontology is discussed.
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Carlo Torniai
ICME2
2005 MPEG-7 Compliant Shot Detection in Sport Videos
abstract
In this paper we propose a system for automatic detection of shots in sport videos. Our work covers two main aspects: the first is robust shot detection in presence of fast object motion and camera operations. To this aim we propose a new algorithm, unique for both cuts and linear transitions detection, which only needs the tuning of two parameters. An extended comparison with four transition detection algorithms, representing the state of the art in literature, is reported. Examples with formula 1, basket, soccer and cycling videos are analyzed. The second aspect is an in depth discussion on the annotation of shots and transitions with the MPEG-7 standard.
Costantino Grana, Giovanni Tardini, Rita Cucchiara
ISM3
2005 On the usefulness of object shape coding with MPEG-4
abstract
This paper reports the results of an in-depth analysis of the degree of usefulness of object shape coding in video compression. In particular, MPEG-4 is used as reference standard. The influence of different coding parameters on the performance is deeply examined and discussions on the results are provided. Object shape coding is compared with classical (MPEG-2) frame-based coding both at an objective level (by comparing PSNR/quality and bitrate/filesize) and at a subjective level (asking to a set of users to express their opinion on overall quality, cognitive effectiveness, and willingness to pay). In conclusion, this paper aims at answering to the question whether it is convenient to use object shape coding instead of frame-based coding or not.
Andrea Prati 0001, Rita Cucchiara
ISM2
2005 An Integrated Framework for Semantic Annotation and Adaptation
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001
Multim. Tools Appl.2
2005 Guest Editorial: Special Issue on Video Segmentation for Semantic Annotation and Transcoding
Rita Cucchiara, Alberto Del Bimbo
Multim. Tools Appl.1
2005 Probabilistic posture classification for Human-behavior analysis
abstract
Computer vision and ubiquitous multimedia access nowadays make feasible the development of a mostly automated system for human-behavior analysis. In this context, our proposal is to analyze human behaviors by classifying the posture of the monitored person and, consequently, detecting corresponding events and alarm situations, like a fall. To this aim, our approach can be divided in two phases: for each frame, the projection histograms (Haritaoglu et al., 1998) of each person are computed and compared with the probabilistic projection maps stored for each posture during the training phase; then, the obtained posture is further validated exploiting the information extracted by a tracking module in order to take into account the reliability of the classification of the first phase. Moreover, the tracking algorithm is used to handle occlusions, making the system particularly robust even in indoors environments. Extensive experimental results demonstrate a promising average accuracy of more than 95% in correctly classifying human postures, even in the case of challenging conditions.
Rita Cucchiara, Costantino Grana, Andrea Prati 0001, Roberto Vezzani
IEEE Trans. Syst. Man Cybern. Part A1
2004 Content-based video adaptation with user's preferences
abstract
We present an integrated system that has been designed to support automatic semantic extraction of highlights in sports video and automatic video adaptation according to user's preferences. To analyze the user's satisfaction, we propose a new performance measure that explicitly takes into account the user's preferences and considers the number and type of errors produced by the annotation engine and the way in which these errors affect the compressed video quality and bandwidth allocation. We provide experimental results with application to soccer and swimming.
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001
ICME2
2004 Neighbor cache prefetching for multimedia image and video processing
abstract
Cache performance is strongly influenced by the type of locality embodied in programs. In particular, multimedia programs handling images and videos are characterized by a bidimensional spatial locality, which is not adequately exploited by standard caches. In this paper we propose novel cache prefetching techniques for image data, called neighbor prefetching, able to improve exploitation of bidimensional spatial locality. A performance comparison is provided against other assessed prefetching techniques on a multimedia workload (with MPEG-2 and MPEG-4 decoding, image processing, and visual object segmentation), including a detailed evaluation of both the miss rate and the memory access time. Results prove that neighbor prefetching achieves a significant reduction in the time due to delayed memory cycles (more than 97% on MPEG-4 with respect to 75% of the second performing technique). This reduction leads to a substantial speedup on the overall memory access time (up to 140% for MPEG-4). Performance has been measured with the PRIMA trace-driven simulator, specifically devised to support cache prefetching.
Rita Cucchiara, Massimo Piccardi, Andrea Prati 0001
IEEE Trans. Multim.1
2003 Object Segmentation in Videos from Moving Camera with MRFs on Color and Motion Features
abstract
In this paper we address the problem of fast segmenting moving objects in video acquired by moving camera or more generally with a moving background. We present an approach based on a color segmentation followed by a region-merging on motion through Markov random fields (MRFs). The technique we propose is inspired by the work of Gelgon and Bouthemy (2000), that has been modified to reduce computational cost in order to achieve a fast segmentation (about ten frame per second). To this aim a modified region matching algorithm (namely partitioned region matching) and an innovative arc-based MRF optimization algorithm with a suitable definition of the motion reliability are proposed. Results on both synthetic and real sequences are reported to confirm validity of our solution.
Rita Cucchiara, Andrea Prati 0001, Roberto Vezzani
CVPR (1)1
2003 Object and event detection for semantic annotation and transcoding
abstract
Video annotation provides a suitable way to describe, organize, and index stored videos. On the other hand, transcoding aims at adapting content to the user/client capabilities and requirements. Both cues are now mandatory, given the tremendous demand of multimedia access from remote clients, in particular nowadays that new terminals with limited resources (PDAs, HCCs, Smart phones) have access to the network. In this paper we propose a unified framework to define event-based and object-based semantic extraction from video to provide both semantic video annotation for video stored and semantic on-line transcoding from live cameras. Two case studies (highlights' extraction from soccer videos for the annotation and people behavior detection in domotic application for transcoding) and corresponding experimental results are reported.
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001
ICME2
2003 Improving Data Prefetching Efficacy in Multimedia Applications
Rita Cucchiara, Andrea Prati 0001, Massimo Piccardi
Multim. Tools Appl.1
2003 Detecting Moving Objects, Ghosts, and Shadows in Video Streams
abstract
Background subtraction methods are widely exploited for moving object detection in videos in many applications, such as traffic monitoring, human motion capture, and video surveillance. How to correctly and efficiently model and update the background model and how to deal with shadows are two of the most distinguishing and challenging aspects of such approaches. The article proposes a general-purpose method that combines statistical assumptions with the object-level knowledge of moving objects, apparent objects (ghosts), and shadows acquired in the processing of the previous frames. Pixels belonging to moving objects, ghosts, and shadows are processed differently in order to supply an object-based selective update. The proposed approach exploits color information for both background subtraction and shadow detection to improve object segmentation and background update. The approach proves fast, flexible, and precise in terms of both pixel accuracy and reactivity to background changes.
Rita Cucchiara, Costantino Grana, Massimo Piccardi, Andrea Prati 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Detecting Moving Shadows: Algorithms and Evaluation
abstract
Moving shadows need careful consideration in the development of robust dynamic scene analysis systems. Moving shadow detection is critical for accurate object detection in video streams since shadow points are often misclassified as object points, causing errors in segmentation and tracking. Many algorithms have been proposed in the literature that deal with shadows. However, a comparative evaluation of the existing approaches is still lacking. In this paper, we present a comprehensive survey of moving shadow detection approaches. We organize contributions reported in the literature in four classes two of them are statistical and two are deterministic. We also present a comparative empirical evaluation of representative algorithms selected from these four classes. Novel quantitative (detection and discrimination rate) and qualitative metrics (scene and object independence, flexibility to shadow situations, and robustness to noise) are proposed to evaluate these classes of algorithms on a benchmark suite of indoor and outdoor video sequences. These video sequences and associated "ground-truth" data are made available at http://cvrr.ucsd.edu/aton/shadow to allow for others in the community to experiment with new algorithms and metrics.
Andrea Prati 0001, Ivana Mikic, Mohan M. Trivedi, Rita Cucchiara
IEEE Trans. Pattern Anal. Mach. Intell.4
2003 A New Algorithm for Border Description of Polarized Light Surface Microscopic Images of Pigmented Skin Lesions
abstract
The aim of this study was to provide mathematical descriptors for the border of pigmented skin lesion images and to assess their efficacy for distinction among different lesion groups. New descriptors such as lesion slope and lesion slope regularity are introduced and mathematically defined. A new algorithm based on the Catmull-Rom spline method and the computation of the gray-level gradient of points extracted by interpolation of normal direction on spline points was employed. The efficacy of these new descriptors was tested on a data set of 510 pigmented skin lesions, composed by 85 melanomas and 425 nevi, by employing statistical methods for discrimination between the two populations.
Costantino Grana, Giovanni Pellacani, Rita Cucchiara, Stefania Seidenari
IEEE Trans. Medical Imaging3
2002 Semantic transcoding for live video server
abstract
In this paper we present transcoding techniques for a video server architecture that enables the user to access live video streams by using different devices with different capabilities. For live videos, annotation methods cannot be exploited. Instead we propose methods of on-the-fly transcoding that adapt the video content with respect to the user resources and the video semantic. Thus we propose an object-based transcoding with classes of relevance (for instance People, Face and Background). To compare the different strategies we propose a metric based on the Weighted Mean Square Error that allows the analysis of different application scenarios by means of a class-wise distortion measure. The obtained results show that the use of semantic can improve the bandwidth to distortion ratio significantly.
Rita Cucchiara, Costantino Grana, Andrea Prati 0001
ACM Multimedia1
2001 Analysis and Detection of Shadows in Video Streams: A Comparative Evaluation
abstract
Robustness to changes in illumination conditions as well as viewing perspectives is an important requirement for many computer vision applications. One of the key factors in enhancing the robustness of dynamic scene analysis is that of accurate and reliable means for shadow detection. Shadow detection is critical for correct object detection in image sequences. Many algorithms have been proposed in the literature that deal with shadows. However, a comparative evaluation of the existing approaches is still lacking. In this paper, the full range of problems underlying the shadow detection is identified and discussed. We classify the proposed solutions to this problem using a taxonomy of four main classes, deterministic model and non-model based, and statistical parametric and nonparametric. Novel quantitative (detection and discrimination accuracy) and qualitative metrics (scene and object independence, flexibility to shadow situations and robustness to noise) are proposed to evaluate these classes of algorithms on a benchmark suite of indoor and outdoor video sequences.
Andrea Prati 0001, Rita Cucchiara, Ivana Mikic, Mohan M. Trivedi
CVPR (2)2
2001 An application of machine learning and statistics to defect detection
Rita Cucchiara, Paola Mello, Massimo Piccardi, Fabrizio Riguzzi
Intell. Data Anal.1
2000 Focus based Feature Extraction for Pallets Recognition
abstract
Visual recognition for object grasping is a well-known challenge for robot automation in industrial applications. A typical example is pallet recognition in industrial environment for pick-and-place automated process. The aim of vision and reasoning algorithms is to help robots in choosing the best pallets holes location. This work proposes an application-based approach, which full all requirements, dealing with every kind of occlusions and light situ-ations possible. Even some meaning noise (or meaning misunderstand-ing) is considered. A pallet model, with limited degrees of freedom, is de-scribed and, starting from it, a complete approach to pallet recognition is out-lined. In the model we dene both virtual and real corners, that are geomet-rical object proprieties computed by different image analysis operators. Real corners are perceived by processing brightness information directly from the image, while virtual corners are inferred at a higher level of abstraction. A nal reasoning stage selects the best solution tting the model. Experimental results and performance are reported in order to demonstrate the suitability of the proposed approach. 1
Rita Cucchiara, Massimo Piccardi, Andrea Prati 0001
BMVC1
2000 Optimal Range Segmentation Parameters through Genetic Algorithms
abstract
A wide number of algorithms for surface segmentation in range images have been recently proposed characterized by different approaches (edge filling, region growing,...), different surface types (either for planar or curved surfaces) and different parameters involved. Optimization of the parameter set is a particularly critical task since the range of parameter variability is often quite large: parameter selection depends on surface type, sensors and the required speed which strongly of affect performance. A framework for parameter optimization is proposed based on genetic algorithms. Such algorithms allow a general approach that has been successfully applied on different state-of-the-art segmenters and different range image databases.
Luigi Cinque, Stefano Levialdi, Gianluca Pignalberi, Rita Cucchiara, Stefano Martinz
ICPR4
2000 Image analysis and rule-based reasoning for a traffic monitoring system
abstract
The paper presents an approach for detecting vehicles in urban traffic scenes by means of rule-based reasoning on visual data. The strength of the approach is its formal separation between the low-level image processing modules and the high-level module, which provides a general-purpose knowledge-based framework for tracking vehicles in the scene. The image-processing modules extract visual data from the scene by spatio-temporal analysis during daytime, and by morphological analysis of headlights at night. The high-level module is designed as a forward chaining production rule system, working on symbolic data, i.e., vehicles and their attributes (area, pattern, direction, and others) and exploiting a set of heuristic rules tuned to urban traffic conditions. The synergy between the artificial intelligence techniques of the high-level and the low-level image analysis techniques provides the system with flexibility and robustness.
Rita Cucchiara, Massimo Piccardi, Paola Mello
IEEE Trans. Intell. Transp. Syst.1
1999 Constraint Propagation and Value Acquisition: Why we should do it Interactively
Evelina Lamma, Paola Mello, Michela Milano, Rita Cucchiara, Marco Gavanelli, Massimo Piccardi
IJCAI4
1999 Eliciting visual primitives for detection of elongated shapes
Rita Cucchiara, Massimo Piccardi
Image Vis. Comput.1
1998 Exploiting image processing locality in cache pre-fetching
abstract
Emerging trends in computer design attempt to include specific solutions for handling images also in general-purpose computers, because of the current spread of multimedia, image processing and computer graphics applications. In this context, we propose hardware pre-fetching techniques specific for caching images: the main issue we state is that most algorithms working on images exhibit a 2D spatial locality that is not taken into account in current cache organization and data access strategies. To this aim we propose an adaptive local pre-fetching for the image data type; this technique, mirroring the two-dimensional spatial locality of image processing algorithms, results in being more efficient than other approaches, such as sequential pre-fetching and adaptive pre-fetching. Performance is evaluated on different classes of image processing algorithms, namely raster-scan and propagative algorithms, common in computer vision and multimedia applications.
Rita Cucchiara, Massimo Piccardi
HiPC1
1998 A real-time hardware implementation of the hough transform
Rita Cucchiara, Giovanni Neri, Massimo Piccardi
J. Syst. Archit.1
1998 Genetic algorithms for clustering in machine vision
Rita Cucchiara
Mach. Vis. Appl.1
1998 The Vector-Gradient Hough Transform
abstract
The paper presents a new transform, called vector-gradient Hough transform, for identifying elongated shapes in gray-scale images. This goal is achieved not only by collecting information on the edges of the objects, but also by reconstructing their transversal profile of luminosity. The main features of the new approach are related to its vector space formulation and the associated capability of exploiting all the vector information of the luminosity gradient.
Rita Cucchiara, Fabio Filicori
IEEE Trans. Pattern Anal. Mach. Intell.1
1997 Exploiting Symbolic Learning in Visual Inspection
Massimo Piccardi, Rita Cucchiara, Michele Bariani, Paola Mello
IDA2
1997 An Interactive Constraint-Based System for Selective Attention in Visual Search
Rita Cucchiara, Evelina Lamma, Paola Mello, Michela Milano
ISMIS1
1997 Block processing on multiprocessor DSPs for multimedia applications
abstract
The paper explores software development for multiprocessor DSPs for data parallel local algorithms. These algorithms are very common in multimedia applications, such as filtering, compression and so on. Multiprocessor DSPs are very attractive for this application since they offer performance typical of parallel machines together with limited cost. The paper provides performance analysis and software design issues according to different data partitioning models. As a case study, performance evaluation has been carried out on the Multimedia Video Processor from Texas Instrument.
Rita Cucchiara, Alessandro Callipo, Massimo Piccardi
MMSP1
1996 Detection of luminosity profiles of elongated shapes
abstract
A novel technique for identifying elongated shapes in grey-scale images is presented. The method provides the detection and identification of elongated shapes not only modelling their principal direction, but also reconstructing the transversal luminosity profile. The approach is proposed starting from the gradient-weighted Hough transform, endowing the Hough space with more complete information about the luminance gradient of the image. This paper presents and discusses the algorithm devised to implement the method on discretized data. As an example of application, we present results on images from mechanical pieces, where real and false defects are discriminated through effective reconstruction of their luminosity profile.
Rita Cucchiara, Massimo Piccardi
ICIP (3)1
1996 The vector-gradient Hough transform for identifying straight-translation generated shapes
abstract
The paper introduces the vector-gradient Hough transform (VGHT), a modified version of the gradient weighted Hough transform (GWHT), defined in vector space and able to exploit all the vector information of the gradient of luminosity. The new formulation, directly derived from the Radon transform, is analyzed and compared with the GWHT, in order to point out the improvement in selectivity provided by the VGHT in a strictly polar parametric space, without any relevant increase in computational complexity. This approach can be very suitable for identifying a specifically defined model of shapes in gray level images, ideally generated by a translation in the 2D space of a 1D luminosity profile. Finally, the suitability of the VGHT in real applications is shown with examples in the area of defect identification for automated visual inspection.
Rita Cucchiara, Fabio Filicori
ICPR1
1995 A Highly Selective HT based Algorithm for Detecting Extended, Almost Rectilinear Shapes
Rita Cucchiara, Fabio Filicori
CAIP1
1995 Detection of Circular Objects by Wave Propagation on a Mesh-Connected Computer
Rita Cucchiara, Luigi Di Stefano, Massimo Piccardi
J. Parallel Distributed Comput.1
1993 Processing of variable size images on a cellular array: Performance analysis with the Abingdon Cross Benchmark
abstract
Handling a continuous flow of variable size images is a requirement for real time computer vision machines. A modular system based on a small size SIMD cellular array of 1-bit processing elements has been developed with this goal in mind and it is now evaluated against the Abingdon Cross Benchmark specifications. The benchmark tests the combination of algorithms and architecture and generates a quality factor expressed as the ratio of the image lateral size and the processing time. The examined machine supports an efficient means to automatically partition, process and reconstruct images larger than the array size. The authors briefly describe the system, discuss the selected algorithms and present performance results and estimates for several system configurations.>
Massimo Piccardi, Luigi Di Stefano, Rita Cucchiara, Tullio Salmon Cinotti
ASAP3