VLDB 2026 Research / reviewers in the wild / expert
Lorenzo Baraldi 0001
dblp:158/5775
· DBLP profile ↗
91ranked-venue papers
10as first author
59since 2021 · last 2026
0000-0001-5125-4957ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 69 · 10 first-author · 41 since 2021Artificial intelligence and machine learning · 60 · 4 first-author · 42 since 2021Systems, architecture and hardware · 4 · 2 since 2021Computer networks · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
Silvia Cappelletti, Tobia Poppi, Samuele Poppi, Diego Garcia-Olano, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR (3) | 7 |
| 2026 | GramSR: Visual Feature Conditioning for Diffusion-Based Super-Resolution
Fabio D'Oronzio, Federico Putamorsi, Leonardo Zini, Marcella Cornia, Lorenzo Baraldi 0001 |
ICPR (7) | 5 |
| 2026 | RaTA-Tool: Retrieval-Based Tool Selection with Multimodal Large Language Models
Gabriele Mattioli, Evelyn Turri, Sara Sarto, Lorenzo Baraldi 0002, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR (7) | 6 |
| 2026 | FG-Tracer: Tracing Information Flow in Multimodal Large Language Models in Free-Form GenerationabstractMultimodal Large Language Models (MLLMs) have achieved impressive performance across a variety of vision–language tasks. However, their internal working mechanisms remain largely underexplored. In this work, we introduce FG-Tracer, a framework designed to analyze the information flow between visual and textual modalities in MLLMs in free-form generation. Notably, our numerically stabilized computational method enables the first systematic analysis of multimodal information flow in underexplored domains such as image captioning and chain-of-thought (CoT) reasoning. We apply FG-Tracer to three state-of-the-art MLLMs—LLaVA 1.5, LLaMA 3.2-Vision, and Qwen 2.5-VL—across three vision–language benchmarks—TextVQA, COCO 2014, and ChartQA—and we conduct a word-level analysis of multimodal integration. Our findings uncover distinct patterns of multimodal fusion across models and tasks, demonstrating that fusion dynamics are both model- and task-dependent. Overall, FG-Tracer offers a robust methodology for probing the internal mechanisms of MLLMs in free-form settings, providing new insights into their multimodal reasoning strategies. Our source code is publicly available at https://github.com/AImageLab-zip/FG-TRACER Alessia Saporita, Vittorio Pipoli, Federico Bolelli, Lorenzo Baraldi 0001, Andrea Acquaviva, Elisa Ficarra |
WACV | 4 |
| 2026 | Hallucination Early Detection in Diffusion Models
Federico Betti 0001, Lorenzo Baraldi 0002, Lorenzo Baraldi 0001, Rita Cucchiara, Nicu Sebe |
Int. J. Comput. Vis. | 3 |
| 2026 | An In-Depth Survey on Multimodal Automatic Fact-Checking DatasetsabstractAbstract The rapid spread of misinformation poses a significant challenge in the digital age, with false claims appearing across multiple modalities, particularly combinations of textual and visual content such as images and videos. While automatic fact-checking plays a crucial role in countering misinformation, traditional approaches predominantly rely on textual data, often neglecting the multimodal nature of modern misinformation. In this survey, we provide a comprehensive evaluation of multimodal datasets designed for automatic fact-checking that combine textual and visual information, systematically analyzing their sources, annotation methodologies, and key statistical properties, such as class distribution, topic diversity, and label availability. Additionally, we assess the usability of these datasets in real-world scenarios, discussing their limitations, biases, and potential risks, such as information leakage. Motivated by the practical difficulties we encountered when attempting to integrate existing datasets into our own multimodal fact-checking pipeline, our work also offers concrete guidance to help researchers choose the most suitable resources. By identifying gaps and challenges in existing datasets, our survey aims to support the development of more reliable and scalable multimodal fact-checking systems. Ian Marco Gallegos Carvajal, Beatrice Portelli, Leonardo Zini, Lorenzo Baraldi 0001, Giuseppe Serra 0001 |
Multim. Tools Appl. | 4 |
| 2025 | Tracing Information Flow in LLaMA Vision: A Step Toward Multimodal Understanding
Alessia Saporita, Vittorio Pipoli, Federico Bolelli, Lorenzo Baraldi 0001, Andrea Acquaviva, Elisa Ficarra |
CAIP (2) | 4 |
| 2025 | Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document RetrievalabstractCross-modal retrieval is gaining increasing efficacy and interest from the research community, thanks to large-scale training, novel architectural and learning designs, and its application in LLMs and multimodal LLMs. In this paper, we move a step forward and design an approach that allows for multimodal queries – composed of both an image and a text – and can search within collections of multi-modal documents, where images and text are interleaved. Our model, ReT, employs multi-level representations extracted from different layers of both visual and textual backbones, both at the query and document side. To allow for multi-level and cross-modal understanding and feature extraction, ReT employs a novel Transformer-based recurrent cell that integrates both textual and visual features at different layers, and leverages sigmoidal gates inspired by the classical design of LSTMs. Extensive experiments on M2KR and M-BEIR benchmarks show that ReT achieves state-of-the-art performance across diverse settings. Our source code and trained models are publicly available at: https://github.com/aimagelab/ReT. Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 4 |
| 2025 | Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question AnsweringabstractMultimodal LLMs (MLLMs) are the natural extension of large language models to handle multimodal inputs, combining text and image data. They have recently garnered attention due to their capability to address complex tasks involving both modalities. However, their effectiveness is limited to the knowledge acquired during training, which restricts their practical utility. In this work, we introduce a novel method to enhance the adaptability of MLLMs by integrating external knowledge sources. Our proposed model, Reflective LLaVA (ReflectiVA), utilizes reflective tokens to dynamically determine the need for external knowledge and predict the relevance of information retrieved from an external database. Tokens are trained following a two-stage two-model training recipe. This ultimately enables the MLLM to manage external knowledge while preserving fluency and performance on tasks where external knowledge is not needed. Through our experiments, we demonstrate the efficacy of ReflectiVA for knowledge-based visual question answering, highlighting its superior performance compared to existing methods. Source code and trained models are publicly available at https://aimagelab.github.io/ReflectiVA. Federico Cocchi, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 4 |
| 2025 | Hyperbolic Safety-Aware Vision-Language ModelsabstractAddressing the retrieval of unsafe content from vision-language models such as CLIP is an important step towards real-world integration. Current efforts have relied on unlearning techniques that try to erase the model’s knowledge of unsafe concepts. While effective in reducing unwanted outputs, unlearning limits the model’s capacity to discern between safe and unsafe content. In this work, we introduce a novel approach that shifts from unlearning to an awareness paradigm by leveraging the inherent hierarchical properties of the hyperbolic space. We propose to encode safe and unsafe content as an entailment hierarchy, where both are placed in different regions of hyperbolic space. Our HySAC, Hyperbolic Safety-Aware CLIP, employs entailment loss functions to model the hierarchical and asymmetrical relations between safe and unsafe image-text pairs. This modelling – ineffective in standard vision-language models due to their reliance on Euclidean embeddings – endows the model with awareness of unsafe content, enabling it to serve as both a multimodal unsafe classifier and a flexible content retriever, with the option to dynamically redirect unsafe queries toward safer alternatives or retain the original output. Extensive experiments show that our approach not only enhances safety recognition but also establishes a more adaptable and interpretable framework for content moderation in vision-language models. Our source code is available at: https://github.com/aimagelab/HySAC Tobia Poppi, Tejaswi Kasarla, Pascal Mettes, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 4 |
| 2025 | Multimodal Emotion Recognition in Conversation via Possible Speaker's Audio and Visual Sequence SelectionabstractMultimodal Emotion Recognition in Conversation (MERC) is an important element in human-machine interaction. It allows machines to automatically identify and track the emotional status of speakers during a conversation in a multimodal setting. However, the conversations involving various audio and visual cues aligned with textual cues are very complex. Recent works have tried integrating the audio and visual modalities with textual to improve the performance of emotion recognition in conversation. Although many MERC models leverage textual, audio, and visual modalities, those models assume that the speaker’s textual utterance, audio speech, and facial sequences are present. However, a conversation may contain multiple parties, among which only one is the speaker. Previous MERC assumed the availability of all modalities, but in many instances, one or more modalities may be unavailable during multiparty conversations. To tackle these issues, we propose the Possible Speaker Informed Multimodal Emotion Recognition in Conversation framework (PSI). PSI is specifically tasked to extract audio (speech) and visual (face) sequences of a possible speaker in the presence of multiple parties. Further, PSI seamlessly extracts the rich unimodal features and fuses them while addressing the unavailability of specific modalities. PSI demonstrates competitive performance with existing state-of-the-art models through experiments with a benchmark dataset. Rahul Singh Maharjan, Niyati Rawal, Marta Romeo, Lorenzo Baraldi 0001, Rita Cucchiara, Angelo Cangelosi |
ICASSP | 4 |
| 2025 | What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language ModelsabstractInstruction-based image editing models offer increased personalization opportunities in generative tasks. However, properly evaluating their results is challenging, and most of the existing metrics lag in terms of alignment with human judgment and explainability. To tackle these issues, we introduce DICE (DIfference Coherence Estimator), a model designed to detect localized differences between the original and the edited image and to assess their relevance to the given modification request. DICE consists of two key components: a difference detector and a coherence estimator, both built on an autoregressive Multimodal Large Language Model (MLLM) and trained using a strategy that leverages self-supervision, distillation from inpainting networks, and full supervision. Through extensive experiments, we evaluate each stage of our pipeline, comparing different MLLMs within the proposed framework. We demonstrate that DICE effectively identifies coherent edits, effectively evaluating images generated by different editing models with a strong correlation with human judgment. We publicly release our source code, models, and data. Lorenzo Baraldi 0001, Davide Bucciarelli, Federico Betti 0001, Marcella Cornia, Lorenzo Baraldi 0002, Nicu Sebe, Rita Cucchiara |
ICCV | 1 |
| 2025 | Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary SegmentationabstractOpen-Vocabulary Segmentation (OVS) aims at segmenting images from free-form textual concepts without predefined training classes. While existing vision-language models such as CLIP can generate segmentation masks by leveraging coarse spatial information from Vision Transformers, they face challenges in spatial localization due to their global alignment of image and text features. Conversely, self-supervised visual models like DINO excel in fine-grained visual encoding but lack integration with language. To bridge this gap, we present Talk2DINO, a novel hybrid approach that combines the spatial accuracy of DINOv2 with the language understanding of CLIP. Our approach aligns the textual embeddings of CLIP to the patch-level features of DINOv2 through a learned mapping function without the need to fine-tune the underlying backbones. At training time, we exploit the attention maps of DINOv2 to selectively align local visual patches with textual embeddings. We show that the powerful semantic and localization abilities of Talk2DINO can enhance the segmentation process, resulting in more natural and less noisy segmentations, and that our approach can also effectively distinguish foreground objects from the background. Experimental results demonstrate that Talk2DINO achieves state-of-the-art performance across several unsupervised OVS benchmarks. Source code and models are publicly available at: https://lorebianchi98.github.io/Talk2DINO/. Luca Barsellotti, Lorenzo Bianchi 0001, Nicola Messina, Fabio Carrara, Marcella Cornia, Lorenzo Baraldi 0001, Fabrizio Falchi, Rita Cucchiara |
ICCV | 6 |
| 2025 | MISSRAG: Addressing the Missing Modality Challenge in Multimodal Large Language Models
Vittorio Pipoli, Alessia Saporita, Federico Bolelli, Marcella Cornia, Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara, Elisa Ficarra |
ICCV | 5 |
| 2025 | Causal Graphical Models for Vision-Language Compositional UnderstandingabstractRecent work has empirically shown that Vision-Language Models (VLMs) struggle
to fully understand the compositional properties of the human language, usually
modeling an image caption as a “bag of words”. As a result, they perform
poorly on compositional tasks, which require a deeper understanding of the different
entities of a sentence (subject, verb, etc.) jointly with their mutual relationships
in order to be solved. In this paper, we model the dependency relations
among textual and visual tokens using a Causal Graphical Model (CGM), built using
a dependency parser, and we train a decoder conditioned by the VLM visual
encoder. Differently from standard autoregressive or parallel predictions, our decoder’s
generative process is partially-ordered following the CGM structure. This
structure encourages the decoder to learn only the main causal dependencies in
a sentence discarding spurious correlations. Using extensive experiments on five
compositional benchmarks, we show that our method significantly outperforms
all the state-of-the-art compositional approaches by a large margin, and it also improves
over methods trained using much larger datasets.
Our model weights and code are publicly available. Fiorenzo Parascandolo, Nicholas Moratelli, Enver Sangineto, Lorenzo Baraldi 0001, Rita Cucchiara |
ICLR | 4 |
| 2025 | vHector and HeisenVec: Scalable Vector Graphics Generation Through Large Language ModelsabstractWe introduce HeisenVec, a large-scale dataset designed to advance research in vector graphics generation from natural language descriptions. Unlike conventional image generation datasets that focus on raster images, HeisenVec targets the structured and symbolic domain of Scalable Vector Graphics (SVG), where images are represented as sequences of drawing commands and style attributes. The dataset comprises 2.2 million SVGs collected from different online sources, each paired with four complementary textual descriptions generated by multi-modal models. To ensure structural consistency and efficiency for autoregressive modeling, all SVGs are standardized through a pre-processing pipeline that unifies geometric primitives as paths, applies affine transformations, and compresses syntax via custom tokens set. HeisenVec exhibits broad coverage among visual styles and sequence lengths, with a substantial portion of samples exceeding 8,000 tokens, making it particularly well-suited for benchmarking long-context language models. Our benchmark enables rigorous evaluation of text-conditioned SVG generation, encourages progress on sequence modeling with symbolic outputs, and bridges the gap between vision, graphics, and language. We release the dataset, tokenization tools, and evaluation pipeline to foster further research in this emerging domain. Leonardo Zini, Elia Frigieri, Sebastiano Aloscari, Lorenzo Baraldi 0001 |
NeurIPS | 4 |
| 2025 | Perceive. Query & Reason: Enhancing Video QA with Question-Guided Temporal QueriesabstractVideo Question Answering (Video QA) is a challenging video understanding task that requires models to compre-hend entire videos, identify the most relevant information based on contextual cues from a given question, and rea-son accurately to provide answers. Recent advancements in Multimodal Large Language Models (MLLMs) have trans-formed video QA by leveraging their exceptional common-sense reasoning capabilities. This progress is largely driven by the effective alignment between visual data and the language space of MLLMs. However, for video QA, an ad-ditional space-time alignment poses a considerable chal-lenge for extracting question-relevant information across frames. In this work, we investigate diverse temporal modeling techniques to integrate with MLLMs, aiming to achieve question-guided temporal modeling that leverages pre-trained visual and textual alignment in MLLMs. We propose T-Former, a novel temporal modeling method that creates a question-guided temporal bridge between frame-wise visual perception and the reasoning capabilities of LLMs. Our evaluation across multiple video QA bench-marks demonstrates that T-Former competes favorably with existing temporal modeling approaches and aligns with re-cent advancements in video QA. Roberto Amoroso, Gengyuan Zhang, Rajat Koner, Lorenzo Baraldi 0001, Rita Cucchiara, Volker Tresp |
WACV | 4 |
| 2025 | Semantically Conditioned Prompts for Visual Recognition Under Missing Modality ScenariosabstractThis paper tackles the domain of multimodal prompting for visual recognition, specifically when dealing with missing modalities through multimodal Transformers. It presents two main contributions: (i) we introduce a novel prompt learning module which is designed to produce sample-specific prompts and (ii) we show that modalityagnostic prompts can effectively adjust to diverse missing modality scenarios. Our model, termed SCP, exploits the semantic representation of available modalities to query a learnable memory bank, which allows the generation of prompts based on the semantics of the input. Notably, SCP distinguishes itself from existing methodologies for its capacity of self-adjusting to both the missing modality scenario and the semantic context of the input, without prior knowledge about the specific missing modality and the number of modalities. Through extensive experiments, we show the effectiveness of the proposed prompt learning framework and demonstrate enhanced performance and robustness across a spectrum of missing modality cases. Our source code is available at https://github.com/vittoriopipoli/SCP_WACV2025. Vittorio Pipoli, Federico Bolelli, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara, Elisa Ficarra |
WACV | 5 |
| 2025 | Learning to mask and permute visual tokens for Vision Transformer pre-trainingabstractThe use of self-supervised pre-training has emerged as a promising approach to enhance the performance of many different visual tasks. In this context, recent approaches have employed the Masked Image Modeling paradigm, which pre-trains a backbone by reconstructing visual tokens associated with randomly masked image patches. This masking approach, however, introduces noise into the input data during pre-training, leading to discrepancies that can impair performance during the fine-tuning phase. Furthermore, input masking neglects the dependencies between corrupted patches, increasing the inconsistencies observed in downstream fine-tuning tasks. To overcome these issues, we propose a new self-supervised pre-training approach, named Masked and Permuted Vision Transformer (MaPeT), that employs autoregressive and permuted predictions to capture intra-patch dependencies. In addition, MaPeT employs auxiliary positional information to reduce the disparity between the pre-training and fine-tuning phases. In our experiments, we employ a fair setting to ensure reliable and meaningful comparisons and conduct investigations on multiple visual tokenizers, including our proposed k -CLIP which directly employs discretized CLIP features. Our results demonstrate that MaPeT achieves competitive performance on ImageNet, compared to baselines and competitors under the same model setting. We release an implementation of our code and models at https://github.com/aimagelab/MaPeT . Lorenzo Baraldi 0002, Roberto Amoroso, Marcella Cornia, Lorenzo Baraldi 0001, Andrea Pilzer, Rita Cucchiara |
Comput. Vis. Image Underst. | 4 |
| 2025 | Positive-Augmented Contrastive Learning for Vision-and-Language Evaluation and Training
Sara Sarto, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
Int. J. Comput. Vis. | 4 |
| 2025 | Augmenting and mixing Transformers with synthetic data for image captioningabstractImage captioning has attracted significant attention within the Computer Vision and Multimedia research domains, resulting in the development of effective methods for generating natural language descriptions of images. Concurrently, the rise of generative models has facilitated the production of highly realistic and high-quality images, particularly through recent advancements in latent diffusion models. In this paper, we propose to leverage the recent advances in Generative AI and create additional training data that can be effectively used to boost the performance of an image captioning model. Specifically, we combine real images with their synthetic counterparts generated by Stable Diffusion using a Mixup data augmentation technique to create novel training examples. Extensive experiments on the COCO dataset demonstrate the effectiveness of our solution in comparison to different baselines and state-of-the-art methods and validate the benefits of using synthetic data to augment the training stage of an image captioning model and improve the quality of the generated captions. Source code and trained models are publicly available at: https://github.com/aimagelab/synthcap_pp . Davide Caffagni, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
Image Vis. Comput. | 3 |
| 2025 | Parents and Children: Distinguishing Multimodal Deepfakes from Natural ImagesabstractRecent advancements in diffusion models have enabled the generation of realistic deepfakes from textual prompts in natural language. While these models have numerous benefits across various sectors, they have also raised concerns about the potential misuse of fake images and cast new pressures on fake image detection. In this work, we pioneer a systematic study on deepfake detection generated by state-of-the-art diffusion models. Firstly, we conduct a comprehensive analysis of the performance of contrastive and classification-based visual features, respectively, extracted from CLIP-based models and ResNet or Vision Transformer (ViT)-based architectures trained on image classification datasets. Our results demonstrate that fake images share common low-level cues, which render them easily recognizable. Further, we devise a multimodal setting wherein fake images are synthesized by different textual captions, which are used as seeds for a generator. Under this setting, we quantify the performance of fake detection strategies and introduce a contrastive-based disentangling method that lets us analyze the role of the semantics of textual descriptions and low-level perceptual cues. Finally, we release a new dataset, called COCOFake, containing about 1.2 million images generated from the original COCO image–caption pairs using two recent text-to-image diffusion models, namely Stable Diffusion v1.4 and v2.0. Roberto Amoroso, Davide Morelli, Marcella Cornia, Lorenzo Baraldi 0001, Alberto Del Bimbo, Rita Cucchiara |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
Nicholas Moratelli, Davide Caffagni, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
BMVC | 4 |
| 2024 | Training-Free Open-Vocabulary Segmentation with Offline Diffusion-Augmented Prototype GenerationabstractOpen-vocabulary semantic segmentation aims at segmenting arbitrary categories expressed in textual form. Pre-vious works have trained over large amounts of image-caption pairs to enforce pixel-level multimodal alignments. However, captions provide global information about the semantics of a given image but lack direct localization of individual concepts. Further, training on large-scale datasets inevitably brings significant computational costs. In this paper, we propose FreeDA, a training-free diffusion-augmented method for open-vocabulary semantic segmentation, which leverages the ability of diffusion models to visually localize generated concepts and local-global similarities to match class-agnostic regions with semantic classes. Our approach involves an offline stage in which textual-visual reference embeddings are collected, starting from a large set of captions and leveraging visual and semantic contexts. At test time, these are queried to support the visual matching process, which is carried out by jointly considering class-agnostic regions and global semantic similarities. Extensive analyses demonstrate that FreeDA achieves state-of-the-art performance on five datasets, surpassing previous methods by more than 7.0 average points in terms of mIoU and without requiring any training. Our source code is available at aimagelab.github. io/freeda. Luca Barsellotti, Roberto Amoroso, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 4 |
| 2024 | Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities
Federico Cocchi, Marcella Cornia, Lorenzo Baraldi 0001, Alessandro Nicolosi, Rita Cucchiara |
ECCV (63) | 3 |
| 2024 | Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models
Samuele Poppi, Tobia Poppi, Federico Cocchi, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ECCV (53) | 5 |
| 2024 | BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues
Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ECCV (78) | 3 |
| 2024 | Adapt to Scarcity: Few-Shot Deepfake Detection via Low-Rank Adaptation
Silvia Cappelletti, Lorenzo Baraldi 0002, Federico Cocchi, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR (21) | 5 |
| 2024 | Fluent and Accurate Image Captioning with a Self-trained Reward Model
Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR (18) | 3 |
| 2024 | Unlearning Vision Transformers Without Retaining Data via Low-Rank Decompositions
Samuele Poppi, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR (3) | 4 |
| 2024 | Mapping High-level Semantic Regions in Indoor Environments without Object RecognitionabstractRobots require a semantic understanding of their surroundings to operate in an efficient and explainable way in human environments. In the literature, there has been an extensive focus on object labeling and exhaustive scene graph generation; less effort has been focused on the task of purely identifying and mapping large semantic regions. The present work proposes a method for semantic region mapping via embodied navigation in indoor environments, generating a high-level representation of the knowledge of the agent. To enable region identification, the method uses a vision-to-language model to provide scene information for mapping. By projecting egocentric scene understanding into the global frame, the proposed method generates a semantic map as a distribution over possible region labels at each location. This mapping procedure is paired with a trained navigation policy to enable autonomous map generation. The proposed method significantly outperforms a variety of baselines, including an object-based system and a pretrained scene classifier, in experiments in a photorealistic simulator. Roberto Bigazzi, Lorenzo Baraldi 0001, Shreyas Kousik, Rita Cucchiara, Marco Pavone 0001 |
ICRA | 2 |
| 2024 | Personalized Instance-based Navigation Toward User-Specific Objects in Realistic EnvironmentsabstractIn the last years, the research interest in visual navigation towards objects in indoor environments has grown significantly. This growth can be attributed to the recent availability of large navigation datasets in photo-realistic simulated environments, like Gibson and Matterport3D. However, the navigation tasks supported by these datasets are often restricted to the objects present in the environment at acquisition time. Also, they fail to account for the realistic scenario in which the target object is a user-specific instance that can be easily confused with similar objects and may be found in multiple locations within the environment. To address these limitations, we propose a new task denominated Personalized Instance-based Navigation (PIN), in which an embodied agent is tasked with locating and reaching a specific personal object by distinguishing it among multiple instances of the same category. The task is accompanied by PInNED, a dedicated new dataset composed of photo-realistic scenes augmented with additional 3D objects. In each episode, the target object is presented to the agent using two modalities: a set of visual reference images on a neutral background and manually annotated textual descriptions. Through comprehensive evaluations and analyses, we showcase the challenges of the PIN task as well as the performance and shortcomings of currently available methods designed for object-driven navigation, considering modular and end-to-end agents. Luca Barsellotti, Roberto Bigazzi, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
NeurIPS | 4 |
| 2024 | FOSSIL: Free Open-Vocabulary Semantic Segmentation through Synthetic References RetrievalabstractUnsupervised Open-Vocabulary Semantic Segmentation aims to segment an image into regions referring to an arbitrary set of concepts described by text, without relying on dense annotations that are available only for a subset of the categories. Previous works rely on inducing pixel-level alignment in a multi-modal space through contrastive training over vast corpora of image-caption pairs. However, representing a semantic category solely through its textual embedding is insufficient to encompass the wide-ranging variability in the visual appearances of the images associated with that category. In this paper, we propose FOSSIL, a pipeline that enables a self-supervised backbone to perform open-vocabulary segmentation relying only on the visual modality. In particular, we decouple the task into two components: (1) we leverage text-conditioned diffusion models to generate a large collection of visual embeddings, starting from a set of captions. These can be retrieved at inference time to obtain a support set of references for the set of textual concepts. Further, (2) we exploit self-supervised dense features to partition the image into semantically coherent regions. We demonstrate that our approach provides strong performance on different semantic segmentation datasets, without requiring any additional training. Luca Barsellotti, Roberto Amoroso, Lorenzo Baraldi 0001, Rita Cucchiara |
WACV | 3 |
| 2024 | What's Outside the Intersection? Fine-grained Error Analysis for Semantic Segmentation Beyond IoUabstractSemantic segmentation represents a fundamental task in computer vision with various application areas such as autonomous driving, medical imaging, or remote sensing. For evaluating and comparing semantic segmentation models, the mean intersection over union (mIoU) is currently the gold standard. However, while mIoU serves as a valuable benchmark, it does not offer insights into the types of errors incurred by a model. Moreover, different types of errors may have different impacts on downstream applications. To address this issue, we propose an intuitive method for the systematic categorization of errors, thereby enabling a fine-grained analysis of semantic segmentation models. Since we assign each erroneous pixel to precisely one error type, our method seamlessly extends the popular IoU-based evaluation by shedding more light on the false positive and false negative predictions. Our approach is model- and dataset-agnostic, as it does not rely on additional information besides the predicted and ground-truth segmentation masks. In our experiments, we demonstrate that our method accurately assesses model strengths and weaknesses on a quantitative basis, thus reducing the dependence on time-consuming qualitative model inspection. We analyze a variety of state-of-the-art semantic segmentation models, revealing systematic differences across various architectural paradigms. Exploiting the gained insights, we showcase that combining two models with complementary strengths in a straightforward way is sufficient to consistently improve mIoU, even for models setting the current state of the art on ADE20K. We release a toolkit for our evaluation method at https://github.com/mxbh/beyond-iou. Maximilian Bernhard, Roberto Amoroso, Yannic Kindermann, Lorenzo Baraldi 0001, Rita Cucchiara, Volker Tresp, Matthias Schubert |
WACV | 4 |
| 2024 | Generating More Pertinent Captions by Leveraging Semantics and Style on Multi-Source Datasets
Marcella Cornia, Lorenzo Baraldi 0001, Giuseppe Fiameni, Rita Cucchiara |
Int. J. Comput. Vis. | 2 |
| 2024 | Towards Retrieval-Augmented Architectures for Image CaptioningabstractThe objective of image captioning models is to bridge the gap between the visual and linguistic modalities by generating natural language descriptions that accurately reflect the content of input images. In recent years, researchers have leveraged deep learning-based models and made advances in the extraction of visual features and the design of multimodal connections to tackle this task. This work presents a novel approach toward developing image captioning models that utilize an externalkNN memory to improve the generation process. Specifically, we propose two model variants that incorporate a knowledge retriever component that is based on visual similarities, a differentiable encoder to represent input images, and akNN-augmented language model to predict tokens based on contextual cues and text retrieved from the external memory. We experimentally validate our approach on COCO and nocaps datasets and demonstrate that incorporating an explicit external memory can significantly enhance the quality of captions, especially with a larger retrieval corpus. This work provides valuable insights into retrieval-augmented captioning models and opens up new avenues for improving image captioning at a larger scale. Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Alessandro Nicolosi, Rita Cucchiara |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Superpixel Positional Encoding to Improve ViT-based Semantic Segmentation Models
Roberto Amoroso, Matteo Tomei, Lorenzo Baraldi 0001, Rita Cucchiara |
BMVC | 3 |
| 2023 | Positive-Augmented Contrastive Learning for Image and Video Captioning EvaluationabstractThe CLIP model has been recently proven to be very effective for a variety of cross-modal tasks, including the evaluation of captions generated from vision-and-language architectures. In this paper, we propose a new recipe for a contrastive-based evaluation metric for image captioning, namely Positive-Augmented Contrastive learning Score (PAC-S), that in a novel way unifies the learning of a contrastive visual-semantic space with the addition of generated images and text on curated data. Experiments spanning several datasets demonstrate that our new metric achieves the highest correlation with human judgments on both images and videos, outperforming existing referencebased metrics like CIDEr and SPICE and reference-free metrics like CLIP-Score. Finally, we test the system-level correlation of the proposed metric when considering popular image captioning approaches, and assess the impact of employing different cross-modal features. Our source code and trained models are publicly available at: https://github.com/aimagelab/pacscore. Sara Sarto, Manuele Barraco, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 4 |
| 2023 | With a Little Help from your own Past: Prototypical Memory Networks for Image CaptioningabstractImage captioning, like many tasks involving vision and language, currently relies on Transformer-based architectures for extracting the semantics in an image and translating it into linguistically coherent descriptions. Although successful, the attention operator only considers a weighted summation of projections of the current input sample, therefore ignoring the relevant semantic information which can come from the joint observation of other samples. In this paper, we devise a network which can perform attention over activations obtained while processing other training samples, through a prototypical memory model. Our memory models the distribution of past keys and values through the definition of prototype vectors which are both discriminative and compact. Experimentally, we assess the performance of the proposed model on the COCO dataset, in comparison with carefully designed baselines and state-of-the-art approaches, and by investigating the role of each of the proposed components. We demonstrate that our proposal can increase the performance of an encoder-decoder Transformer by 3.7 CIDEr points both when training in cross-entropy only and when fine-tuning with self-critical sequence training. Source code and trained models are available at: https://github.com/aimagelab/PMA-Net. Manuele Barraco, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICCV | 4 |
| 2023 | Embodied Agents for Efficient Exploration and Smart Scene DescriptionabstractThe development of embodied agents that can communicate with humans in natural language has gained increasing interest over the last years, as it facilitates the diffusion of robotic platforms in human-populated environments. As a step towards this objective, in this work, we tackle a setting for visual navigation in which an autonomous agent needs to explore and map an unseen indoor environment while portraying interesting scenes with natural language descriptions. To this end, we propose and evaluate an approach that combines recent advances in visual robotic exploration and image captioning on images generated through agent-environment interaction. Our approach can generate smart scene descriptions that maximize semantic knowledge of the environment and avoid repetitions. Further, such descriptions offer user-understandable insights into the robot's representation of the environment by high-lighting the prominent objects and the correlation between them as encountered during the exploration. To quantitatively assess the performance of the proposed approach, we also devise a specific score that takes into account both exploration and description skills. The experiments carried out on both photorealistic simulated environments and real-world ones demonstrate that our approach can effectively describe the robot's point of view during exploration, improving the human-friendly interpretability of its observations. Roberto Bigazzi, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara |
ICRA | 4 |
| 2023 | Let's ViCE! Mimicking Human Cognitive Behavior in Image Generation EvaluationabstractResearch in Image Generation has recently made significant progress, particularly boosted by the introduction of VisionLanguage models which are able to produce high-quality visual content based on textual inputs. Despite ongoing advancements in terms of generation quality and realism, no methodical frameworks have been defined yet to quantitatively measure the quality of the generated content and the adherence with the prompted requests: so far, only humanbased evaluations have been adopted for quality satisfaction and for comparing different generative methods. We introduce a novel automated method for Visual Concept Evaluation (ViCE), i.e. to assess consistency between a generated/edited image and the corresponding prompt/instructions, with a process inspired by the human cognitive behaviour. ViCE combines the strengths of Large Language Models (LLMs) and Visual Question Answering (VQA) into a unified pipeline, aiming to replicate the human cognitive process in quality assessment. This method outlines visual concepts, formulates image-specific verification questions, utilizes the Q&A system to investigate the image, and scores the combined outcome. Although this brave new hypothesis of mimicking humans in the image evaluation process is in its preliminary assessment stage, results are promising and open the door to a new form of automatic evaluation which could have significant impact as the image generation or the image target editing tasks become more and more sophisticated. Federico Betti 0001, Jacopo Staiano, Lorenzo Baraldi 0002, Lorenzo Baraldi 0001, Rita Cucchiara, Nicu Sebe |
ACM Multimedia | 4 |
| 2023 | Fully-attentive iterative networks for region-based controllable image and video captioningabstractControllable image captioning has recently gained attention as a way to increase the diversity and the applicability to real-world scenarios of image captioning algorithms. In this task, a captioner is conditioned on an external control signal, which needs to be followed during the generation of the caption. We aim to overcome the limitations of current controllable captioning methods by proposing a fully-attentive and iterative network that can generate grounded and controllable captions from a control signal given as a sequence of visual regions from the image. Our architecture is based on a set of novel attention operators, which take into account the hierarchical nature of the control signal, and is endowed with a decoder which explicitly focuses on each part of the control signal. We demonstrate the effectiveness of the proposed approach by conducting experiments on three datasets, where our model surpasses the performances of previous methods and achieves a new state of the art on both image and video controllable captioning. Marcella Cornia, Lorenzo Baraldi 0001, Ayellet Tal, Rita Cucchiara |
Comput. Vis. Image Underst. | 2 |
| 2023 | From Show to Tell: A Survey on Deep Learning-Based Image CaptioningabstractConnecting Vision and Language plays an essential role in Generative Intelligence. For this reason, large research efforts have been devoted to image captioning, i.e. describing images with syntactically and semantically meaningful sentences. Starting from 2015 the task has generally been addressed with pipelines composed of a visual encoder and a language model for text generation. During these years, both components have evolved considerably through the exploitation of object regions, attributes, the introduction of multi-modal connections, fully-attentive approaches, and BERT-like early-fusion strategies. However, regardless of the impressive results, research in image captioning has not reached a conclusive answer yet. This work aims at providing a comprehensive overview of image captioning approaches, from visual encoding and text generation to training strategies, datasets, and evaluation metrics. In this respect, we quantitatively compare many relevant state-of-the-art approaches to identify the most impactful technical innovations in architectures and training strategies. Moreover, many variants of the problem and its open challenges are discussed. The final goal of this work is to serve as a tool for understanding the existing literature and highlighting the future directions for a research area where Computer Vision and Natural Language Processing can find an optimal synergy. Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi 0001, Silvia Cascianelli, Giuseppe Fiameni, Rita Cucchiara |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Evaluating synthetic pre-Training for handwriting processing tasksabstractIn this work, we explore massive pre-training on synthetic word images for enhancing the performance on four benchmark downstream handwriting analysis tasks. To this end, we build a large synthetic dataset of word images rendered in several handwriting fonts, which offers a complete supervision signal. We use it to train a simple convolutional neural network (ConvNet) with a fully supervised objective. The vector representations of the images obtained from the pre-trained ConvNet can then be considered as encodings of the handwriting style. We exploit such representations for Writer Retrieval, Writer Identification, Writer Verification, and Writer Classification and demonstrate that our pre-training strategy allows extracting rich representations of the writers’ style that enable the aforementioned tasks with competitive results with respect to task-specific State-of-the-Art approaches. Vittorio Pippi, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara |
Pattern Recognit. Lett. | 3 |
| 2022 | ALADIN: Distilling Fine-grained Alignment Scores for Efficient Image-Text Matching and RetrievalabstractImage-text matching is gaining a leading role among tasks involving the joint understanding of vision and language. In literature, this task is often used as a pre-training objective to forge architectures able to jointly deal with images and texts. Nonetheless, it has a direct downstream application: cross-modal retrieval, which consists in finding images related to a given query text or vice-versa. Solving this task is of critical importance in cross-modal search engines. Many recent methods proposed effective solutions to the image-text matching problem, mostly using recent large vision-language (VL) Transformer networks. However, these models are often computationally expensive, especially at inference time. This prevents their adoption in large-scale cross-modal retrieval scenarios, where results should be provided to the user almost instantaneously. In this paper, we propose to fill in the gap between effectiveness and efficiency by proposing an ALign And DIstill Network (ALADIN). ALADIN first produces high-effective scores by aligning at fine-grained level images and texts. Then, it learns a shared embedding space – where an efficient kNN search can be performed – by distilling the relevance scores obtained from the fine-grained alignments. We obtained remarkable results on MS-COCO, showing that our method can compete with state-of-the-art VL Transformers while being almost 90 times faster. The code for reproducing our results is available at https://github.com/mesnico/ALADIN. Nicola Messina, Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi 0001, Fabrizio Falchi, Giuseppe Amato 0001, Rita Cucchiara |
CBMI | 4 |
| 2022 | Retrieval-Augmented Transformer for Image CaptioningabstractImage captioning models aim at connecting Vision and Language by providing natural language descriptions of input images. In the past few years, the task has been tackled by learning parametric models and proposing visual feature extraction advancements or by modeling better multi-modal connections. In this paper, we investigate the development of an image captioning approach with a kNN memory, with which knowledge can be retrieved from an external corpus to aid the generation process. Our architecture combines a knowledge retriever based on visual similarities, a differentiable encoder, and a kNN-augmented attention layer to predict tokens based on the past context and on text retrieved from the external memory. Experimental results, conducted on the COCO dataset, demonstrate that employing an explicit external memory can aid the generation process and increase caption quality. Our work opens up new avenues for improving image captioning models at larger scale. Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CBMI | 3 |
| 2022 | CaMEL: Mean Teacher Learning for Image CaptioningabstractDescribing images in natural language is a fundamental step towards the automatic modeling of connections between the visual and textual modalities. In this paper we present CaMEL, a novel Transformer-based architecture for image captioning. Our proposed approach leverages the interaction of two interconnected language models that learn from each other during the training phase. The interplay between the two language models follows a mean teacher learning paradigm with knowledge distillation. Experimentally, we assess the effectiveness of the proposed solution on the COCO dataset and in conjunction with different visual feature extractors. When comparing with existing proposals, we demonstrate that our model provides state-of-the-art caption quality with a significantly reduced number of parameters. According to the CIDEr metric, we obtain a new state of the art on COCO when training without using external data. The source code and trained models will be made publicly available at: https://github.com/aimagelab/camel. Manuele Barraco, Matteo Stefanini, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR | 5 |
| 2022 | The LAM Dataset: A Novel Benchmark for Line-Level Handwritten Text RecognitionabstractHandwritten Text Recognition (HTR) is an open problem at the intersection of Computer Vision and Natural Language Processing. The main challenges, when dealing with historical manuscripts, are due to the preservation of the paper support, the variability of the handwriting – even of the same author over a wide time-span – and the scarcity of data from ancient, poorly represented languages. With the aim of fostering the research on this topic, in this paper we present the Ludovico Antonio Muratori (LAM) dataset, a large line-level HTR dataset of Italian ancient manuscripts edited by a single author over 60 years. The dataset comes in two configurations: a basic splitting and a date-based splitting which takes into account the age of the author. The first setting is intended to study HTR on ancient documents in Italian, while the second focuses on the ability of HTR systems to recognize text written by the same writer in time periods for which training data are not available. For both configurations, we analyze quantitative and qualitative characteristics, also with respect to other line-level HTR benchmarks, and present the recognition performance of state-of-the-art HTR architectures. The dataset is available for download at https://aimagelab.ing.unimore.it/go/lam. Silvia Cascianelli, Vittorio Pippi, Martin Maarand, Marcella Cornia, Lorenzo Baraldi 0001, Christopher Kermorvant, Rita Cucchiara |
ICPR | 5 |
| 2022 | Spot the Difference: A Novel Task for Embodied Agents in Changing EnvironmentsabstractEmbodied AI is a recent research area that aims at creating intelligent agents that can move and operate inside an environment. Existing approaches in this field demand the agents to act in completely new and unexplored scenes. However, this setting is far from realistic use cases that instead require executing multiple tasks in the same environment. Even if the environment changes over time, the agent could still count on its global knowledge about the scene while trying to adapt its internal representation to the current state of the environment. To make a step towards this setting, we propose Spot the Difference: a novel task for Embodied AI where the agent has access to an outdated map of the environment and needs to recover the correct layout in a fixed time budget. To this end, we collect a new dataset of occupancy maps starting from existing datasets of 3D spaces and generating a number of possible layouts for a single environment. This dataset can be employed in the popular Habitat simulator and is fully compliant with existing methods that employ reconstructed occupancy maps during navigation. Furthermore, we propose an exploration policy that can take advantage of previous knowledge of the environment and identify changes in the scene faster and more effectively than existing agents. Experimental results show that the proposed architecture outperforms existing state-of-the-art models for exploration on this new setting. Federico Landi, Roberto Bigazzi, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR | 5 |
| 2022 | Boosting modern and historical handwritten text recognition with deformable convolutions
Silvia Cascianelli, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
Int. J. Document Anal. Recognit. | 3 |
| 2022 | A computational approach for progressive architecture shrinkage in action recognitionabstractAbstract Efficiency plays a key role in video understanding modeling, and developing more efficient spatiotemporal deep networks is a key ingredient for enabling their usage in production scenarios. In this work, we propose a methodology for reducing the computational complexity of a video understanding backbone while limiting the drop in accuracy caused by architectural changes. Our approach, named, Progressive Architecture Shrinkage, applies a sequence of reduction operators to the hyperparameters of a network to reduce its computational footprint. The choice of the sequence of operations is automatically optimized in a coordinate‐descent schema, and the approach transfers knowledge from both the initial network and previous stages of the shrinking process by employing a Knowledge Distillation and an adaptive fine‐tuning strategy. As each iteration of the shrinking algorithm requires to train a large‐scale video understanding network, we perform experiments on MARCONI 100—a supercomputer equipped with an IBM Power9 architecture and Volta NVIDIA GPUs. Experimental evaluations are conducted using two backbones and three different action recognition benchmarks. We show that, through our approach, high accuracy levels can be maintained while reducing the number of multiply–adds operations by four times with respect to the original architectures. Code will be made available. Matteo Tomei, Lorenzo Baraldi 0001, Giuseppe Fiameni, Simone Bronzin, Rita Cucchiara |
Softw. Pract. Exp. | 2 |
| 2022 | Matching Faces and Attributes Between the Artistic and the Real Domain: the PersonArt ApproachabstractIn this article, we present an approach for retrieving similar faces between the artistic and the real domain. The application we refer to is an interactive exhibition inside a museum, in which a visitor can take a photo of himself and search for a lookalike in the collection of paintings. The task requires not only to identify faces but also to extract discriminative features from artistic and photo-realistic images, tackling a significant domain shift. Our method integrates feature extraction networks which account for the aesthetic similarity of two faces and their correspondences in terms of semantic attributes. Also, it addresses the domain shift between realistic images and paintings by translating photo-realistic images into the artistic domain. Noticeably, by exploiting the same technique, our model does not need to rely on annotated data in the artistic domain. Experimental results are conducted on different paired datasets to show the effectiveness of the proposed solution in terms of identity and attribute preservation. The approach is also evaluated on unpaired settings and in combination with an interactive relevance feedback strategy. Finally, we show how the proposed algorithm has been implemented in a real showcase at the Gallerie Estensi museum in Italy, with the participation of more than 1,100 visitors in just three days. Marcella Cornia, Matteo Tomei, Lorenzo Baraldi 0001, Rita Cucchiara |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Assessing the Role of Boundary-Level Objectives in Indoor Semantic Segmentation
Roberto Amoroso, Lorenzo Baraldi 0001, Rita Cucchiara |
CAIP (1) | 2 |
| 2021 | Out of the Box: Embodied Navigation in the Real World
Roberto Bigazzi, Federico Landi, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara |
CAIP (1) | 5 |
| 2021 | Learning to Read L'Infinito: Handwritten Text Recognition with Synthetic Training Data
Silvia Cascianelli, Marcella Cornia, Lorenzo Baraldi 0001, Maria Ludovica Piazzi, Rosiana Schiuma, Rita Cucchiara |
CAIP (2) | 3 |
| 2021 | Learning to Select: A Fully Attentive Approach for Novel Object CaptioningabstractImage captioning models have lately shown impressive results when applied to standard datasets. Switching to real-life scenarios, however, constitutes a challenge due to the larger variety of visual concepts which are not covered in existing training sets. For this reason, novel object captioning (NOC) has recently emerged as a paradigm to test captioning models on objects which are unseen during the training phase. In this paper, we present a novel approach for NOC that learns to select the most relevant objects of an image, regardless of their adherence to the training set, and to constrain the generative process of a language model accordingly. Our architecture is fully-attentive and end-to-end trainable, also when incorporating constraints. We perform experiments on the held-out COCO dataset, where we demonstrate improvements over the state of the art, both in terms of adaptability to novel objects and caption quality. Marco Cagrandi, Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi 0001, Rita Cucchiara |
ICMR | 4 |
| 2021 | Multimodal attention networks for low-level vision-and-language navigation
Federico Landi, Lorenzo Baraldi 0001, Marcella Cornia, Massimiliano Corsini, Rita Cucchiara |
Comput. Vis. Image Underst. | 2 |
| 2021 | Video action detection by learning graph-based spatio-temporal interactionsabstractAction Detection is a complex task that aims to detect and classify human actions in video clips. Typically, it has been addressed by processing fine-grained features extracted from a video classification backbone. Recently, thanks to the robustness of object and people detectors, a deeper focus has been added on relationship modeling. Following this line, we propose a graph-based framework to learn high-level interactions between people and objects, in both space and time. In our formulation, spatio-temporal relationships are learned through self-attention on a multi-layer graph structure which can connect entities from consecutive clips, thus considering long-range spatial and temporal dependencies. The proposed module is backbone independent by design and does not require end-to-end training. Extensive experiments are conducted on the AVA dataset, where our model demonstrates state-of-the-art results and consistent improvements over baselines built with different backbones. Code is publicly available at https://github.com/aimagelab/STAGE_action_detection. Matteo Tomei, Lorenzo Baraldi 0001, Simone Calderara, Simone Bronzin, Rita Cucchiara |
Comput. Vis. Image Underst. | 2 |
| 2021 | Working Memory Connections for LSTM
Federico Landi, Lorenzo Baraldi 0001, Marcella Cornia, Rita Cucchiara |
Neural Networks | 2 |
| 2020 | Meshed-Memory Transformer for Image CaptioningabstractTransformer-based architectures represent the state of the art in sequence modeling tasks like machine translation and language understanding. Their applicability to multi-modal contexts like image captioning, however, is still largely under-explored. With the aim of filling this gap, we present M2- a Meshed Transformer with Memory for Image Captioning. The architecture improves both the image encoding and the language generation steps: it learns a multi-level representation of the relationships between image regions integrating learned a priori knowledge, and uses a mesh-like connectivity at decoding stage to exploit low- and high-level features. Experimentally, we investigate the performance of the M2Transformer and different fully-attentive models in comparison with recurrent ones. When tested on COCO, our proposal achieves a new state of the art in single-model and ensemble configurations on the "Karpathy" test split and on the online test server. We also assess its performances when describing objects unseen in the training set. Trained models and code for reproducing the experiments are publicly available at: https://github.com/aimagelab/meshed-memory-transformer. Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 3 |
| 2020 | Explore and Explain: Self-supervised Navigation and RecountingabstractEmbodied AI has been recently gaining attention as it aims to foster the development of autonomous and intelligent agents. In this paper, we devise a novel embodied setting in which an agent needs to explore a previously unknown environment while recounting what it sees during the path. In this context, the agent needs to navigate the environment driven by an exploration goal, select proper moments for description, and output natural language descriptions of relevant objects and scenes. Our model integrates a novel self-supervised exploration module with penalty, and a fully-attentive captioning model for explanation. Also, we investigate different policies for selecting proper moments for explanation, driven by information coming from both the environment and the navigation. Experiments are conducted on photorealistic environments from the Matterport3D dataset and investigate the navigation and explanation capabilities of the agent as well as the role of their interactions. Roberto Bigazzi, Federico Landi, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR | 5 |
| 2020 | Watch Your Strokes: Improving Handwritten Text Recognition with Deformable ConvolutionsabstractHandwritten Text Recognition (HTR) in free-layout pages is a valuable yet challenging task which aims to automatically understand handwritten texts. State-of-the-art approaches in this field usually encode input images with Convolutional Neural Networks, whose kernels are typically defined on a fixed grid and focus on all input pixels independently. However, this is in contrast with the sparse nature of handwritten pages, in which only pixels representing the ink of the writing are useful for the recognition task. Furthermore, the standard convolution operator is not explicitly designed to take into account the great variability in shape, scale, and orientation of handwritten characters. To overcome these limitations, we investigate the use of deformable convolutions for handwriting recognition. The kernel of this type of convolution deforms according to the content of the neighborhood, and can therefore be more adaptable to geometric variations and other deformations of the text. Experiments conducted on the IAM and RIMES datasets demonstrate that the use of deformable convolutions is a promising direction for the design of novel architectures for handwritten text recognition. Iulian Cojocaru, Silvia Cascianelli, Lorenzo Baraldi 0001, Massimiliano Corsini, Rita Cucchiara |
ICPR | 3 |
| 2020 | A Novel Attention-based Aggregation Function to Combine Vision and LanguageabstractThe joint understanding of vision and language has been recently gaining a lot of attention in both the Computer Vision and Natural Language Processing communities, with the emergence of tasks such as image captioning, image-text matching, and visual question answering. As both images and text can be encoded as sets or sequences of elements - like regions and words - proper reduction functions are needed to transform a set of encoded elements into a single response, like a classification or similarity score. In this paper, we propose a novel fully-attentive reduction method for vision and language. Specifically, our approach computes a set of scores for each element of each modality employing a novel variant of cross-attention, and performs a learnable and cross-modal reduction, which can be used for both classification and ranking. We test our approach on image-text matching and visual question answering, building fair comparisons with other reduction choices, on both COCO and VQA 2.0 datasets. Experimentally, we demonstrate that our approach leads to a performance increase on both tasks. Further, we conduct ablation studies to validate the role of each component of the approach. Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR | 3 |
| 2020 | RMS-Net: Regression and Masking for Soccer Event SpottingabstractThe recently proposed action spotting task consists in finding the exact timestamp in which an event occurs. This task fits particularly well for soccer videos, where events correspond to salient actions strictly defined by soccer rules (a goal occurs when the ball crosses the goal line). In this paper, we devise a lightweight and modular network for action spotting, which can simultaneously predict the event label and its temporal offset using the same underlying features. We enrich our model with two training strategies: the first one for data balancing and uniform sampling, the second for masking ambiguous frames and keeping the most discriminative visual cues. When tested on the SoccerNet dataset and using standard features, our full proposal exceeds the current state of the art by 3 Average-mAP points. Additionally, it reaches a gain of more than 10 Average-mAP points on the test set when fine-tuned in combination with a strong 2D backbone. Matteo Tomei, Lorenzo Baraldi 0001, Simone Calderara, Simone Bronzin, Rita Cucchiara |
ICPR | 2 |
| 2020 | SMArT: Training Shallow Memory-aware Transformers for Robotic ExplainabilityabstractThe ability to generate natural language explanations conditioned on the visual perception is a crucial step towards autonomous agents which can explain themselves and communicate with humans. While the research efforts in image and video captioning are giving promising results, this is often done at the expense of the computational requirements of the approaches, limiting their applicability to real contexts. In this paper, we propose a fully-attentive captioning algorithm which can provide state-of-the-art performances on language generation while restricting its computational demands. Our model is inspired by the Transformer model and employs only two Transformer layers in the encoding and decoding stages. Further, it incorporates a novel memory-aware encoding of image regions. Experiments demonstrate that our approach achieves competitive results in terms of caption quality while featuring reduced computational demands. Further, to evaluate its applicability on autonomous agents, we conduct experiments on simulated scenes taken from the perspective of domestic robots. Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICRA | 2 |
| 2020 | A unified cycle-consistent neural model for text and image retrieval
Marcella Cornia, Lorenzo Baraldi 0001, Hamed Rezazadegan Tavakoli, Rita Cucchiara |
Multim. Tools Appl. | 2 |
| 2020 | Explaining digital humanities by aligning images and textual descriptions
Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi 0001, Massimiliano Corsini, Rita Cucchiara |
Pattern Recognit. Lett. | 3 |
| 2020 | Spaghetti Labeling: Directed Acyclic Graphs for Block-Based Connected Components LabelingabstractConnected Components Labeling is an essential step of many Image Processing and Computer Vision tasks. Since the first proposal of a labeling algorithm, which dates back to the sixties, many approaches have optimized the computational load needed to label an image. In particular, the use of decision forests and state prediction have recently appeared as valuable strategies to improve performance. However, due to the overhead of the manual construction of prediction states and the size of the resulting machine code, the application of these strategies has been restricted to small masks, thus ignoring the benefit of using a block-based approach. In this paper, we combine a block-based mask with state prediction and code compression: the resulting algorithm is modeled as a Directed Rooted Acyclic Graph with multiple entry points, which is automatically generated without manual intervention. When tested on synthetic and real datasets, in comparison with optimized implementations of state-of-the-art algorithms, the proposed approach shows superior performance, surpassing the results obtained by all compared approaches in all settings. Federico Bolelli, Stefano Allegretti, Lorenzo Baraldi 0001, Costantino Grana |
IEEE Trans. Image Process. | 3 |
| 2019 | Embodied Vision-and-Language Navigation with Dynamic Convolutional Filters
Federico Landi, Lorenzo Baraldi 0001, Massimiliano Corsini, Rita Cucchiara |
BMVC | 2 |
| 2019 | A Deep-learning-based approach to VM behavior Identification in Cloud SystemsabstractCloud computing data centers are growing in size and complexity to the point where monitoring and management of the infrastructure become a challenge due to scalability issues. A possible approach to cope with the size of such data centers is to identify VMs exhibiting a similar behavior. Existing literature demonstrated that clustering together VMs that show a similar behavior may improve the scalability of both monitoring andmanagement of a data center. However, available techniques suffer from a trade-off between accuracy and time to achieve this result. Throughout this paper we propose a different approach where, instead of an unsupervised clustering, we rely on classifiers based on deep learning techniques to assigna newly deployed VMs to a cluster of already-known VMs. The two proposed classifiers, namely DeepConv and DeepFFT use a convolution neural network and (in the latter model) exploits Fast Fourier Transformation to classify the VMs. Our proposal is validated using a set of traces describing the behavior of VMs from a realcloud data center. The experiments compare our proposal with state-of-the-art solutions available in literature, demonstrating that our proposal achieve better performance. Furthermore, we show that our solution issignificantly faster than the alternatives as it can produce a perfect classification even with just a few samples of data, making our proposal viable also toclassify on-demand VMs that are characterized by a short life span. Matteo Stefanini, Riccardo Lancellotti, Lorenzo Baraldi 0001, Simone Calderara |
CLOSER | 3 |
| 2019 | Show, Control and Tell: A Framework for Generating Controllable and Grounded CaptionsabstractCurrent captioning approaches can describe images using black-box architectures whose behavior is hardly controllable and explainable from the exterior. As an image can be described in infinite ways depending on the goal and the context at hand, a higher degree of controllability is needed to apply captioning algorithms in complex scenarios. In this paper, we introduce a novel framework for image captioning which can generate diverse descriptions by allowing both grounding and controllability. Given a control signal in the form of a sequence or set of image regions, we generate the corresponding caption through a recurrent architecture which predicts textual chunks explicitly grounded on regions, following the constraints of the given control. Experiments are conducted on Flickr30k Entities and on COCO Entities, an extended version of COCO in which we add grounding annotations collected in a semi-automatic manner. Results demonstrate that our method achieves state of the art performances on controllable image captioning, in terms of caption quality and diversity. Code and annotations are publicly available at: https://github.com/aimagelab/show-control-and-tell. Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 2 |
| 2019 | Art2Real: Unfolding the Reality of Artworks via Semantically-Aware Image-To-Image TranslationabstractThe applicability of computer vision to real paintings and artworks has been rarely investigated, even though a vast heritage would greatly benefit from techniques which can understand and process data from the artistic domain. This is partially due to the small amount of annotated artistic data, which is not even comparable to that of natural images captured by cameras. In this paper, we propose a semantic-aware architecture which can translate artworks to photo-realistic visualizations, thus reducing the gap between visual features of artistic and realistic data. Our architecture can generate natural images by retrieving and learning details from real photos through a similarity matching strategy which leverages a weakly-supervised semantic understanding of the scene. Experimental results show that the proposed technique leads to increased realism and to a reduction in domain shift, which improves the performance of pre-trained architectures for classification, detection, and segmentation. Code is publicly available at: https://github.com/aimagelab/art2real. Matteo Tomei, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 3 |
| 2019 | M-VAD names: a dataset for video captioning with naming
Stefano Pini, Marcella Cornia, Federico Bolelli, Lorenzo Baraldi 0001, Rita Cucchiara |
Multim. Tools Appl. | 4 |
| 2018 | LAMV: Learning to Align and Match Videos With Kernelized Temporal LayersabstractThis paper considers a learnable approach for comparing and aligning videos. Our architecture builds upon and revisits temporal match kernels within neural networks: we propose a new temporal layer that finds temporal alignments by maximizing the scores between two sequences of vectors, according to a time-sensitive similarity metric parametrized in the Fourier domain. We learn this layer with a temporal proposal strategy, in which we minimize a triplet loss that takes into account both the localization accuracy and the recognition rate. We evaluate our approach on video alignment, copy detection and event retrieval. Our approach outperforms the state on the art on temporal video alignment and video copy detection datasets in comparable setups. It also attains the best reported results for particular event search, while precisely aligning videos. Lorenzo Baraldi 0001, Matthijs Douze, Rita Cucchiara, Hervé Jégou |
CVPR | 1 |
| 2018 | Aligning Text and Document Illustrations: Towards Visually Explainable Digital HumanitiesabstractWhile several approaches to bring vision and language together are emerging, none of them has yet addressed the digital humanities domain, which, nevertheless, is a rich source of visual and textual data. To foster research in this direction, we investigate the learning of visual-semantic embeddings for historical document illustrations, devising both supervised and semi-supervised approaches. We exploit the joint visual-semantic embeddings to automatically align illustrations and textual elements, thus providing an automatic annotation of the visual content of a manuscript. Experiments are performed on the Borso d'Este Holy Bible, one of the most sophisticated illuminated manuscript from the Renaissance, which we manually annotate aligning every illustration with textual commentaries written by experts. Experimental results quantify the domain shift between ordinary visual-semantic datasets and the proposed one, validate the proposed strategies, and devise future works on the same line. Lorenzo Baraldi 0001, Marcella Cornia, Costantino Grana, Rita Cucchiara |
ICPR | 1 |
| 2018 | Connected Components Labeling on DRAGsabstractIn this paper we introduce a new Connected Components Labeling (CCL) algorithm which exploits a novel approach to model decision problems as Directed Acyclic Graphs with a root, which will be called Directed Rooted Acyclic Graphs (DRAGs). This structure supports the use of sets of equivalent actions, as required by CCL, and optimally leverages these equivalences to reduce the number of nodes (decision points). The advantage of this representation is that a DRAG, differently from decision trees usually exploited by the state-of-the-art algorithms, will contain only the minimum number of nodes required to reach the leaf corresponding to a set of condition values. This combines the benefits of using binary decision trees with a reduction of the machine code size. Experiments show a consistent improvement of the execution time when the model is applied to CCL. Federico Bolelli, Lorenzo Baraldi 0001, Michele Cancilla, Costantino Grana |
ICPR | 2 |
| 2018 | A Hierarchical Quasi-Recurrent approach to Video CaptioningabstractVideo captioning has picked up a considerable attention thanks to the ability of Recurrent Neural Networks to extrapolate an encoded representation of the input video, and then use it to generate a description. We propose a recurrent encoding approach able to find and exploit the layered design of the video. Differently from the established encoder-decoder procedure, in which a video is repeatedly encoded by a recurrent layer, we employ revised Quasi-Recurrent Neural Networks. We further extend their basic cell with a boundary detector in order to recognize discontinuous segments boundaries and likewise correct the temporal connections of the encoding layer accordingly. Experiments, on the Montreal Video Annotation dataset, demonstrate that our approach can find suitable levels of representation of the input information, while reducing the computational requirements. Federico Bolelli, Lorenzo Baraldi 0001, Federico Pollastri, Costantino Grana |
IPAS | 2 |
| 2018 | Predicting Human Eye Fixations via an LSTM-Based Saliency Attentive ModelabstractData-driven saliency has recently gained a lot of attention thanks to the use of Convolutional Neural Networks for predicting gaze fixations. In this paper we go beyond standard approaches to saliency prediction, in which gaze maps are computed with a feed-forward network, and present a novel model which can predict accurate saliency maps by incorporating neural attentive mechanisms. The core of our solution is a Convolutional LSTM that focuses on the most salient regions of the input image to iteratively refine the predicted saliency map. Additionally, to tackle the center bias typical of human eye fixations, our model can learn a set of prior maps generated with Gaussian functions. We show, through an extensive evaluation, that the proposed architecture outperforms the current state of the art on public saliency prediction datasets. We further study the contribution of each key component to demonstrate their robustness on different scenarios. Marcella Cornia, Lorenzo Baraldi 0001, Giuseppe Serra 0001, Rita Cucchiara |
IEEE Trans. Image Process. | 2 |
| 2018 | Paying More Attention to Saliency: Image Captioning with Saliency and Context AttentionabstractImage captioning has been recently gaining a lot of attention thanks to the impressive achievements shown by deep captioning architectures, which combine Convolutional Neural Networks to extract image representations and Recurrent Neural Networks to generate the corresponding captions. At the same time, a significant research effort has been dedicated to the development of saliency prediction models, which can predict human eye fixations. Even though saliency information could be useful to condition an image captioning architecture, by providing an indication of what is salient and what is not, research is still struggling to incorporate these two techniques. In this work, we propose an image captioning approach in which a generative recurrent neural network can focus on different parts of the input image during the generation of the caption, by exploiting the conditioning given by a saliency prediction model on which parts of the image are salient and which are contextual. We show, through extensive quantitative and qualitative experiments on large-scale datasets, that our model achieves superior performance with respect to captioning baselines with and without saliency and to different state-of-the-art approaches combining saliency and captioning. Marcella Cornia, Lorenzo Baraldi 0001, Giuseppe Serra 0001, Rita Cucchiara |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | Hierarchical Boundary-Aware Neural Encoder for Video CaptioningabstractThe use of Recurrent Neural Networks for video captioning has recently gained a lot of attention, since they can be used both to encode the input video and to generate the corresponding description. In this paper, we present a recurrent video encoding scheme which can discover and leverage the hierarchical structure of the video. Unlike the classical encoder-decoder approach, in which a video is encoded continuously by a recurrent layer, we propose a novel LSTM cell which can identify discontinuity points between frames or segments and modify the temporal connections of the encoding layer accordingly. We evaluate our approach on three large-scale datasets: the Montreal Video Annotation dataset, the MPII Movie Description dataset and the Microsoft Video Description Corpus. Experiments show that our approach can discover appropriate hierarchical representations of input videos and improve the state of the art results on movie description datasets. Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
CVPR | 1 |
| 2017 | Modeling multimodal cues in a deep learning-based framework for emotion recognition in the wildabstractIn this paper, we propose a multimodal deep learning architecture for emotion recognition in video regarding our participation to the audio-video based sub-challenge of the Emotion Recognition in the Wild 2017 challenge. Our model combines cues from multiple video modalities, including static facial features, motion patterns related to the evolution of the human expression over time, and audio information. Specifically, it is composed of three sub-networks trained separately: the first and second ones extract static visual features and dynamic patterns through 2D and 3D Convolutional Neural Networks (CNN), while the third one consists in a pretrained audio network which is used to extract useful deep acoustic signals from video. In the audio branch, we also apply Long Short Term Memory (LSTM) networks in order to capture the temporal evolution of the audio features. To identify and exploit possible relationships among different modalities, we propose a fusion network that merges cues from the different modalities in one representation. The proposed architecture outperforms the challenge baselines (38.81 % and 40.47 %): we achieve an accuracy of 50.39 % and 49.92 % respectively on the validation and the testing data. Stefano Pini, Olfa Ben Ahmed, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara, Benoit Huet |
ICMI | 4 |
| 2017 | Recognizing and Presenting the Storytelling Video Structure With Deep Multimodal NetworksabstractIn this paper, we propose a novel scene detection algorithm which employs semantic, visual, textual, and audio cues. We also show how the hierarchical decomposition of the storytelling video structure can improve retrieval results presentation with semantically and aesthetically effective thumbnails. Our method is built upon two advancements of the state of the art: first is semantic feature extraction which builds video-specific concept detectors; and second is multimodal feature embedding learning that maps the feature vector of a shot to a space in which the Euclidean distance has task specific semantic properties. The proposed method is able to decompose the video in annotated temporal segments which allow us for a query specific thumbnail extraction. Extensive experiments are performed on different data sets to demonstrate the effectiveness of our algorithm. An in-depth discussion on how to deal with the subjectivity of the task is conducted and a strategy to overcome the problem is suggested. Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
IEEE Trans. Multim. | 1 |
| 2016 | Optimized Connected Components Labeling with Pixel Prediction
Costantino Grana, Lorenzo Baraldi 0001, Federico Bolelli |
ACIVS | 2 |
| 2016 | Historical document digitization through layout analysis and deep content classificationabstractDocument layout segmentation and recognition is an important task in the creation of digitized documents collections, especially when dealing with historical documents. This paper presents an hybrid approach to layout segmentation as well as a strategy to classify document regions, which is applied to the process of digitization of an historical encyclopedia. Our layout analysis method merges a classic top-down approach and a bottom-up classification process based on local geometrical features, while regions are classified by means of features extracted from a Convolutional Neural Network merged in a Random Forest classifier. Experiments are conducted on the first volume of the “Enciclopedia Treccani”, a large dataset containing 999 manually annotated pages from the historical Italian encyclopedia. Andrea Corbelli, Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
ICPR | 2 |
| 2016 | A deep multi-level network for saliency predictionabstractThis paper presents a novel deep architecture for saliency prediction. Current state of the art models for saliency prediction employ Fully Convolutional networks that perform a non-linear combination of features extracted from the last convolutional layer to predict saliency maps. We propose an architecture which, instead, combines features extracted at different levels of a Convolutional Neural Network (CNN). Our model is composed of three main blocks: a feature extraction CNN, a feature encoding network, that weights low and high level feature maps, and a prior learning network. We compare our solution with state of the art saliency models on two public benchmarks datasets. Results show that our model outperforms under all evaluation metrics on the SALICON dataset, which is currently the largest public dataset for saliency prediction, and achieves competitive results on the MIT300 benchmark. Code is available at https://github.com/marcellacornia/mlnet. Marcella Cornia, Lorenzo Baraldi 0001, Giuseppe Serra 0001, Rita Cucchiara |
ICPR | 2 |
| 2016 | YACCLAB - Yet Another Connected Components Labeling BenchmarkabstractThe problem of labeling the connected components (CCL) of a binary image is well-defined and several proposals have been presented in the past. Since an exact solution to the problem exists and should be mandatory provided as output, algorithms mainly differ on their execution speed. In this paper, we propose and describe YACCLAB, Yet Another Connected Components Labeling Benchmark. Together with a rich and varied dataset, YACCLAB contains an open source platform to test new proposals and to compare them with publicly available competitors. Textual and graphical outputs are automatically generated for three kinds of test, which analyze the methods from different perspectives. The fairness of the comparisons is guaranteed by running on the same system and over the same datasets. Examples of usage and the corresponding comparisons among state-of-the-art techniques are reported to confirm the potentiality of the benchmark. Costantino Grana, Federico Bolelli, Lorenzo Baraldi 0001, Roberto Vezzani |
ICPR | 3 |
| 2016 | Scene-driven Retrieval in Edited Videos using Aesthetic and Semantic Deep FeaturesabstractThis paper presents a novel retrieval pipeline for video collections, which aims to retrieve the most significant parts of an edited video for a given query, and represent them with thumbnails which are at the same time semantically meaningful and aesthetically remarkable. Videos are first segmented into coherent and story-telling scenes, then a retrieval algorithm based on deep learning is proposed to retrieve the most significant scenes for a textual query. A ranking strategy based on deep features is finally used to tackle the problem of visualizing the best thumbnail. Qualitative and quantitative experiments are conducted on a collection of edited videos to demonstrate the effectiveness of our approach. Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
ICMR | 1 |
| 2016 | A Browsing and Retrieval System for Broadcast Videos using Scene Detection and Automatic AnnotationabstractThis paper presents a novel video access and retrieval system for edited videos. The key element of the proposal is that videos are automatically decomposed into semantically coherent parts (called scenes) to provide a more manageable unit for browsing, tagging and searching. The system features an automatic annotation pipeline, with which videos are tagged by exploiting both the transcript and the video itself. Scenes can also be retrieved with textual queries; the best thumbnail for a query is selected according to both semantics and aesthetics criteria. Lorenzo Baraldi 0001, Costantino Grana, Alberto Messina, Rita Cucchiara |
ACM Multimedia | 1 |
| 2015 | Shot and Scene Detection via Hierarchical Clustering for Re-using Broadcast Video
Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
CAIP (1) | 1 |
| 2015 | Scene segmentation using temporal clustering for accessing and re-using broadcast videoabstractScene detection is a fundamental tool for allowing effective video browsing and re-using. In this paper we present a model that automatically divides videos into coherent scenes, which is based on a novel combination of local image descriptors and temporal clustering techniques. Experiments are performed to demonstrate the effectiveness of our approach, by comparing our algorithm against two recent proposals for automatic scene segmentation. We also propose improved performance measures that aim to reduce the gap between numerical evaluation and expected results. Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
ICME | 1 |
| 2015 | A Deep Siamese Network for Scene Detection in Broadcast VideosabstractWe present a model that automatically divides broadcast videos into coherent scenes by learning a distance measure between shots. Experiments are performed to demonstrate the effectiveness of our approach by comparing our algorithm against recent proposals for automatic scene segmentation. We also propose an improved performance measure that aims to reduce the gap between numerical evaluation and expected results, and propose and release a new benchmark dataset. Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
ACM Multimedia | 1 |