VLDB 2026 Research / reviewers in the wild / expert
Marcella Cornia
dblp:186/8196
· DBLP profile ↗
70ranked-venue papers
12as first author
55since 2021 · last 2026
0000-0001-9640-9385ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 8 first-author · 42 since 2021Artificial intelligence and machine learning · 50 · 7 first-author · 40 since 2021Computer networks · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
Silvia Cappelletti, Tobia Poppi, Samuele Poppi, Diego Garcia-Olano, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR (3) | 6 |
| 2026 | GramSR: Visual Feature Conditioning for Diffusion-Based Super-Resolution
Fabio D'Oronzio, Federico Putamorsi, Leonardo Zini, Marcella Cornia, Lorenzo Baraldi 0001 |
ICPR (7) | 4 |
| 2026 | RaTA-Tool: Retrieval-Based Tool Selection with Multimodal Large Language Models
Gabriele Mattioli, Evelyn Turri, Sara Sarto, Lorenzo Baraldi 0002, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR (7) | 5 |
| 2026 | From Captioning to Multimodal Reasoning: The Evolution of Vision-Language Research in the Era of Multimodal LLMsabstractOver the past decade, vision-language research has undergone a major evolution. Early image captioning systems primarily focused on generating concise textual descriptions of visual content through task-specific architectures trained on relatively constrained datasets. While successful in narrow settings, these approaches provided only limited reasoning capabilities and shallow alignment between visual and linguistic representations. The emergence of large-scale multimodal pretraining and Multimodal Large Language Models (MLLMs) has fundamentally reshaped this landscape. Modern multimodal systems are no longer restricted to visual description, but instead support open-ended interaction, multimodal generation, and increasingly sophisticated reasoning capabilities across images, text, video, and diverse modalities. This transition has progressively shifted the focus of the field from perception-oriented architectures toward general-purpose multimodal foundation models. Despite these advances, current MLLMs still exhibit significant limitations in knowledge-intensive and reasoning-heavy scenarios. In particular, tasks requiring external, long-tail, or continuously evolving knowledge remain highly challenging, as model predictions are often constrained by information implicitly encoded within model parameters. These limitations have recently motivated growing interest in knowledge-intensive visual question answering, retrieval-augmented multimodal architectures, adaptive knowledge integration, and reasoning-aware generation strategies. At the same time, the increasing scale and complexity of multimodal foundation models raise important challenges concerning hallucination mitigation, interpretability, and trustworthiness. This talk will discuss the major conceptual and technical paradigm shifts that have shaped the evolution of vision-language research, from classical captioning systems to retrieval-augmented MLLMs, highlighting emerging directions toward knowledge-aware, reasoning-centric, and trustworthy multimodal AI systems. Marcella Cornia |
ICMR | 1 |
| 2026 | Sketch2Stitch: GANs for Abstract Sketch-Based Dress SynthesisabstractIn the realm of creative expression, not everyone possesses the gift of effortlessly translating their imaginative visions into flawless sketches. More often than not, the outcome resembles an abstract, perhaps even slightly distorted representation. The art of producing impeccable sketches is not only challenging but also a time-consuming process. Our work is the first of this kind in transforming abstract, sometimes deformed garment sketches into photorealistic catalog images, to empower the everyday individual to become their own fashion designer. We create Sketch2Stitch, a dataset featuring over 65,000 abstract sketch images generated from garments of Dress Code [40] and VITON HD [6], two benchmark datasets in the virtual try-on task. Sketch2Stitch is the first dataset in the literature to provide abstract sketches in the fashion domain. We propose a StyleGAN-based generative framework that bridges freehand sketching with photorealistic garment synthesis. We demonstrate that our framework allows users to sketch rough outlines and optionally provide color hints, producing realistic designs in seconds. Experimental results demonstrate, both quantitatively and qualitatively, that the proposed framework achieves superior performance against various baselines and existing methods on both subsets of our dataset. Our work highlights a pathway toward AI-assisted fashion design tools, democratizing garment ideation for students, independent designers, and casual creators. Faizan Farooq Khan, Eslam Abdelrahman, Davide Morelli, Marcella Cornia, Rita Cucchiara, Mohamed Elhoseiny 0001 |
WACV | 4 |
| 2026 | Multimodal-Conditioned Latent Diffusion Models for Fashion Image EditingabstractFashion illustration is a crucial medium for designers to convey their creative vision and transform design concepts into tangible representations that showcase the interplay between clothing and the human body. In the context of fashion design, computer vision techniques have the potential to enhance and streamline the design process. Departing from prior research primarily focused on virtual try-on, this article tackles the task of multimodal-conditioned fashion image editing. Our approach aims to generate human-centric fashion images guided by multimodal prompts, including text, human body poses, garment sketches, and fabric textures. To address this problem, we propose extending latent diffusion models to incorporate these multiple modalities and modifying the structure of the denoising network, taking multimodal prompts as input. To condition the proposed architecture on fabric textures, we employ textual inversion techniques and let diverse cross-attention layers of the denoising network attend to textual and texture information, thus incorporating different granularity conditioning details. Given the lack of datasets for the task, we extend two existing fashion datasets, Dress Code and VITON-HD, with multimodal annotations. Experimental evaluations demonstrate the effectiveness of our proposed approach in terms of realism and coherence concerning the provided multimodal inputs. Alberto Baldrati, Davide Morelli, Marcella Cornia, Marco Bertini 0001, Rita Cucchiara |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document RetrievalabstractCross-modal retrieval is gaining increasing efficacy and interest from the research community, thanks to large-scale training, novel architectural and learning designs, and its application in LLMs and multimodal LLMs. In this paper, we move a step forward and design an approach that allows for multimodal queries – composed of both an image and a text – and can search within collections of multi-modal documents, where images and text are interleaved. Our model, ReT, employs multi-level representations extracted from different layers of both visual and textual backbones, both at the query and document side. To allow for multi-level and cross-modal understanding and feature extraction, ReT employs a novel Transformer-based recurrent cell that integrates both textual and visual features at different layers, and leverages sigmoidal gates inspired by the classical design of LSTMs. Extensive experiments on M2KR and M-BEIR benchmarks show that ReT achieves state-of-the-art performance across diverse settings. Our source code and trained models are publicly available at: https://github.com/aimagelab/ReT. Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 3 |
| 2025 | Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question AnsweringabstractMultimodal LLMs (MLLMs) are the natural extension of large language models to handle multimodal inputs, combining text and image data. They have recently garnered attention due to their capability to address complex tasks involving both modalities. However, their effectiveness is limited to the knowledge acquired during training, which restricts their practical utility. In this work, we introduce a novel method to enhance the adaptability of MLLMs by integrating external knowledge sources. Our proposed model, Reflective LLaVA (ReflectiVA), utilizes reflective tokens to dynamically determine the need for external knowledge and predict the relevance of information retrieved from an external database. Tokens are trained following a two-stage two-model training recipe. This ultimately enables the MLLM to manage external knowledge while preserving fluency and performance on tasks where external knowledge is not needed. Through our experiments, we demonstrate the efficacy of ReflectiVA for knowledge-based visual question answering, highlighting its superior performance compared to existing methods. Source code and trained models are publicly available at https://aimagelab.github.io/ReflectiVA. Federico Cocchi, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 3 |
| 2025 | Generating Synthetic Data with Large Language Models for Low-Resource Sentence Retrieval
Davide Caffagni, Federico Cocchi, Anna Mambelli, Fabio Tutrone, Marco Zanella, Marcella Cornia, Rita Cucchiara |
TPDL | 6 |
| 2025 | What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language ModelsabstractInstruction-based image editing models offer increased personalization opportunities in generative tasks. However, properly evaluating their results is challenging, and most of the existing metrics lag in terms of alignment with human judgment and explainability. To tackle these issues, we introduce DICE (DIfference Coherence Estimator), a model designed to detect localized differences between the original and the edited image and to assess their relevance to the given modification request. DICE consists of two key components: a difference detector and a coherence estimator, both built on an autoregressive Multimodal Large Language Model (MLLM) and trained using a strategy that leverages self-supervision, distillation from inpainting networks, and full supervision. Through extensive experiments, we evaluate each stage of our pipeline, comparing different MLLMs within the proposed framework. We demonstrate that DICE effectively identifies coherent edits, effectively evaluating images generated by different editing models with a strong correlation with human judgment. We publicly release our source code, models, and data. Lorenzo Baraldi 0001, Davide Bucciarelli, Federico Betti 0001, Marcella Cornia, Lorenzo Baraldi 0002, Nicu Sebe, Rita Cucchiara |
ICCV | 4 |
| 2025 | Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary SegmentationabstractOpen-Vocabulary Segmentation (OVS) aims at segmenting images from free-form textual concepts without predefined training classes. While existing vision-language models such as CLIP can generate segmentation masks by leveraging coarse spatial information from Vision Transformers, they face challenges in spatial localization due to their global alignment of image and text features. Conversely, self-supervised visual models like DINO excel in fine-grained visual encoding but lack integration with language. To bridge this gap, we present Talk2DINO, a novel hybrid approach that combines the spatial accuracy of DINOv2 with the language understanding of CLIP. Our approach aligns the textual embeddings of CLIP to the patch-level features of DINOv2 through a learned mapping function without the need to fine-tune the underlying backbones. At training time, we exploit the attention maps of DINOv2 to selectively align local visual patches with textual embeddings. We show that the powerful semantic and localization abilities of Talk2DINO can enhance the segmentation process, resulting in more natural and less noisy segmentations, and that our approach can also effectively distinguish foreground objects from the background. Experimental results demonstrate that Talk2DINO achieves state-of-the-art performance across several unsupervised OVS benchmarks. Source code and models are publicly available at: https://lorebianchi98.github.io/Talk2DINO/. Luca Barsellotti, Lorenzo Bianchi 0001, Nicola Messina, Fabio Carrara, Marcella Cornia, Lorenzo Baraldi 0001, Fabrizio Falchi, Rita Cucchiara |
ICCV | 5 |
| 2025 | Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
Giuseppe Cartella, Vittorio Cuculo, Alessandro D'Amelio, Marcella Cornia, Giuseppe Boccignone, Rita Cucchiara |
ICCV | 4 |
| 2025 | MISSRAG: Addressing the Missing Modality Challenge in Multimodal Large Language Models
Vittorio Pipoli, Alessia Saporita, Federico Bolelli, Marcella Cornia, Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara, Elisa Ficarra |
ICCV | 4 |
| 2025 | Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future PerspectivesabstractThe evaluation of machine-generated captions is a complex and evolving challenge. With the advent of Multimodal Large Language Models (MLLMs), image captioning has become a core task, increasing the need for robust and reliable evaluation metrics. This survey provides a comprehensive overview of advancements in image captioning evaluation, analyzing the evolution, strengths, and limitations of existing metrics. We assess these metrics across multiple dimensions, including correlation with human judgment, ranking accuracy, and sensitivity to hallucinations. Additionally, we explore the challenges posed by the longer and more detailed captions generated by MLLMs and examine the adaptability of current metrics to these stylistic variations. Our analysis highlights some limitations of standard evaluation approaches and suggests promising directions for future research in image captioning assessment. For a comprehensive overview of captioning evaluation refer to our project page available at https://github.com/aimagelab/awesome-captioning-evaluation. Sara Sarto, Marcella Cornia, Rita Cucchiara |
IJCAI | 2 |
| 2025 | Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented GenerationabstractIn recent years, the fashion industry has increasingly adopted AI technologies to enhance customer experience, driven by the proliferation of e-commerce platforms and virtual applications. Among the various tasks, virtual try-on and multimodal fashion image editing–which utilizes diverse input modalities such as text, garment sketches, and body poses–have become a key area of research. Diffusion models have emerged as a leading approach for such generative tasks, offering superior image quality and diversity. However, most existing virtual try-on methods rely on having a specific garment input, which is often impractical in real-world scenarios where users may only provide textual specifications. To address this limitation, in this work we introduce Fashion Retrieval-Augmented Generation (Fashion-RAG), a novel method that enables the customization of fashion items based on user preferences provided in textual form. Our approach retrieves multiple garments that match the input specifications and generates a personalized image by incorporating attributes from the retrieved items. To achieve this, we employ textual inversion techniques, where retrieved garment images are projected into the textual embedding space of the Stable Diffusion text encoder, allowing seamless integration of retrieved elements into the generative process. Experimental results on the Dress Code dataset demonstrate that Fashion-RAG outperforms existing methods both qualitatively and quantitatively, effectively capturing fine-grained visual details from retrieved garments. To the best of our knowledge, this is the first work to introduce a retrieval-augmented generation approach specifically tailored for multimodal fashion image editing. Fulvio Sanguigni, Davide Morelli, Marcella Cornia, Rita Cucchiara |
IJCNN | 3 |
| 2025 | TPP-Gaze: Modelling Gaze Dynamics in Space and Time with Neural Temporal Point ProcessesabstractAttention guides our gaze to fixate the proper location of the scene and holds it in that location for the de-served amount of time given current processing demands, before shifting to the next one. As such, gaze deploy-ment crucially is a temporal process. Existing computational models have made significant strides in predicting spatial aspects of observer's visual scanpaths (where to look), while often putting on the background the tempo-ral facet of attention dynamics (when). In this paper we present TPP-Gaze, a novel and principled approach to model scanpath dynamics based on Neural Temporal Point Process (TPP), that Jointly learns the temporal dynamics of fixations position and duration, integrating deep learning methodologies with point process theory. We conduct ex-tensive experiments across five publicly available datasets. Our results show the overall superior performance of the proposed model compared to state-of-the-art approaches. Source code and trained models are publicly available at: https://github.com/phuselab/tppgaze. Alessandro D'Amelio, Giuseppe Cartella, Vittorio Cuculo, Manuele Lucchi, Marcella Cornia, Rita Cucchiara, Giuseppe Boccignone |
WACV | 5 |
| 2025 | Semantically Conditioned Prompts for Visual Recognition Under Missing Modality ScenariosabstractThis paper tackles the domain of multimodal prompting for visual recognition, specifically when dealing with missing modalities through multimodal Transformers. It presents two main contributions: (i) we introduce a novel prompt learning module which is designed to produce sample-specific prompts and (ii) we show that modalityagnostic prompts can effectively adjust to diverse missing modality scenarios. Our model, termed SCP, exploits the semantic representation of available modalities to query a learnable memory bank, which allows the generation of prompts based on the semantics of the input. Notably, SCP distinguishes itself from existing methodologies for its capacity of self-adjusting to both the missing modality scenario and the semantic context of the input, without prior knowledge about the specific missing modality and the number of modalities. Through extensive experiments, we show the effectiveness of the proposed prompt learning framework and demonstrate enhanced performance and robustness across a spectrum of missing modality cases. Our source code is available at https://github.com/vittoriopipoli/SCP_WACV2025. Vittorio Pipoli, Federico Bolelli, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara, Elisa Ficarra |
WACV | 4 |
| 2025 | Learning to mask and permute visual tokens for Vision Transformer pre-trainingabstractThe use of self-supervised pre-training has emerged as a promising approach to enhance the performance of many different visual tasks. In this context, recent approaches have employed the Masked Image Modeling paradigm, which pre-trains a backbone by reconstructing visual tokens associated with randomly masked image patches. This masking approach, however, introduces noise into the input data during pre-training, leading to discrepancies that can impair performance during the fine-tuning phase. Furthermore, input masking neglects the dependencies between corrupted patches, increasing the inconsistencies observed in downstream fine-tuning tasks. To overcome these issues, we propose a new self-supervised pre-training approach, named Masked and Permuted Vision Transformer (MaPeT), that employs autoregressive and permuted predictions to capture intra-patch dependencies. In addition, MaPeT employs auxiliary positional information to reduce the disparity between the pre-training and fine-tuning phases. In our experiments, we employ a fair setting to ensure reliable and meaningful comparisons and conduct investigations on multiple visual tokenizers, including our proposed k -CLIP which directly employs discretized CLIP features. Our results demonstrate that MaPeT achieves competitive performance on ImageNet, compared to baselines and competitors under the same model setting. We release an implementation of our code and models at https://github.com/aimagelab/MaPeT . Lorenzo Baraldi 0002, Roberto Amoroso, Marcella Cornia, Lorenzo Baraldi 0001, Andrea Pilzer, Rita Cucchiara |
Comput. Vis. Image Underst. | 3 |
| 2025 | Positive-Augmented Contrastive Learning for Vision-and-Language Evaluation and Training
Sara Sarto, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
Int. J. Comput. Vis. | 3 |
| 2025 | Augmenting and mixing Transformers with synthetic data for image captioningabstractImage captioning has attracted significant attention within the Computer Vision and Multimedia research domains, resulting in the development of effective methods for generating natural language descriptions of images. Concurrently, the rise of generative models has facilitated the production of highly realistic and high-quality images, particularly through recent advancements in latent diffusion models. In this paper, we propose to leverage the recent advances in Generative AI and create additional training data that can be effectively used to boost the performance of an image captioning model. Specifically, we combine real images with their synthetic counterparts generated by Stable Diffusion using a Mixup data augmentation technique to create novel training examples. Extensive experiments on the COCO dataset demonstrate the effectiveness of our solution in comparison to different baselines and state-of-the-art methods and validate the benefits of using synthetic data to augment the training stage of an image captioning model and improve the quality of the generated captions. Source code and trained models are publicly available at: https://github.com/aimagelab/synthcap_pp . Davide Caffagni, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
Image Vis. Comput. | 2 |
| 2025 | Parents and Children: Distinguishing Multimodal Deepfakes from Natural ImagesabstractRecent advancements in diffusion models have enabled the generation of realistic deepfakes from textual prompts in natural language. While these models have numerous benefits across various sectors, they have also raised concerns about the potential misuse of fake images and cast new pressures on fake image detection. In this work, we pioneer a systematic study on deepfake detection generated by state-of-the-art diffusion models. Firstly, we conduct a comprehensive analysis of the performance of contrastive and classification-based visual features, respectively, extracted from CLIP-based models and ResNet or Vision Transformer (ViT)-based architectures trained on image classification datasets. Our results demonstrate that fake images share common low-level cues, which render them easily recognizable. Further, we devise a multimodal setting wherein fake images are synthesized by different textual captions, which are used as seeds for a generator. Under this setting, we quantify the performance of fake detection strategies and introduce a contrastive-based disentangling method that lets us analyze the role of the semantics of textual descriptions and low-level perceptual cues. Finally, we release a new dataset, called COCOFake, containing about 1.2 million images generated from the original COCO image–caption pairs using two recent text-to-image diffusion models, namely Stable Diffusion v1.4 and v2.0. Roberto Amoroso, Davide Morelli, Marcella Cornia, Lorenzo Baraldi 0001, Alberto Del Bimbo, Rita Cucchiara |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
Nicholas Moratelli, Davide Caffagni, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
BMVC | 3 |
| 2024 | Training-Free Open-Vocabulary Segmentation with Offline Diffusion-Augmented Prototype GenerationabstractOpen-vocabulary semantic segmentation aims at segmenting arbitrary categories expressed in textual form. Pre-vious works have trained over large amounts of image-caption pairs to enforce pixel-level multimodal alignments. However, captions provide global information about the semantics of a given image but lack direct localization of individual concepts. Further, training on large-scale datasets inevitably brings significant computational costs. In this paper, we propose FreeDA, a training-free diffusion-augmented method for open-vocabulary semantic segmentation, which leverages the ability of diffusion models to visually localize generated concepts and local-global similarities to match class-agnostic regions with semantic classes. Our approach involves an offline stage in which textual-visual reference embeddings are collected, starting from a large set of captions and leveraging visual and semantic contexts. At test time, these are queried to support the visual matching process, which is carried out by jointly considering class-agnostic regions and global semantic similarities. Extensive analyses demonstrate that FreeDA achieves state-of-the-art performance on five datasets, surpassing previous methods by more than 7.0 average points in terms of mIoU and without requiring any training. Our source code is available at aimagelab.github. io/freeda. Luca Barsellotti, Roberto Amoroso, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 3 |
| 2024 | Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities
Federico Cocchi, Marcella Cornia, Lorenzo Baraldi 0001, Alessandro Nicolosi, Rita Cucchiara |
ECCV (63) | 2 |
| 2024 | Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models
Samuele Poppi, Tobia Poppi, Federico Cocchi, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ECCV (53) | 4 |
| 2024 | BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues
Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ECCV (78) | 2 |
| 2024 | Adapt to Scarcity: Few-Shot Deepfake Detection via Low-Rank Adaptation
Silvia Cappelletti, Lorenzo Baraldi 0002, Federico Cocchi, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR (21) | 4 |
| 2024 | Fluent and Accurate Image Captioning with a Self-trained Reward Model
Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR (18) | 2 |
| 2024 | Unlearning Vision Transformers Without Retaining Data via Low-Rank Decompositions
Samuele Poppi, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR (3) | 3 |
| 2024 | Trends, Applications, and Challenges in Human Attention Modelling
Giuseppe Cartella, Marcella Cornia, Vittorio Cuculo, Alessandro D'Amelio, Dario Zanca, Giuseppe Boccignone, Rita Cucchiara |
IJCAI | 2 |
| 2024 | Personalized Instance-based Navigation Toward User-Specific Objects in Realistic EnvironmentsabstractIn the last years, the research interest in visual navigation towards objects in indoor environments has grown significantly. This growth can be attributed to the recent availability of large navigation datasets in photo-realistic simulated environments, like Gibson and Matterport3D. However, the navigation tasks supported by these datasets are often restricted to the objects present in the environment at acquisition time. Also, they fail to account for the realistic scenario in which the target object is a user-specific instance that can be easily confused with similar objects and may be found in multiple locations within the environment. To address these limitations, we propose a new task denominated Personalized Instance-based Navigation (PIN), in which an embodied agent is tasked with locating and reaching a specific personal object by distinguishing it among multiple instances of the same category. The task is accompanied by PInNED, a dedicated new dataset composed of photo-realistic scenes augmented with additional 3D objects. In each episode, the target object is presented to the agent using two modalities: a set of visual reference images on a neutral background and manually annotated textual descriptions. Through comprehensive evaluations and analyses, we showcase the challenges of the PIN task as well as the performance and shortcomings of currently available methods designed for object-driven navigation, considering modular and end-to-end agents. Luca Barsellotti, Roberto Bigazzi, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
NeurIPS | 3 |
| 2024 | Generating More Pertinent Captions by Leveraging Semantics and Style on Multi-Source Datasets
Marcella Cornia, Lorenzo Baraldi 0001, Giuseppe Fiameni, Rita Cucchiara |
Int. J. Comput. Vis. | 1 |
| 2024 | Unveiling the Truth: Exploring Human Gaze Patterns in Fake ImagesabstractCreating high-quality and realistic images is now possible thanks to the impressive advancements in image generation. A description in natural language of your desired output is all you need to obtain breathtaking results. However, as the use of generative models grows, so do concerns about the propagation of malicious content and misinformation. Consequently, the research community is actively working on the development of novel fake detection techniques, primarily focusing on low-level features and possible fingerprints left by generative models during the image generation process. In a different vein, in our work, we leverage human semantic knowledge to investigate the possibility of being included in frameworks of fake image detection. To achieve this, we collect a novel dataset of partially manipulated images using diffusion models and conduct an eye-tracking experiment to record the eye movements of different observers while viewing real and fake stimuli. A preliminary statistical analysis is conducted to explore the distinctive patterns in how humans perceive genuine and altered images. Statistical findings reveal that, when perceiving counterfeit samples, humans tend to focus on more confined regions of the image, in contrast to the more dispersed observational pattern observed when viewing genuine images. Our dataset is publicly available at:https://github.com/aimagelab/unveiling-the-truth. Giuseppe Cartella, Vittorio Cuculo, Marcella Cornia, Rita Cucchiara |
IEEE Signal Process. Lett. | 3 |
| 2024 | Towards Retrieval-Augmented Architectures for Image CaptioningabstractThe objective of image captioning models is to bridge the gap between the visual and linguistic modalities by generating natural language descriptions that accurately reflect the content of input images. In recent years, researchers have leveraged deep learning-based models and made advances in the extraction of visual features and the design of multimodal connections to tackle this task. This work presents a novel approach toward developing image captioning models that utilize an externalkNN memory to improve the generation process. Specifically, we propose two model variants that incorporate a knowledge retriever component that is based on visual similarities, a differentiable encoder to represent input images, and akNN-augmented language model to predict tokens based on contextual cues and text retrieved from the external memory. We experimentally validate our approach on COCO and nocaps datasets and demonstrate that incorporating an explicit external memory can significantly enhance the quality of captions, especially with a larger retrieval corpus. This work provides valuable insights into retrieval-augmented captioning models and opens up new avenues for improving image captioning at a larger scale. Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Alessandro Nicolosi, Rita Cucchiara |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Positive-Augmented Contrastive Learning for Image and Video Captioning EvaluationabstractThe CLIP model has been recently proven to be very effective for a variety of cross-modal tasks, including the evaluation of captions generated from vision-and-language architectures. In this paper, we propose a new recipe for a contrastive-based evaluation metric for image captioning, namely Positive-Augmented Contrastive learning Score (PAC-S), that in a novel way unifies the learning of a contrastive visual-semantic space with the addition of generated images and text on curated data. Experiments spanning several datasets demonstrate that our new metric achieves the highest correlation with human judgments on both images and videos, outperforming existing referencebased metrics like CIDEr and SPICE and reference-free metrics like CLIP-Score. Finally, we test the system-level correlation of the proposed metric when considering popular image captioning approaches, and assess the impact of employing different cross-modal features. Our source code and trained models are publicly available at: https://github.com/aimagelab/pacscore. Sara Sarto, Manuele Barraco, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 3 |
| 2023 | Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image EditingabstractFashion illustration is used by designers to communicate their vision and to bring the design idea from conceptualization to realization, showing how clothes interact with the human body. In this context, computer vision can thus be used to improve the fashion design process. Differently from previous works that mainly focused on the virtual try-on of garments, we propose the task of multimodal-conditioned fashion image editing, guiding the generation of human-centric fashion images by following multimodal prompts, such as text, human body poses, and garment sketches. We tackle this problem by proposing a new architecture based on latent diffusion models, an approach that has not been used before in the fashion domain. Given the lack of existing datasets suitable for the task, we also extend two existing fashion datasets, namely Dress Code and VITON-HD, with multimodal annotations collected in a semi-automatic manner. Experimental results on these new datasets demonstrate the effectiveness of our proposal, both in terms of realism and coherence with the given multimodal inputs. Source code and collected multimodal annotations are publicly available at: https://github.com/aimagelab/multimodal-garment-designer. Alberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia, Marco Bertini 0001, Rita Cucchiara |
ICCV | 4 |
| 2023 | With a Little Help from your own Past: Prototypical Memory Networks for Image CaptioningabstractImage captioning, like many tasks involving vision and language, currently relies on Transformer-based architectures for extracting the semantics in an image and translating it into linguistically coherent descriptions. Although successful, the attention operator only considers a weighted summation of projections of the current input sample, therefore ignoring the relevant semantic information which can come from the joint observation of other samples. In this paper, we devise a network which can perform attention over activations obtained while processing other training samples, through a prototypical memory model. Our memory models the distribution of past keys and values through the definition of prototype vectors which are both discriminative and compact. Experimentally, we assess the performance of the proposed model on the COCO dataset, in comparison with carefully designed baselines and state-of-the-art approaches, and by investigating the role of each of the proposed components. We demonstrate that our proposal can increase the performance of an encoder-decoder Transformer by 3.7 CIDEr points both when training in cross-entropy only and when fine-tuning with self-critical sequence training. Source code and trained models are available at: https://github.com/aimagelab/PMA-Net. Manuele Barraco, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICCV | 3 |
| 2023 | Embodied Agents for Efficient Exploration and Smart Scene DescriptionabstractThe development of embodied agents that can communicate with humans in natural language has gained increasing interest over the last years, as it facilitates the diffusion of robotic platforms in human-populated environments. As a step towards this objective, in this work, we tackle a setting for visual navigation in which an autonomous agent needs to explore and map an unseen indoor environment while portraying interesting scenes with natural language descriptions. To this end, we propose and evaluate an approach that combines recent advances in visual robotic exploration and image captioning on images generated through agent-environment interaction. Our approach can generate smart scene descriptions that maximize semantic knowledge of the environment and avoid repetitions. Further, such descriptions offer user-understandable insights into the robot's representation of the environment by high-lighting the prominent objects and the correlation between them as encountered during the exploration. To quantitatively assess the performance of the proposed approach, we also devise a specific score that takes into account both exploration and description skills. The experiments carried out on both photorealistic simulated environments and real-world ones demonstrate that our approach can effectively describe the robot's point of view during exploration, improving the human-friendly interpretability of its observations. Roberto Bigazzi, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara |
ICRA | 2 |
| 2023 | LaDI-VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-OnabstractThe rapidly evolving fields of e-commerce and metaverse continue to seek innovative approaches to enhance the consumer experience. At the same time, recent advancements in the development of diffusion models have enabled generative networks to create remarkably realistic images. In this context, image-based virtual try-on, which consists in generating a novel image of a target model wearing a given in-shop garment, has yet to capitalize on the potential of these powerful generative solutions. This work introduces LaDI-VTON, the first Latent Diffusion textual Inversion-enhanced model for the Virtual Try-ON task. The proposed architecture relies on a latent diffusion model extended with a novel additional autoencoder module that exploits learnable skip connections to enhance the generation process preserving the model's characteristics. To effectively maintain the texture and details of the in-shop garment, we propose a textual inversion component that can map the visual features of the garment to the CLIP token embedding space and thus generate a set of pseudo-word token embeddings capable of conditioning the generation process. Experimental results on Dress Code and VITON-HD datasets demonstrate that our approach outperforms the competitors by a consistent margin, achieving a significant milestone for the task. Source code and trained models are publicly available at: https://github.com/miccunifi/ladi-vton. Davide Morelli, Alberto Baldrati, Giuseppe Cartella, Marcella Cornia, Marco Bertini 0001, Rita Cucchiara |
ACM Multimedia | 4 |
| 2023 | Fully-attentive iterative networks for region-based controllable image and video captioningabstractControllable image captioning has recently gained attention as a way to increase the diversity and the applicability to real-world scenarios of image captioning algorithms. In this task, a captioner is conditioned on an external control signal, which needs to be followed during the generation of the caption. We aim to overcome the limitations of current controllable captioning methods by proposing a fully-attentive and iterative network that can generate grounded and controllable captions from a control signal given as a sequence of visual regions from the image. Our architecture is based on a set of novel attention operators, which take into account the hierarchical nature of the control signal, and is endowed with a decoder which explicitly focuses on each part of the control signal. We demonstrate the effectiveness of the proposed approach by conducting experiments on three datasets, where our model surpasses the performances of previous methods and achieves a new state of the art on both image and video controllable captioning. Marcella Cornia, Lorenzo Baraldi 0001, Ayellet Tal, Rita Cucchiara |
Comput. Vis. Image Underst. | 1 |
| 2023 | From Show to Tell: A Survey on Deep Learning-Based Image CaptioningabstractConnecting Vision and Language plays an essential role in Generative Intelligence. For this reason, large research efforts have been devoted to image captioning, i.e. describing images with syntactically and semantically meaningful sentences. Starting from 2015 the task has generally been addressed with pipelines composed of a visual encoder and a language model for text generation. During these years, both components have evolved considerably through the exploitation of object regions, attributes, the introduction of multi-modal connections, fully-attentive approaches, and BERT-like early-fusion strategies. However, regardless of the impressive results, research in image captioning has not reached a conclusive answer yet. This work aims at providing a comprehensive overview of image captioning approaches, from visual encoding and text generation to training strategies, datasets, and evaluation metrics. In this respect, we quantitatively compare many relevant state-of-the-art approaches to identify the most impactful technical innovations in architectures and training strategies. Moreover, many variants of the problem and its open challenges are discussed. The final goal of this work is to serve as a tool for understanding the existing literature and highlighting the future directions for a research area where Computer Vision and Natural Language Processing can find an optimal synergy. Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi 0001, Silvia Cascianelli, Giuseppe Fiameni, Rita Cucchiara |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | ALADIN: Distilling Fine-grained Alignment Scores for Efficient Image-Text Matching and RetrievalabstractImage-text matching is gaining a leading role among tasks involving the joint understanding of vision and language. In literature, this task is often used as a pre-training objective to forge architectures able to jointly deal with images and texts. Nonetheless, it has a direct downstream application: cross-modal retrieval, which consists in finding images related to a given query text or vice-versa. Solving this task is of critical importance in cross-modal search engines. Many recent methods proposed effective solutions to the image-text matching problem, mostly using recent large vision-language (VL) Transformer networks. However, these models are often computationally expensive, especially at inference time. This prevents their adoption in large-scale cross-modal retrieval scenarios, where results should be provided to the user almost instantaneously. In this paper, we propose to fill in the gap between effectiveness and efficiency by proposing an ALign And DIstill Network (ALADIN). ALADIN first produces high-effective scores by aligning at fine-grained level images and texts. Then, it learns a shared embedding space – where an efficient kNN search can be performed – by distilling the relevance scores obtained from the fine-grained alignments. We obtained remarkable results on MS-COCO, showing that our method can compete with state-of-the-art VL Transformers while being almost 90 times faster. The code for reproducing our results is available at https://github.com/mesnico/ALADIN. Nicola Messina, Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi 0001, Fabrizio Falchi, Giuseppe Amato 0001, Rita Cucchiara |
CBMI | 3 |
| 2022 | Retrieval-Augmented Transformer for Image CaptioningabstractImage captioning models aim at connecting Vision and Language by providing natural language descriptions of input images. In the past few years, the task has been tackled by learning parametric models and proposing visual feature extraction advancements or by modeling better multi-modal connections. In this paper, we investigate the development of an image captioning approach with a kNN memory, with which knowledge can be retrieved from an external corpus to aid the generation process. Our architecture combines a knowledge retriever based on visual similarities, a differentiable encoder, and a kNN-augmented attention layer to predict tokens based on the past context and on text retrieved from the external memory. Experimental results, conducted on the COCO dataset, demonstrate that employing an explicit external memory can aid the generation process and increase caption quality. Our work opens up new avenues for improving image captioning models at larger scale. Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CBMI | 2 |
| 2022 | Dress Code: High-Resolution Multi-category Virtual Try-On
Davide Morelli, Matteo Fincato, Marcella Cornia, Federico Landi, Fabio Cesari, Rita Cucchiara |
ECCV (8) | 3 |
| 2022 | CaMEL: Mean Teacher Learning for Image CaptioningabstractDescribing images in natural language is a fundamental step towards the automatic modeling of connections between the visual and textual modalities. In this paper we present CaMEL, a novel Transformer-based architecture for image captioning. Our proposed approach leverages the interaction of two interconnected language models that learn from each other during the training phase. The interplay between the two language models follows a mean teacher learning paradigm with knowledge distillation. Experimentally, we assess the effectiveness of the proposed solution on the COCO dataset and in conjunction with different visual feature extractors. When comparing with existing proposals, we demonstrate that our model provides state-of-the-art caption quality with a significantly reduced number of parameters. According to the CIDEr metric, we obtain a new state of the art on COCO when training without using external data. The source code and trained models will be made publicly available at: https://github.com/aimagelab/camel. Manuele Barraco, Matteo Stefanini, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR | 3 |
| 2022 | The LAM Dataset: A Novel Benchmark for Line-Level Handwritten Text RecognitionabstractHandwritten Text Recognition (HTR) is an open problem at the intersection of Computer Vision and Natural Language Processing. The main challenges, when dealing with historical manuscripts, are due to the preservation of the paper support, the variability of the handwriting – even of the same author over a wide time-span – and the scarcity of data from ancient, poorly represented languages. With the aim of fostering the research on this topic, in this paper we present the Ludovico Antonio Muratori (LAM) dataset, a large line-level HTR dataset of Italian ancient manuscripts edited by a single author over 60 years. The dataset comes in two configurations: a basic splitting and a date-based splitting which takes into account the age of the author. The first setting is intended to study HTR on ancient documents in Italian, while the second focuses on the ability of HTR systems to recognize text written by the same writer in time periods for which training data are not available. For both configurations, we analyze quantitative and qualitative characteristics, also with respect to other line-level HTR benchmarks, and present the recognition performance of state-of-the-art HTR architectures. The dataset is available for download at https://aimagelab.ing.unimore.it/go/lam. Silvia Cascianelli, Vittorio Pippi, Martin Maarand, Marcella Cornia, Lorenzo Baraldi 0001, Christopher Kermorvant, Rita Cucchiara |
ICPR | 4 |
| 2022 | Spot the Difference: A Novel Task for Embodied Agents in Changing EnvironmentsabstractEmbodied AI is a recent research area that aims at creating intelligent agents that can move and operate inside an environment. Existing approaches in this field demand the agents to act in completely new and unexplored scenes. However, this setting is far from realistic use cases that instead require executing multiple tasks in the same environment. Even if the environment changes over time, the agent could still count on its global knowledge about the scene while trying to adapt its internal representation to the current state of the environment. To make a step towards this setting, we propose Spot the Difference: a novel task for Embodied AI where the agent has access to an outdated map of the environment and needs to recover the correct layout in a fixed time budget. To this end, we collect a new dataset of occupancy maps starting from existing datasets of 3D spaces and generating a number of possible layouts for a single environment. This dataset can be employed in the popular Habitat simulator and is fully compliant with existing methods that employ reconstructed occupancy maps during navigation. Furthermore, we propose an exploration policy that can take advantage of previous knowledge of the environment and identify changes in the scene faster and more effectively than existing agents. Experimental results show that the proposed architecture outperforms existing state-of-the-art models for exploration on this new setting. Federico Landi, Roberto Bigazzi, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR | 3 |
| 2022 | Boosting modern and historical handwritten text recognition with deformable convolutions
Silvia Cascianelli, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
Int. J. Document Anal. Recognit. | 2 |
| 2022 | Matching Faces and Attributes Between the Artistic and the Real Domain: the PersonArt ApproachabstractIn this article, we present an approach for retrieving similar faces between the artistic and the real domain. The application we refer to is an interactive exhibition inside a museum, in which a visitor can take a photo of himself and search for a lookalike in the collection of paintings. The task requires not only to identify faces but also to extract discriminative features from artistic and photo-realistic images, tackling a significant domain shift. Our method integrates feature extraction networks which account for the aesthetic similarity of two faces and their correspondences in terms of semantic attributes. Also, it addresses the domain shift between realistic images and paintings by translating photo-realistic images into the artistic domain. Noticeably, by exploiting the same technique, our model does not need to rely on annotated data in the artistic domain. Experimental results are conducted on different paired datasets to show the effectiveness of the proposed solution in terms of identity and attribute preservation. The approach is also evaluated on unpaired settings and in combination with an interactive relevance feedback strategy. Finally, we show how the proposed algorithm has been implemented in a real showcase at the Gallerie Estensi museum in Italy, with the participation of more than 1,100 visitors in just three days. Marcella Cornia, Matteo Tomei, Lorenzo Baraldi 0001, Rita Cucchiara |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Transform, Warp, and Dress: A New Transformation-guided Model for Virtual Try-onabstractVirtual try-on has recently emerged in computer vision and multimedia communities with the development of architectures that can generate realistic images of a target person wearing a custom garment. This research interest is motivated by the large role played by e-commerce and online shopping in our society. Indeed, the virtual try-on task can offer many opportunities to improve the efficiency of preparing fashion catalogs and to enhance the online user experience. The problem is far to be solved: current architectures do not reach sufficient accuracy with respect to manually generated images and can only be trained on image pairs with a limited variety. Existing virtual try-on datasets have two main limits: they contain only female models, and all the images are available only in low resolution. This not only affects the generalization capabilities of the trained architectures but makes the deployment to real applications impractical. To overcome these issues, we present Dress Code , a new dataset for virtual try-on that contains high-resolution images of a large variety of upper-body clothes and both male and female models. Leveraging this enriched dataset, we propose a new model for virtual try-on capable of generating high-quality and photo-realistic images using a three-stage pipeline. The first two stages perform two different geometric transformations to warp the desired garment and make it fit into the target person’s body pose and shape. Then, we generate the new image of that same person wearing the try-on garment using a generative network. We test the proposed solution on the most widely used dataset for this task as well as on our newly collected dataset and demonstrate its effectiveness when compared to current state-of-the-art methods. Through extensive analyses on our Dress Code dataset, we show the adaptability of our model, which can generate try-on images even with a higher resolution. Matteo Fincato, Marcella Cornia, Federico Landi, Fabio Cesari, Rita Cucchiara |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Out of the Box: Embodied Navigation in the Real World
Roberto Bigazzi, Federico Landi, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara |
CAIP (1) | 3 |
| 2021 | Learning to Read L'Infinito: Handwritten Text Recognition with Synthetic Training Data
Silvia Cascianelli, Marcella Cornia, Lorenzo Baraldi 0001, Maria Ludovica Piazzi, Rosiana Schiuma, Rita Cucchiara |
CAIP (2) | 2 |
| 2021 | Learning to Select: A Fully Attentive Approach for Novel Object CaptioningabstractImage captioning models have lately shown impressive results when applied to standard datasets. Switching to real-life scenarios, however, constitutes a challenge due to the larger variety of visual concepts which are not covered in existing training sets. For this reason, novel object captioning (NOC) has recently emerged as a paradigm to test captioning models on objects which are unseen during the training phase. In this paper, we present a novel approach for NOC that learns to select the most relevant objects of an image, regardless of their adherence to the training set, and to constrain the generative process of a language model accordingly. Our architecture is fully-attentive and end-to-end trainable, also when incorporating constraints. We perform experiments on the held-out COCO dataset, where we demonstrate improvements over the state of the art, both in terms of adaptability to novel objects and caption quality. Marco Cagrandi, Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi 0001, Rita Cucchiara |
ICMR | 2 |
| 2021 | Multimodal attention networks for low-level vision-and-language navigation
Federico Landi, Lorenzo Baraldi 0001, Marcella Cornia, Massimiliano Corsini, Rita Cucchiara |
Comput. Vis. Image Underst. | 3 |
| 2021 | Working Memory Connections for LSTM
Federico Landi, Lorenzo Baraldi 0001, Marcella Cornia, Rita Cucchiara |
Neural Networks | 3 |
| 2020 | Meshed-Memory Transformer for Image CaptioningabstractTransformer-based architectures represent the state of the art in sequence modeling tasks like machine translation and language understanding. Their applicability to multi-modal contexts like image captioning, however, is still largely under-explored. With the aim of filling this gap, we present M2- a Meshed Transformer with Memory for Image Captioning. The architecture improves both the image encoding and the language generation steps: it learns a multi-level representation of the relationships between image regions integrating learned a priori knowledge, and uses a mesh-like connectivity at decoding stage to exploit low- and high-level features. Experimentally, we investigate the performance of the M2Transformer and different fully-attentive models in comparison with recurrent ones. When tested on COCO, our proposal achieves a new state of the art in single-model and ensemble configurations on the "Karpathy" test split and on the online test server. We also assess its performances when describing objects unseen in the training set. Trained models and code for reproducing the experiments are publicly available at: https://github.com/aimagelab/meshed-memory-transformer. Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 1 |
| 2020 | Explore and Explain: Self-supervised Navigation and RecountingabstractEmbodied AI has been recently gaining attention as it aims to foster the development of autonomous and intelligent agents. In this paper, we devise a novel embodied setting in which an agent needs to explore a previously unknown environment while recounting what it sees during the path. In this context, the agent needs to navigate the environment driven by an exploration goal, select proper moments for description, and output natural language descriptions of relevant objects and scenes. Our model integrates a novel self-supervised exploration module with penalty, and a fully-attentive captioning model for explanation. Also, we investigate different policies for selecting proper moments for explanation, driven by information coming from both the environment and the navigation. Experiments are conducted on photorealistic environments from the Matterport3D dataset and investigate the navigation and explanation capabilities of the agent as well as the role of their interactions. Roberto Bigazzi, Federico Landi, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR | 3 |
| 2020 | VITON-GT: An Image-based Virtual Try-On Model with Geometric TransformationsabstractThe large spread of online shopping has led computer vision researchers to develop different solutions for the fashion domain to potentially increase the online user experience and improve the efficiency of preparing fashion catalogs. Among them, image-based virtual try-on has recently attracted a lot of attention resulting in several architectures that can generate a new image of a person wearing an input try-on garment in a plausible and realistic way. In this paper, we present VITON-G T, a new model for virtual try-on that generates high-quality and photo-realistic images thanks to multiple geometric transformations. In particular, our model is composed of a two-stage geometric transformation module that performs two different projections on the input garment, and a transformation-guided try-on module that synthesizes the new image. We experimentally validate the proposed solution on the most common dataset for this task, containing mainly t-shirts, and we demonstrate its effectiveness compared to different baselines and previous methods. Additionally, we assess the generalization capabilities of our model on a new set of fashion items composed of upper-body clothes from different categories. To the best of our knowledge, we are the first to test virtual try-on architectures in this challenging experimental setting. Matteo Fincato, Federico Landi, Marcella Cornia, Fabio Cesari, Rita Cucchiara |
ICPR | 3 |
| 2020 | A Novel Attention-based Aggregation Function to Combine Vision and LanguageabstractThe joint understanding of vision and language has been recently gaining a lot of attention in both the Computer Vision and Natural Language Processing communities, with the emergence of tasks such as image captioning, image-text matching, and visual question answering. As both images and text can be encoded as sets or sequences of elements - like regions and words - proper reduction functions are needed to transform a set of encoded elements into a single response, like a classification or similarity score. In this paper, we propose a novel fully-attentive reduction method for vision and language. Specifically, our approach computes a set of scores for each element of each modality employing a novel variant of cross-attention, and performs a learnable and cross-modal reduction, which can be used for both classification and ranking. We test our approach on image-text matching and visual question answering, building fair comparisons with other reduction choices, on both COCO and VQA 2.0 datasets. Experimentally, we demonstrate that our approach leads to a performance increase on both tasks. Further, we conduct ablation studies to validate the role of each component of the approach. Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICPR | 2 |
| 2020 | SMArT: Training Shallow Memory-aware Transformers for Robotic ExplainabilityabstractThe ability to generate natural language explanations conditioned on the visual perception is a crucial step towards autonomous agents which can explain themselves and communicate with humans. While the research efforts in image and video captioning are giving promising results, this is often done at the expense of the computational requirements of the approaches, limiting their applicability to real contexts. In this paper, we propose a fully-attentive captioning algorithm which can provide state-of-the-art performances on language generation while restricting its computational demands. Our model is inspired by the Transformer model and employs only two Transformer layers in the encoding and decoding stages. Further, it incorporates a novel memory-aware encoding of image regions. Experiments demonstrate that our approach achieves competitive results in terms of caption quality while featuring reduced computational demands. Further, to evaluate its applicability on autonomous agents, we conduct experiments on simulated scenes taken from the perspective of domestic robots. Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
ICRA | 1 |
| 2020 | A unified cycle-consistent neural model for text and image retrieval
Marcella Cornia, Lorenzo Baraldi 0001, Hamed Rezazadegan Tavakoli, Rita Cucchiara |
Multim. Tools Appl. | 1 |
| 2020 | Explaining digital humanities by aligning images and textual descriptions
Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi 0001, Massimiliano Corsini, Rita Cucchiara |
Pattern Recognit. Lett. | 1 |
| 2019 | Show, Control and Tell: A Framework for Generating Controllable and Grounded CaptionsabstractCurrent captioning approaches can describe images using black-box architectures whose behavior is hardly controllable and explainable from the exterior. As an image can be described in infinite ways depending on the goal and the context at hand, a higher degree of controllability is needed to apply captioning algorithms in complex scenarios. In this paper, we introduce a novel framework for image captioning which can generate diverse descriptions by allowing both grounding and controllability. Given a control signal in the form of a sequence or set of image regions, we generate the corresponding caption through a recurrent architecture which predicts textual chunks explicitly grounded on regions, following the constraints of the given control. Experiments are conducted on Flickr30k Entities and on COCO Entities, an extended version of COCO in which we add grounding annotations collected in a semi-automatic manner. Results demonstrate that our method achieves state of the art performances on controllable image captioning, in terms of caption quality and diversity. Code and annotations are publicly available at: https://github.com/aimagelab/show-control-and-tell. Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 1 |
| 2019 | Art2Real: Unfolding the Reality of Artworks via Semantically-Aware Image-To-Image TranslationabstractThe applicability of computer vision to real paintings and artworks has been rarely investigated, even though a vast heritage would greatly benefit from techniques which can understand and process data from the artistic domain. This is partially due to the small amount of annotated artistic data, which is not even comparable to that of natural images captured by cameras. In this paper, we propose a semantic-aware architecture which can translate artworks to photo-realistic visualizations, thus reducing the gap between visual features of artistic and realistic data. Our architecture can generate natural images by retrieving and learning details from real photos through a similarity matching strategy which leverages a weakly-supervised semantic understanding of the scene. Experimental results show that the proposed technique leads to increased realism and to a reduction in domain shift, which improves the performance of pre-trained architectures for classification, detection, and segmentation. Code is publicly available at: https://github.com/aimagelab/art2real. Matteo Tomei, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara |
CVPR | 2 |
| 2019 | M-VAD names: a dataset for video captioning with naming
Stefano Pini, Marcella Cornia, Federico Bolelli, Lorenzo Baraldi 0001, Rita Cucchiara |
Multim. Tools Appl. | 2 |
| 2018 | Aligning Text and Document Illustrations: Towards Visually Explainable Digital HumanitiesabstractWhile several approaches to bring vision and language together are emerging, none of them has yet addressed the digital humanities domain, which, nevertheless, is a rich source of visual and textual data. To foster research in this direction, we investigate the learning of visual-semantic embeddings for historical document illustrations, devising both supervised and semi-supervised approaches. We exploit the joint visual-semantic embeddings to automatically align illustrations and textual elements, thus providing an automatic annotation of the visual content of a manuscript. Experiments are performed on the Borso d'Este Holy Bible, one of the most sophisticated illuminated manuscript from the Renaissance, which we manually annotate aligning every illustration with textual commentaries written by experts. Experimental results quantify the domain shift between ordinary visual-semantic datasets and the proposed one, validate the proposed strategies, and devise future works on the same line. Lorenzo Baraldi 0001, Marcella Cornia, Costantino Grana, Rita Cucchiara |
ICPR | 2 |
| 2018 | Predicting Human Eye Fixations via an LSTM-Based Saliency Attentive ModelabstractData-driven saliency has recently gained a lot of attention thanks to the use of Convolutional Neural Networks for predicting gaze fixations. In this paper we go beyond standard approaches to saliency prediction, in which gaze maps are computed with a feed-forward network, and present a novel model which can predict accurate saliency maps by incorporating neural attentive mechanisms. The core of our solution is a Convolutional LSTM that focuses on the most salient regions of the input image to iteratively refine the predicted saliency map. Additionally, to tackle the center bias typical of human eye fixations, our model can learn a set of prior maps generated with Gaussian functions. We show, through an extensive evaluation, that the proposed architecture outperforms the current state of the art on public saliency prediction datasets. We further study the contribution of each key component to demonstrate their robustness on different scenarios. Marcella Cornia, Lorenzo Baraldi 0001, Giuseppe Serra 0001, Rita Cucchiara |
IEEE Trans. Image Process. | 1 |
| 2018 | Paying More Attention to Saliency: Image Captioning with Saliency and Context AttentionabstractImage captioning has been recently gaining a lot of attention thanks to the impressive achievements shown by deep captioning architectures, which combine Convolutional Neural Networks to extract image representations and Recurrent Neural Networks to generate the corresponding captions. At the same time, a significant research effort has been dedicated to the development of saliency prediction models, which can predict human eye fixations. Even though saliency information could be useful to condition an image captioning architecture, by providing an indication of what is salient and what is not, research is still struggling to incorporate these two techniques. In this work, we propose an image captioning approach in which a generative recurrent neural network can focus on different parts of the input image during the generation of the caption, by exploiting the conditioning given by a saliency prediction model on which parts of the image are salient and which are contextual. We show, through extensive quantitative and qualitative experiments on large-scale datasets, that our model achieves superior performance with respect to captioning baselines with and without saliency and to different state-of-the-art approaches combining saliency and captioning. Marcella Cornia, Lorenzo Baraldi 0001, Giuseppe Serra 0001, Rita Cucchiara |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2017 | Modeling multimodal cues in a deep learning-based framework for emotion recognition in the wildabstractIn this paper, we propose a multimodal deep learning architecture for emotion recognition in video regarding our participation to the audio-video based sub-challenge of the Emotion Recognition in the Wild 2017 challenge. Our model combines cues from multiple video modalities, including static facial features, motion patterns related to the evolution of the human expression over time, and audio information. Specifically, it is composed of three sub-networks trained separately: the first and second ones extract static visual features and dynamic patterns through 2D and 3D Convolutional Neural Networks (CNN), while the third one consists in a pretrained audio network which is used to extract useful deep acoustic signals from video. In the audio branch, we also apply Long Short Term Memory (LSTM) networks in order to capture the temporal evolution of the audio features. To identify and exploit possible relationships among different modalities, we propose a fusion network that merges cues from the different modalities in one representation. The proposed architecture outperforms the challenge baselines (38.81 % and 40.47 %): we achieve an accuracy of 50.39 % and 49.92 % respectively on the validation and the testing data. Stefano Pini, Olfa Ben Ahmed, Marcella Cornia, Lorenzo Baraldi 0001, Rita Cucchiara, Benoit Huet |
ICMI | 3 |
| 2016 | A deep multi-level network for saliency predictionabstractThis paper presents a novel deep architecture for saliency prediction. Current state of the art models for saliency prediction employ Fully Convolutional networks that perform a non-linear combination of features extracted from the last convolutional layer to predict saliency maps. We propose an architecture which, instead, combines features extracted at different levels of a Convolutional Neural Network (CNN). Our model is composed of three main blocks: a feature extraction CNN, a feature encoding network, that weights low and high level feature maps, and a prior learning network. We compare our solution with state of the art saliency models on two public benchmarks datasets. Results show that our model outperforms under all evaluation metrics on the SALICON dataset, which is currently the largest public dataset for saliency prediction, and achieves competitive results on the MIT300 benchmark. Code is available at https://github.com/marcellacornia/mlnet. Marcella Cornia, Lorenzo Baraldi 0001, Giuseppe Serra 0001, Rita Cucchiara |
ICPR | 1 |