VLDB 2026 Research / reviewers in the wild / expert
Elisa Ricci 0001
dblp:88/397
· DBLP profile ↗
189ranked-venue papers
10as first author
91since 2021 · last 2026
0000-0002-0228-1147ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 144 · 6 first-author · 74 since 2021Graphics, computer vision, multimedia, augmented reality and games · 128 · 4 first-author · 62 since 2021Systems, architecture and hardware · 8 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training-Free Semantic Multi-Object Tracking with Vision-Language Models
Laurence Bonat, Francesco Tonini, Elisa Ricci 0001, Lorenzo Vaquero |
FG | 3 |
| 2026 | Zero-Shot Temporal Action Localization Through Textual Guidance
Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero, Paolo Rota, Yiming Wang 0002, Elisa Ricci 0001 |
FG | 6 |
| 2026 | Towards Unconstrained Human-Object Interaction
Francesco Tonini, Alessandro Conti, Lorenzo Vaquero, Cigdem Beyan, Elisa Ricci 0001 |
FG | 5 |
| 2026 | Vocabulary-Free Image Classification and Semantic SegmentationabstractLarge vision-language models revolutionized image classification and semantic segmentation paradigms. However, they typically assume a pre-defined set of categories, or vocabulary, at test time for composing textual prompts. This assumption is impractical in scenarios with unknown or evolving semantic context. Here, we address this issue and introduce the Vocabulary-free Image Classification (VIC) task, which aims to assign a class from an unconstrained language-induced semantic space to an input image without needing a known vocabulary. VIC is challenging due to the vastness of the semantic space, which contains millions of concepts, including fine-grained categories. To address VIC, we propose Category Search from External Databases (CaSED), a training-free method that leverages a pre-trained vision-language model and an external database. CaSED first extracts the set of candidate categories from the most semantically similar captions in the database and then assigns the image to the best-matching candidate category according to the same vision-language model. Furthermore, we demonstrate that CaSED can be applied locally to generate a coarse segmentation mask that classifies image regions, introducing the task of Vocabulary-free Semantic Segmentation. CaSED and its variants outperform other more complex vision-language models, on classification and semantic segmentation benchmarks, while using much fewer parameters. Alessandro Conti, Enrico Fini, Massimiliano Mancini, Paolo Rota, Yiming Wang 0002, Elisa Ricci 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language ModelsabstractVision-Language Models (VLMs) learn a shared feature space for text and images, enabling the comparison of inputs of different modalities. While prior works demonstrated that VLMs organize natural language representations into regular structures encoding composite meanings, it remains unclear if compositional patterns also emerge in the visual embedding space. In this work, we investigate compositionality in the image domain, where the analysis of compositional properties is challenged by noise and sparsity of visual data. We address these problems and propose a framework, called Geodesically Decomposable Embeddings (GDE), that approximates image representations with geometry-aware compositional structures in the latent space. We demonstrate that visual embeddings of pre-trained VLMs exhibit a compositional arrangement, and evaluate the effectiveness of this property in the tasks of compositional classification and group robustness. GDE achieves stronger performance in compositional classification compared to its counterpart method that assumes linear geometry of the latent space. Notably, it is particularly effective for group robustness, where we achieve higher results than task-specific solutions. Our results indicate that VLMs can automatically develop a human-like form of compositional reasoning in the visual domain, making their underlying processes more interpretable. Code is available at https://github.com/BerasiDavide/vlm_image_compositionality. Davide Berasi, Matteo Farina, Massimiliano Mancini, Elisa Ricci 0001, Nicola Strisciuglio |
CVPR | 4 |
| 2025 | Rethinking Few-Shot Adaptation of Vision-Language Models in Two StagesabstractAn old-school recipe for training a classifier is to (i) learn a good feature extractor and (ii) optimize a linear layer atop. When only a handful of samples are available per category, as in Few-Shot Adaptation (FSA), data are insufficient to fit a large number of parameters, rendering the above impractical. This is especially true with large pre-trained Vision-Language Models (VLMs), which motivated successful research at the intersection of Parameter-Efficient Fine-tuning (PEFT) and FSA. In this work, we start by analyzing the learning dynamics of PEFT techniques when trained on few-shot data from only a subset of categories, referred to as the "base" classes. We show that such dynamics naturally splits into two distinct phases: (i) task-level feature extraction and (ii) specialization to the available concepts. To accommodate this dynamic, we then depart from prompt- or adapter-based methods and tackle FSA differently. Specifically, given a fixed computational budget, we split it to (i) learn a task-specific feature extractor via PEFT and (ii) train a linear classifier on top. We call this scheme Two-Stage Few-Shot Adaptation (2SFS). Differently from established methods, our scheme enables a novel form of selective inference at a category level, i.e., at test time, only novel categories are embedded by the adapted text encoder, while embeddings of base categories are available within the classifier. Results with fixed hyperparameters across two settings, three backbones, and eleven datasets, show that 2SFS matches or surpasses the state-of-the-art, while established methods degrade significantly across settings. . Matteo Farina, Massimiliano Mancini, Giovanni Iacca, Elisa Ricci 0001 |
CVPR | 4 |
| 2025 | Compositional Caching for Training-free Open-vocabulary Attribute DetectionabstractAttribute detection is crucial for many computer vision tasks, as it enables systems to describe properties such as color, texture, and material. Current approaches often rely on labor-intensive annotation processes which are inherently limited: objects can be described at an arbitrary level of detail (e.g., color vs. color shades), leading to ambiguities when the annotators are not instructed carefully. Furthermore, they operate within a predefined set of attributes, reducing scalability and adaptability to unforeseen downstream applications. We present Compositional Caching (ComCa), a training-free method for open-vocabulary attribute detection that overcomes these constraints. ComCa requires only the list of target attributes and objects as input, using them to populate an auxiliary cache of images by leveraging web-scale databases and Large Language Models to determine attribute-object compatibility. To account for the compositional nature of attributes, cache images receive soft attribute labels. Those are aggregated at inference time based on the similarity between the input and cache images, refining the predictions of underlying Vision-Language Models (VLMs). Importantly, our approach is model-agnostic, compatible with various VLMs. Experiments on public datasets demonstrate that ComCa significantly outperforms zero-shot and cache-based baselines, competing with recent training-based methods, proving that a carefully designed training-free approach can successfully address open-vocabulary attribute detection. Marco Garosi, Alessandro Conti, Gaowen Liu, Elisa Ricci 0001, Massimiliano Mancini |
CVPR | 4 |
| 2025 | Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual ClassifiersabstractA person downloading a pre-trained model from the web should be aware of its biases. Existing approaches for bias identification rely on datasets containing labels for the task of interest, something that a non-expert may not have access to, or may not have the necessary resources to collect: this greatly limits the number of tasks where model biases can be identified. In this work, we present Classifier-to-Bias (C2B), the first bias discovery framework that works without access to any labeled data: it only relies on a textual description of the classification task to identify biases in the target classification model. This description is fed to a large language model to generate bias proposals and corresponding captions depicting biases together with task-specific target labels. A retrieval model collects images for those captions, which are then used to assess the accuracy of the model w.r.t. the given biases. C2B is training-free, does not require any annotations, has no constraints on the list of biases, and can be applied to any pre-trained model on any classification task. Experiments on two publicly available datasets show that C2B discovers biases beyond those of the original datasets and outperforms a recent state-of-the-art bias detection baseline that relies on task-specific annotations, being a promising first step toward addressing task-agnostic unsupervised bias detection. Quentin Guimard, Moreno D'Incà, Massimiliano Mancini, Elisa Ricci 0001 |
CVPR | 4 |
| 2025 | Can Text-to-Video Generation help Video-Language Alignment?abstractRecent video-language alignment models are trained on sets of videos, each with an associated positive caption and a negative caption generated by large language models. A problem with this procedure is that negative captions may introduce linguistic biases, i.e., concepts are seen only as negatives and never associated with a video. While a solution would be to collect videos for the negative captions, existing databases lack the fine-grained variations needed to cover all possible negatives. In this work, we study whether synthetic videos can help to overcome this issue. Our preliminary analysis with multiple generators shows that, while promising on some tasks, synthetic videos harm the performance of the model on others. We hypothesize this issue is linked to noise (semantic and visual) in the generated videos and develop a method, SynViTa, that accounts for those. SynViTa dynamically weights the contribution of each synthetic video based on how similar its target caption is w.r.t. the real counterpart. Moreover, a semantic consistency loss makes the model focus on fine-grained differences across captions, rather than differences in video appearance. Experiments show that, on average, SynViTa improves over existing methods on VideoCon test sets and SSv2-Temporal, SSv2-Events, and ATP-Hard benchmarks, being a first promising step for using synthetic videos when learning video-language models. Luca Zanella, Massimiliano Mancini, Willi Menapace, Sergey Tulyakov, Yiming Wang 0002, Elisa Ricci 0001 |
CVPR | 6 |
| 2025 | On Large Multimodal Models as Open-World Image ClassifiersabstractTraditional image classification requires a predefined list of semantic categories. In contrast, Large Multimodal Models (LMMs) can sidestep this requirement by classifying images directly using natural language (e.g., answering the prompt "What is the main object in the image?"). Despite this remarkable capability, most existing studies on LMM classification performance are surprisingly limited in scope, often assuming a closed-world setting with a predefined set of categories. In this work, we address this gap by thoroughly evaluating LMM classification performance in a truly open-world setting. We first formalize the task and introduce an evaluation protocol, defining various metrics to assess the alignment between predicted and ground truth classes. We then evaluate 13 models across 10 benchmarks, encompassing prototypical, non-prototypical, fine-grained, and very fine-grained classes, demonstrating the challenges LMMs face in this task. Further analyses based on the proposed metrics reveal the types of errors LMMs make, highlighting challenges related to granularity and fine-grained capabilities, showing how tailored prompting and reasoning can alleviate them. Alessandro Conti, Massimiliano Mancini, Enrico Fini, Yiming Wang 0002, Paolo Rota, Elisa Ricci 0001 |
ICCV | 6 |
| 2025 | Training-Free Personalization via Retrieval and Reasoning on Fingerprints
Deepayan Das, Davide Talon, Yiming Wang 0002, Massimiliano Mancini, Elisa Ricci 0001 |
ICCV | 5 |
| 2025 | Superpowering Open-Vocabulary Object Detectors for X-ray Vision
Pablo Garcia-Fernandez, Lorenzo Vaquero, Feng Xue 0001, Daniel Cores, Nicu Sebe, Manuel Mucientes, Elisa Ricci 0001 |
ICCV | 8 |
| 2025 | FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language ModelsabstractIn federated learning, textual prompt tuning adapts Vision-Language Models (e.g., CLIP) by tuning lightweight input tokens (or prompts) on local client data, while keeping network weights frozen. After training, only the prompts are shared by the clients with the central server for aggregation. However, textual prompt tuning suffers from overfitting to known concepts, limiting its generalizability to unseen concepts. To address this limitation, we propose Multimodal Visual Prompt Tuning (FedMVP) that conditions the prompts on multimodal contextual information - derived from the input image and textual attribute features of a class. At the core of FedMVP is a PromptFormer module that synergistically aligns textual and visual features through a cross-attention mechanism. The dynamically generated multimodal visual prompts are then input to the frozen vision encoder of CLIP, and trained with a combination of CLIP similarity loss and a consistency loss. Extensive evaluation on 20 datasets, spanning three generalization settings, demonstrates that FedMVP not only preserves performance on in-distribution classes and domains, but also displays higher generalizability to unseen classes and domains, surpassing state-of-the-art methods by a notable margin of +1.57% - 2.26%. Code is available at https://github.com/mainaksingha01/FedMVP. Mainak Singha, Subhankar Roy, Sarthak Mehrotra, Ankit Jha, Moloud Abdar, Biplab Banerjee, Elisa Ricci 0001 |
ICCV | 7 |
| 2025 | Test-Time Vocabulary Adaptation for Language-Driven Object DetectionabstractOpen-Vocabulary object detection models allow users to freely specify a class vocabulary in natural language at test time, guiding the detection of desired objects. However, vocabularies can be overly broad or even mis-specified, hampering the overall performance of the detector. In this work, we propose a plug-and-play Vocabulary Adapter (VocAda) to refine the user-defined vocabulary, automatically tailoring it to categories that are relevant for a given image. VocAda does not require any training, it operates at inference time in three steps: i) it uses an image captionner to describe visible objects, ii) it parses nouns from those captions, and iii) it selects relevant classes from the user-defined vocabulary, discarding irrelevant ones. Experiments on COCO and Objects365 with three state-of-the-art detectors show that VocAda consistently improves performance, proving its versatility. The code is open source. Tyler L. Hayes, Massimiliano Mancini, Elisa Ricci 0001, Riccardo Volpi, Gabriela Csurka |
ICIP | 4 |
| 2025 | Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection aims to identify humans and objects within images and interpret their interactions. Existing HOI methods rely heavily on large datasets with manual annotations to learn interactions from visual cues. These annotations are labor-intensive to create, prone to inconsistency, and limit scalability to new domains and rare interactions. We argue that recent advances in Vision-Language Models (VLMs) offer untapped potential, particularly in enhancing interaction representation. While prior work has injected such potential and even proposed training-free methods, there remain key gaps. Consequently, we propose a novel training-free HOI detection framework for Dynamic Scoring with enhanced semantics (dysco) that effectively utilizes textual and visual interaction representations within a multimodal registry, enabling robust and nuanced interaction understanding. This registry incorporates a small set of visual cues and uses innovative interaction signatures to improve the semantic alignment of verbs, facilitating effective generalization to rare interactions. Additionally, we propose a unique multi-head attention mechanism that adaptively weights the contributions of the visual and textual features. Experimental results demonstrate that our dysco surpasses training-free state-of-the-art models and is competitive with training-based approaches, particularly excelling in rare interactions. Code is available at https://github.com/francescotonini/dysco. Francesco Tonini, Lorenzo Vaquero, Alessandro Conti, Cigdem Beyan, Elisa Ricci 0001 |
ACM Multimedia | 5 |
| 2025 | LT-Soups: Bridging Head and Tail Classes via Subsampled Model SoupsabstractReal-world datasets typically exhibit long-tailed (LT) distributions, where a few head classes dominate and many tail classes are severely underrepresented. While recent work shows that parameter-efficient fine-tuning (PEFT) methods like LoRA and AdaptFormer preserve tail-class performance on foundation models such as CLIP, we find that they do so at the cost of head-class accuracy. We identify the head-tail ratio, the proportion of head to tail classes, as a crucial but overlooked factor influencing this trade-off. Through controlled experiments on CIFAR100 with varying imbalance ratio ($\rho$) and head-tail ratio ($\eta$), we show that PEFT excels in tail-heavy scenarios but degrades in more balanced and head-heavy distributions. To overcome these limitations, we propose LT-Soups, a two-stage model soups framework designed to generalize across diverse LT regimes. In the first stage, LT-Soups averages models fine-tuned on balanced subsets to reduce head-class bias; in the second, it fine-tunes only the classifier on the full dataset to restore head-class accuracy. Experiments across six benchmark datasets show that LT-Soups achieves superior trade-offs compared to both PEFT and traditional model soups across a wide range of imbalance regimes. Masih Aminbeidokhti, Subhankar Roy, Eric Granger, Elisa Ricci 0001, Marco Pedersoli |
NeurIPS | 4 |
| 2025 | ConViS-Bench: Estimating Video Similarity Through Semantic ConceptsabstractWhat does it mean for two videos to be similar? Videos may appear similar when judged by the actions they depict, yet entirely different if evaluated based on the locations where they were filmed. While humans naturally compare videos by taking different aspects into account, this ability has not been thoroughly studied and presents a challenge for models that often depend on broad global similarity scores. Large Multimodal Models (LMMs) with video understanding capabilities open new opportunities for leveraging natural language in comparative video tasks. We introduce Concept-based Video Similarity estimation (ConViS), a novel task that compares pairs of videos by computing interpretable similarity scores across a predefined set of key semantic concepts. ConViS allows for human-like reasoning about video similarity and enables new applications such as concept-conditioned video retrieval. To support this task, we also introduce ConViS-Bench, a new benchmark comprising carefully annotated video pairs spanning multiple domains. Each pair comes with concept-level similarity scores and textual descriptions of both differences and similarities. Additionally, we benchmark several state-of-the-art models on ConViS, providing insights into their alignment with human judgments. Our results reveal significant performance differences on ConViS, indicating that some concepts present greater challenges for estimating video similarity. We believe that ConViS-Bench will serve as a valuable resource for advancing research in language-driven video understanding. Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero, Yiming Wang 0002, Elisa Ricci 0001, Paolo Rota |
NeurIPS | 5 |
| 2025 | Training-free Online Video Step GroundingabstractGiven a task and a set of steps composing it, Video Step Grounding (VSG) aims to detect which steps are performed in a video. Standard approaches for this task require a labeled training set (e.g., with step-level annotations or narrations), which may be costly to collect. Moreover, they process the full video offline, limiting their applications for scenarios requiring online decisions. Thus, in this work, we explore how to perform VSG online and without training. We achieve this by exploiting the zero-shot capabilities of recent Large Multimodal Models (LMMs). In particular, we use LMMs to predict the step associated with a restricted set of frames, without access to the whole video. We show that this online strategy without task-specific tuning outperforms offline and training-based models. Motivated by this finding, we develop Bayesian Grounding with Large Multimodal Models (BaGLM), further injecting knowledge of past frames into the LMM-based predictions. BaGLM exploits Bayesian filtering principles, modeling step transitions via (i) a dependency matrix extracted through large language models and (ii) an estimation of step progress. Experiments on three datasets show superior performance of BaGLM over state-of-the-art training-based offline methods. Luca Zanella, Massimiliano Mancini, Yiming Wang 0002, Alessio Tonioni, Elisa Ricci 0001 |
NeurIPS | 5 |
| 2025 | One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question AnsweringabstractVision-Language Models (VLMs) have shown significant promise in Visual Question Answering (VQA) tasks by leveraging web-scale multimodal datasets. However, these models often struggle with continual learning due to catastrophic forgetting when adapting to new tasks. As an effective remedy to mitigate catastrophic forgetting, rehearsal strategy uses the data of past tasks upon learning new task. However, such strategy incurs the need of storing past data, which might not be feasible due to hardware constraints or privacy concerns. In this work, we propose the first data-free method that leverages the language generation capability of a VLM, instead of relying on external models, to produce pseudo-rehearsal data for addressing continual VQA. Our proposal, named as GaB, generates pseudo-rehearsal data by posing previous task questions on new task data. Yet, despite being effective, the distribution of generated questions skews towards the most frequently posed questions due to the limited and task-specific training data. To mitigate this issue, we introduce a pseudo-rehearsal balancing module that aligns the generated data towards the ground-truth data distribution using either the question meta-statistics or an unsupervised clustering method. We evaluate our proposed method on two recent benchmarks, i.e. VQACL- VQAv2 and CLOVE-function benchmarks. GaB outperforms all the data-free baselines with substantial improvement in maintaining VQA performance across evolving tasks, while being on-par with methods with access to the past data. Code and models are available at https://github.com/Deepayan137/GaB. Deepayan Das, Davide Talon, Massimiliano Mancini, Yiming Wang 0002, Elisa Ricci 0001 |
WACV | 5 |
| 2025 | Collaborative Neural Painting
Nicola Dall'Asen, Willi Menapace, Elia Peruzzo, Enver Sangineto, Yiming Wang 0002, Elisa Ricci 0001 |
Comput. Vis. Image Underst. | 6 |
| 2025 | Novel Class Discovery Meets Foundation Models for 3D Semantic Segmentation
Luigi Riz, Cristiano Saltori, Yiming Wang 0002, Elisa Ricci 0001, Fabio Poiesi |
Int. J. Comput. Vis. | 4 |
| 2025 | Revisiting Viewing Graph Solvability: An Effective Approach Based on Cycle ConsistencyabstractIn the structure from motion, the viewing graph is a graph where the vertices correspond to cameras (or images) and the edges represent the fundamental matrices. We provide a new formulation and an algorithm for determining whether a viewing graph is solvable, i.e., uniquely determines a set of projective cameras. The known theoretical conditions either do not fully characterize the solvability of all viewing graphs, or are extremely difficult to compute because they involve solving a system of polynomial equations with a large number of unknowns. The main result of this paper is a method to reduce the number of unknowns by exploiting cycle consistency. We advance the understanding of solvability by (i) finishing the classification of all minimal graphs up to 9 nodes, (ii) extending the practical verification of solvability to minimal graphs with up to 90 nodes, (iii) finally answering an open research question by showing that finite solvability is not equivalent to solvability, and (iv) formally drawing the connection with the calibrated case (i.e., parallel rigidity). Finally, we present an experiment on real data that shows that unsolvable graphs may appear in practice. Federica Arrigoni, Andrea Fusiello, Romeo Rizzi, Elisa Ricci 0001, Tomás Pajdla |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Swarmchestrate: Towards a Fully Decentralised Framework for Orchestrating Applications in the Cloud-to-Edge Continuum
Tamás Kiss, Amjad Ullah, Gábor Terstyánszky, Odej Kao, Sören Becker 0001, Giannis Verginadis, Antonis Michalas, Vlado Stankovski, Attila Kertész, Elisa Ricci 0001, Jörn Altmann, Bernhard Egger 0002, Francesco Tusa, József Kovács, Róbert Lovas |
AINA (5) | 10 |
| 2024 | MULTIFLOW: Shifting Towards Task-Agnostic Vision-Language PruningabstractWhile excellent in transfer learning, Vision-Language models (VLMs) come with high computational costs due to their large number of parameters. To address this issue, removing parameters via model pruning is a viable solution. However, existing techniques for VLMs are task-specific, and thus require pruning the network from scratch for each new task of interest. In this work, we explore a new direction: Task-Agnostic Vision-Language Pruning (TA-VLP). Given a pretrained VLM, the goal is to find a unique pruned counterpart transferable to multiple unknown downstream tasks. In this challenging setting, the transferable representations already encoded in the pretrained model are a key aspect to preserve. Thus, we propose Multimodal Flow Pruning (MULTIFLOW), a first, gradient-free, pruning framework for TA-VLP where: (i) the importance of a parameter is expressed in terms of its magnitude and its information flow, by incorporating the saliency of the neurons it connects; and (ii) pruning is driven by the emergent (multimodal) distribution of the VLM parameters after pretraining. We benchmark eight state-of-the-art pruning algorithms in the context of TA-VLP, experimenting with two VLMs, three vision-language tasks, and three pruning ratios. Our experimental results show that MULTIFLOW outperforms recent sophisticated, combinatorial competitors in the vast majority of the cases, paving the way towards addressing TA- VLP. The code is publicly available at https://github.com/FarinaMatteo/multiflow. Matteo Farina, Massimiliano Mancini, Elia Cunegatti, Gaowen Liu, Giovanni Iacca, Elisa Ricci 0001 |
CVPR | 6 |
| 2024 | Test-Time Zero-Shot Temporal Action LocalizationabstractZero-Shot Temporal Action Localization (ZS-TAL) seeks to identify and locate actions in untrimmed videos unseen during training. Existing ZS-TAL methods involve fine-tuning a model on a large amount of annotated training data. While effective, training-based ZS-TAL approaches assume the availability of labeled data for supervised learning, which can be impractical in some applications. Furthermore, the training process naturally induces a domain bias into the learned model, which may adversely affect the model's generalization ability to arbitrary videos. These considerations prompt us to approach the ZS-TAL problem from a radically novel perspective, relaxing the requirement for training data. To this aim, we introduce a novel method that performs Test-Time adaptation for Temporal Action Localization (T3AL). In a nutshell, T3AL adapts a pre-trained Vision and Language Model (VLM). T3AL operates in three steps. First, a video-level pseudo-label of the action category is computed by aggregating information from the entire video. Then, action localization is performed adopting a novel procedure inspired by self-supervised learning. Finally, frame-level textual descriptions extracted with a state-of-the-art captioning model are employed for refining the action region proposals. We validate the effectiveness of T3AL by conducting experiments on the THUMOS14 and the ActivityNet-v1.3 datasets. Our results demonstrate that T3AL significantly outperforms zero-shot baselines based on state-of-the-art VLMs, confirming the benefit of a test-time adaptation approach. Benedetta Liberatori, Alessandro Conti, Paolo Rota, Yiming Wang 0002, Elisa Ricci 0001 |
CVPR | 5 |
| 2024 | SHiNe: Semantic Hierarchy Nexus for Open-Vocabulary Object DetectionabstractOpen-vocabulary object detection (OvOD) has transformed detection into a language-guided task, empowering users to freely define their class vocabularies of interest during inference. However, our initial investigation indicates that existing OvOD detectors exhibit significant vari-ability when dealing with vocabularies across various semantic granularities, posing a concern for real-world deployment. To this end, we introduce Semantic Hierarchy Nexus (SHiNe), a novel classifier that uses semantic knowledge from class hierarchies. It runs offline in three steps: i) it retrieves relevant super-/sub-categories from a hierar-chy for each target class; ii) it integrates these categories into hierarchy-aware sentences; iii) it fuses these sentence embeddings to generate the nexus classifier vector. Our evaluation on various detection benchmarks demonstrates that SHiNe enhances robustness across diverse vocabulary granularities, achieving up to +31.9% mAP50 with ground truth hierarchies, while retaining improvements using hierarchies generated by large language models. Moreover, when applied to open-vocabulary classification on ImageNet-1k, SHiNe improves the CLIP zero-shot baseline by +2.8% accuracy. SHiNe is training-free and can be seamlessly integrated with any off-the-shelf OvOD detector, without incurring additional computational overhead during inference. The code is open source. Tyler L. Hayes, Elisa Ricci 0001, Gabriela Csurka, Riccardo Volpi |
CVPR | 3 |
| 2024 | Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video SynthesisabstractContemporary models for generating images show remarkable quality and versatility. Swayed by these advantages, the research community repurposes them to generate videos. Since video content is highly redundant, we argue that naively bringing advances of image models to the video generation domain reduces motion fidelity, visual quality and impairs scalability. In this work, we build Snap Video, a video-first model that systematically addresses these challenges. To do that, we first extend the EDM framework to take into account spatially and temporally redundant pixels and naturally support video generation. Second, we show that a U-Net—a workhorse behind image generation—scales poorly when generating videos, requiring significant computational overhead. Hence, we propose a new transformer-based architecture that trains 3.31 times faster than U-Nets (and is ∼4.5 faster at inference). This allows us to efficiently train a text-to-video model with billions of parameters for the first time, reach state-of-the-art results on a number of benchmarks, and generate videos with substantially higher quality, temporal consistency, and motion complexity. The user studies showed that our model was favored by a large margin over the most recent methods. Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov, Ekaterina Deyneka, Tsai-Shien Chen, Anil Kag, Yuwei Fang, Aleksei Stoliar, Elisa Ricci 0001, Jian Ren 0005, Sergey Tulyakov |
CVPR | 9 |
| 2024 | Harnessing Large Language Models for Training-Free Video Anomaly DetectionabstractVideo anomaly detection (VAD) aims to temporally locate abnormal events in a video. Existing works mostly rely on training deep models to learn the distribution of normality with either video-level supervision, one-class supervision, or in an unsupervised setting. Training-based methods are prone to be domain-specific, thus being costly for practical deployment as any domain change will involve data collection and model training. In this paper, we radically depart from previous efforts and propose LAnguage-based VAD (LAVAD), a method tackling VAD in a novel, training-free paradigm, exploiting the capabilities of pre-trained large language models (LLMs) and existing vision-language models (VLMs). We leverage VLM-based captioning models to generate textual descriptions for each frame of any test video. With the textual scene description, we then devise a prompting mechanism to unlock the capability of LLMs in terms of temporal aggregation and anomaly score estimation, turning LLMs into an effective video anomaly detector. We further leverage modality-aligned VLMs and propose effective techniques based on cross-modal similarity for cleaning noisy captions and refining the LLM-based anomaly scores. We evaluate LAVAD on two large datasets featuring real-world surveillance scenarios (UCF-Crime and XD- Violence), showing that it outperforms both unsupervised and one-class methods without requiring any training or data collection. Luca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang 0002, Elisa Ricci 0001 |
CVPR | 5 |
| 2024 | Retrieval-enriched zero-shot image classification in low-resource domainsabstractLow-resource domains, characterized by scarce data and annotations, present significant challenges for language and visual understanding tasks, with the latter much under-explored in the literature.Recent advancements in Vision-Language Models (VLM) have shown promising results in high-resource domains but fall short in low-resource concepts that are underrepresented (e.g.only a handful of images per category) in the pre-training set.We tackle the challenging task of zero-shot low-resource image classification from a novel perspective.By leveraging a retrieval-based strategy, we achieve this in a training-free fashion.Specifically, our method, named CORE (Combination of Retrieval Enrichment), enriches the representation of both query images and class prototypes by retrieving relevant textual information from large web-crawled databases.This retrieval-based enrichment significantly boosts classification performance by incorporating the broader contextual information relevant to the specific class.We validate our method on a newly established benchmark covering diverse low-resource domains, including medical imaging, rare plants, and circuits.Our experiments demonstrate that CORE outperforms existing state-of-the-art methods that rely on synthetic data generation and model fine-tuning. Nicola Dall'Asen, Yiming Wang 0002, Enrico Fini, Elisa Ricci 0001 |
EMNLP | 4 |
| 2024 | Enhancing the Domain Robustness of Self-Supervised pre-Training with Synthetic ImagesabstractWe present a novel method for improving the adaptability of self-supervised (SSL) pre-trained models across different domains. Our approach uses synthetic images that are generated using an auxiliary diffusion model, namely InstructPix2Pix. More specifically, starting from a real image, we prompt the diffusion model to generate synthetic versions of that image in the style of the target domains. This allows us to generate a diverse set of multi-domain images that share the same semantics as real images. Integrating these synthetic images into the training dataset enhances the model’s capacity to generalize to other domains. We pre-trained different SSL methods on Imagenet-100 with and without the synthetic images and evaluated their performance on three multi-domain datasets, DomainNet, PACS, and Office-Home. Our results show significant improvements in all datasets and methods, encouraging new research in the direction of leveraging synthetic data to improve the robustness of pre-trained models. Code is available at https://github.com/has97/Diffusion_pre-training. Mohamad Hassan N C, Avigyan Bhattacharya, Victor G. T. da Costa, Biplab Banerjee, Elisa Ricci 0001 |
ICASSP | 5 |
| 2024 | Democratizing Fine-grained Visual Recognition with Large Language ModelsabstractIdentifying subordinate-level categories from images is a longstanding task in computer vision and is referred to as fine-grained visual recognition (FGVR). It has tremendous significance in real-world applications since an average layperson does not excel at differentiating species of birds or mushrooms due to subtle differences among the species. A major bottleneck in developing FGVR systems is caused by the need of high-quality paired expert annotations. To circumvent the need of expert knowledge we propose Fine-grained Semantic Category Reasoning (FineR) that internally leverages the world knowledge of large language models (LLMs) as a proxy in order to reason about fine-grained category names. In detail, to bridge the modality gap between images and LLM, we extract part-level visual attributes from images as text and feed that information to a LLM. Based on the visual attributes and its internal world knowledge the LLM reasons about the subordinate-level category names. Our training-free FineR outperforms several state-of-the-art FGVR and language and vision assistant models and shows promise in working in the wild and in new domains where gathering expert annotation is arduous. Subhankar Roy, Wenjing Li 0005, Zhun Zhong, Nicu Sebe, Elisa Ricci 0001 |
ICLR | 6 |
| 2024 | Text-Enhanced Zero-Shot Action Recognition: A Training-Free Approach
Massimo Bosetti, Shibingfeng Zhang, Benedetta Liberatori, Giacomo Zara, Elisa Ricci 0001, Paolo Rota |
ICPR (15) | 5 |
| 2024 | Conditioned Prompt-Optimization for Continual Deepfake Detection
Francesco Laiti, Benedetta Liberatori, Thomas De Min, Elisa Ricci 0001 |
ICPR (9) | 4 |
| 2024 | Large-Scale Pre-trained Models are Surprisingly Strong in Incremental Novel Class Discovery
Subhankar Roy, Zhun Zhong, Nicu Sebe, Elisa Ricci 0001 |
ICPR (16) | 5 |
| 2024 | AL-GTD: Deep Active Learning for Gaze Target DetectionabstractGaze target detection aims at determining the image location where a person is looking. While existing studies have made significant progress in this area by regressing accurate gaze heatmaps, these achievements have largely relied on access to extensive labeled datasets, which demands substantial human labor. In this paper, our goal is to reduce the reliance on the size of labeled training data for gaze target detection. To achieve this, we propose AL-GTD, an innovative approach that integrates supervised and self-supervised losses within a novel sample acquisition function to perform active learning (AL). Additionally, it utilizes pseudo-labeling to mitigate distribution shifts during the training phase. AL-GTD achieves the best of all AUC results by utilizing only 40-50% of the training data, in contrast to state-of-the-art (SOTA) gaze target detectors requiring the entire training dataset to achieve the same performance. Importantly, AL-GTD quickly reaches satisfactory performance with 10-20% of the training data, showing the effectiveness of our acquisition function, which is able to acquire the most informative samples. We provide a comprehensive experimental analysis by adapting several AL methods for the task. AL-GTD outperforms AL competitors, simultaneously exhibiting superior performance compared to SOTA gaze target detectors when all are trained within a low-data regime. Code is available at: https://github.com/francescotonini/al-gtd. Francesco Tonini, Nicola Dall'Asen, Lorenzo Vaquero, Cigdem Beyan, Elisa Ricci 0001 |
ACM Multimedia | 5 |
| 2024 | Frustratingly Easy Test-Time Adaptation of Vision-Language ModelsabstractVision-Language Models seamlessly discriminate among arbitrary semantic categories, yet they still suffer from poor generalization when presented with challenging examples. For this reason, Episodic Test-Time Adaptation (TTA) strategies have recently emerged as powerful techniques to adapt VLMs in the presence of a single unlabeled image. The recent literature on TTA is dominated by the paradigm of prompt tuning by Marginal Entropy Minimization, which, relying on online backpropagation, inevitably slows down inference while increasing memory. In this work, we theoretically investigate the properties of this approach and unveil that a surprisingly strong TTA method lies dormant and hidden within it. We term this approach ZERO (TTA with “zero” temperature), whose design is both incredibly effective and frustratingly simple: augment N times, predict, retain the most confident predictions, and marginalize after setting the Softmax temperature to zero. Remarkably, ZERO requires a single batched forward pass through the vision encoder only and no backward passes. We thoroughly evaluate our approach following the experimental protocol established in the literature and show that ZERO largely surpasses or compares favorably w.r.t. the state-of-the-art while being almost 10× faster and 13× more memory friendly than standard Test-Time Prompt Tuning. Thanks to its simplicity and comparatively negligible computation, ZERO can serve as a strong baseline for future work in this field. Code will be available. Matteo Farina, Gianni Franchi, Giovanni Iacca, Massimiliano Mancini, Elisa Ricci 0001 |
NeurIPS | 5 |
| 2024 | StyLIP: Multi-Scale Style-Conditioned Prompt Learning for CLIP-based Domain GeneralizationabstractLarge-scale foundation models, such as CLIP, have demonstrated impressive zero-shot generalization performance on downstream tasks, leveraging well-designed language prompts. However, these prompt learning techniques often struggle with domain shift, limiting their generalization capabilities. In our study, we tackle this issue by proposing StyLIP, a novel approach for Domain Generalization (DG) that enhances CLIP’s classification performance across domains. Our method focuses on a domain-agnostic prompt learning strategy, aiming to disentangle the visual style and content information embedded in CLIP’s pre-trained vision encoder, enabling effortless adaptation to novel domains during inference. To achieve this, we introduce a set of style projectors that directly learn the domain-specific prompt tokens from the extracted multi-scale style features. These generated prompt embeddings are subsequently combined with the multi-scale visual content features learned by a content projector. The projectors are trained in a contrastive manner, utilizing CLIP’s fixed vision and text backbones. Through extensive experiments conducted in five different DG settings on multiple benchmark datasets, we consistently demonstrate that StyLIP outperforms the current state-of-the-art (SOTA) methods. Shirsha Bose, Ankit Jha, Enrico Fini, Mainak Singha, Elisa Ricci 0001, Biplab Banerjee |
WACV | 5 |
| 2024 | Delving into CLIP latent space for Video Anomaly Recognition
Luca Zanella, Benedetta Liberatori, Willi Menapace, Fabio Poiesi, Yiming Wang 0002, Elisa Ricci 0001 |
Comput. Vis. Image Underst. | 6 |
| 2024 | Simplifying open-set video domain adaptation with contrastive learning
Giacomo Zara, Victor G. T. da Costa, Subhankar Roy, Paolo Rota, Elisa Ricci 0001 |
Comput. Vis. Image Underst. | 5 |
| 2024 | Unsupervised Point Cloud Representation Learning by Clustering and Neural RenderingabstractAbstract Data augmentation has contributed to the rapid advancement of unsupervised learning on 3D point clouds. However, we argue that data augmentation is not ideal, as it requires a careful application-dependent selection of the types of augmentations to be performed, thus potentially biasing the information learned by the network during self-training. Moreover, several unsupervised methods only focus on uni-modal information, thus potentially introducing challenges in the case of sparse and textureless point clouds. To address these issues, we propose an augmentation-free unsupervised approach for point clouds, named CluRender, to learn transferable point-level features by leveraging uni-modal information for soft clustering and cross-modal information for neural rendering. Soft clustering enables self-training through a pseudo-label prediction task, where the affiliation of points to their clusters is used as a proxy under the constraint that these pseudo-labels divide the point cloud into approximate equal partitions. This allows us to formulate a clustering loss to minimize the standard cross-entropy between pseudo and predicted labels. Neural rendering generates photorealistic renderings from various viewpoints to transfer photometric cues from 2D images to the features. The consistency between rendered and real images is then measured to form a fitting loss, combined with the cross-entropy loss to self-train networks. Experiments on downstream applications, including 3D object detection, semantic segmentation, classification, part segmentation, and few-shot learning, demonstrate the effectiveness of our framework in outperforming state-of-the-art techniques. Guofeng Mei, Cristiano Saltori, Elisa Ricci 0001, Nicu Sebe, Qiang Wu 0001, Jian Zhang 0002, Fabio Poiesi |
Int. J. Comput. Vis. | 3 |
| 2024 | 3D Object Detection From Images for Autonomous Driving: A Surveyabstract3D object detection from images, one of the fundamental and challenging problems in autonomous driving, has received increasing attention from both industry and academia in recent years. Benefiting from the rapid development of deep learning technologies, image-based 3D detection has achieved remarkable progress. Particularly, more than 200 works have studied this problem from 2015 to 2021, encompassing a broad spectrum of theories, algorithms, and applications. However, to date no recent survey exists to collect and organize this knowledge. In this paper, we fill this gap in the literature and provide the first comprehensive survey of this novel and continuously growing research field, summarizing the most commonly used pipelines for image-based 3D detection and deeply analyzing each of their components. Additionally, we also propose two new taxonomies to organize the state-of-the-art methods into different categories, with the intent of providing a more systematic review of existing methods and facilitating fair comparisons with future works. In retrospect of what has been achieved so far, we also analyze the current challenges in the field and discuss future directions for image-based 3D detection research. Xinzhu Ma, Wanli Ouyang, Andrea Simonelli, Elisa Ricci 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Promptable Game Models: Text-guided Game Simulation via Masked Diffusion ModelsabstractNeural video game simulators emerged as powerful tools to generate and edit videos. Their idea is to represent games as the evolution of an environment’s state driven by the actions of its agents. While such a paradigm enables users toplaya game action-by-action, its rigidity precludes more semantic forms of control. To overcome this limitation, we augment game models withpromptsspecified as a set ofnatural languageactions anddesired states. The result—a Promptable Game Model (PGM)—makes it possible for a user toplaythe game by prompting it with high- and low-level action sequences. Most captivatingly, our PGM unlocks thedirector’s mode, where the game is played by specifying goals for the agents in the form of a prompt. This requires learning “game AI,” encapsulated by our animation model, to navigate the scene using high-level constraints, play against an adversary, and devise a strategy to win a point. To render the resulting state, we use a compositional NeRF representation encapsulated in our synthesis model. To foster future research, we present newly collected, annotated and calibrated Tennis and Minecraft datasets. Our method significantly outperforms existing neural video game simulators in terms of rendering quality and unlocks applications beyond the capabilities of the current state-of-the-art. Our framework, data, and models are available at snap-research.github.io/promptable-game-models. Willi Menapace, Aliaksandr Siarohin, Stéphane Lathuilière, Panos Achlioptas, Vladislav Golyanik, Sergey Tulyakov, Elisa Ricci 0001 |
ACM Trans. Graph. | 7 |
| 2023 | Quantum Multi-Model FittingabstractGeometric model fitting is a challenging but fundamental computer vision problem. Recently, quantum optimization has been shown to enhance robust fitting for the case of a single model, while leaving the question of multi-model fitting open. In response to this challenge, this paper shows that the latter case can significantly benefit from quantum hardware and proposes the first quantum approach to multimodel fitting (MMF). We formulate MMF as a problem that can be efficiently sampled by modern adiabatic quantum computers without the relaxation of the objective function. We also propose an iterative and decomposed version of our method, which supports real-world-sized problems. The experimental evaluation demonstrates promising results on a variety of datasets. The source code is available at: https://github.com/FarinaMatteo/qmmf. Matteo Farina, Luca Magri 0002, Willi Menapace, Elisa Ricci 0001, Vladislav Golyanik, Federica Arrigoni |
CVPR | 4 |
| 2023 | Semi-supervised learning made simple with self-supervised clusteringabstractSelf-supervised learning models have been shown to learn rich visual representations without requiring human annotations. However, in many real-world scenarios, labels are partially available, motivating a recent line of work on semi-supervised methods inspired by self-supervised principles. In this paper, we propose a conceptually simple yet empirically powerful approach to turn clustering-based self-supervised methods such as SwAV or DINO into semi-supervised learners. More precisely, we introduce a multi-task framework merging a supervised objective using ground-truth labels and a self-supervised objective relying on clustering assignments with a single cross-entropy loss. This approach may be interpreted as imposing the cluster centroids to be class prototypes. Despite its simplicity, we provide empirical evidence that our approach is highly effective and achieves state-of-the-art performance on CI-FAR100 and ImageNet. Enrico Fini, Pietro Astolfi, Karteek Alahari, Xavier Alameda-Pineda, Julien Mairal, Moin Nabi, Elisa Ricci 0001 |
CVPR | 7 |
| 2023 | Novel Class Discovery for 3D Point Cloud Semantic SegmentationabstractNovel class discovery (NCD) for semantic segmentation is the task of learning a model that can segment unlabelled (novel) classes using only the supervision from labelled (base) classes. This problem has recently been pioneered for 2D image data, but no work exists for 3D point cloud data. In fact, the assumptions made for 2D are loosely applicable to 3D in this case. This paper is presented to advance the state of the art on point cloud data analysis in four directions. Firstly, we address the new problem of NCD for point cloud semantic segmentation. Secondly, we show that the transposition of the only existing NCD method for 2D semantic segmentation to 3D data is sub-optimal. Thirdly, we present a new method for NCD based on online clustering that exploits uncertainty quantification to produce prototypes for pseudo-labelling the points of the novel classes. Lastly, we introduce a new evaluation protocol to assess the performance of NCD for point cloud semantic segmentation. We thoroughly evaluate our method on SemanticKITTI and SemanticPOSS datasets, showing that it can significantly outperform the baseline. Project page: https://github.com/LuigiRiz/NOPS. Luigi Riz, Cristiano Saltori, Elisa Ricci 0001, Fabio Poiesi |
CVPR | 3 |
| 2023 | AutoLabel: CLIP-based framework for Open-Set Video Domain AdaptationabstractOpen-set Unsupervised Video Domain Adaptation (OU-VDA) deals with the task of adapting an action recognition model from a labelled source domain to an unlabelled target domain that contains “target-private” categories, which are present in the target but absent in the source. In this work we deviate from the prior work of training a specialized open-set classifier or weighted adversarial learning by proposing to use pre-trained Language and Vision Models (CLIP). The CLIP is well suited for OUVDA due to its rich representation and the zero-shot recognition capabilities. However, rejecting target-private instances with the CLIP's zero-shot protocol requires oracle knowledge about the target-private label names. To circumvent the impossibility of the knowledge of label names, we propose AutoLabel that automatically discovers and generates object-centric compositional candidate target-private class names. Despite its simplicity, we show that CLIP when equipped with AutoLabel can satisfactorily reject the target-private instances, thereby facilitating better alignment between the shared classes of the two domains. The code is available11https://github.com/gzaraunitn/autolabel. Giacomo Zara, Subhankar Roy, Paolo Rota, Elisa Ricci 0001 |
CVPR | 4 |
| 2023 | A soft nearest-neighbor framework for continual semi-supervised learningabstractDespite significant advances, the performance of state-of-the-art continual learning approaches hinges on the unrealistic scenario of fully labeled data. In this paper, we tackle this challenge and propose an approach for continual semi-supervised learning—a setting where not all the data samples are labeled. A primary issue in this scenario is the model forgetting representations of unlabeled data and overfitting the labeled samples. We leverage the power of nearest-neighbor classifiers to nonlinearly partition the feature space and flexibly model the underlying data distribution thanks to its non-parametric nature. This enables the model to learn a strong representation for the current task, and distill relevant information from previous tasks. We perform a thorough experimental evaluation and show that our method outperforms all the existing approaches by large margins, setting a solid state of the art on the continual semi-supervised learning paradigm. For example, on CIFAR-100 we surpass several others even when using at least 30 times less supervision (0.8% vs. 25% of annotations). Finally, our method works well on both low and high resolution images and scales seamlessly to more complex datasets such as ImageNet-100. Our source code is publicly available at https://github.com/kangzhiq/NNCSL. Zhiqi Kang, Enrico Fini, Moin Nabi, Elisa Ricci 0001, Karteek Alahari |
ICCV | 4 |
| 2023 | Walking Your LiDOG: A Journey Through Multiple Domains for LiDAR Semantic SegmentationabstractThe ability to deploy robots that can operate safely in diverse environments is crucial for developing embodied intelligent agents. As a community, we have made tremendous progress in within-domain LiDAR semantic segmentation. However, do these methods generalize across domains? To answer this question, we design the first experimental setup for studying domain generalization (DG) for LiDAR semantic segmentation (DG-LSS). Our results confirm a significant gap between methods, evaluated in a cross-domain setting: for example, a model trained on the source dataset (SemanticKITTI) obtains 26.53 mIoU on the target data, compared to 48.49 mIoU obtained by the model trained on the target domain (nuScenes). To tackle this gap, we propose the first method specifically designed for DG-LSS, which obtains 34.88 mIoU on the target domain, outperforming all baselines. Our method augments a sparse-convolutional encoder-decoder 3D segmentation network with an additional, dense 2D convolutional decoder that learns to classify a birds-eye view of the point cloud. This simple auxiliary task encourages the 3D network to learn features that are robust to sensor placement shifts and resolution, and are transferable across domains. With this work, we aim to in spire the community to develop and evaluate future models in such cross-domain conditions. Cristiano Saltori, Aljosa Osep, Elisa Ricci 0001, Laura Leal-Taixé |
ICCV | 3 |
| 2023 | Object-aware Gaze Target DetectionabstractGaze target detection aims to predict the image location where the person is looking and the probability that a gaze is out of the scene. Several works have tackled this task by regressing a gaze heatmap centered on the gaze location, however, they overlooked decoding the relationship between the people and the gazed objects. This paper proposes a Transformer-based architecture that automatically detects objects (including heads) in the scene to build associations between every head and the gazed-head/object, resulting in a comprehensive, explainable gaze analysis composed of: gaze target area, gaze pixel point, the class and the image location of the gazed-object. Upon evaluation of the in-the-wild benchmarks, our method achieves state-of-the-art results on all metrics (up to 2.91% gain in AUC, 50% reduction in gaze distance, and 9% gain in out-of-frame average precision) for gaze target detection and 11-13% improvement in average precision for the classification and the localization of the gazed-objects. The code of the proposed method is publicly available1. Francesco Tonini, Nicola Dall'Asen, Cigdem Beyan, Elisa Ricci 0001 |
ICCV | 4 |
| 2023 | The Unreasonable Effectiveness of Large Language-Vision Models for Source-free Video Domain AdaptationabstractSource-Free Video Unsupervised Domain Adaptation (SFVUDA) task consists in adapting an action recognition model, trained on a labelled source dataset, to an unlabelled target dataset, without accessing the actual source data. The previous approaches have attempted to address SFVUDA by leveraging self-supervision (e.g., enforcing temporal consistency) derived from the target data itself. In this work, we take an orthogonal approach by exploiting "web-supervision" from Large Language-Vision Models (LLVMs), driven by the rationale that LLVMs contain a rich world prior surprisingly robust to domain-shift. We showcase the unreasonable effectiveness of integrating LLVMs for SFVUDA by devising an intuitive and parameter-efficient method, which we name Domain Adaptation with Large Language-Vision models (DALL-V), that distills the world prior and complementary source model information into a student network tailored for the target. Despite the simplicity, DALL-V1achieves significant improvement over state-of-the-art SFVUDA methods. Giacomo Zara, Alessandro Conti, Subhankar Roy, Stéphane Lathuilière, Paolo Rota, Elisa Ricci 0001 |
ICCV | 6 |
| 2023 | Exploring Diffusion Models for Unsupervised Video Anomaly DetectionabstractThis paper investigates the performance of diffusion models for video anomaly detection (VAD) within the most challenging but also the most operational scenario in which the data annotations are not used. As being sparse, diverse, contextual, and often ambiguous, detecting abnormal events precisely is a very ambitious task. To this end, we rely only on the information-rich spatio-temporal data, and the reconstruction power of the diffusion models such that a high reconstruction error is utilized to decide the abnormality. Experiments performed on two large-scale video anomaly detection datasets demonstrate the consistent improvement of the proposed method over the state-of-the-art generative models while in some cases our method achieves better scores than the more complex models. This is the first study using a diffusion model and examining its parameters’ influence to present guidance for VAD in surveillance scenarios. Anil Osman Tur, Nicola Dall'Asen, Cigdem Beyan, Elisa Ricci 0001 |
ICIP | 4 |
| 2023 | Rotation Synchronization via Deep Matrix FactorizationabstractIn this paper we address the rotation synchronization problem, where the objective is to recover absolute rotations starting from pairwise ones, where the unknowns and the measures are represented as nodes and edges of a graph, respectively. This problem is an essential task for structure from motion and simultaneous localization and mapping. We focus on the formulation of synchronization via neural networks, which has only recently begun to be explored in the literature. Inspired by deep matrix completion, we express rotation synchronization in terms of matrix factorization with a deep neural network. Our formulation exhibits implicit regularization properties and, more importantly, is unsupervised, whereas previous deep approaches are supervised. Our experiments show that we achieve comparable accuracy to the closest competitors in most scenes, while working under weaker assumptions. GK Tejus, Giacomo Zara, Paolo Rota, Andrea Fusiello, Elisa Ricci 0001, Federica Arrigoni |
ICRA | 5 |
| 2023 | Recognizing Actions in Videos under Domain ShiftabstractAction recognition, which consists in automatically recognizing the action being performed in a video sequence, is a fundamental task in computer vision and multimedia. Supervised action recognition has been widely studied because of the growing need for automatically categorizing video content that are being generated everyday. However, it is nearly impossible for human annotators to keep pace with the enormous volumes of online videos, and thus supervised training becomes infeasible. A cheaper way of leveraging the massive pool of unlabelled data is by exploiting an already trained model to infer the labels on such data and then re-using them to build an improved model. Such an approach is also prone to failure because the unlabelled data may belong to a data distribution that is different from the annotated one. This is often referred to as the domain-shift problem. To address the domain-shift, recently Unsupervised Video Domain Adaptation (UVDA) methods have been proposed. However, these methods typically make strong and unrealistic assumptions. In this talk I will present some recent works of my research group on UVDA, showing that, thanks to recent advances in deep architectures and to the advent of foundation models, it is possible to deal with more challenging and realistic settings and recognize out-of-distribution classes. Elisa Ricci 0001 |
ICMR | 1 |
| 2023 | Vocabulary-free Image ClassificationabstractRecent advances in large vision-language models have revolutionized the image classification paradigm. Despite showing impressive zero-shot capabilities, a pre-defined set of categories, a.k.a. the vocabulary, is assumed at test time for composing the textual prompts. However, such assumption can be impractical when the semantic context is unknown and evolving. We thus formalize a novel task, termed as Vocabulary-free Image Classification (VIC), where we aim to assign to an input image a class that resides in an unconstrained language-induced semantic space, without the prerequisite of a known vocabulary. VIC is a challenging task as the semantic space is extremely large, containing millions of concepts, with hard-to-discriminate fine-grained categories. In this work, we first empirically verify that representing this semantic space by means of an external vision-language database is the most effective way to obtain semantically relevant content for classifying the image. We then propose Category Search from External Databases (CaSED), a method that exploits a pre-trained vision-language model and an external vision-language database to address VIC in a training-free manner. CaSED first extracts a set of candidate categories from captions retrieved from the database based on their semantic similarity to the image, and then assigns to the image the best matching candidate category according to the same vision-language model. Experiments on benchmark datasets validate that CaSED outperforms other complex vision-language frameworks, while being efficient with much fewer parameters, paving the way for future research in this direction. Alessandro Conti, Enrico Fini, Massimiliano Mancini, Paolo Rota, Yiming Wang 0002, Elisa Ricci 0001 |
NeurIPS | 6 |
| 2023 | ConfMix: Unsupervised Domain Adaptation for Object Detection via Confidence-based MixingabstractUnsupervised Domain Adaptation (UDA) for object detection aims to adapt a model trained on a source domain to detect instances from a new target domain for which annotations are not available. Different from traditional approaches, we propose ConfMix, the first method that introduces a sample mixing strategy based on region-level detection confidence for adaptive object detector learning. We mix the local region of the target sample that corresponds to the most confident pseudo detections with a source image, and apply an additional consistency loss term to gradually adapt towards the target data distribution. In order to robustly define a confidence score for a region, we exploit the confidence score per pseudo detection that accounts for both the detector-dependent confidence and the bounding box uncertainty. Moreover, we propose a novel pseudo labelling scheme that progressively filters the pseudo target detections using the confidence metric that varies from a loose to strict manner along the training. We perform extensive experiments with three datasets, achieving state-of-the-art performance in two of them and approaching the supervised target model performance in the other. Code is available at https://github.com/giuliomattolin/ConfMix. Giulio Mattolin, Luca Zanella, Elisa Ricci 0001, Yiming Wang 0002 |
WACV | 3 |
| 2023 | Overlap-guided Gaussian Mixture Models for Point Cloud RegistrationabstractProbabilistic 3D point cloud registration methods have shown competitive performance in overcoming noise, outliers, and density variations. However, registering point cloud pairs in the case of partial overlap is still a challenge. This paper proposes a novel overlap-guided probabilistic registration approach that computes the optimal transformation from matched Gaussian Mixture Model (GMM) parameters. We reformulate the registration problem as the problem of aligning two Gaussian mixtures such that a statistical discrepancy measure between the two corresponding mixtures is minimized. We introduce a Transformer-based detection module to detect overlapping regions, and represent the input point clouds using GMMs by guiding their alignment through overlap scores computed by this detection module. Experiments show that our method achieves superior registration accuracy and efficiency than state-of-the-art methods when handling point clouds with partial overlap and different densities on synthetic and real-world datasets. https://github.com/gfmei/ogmm Guofeng Mei, Fabio Poiesi, Cristiano Saltori, Jian Zhang 0002, Elisa Ricci 0001, Nicu Sebe |
WACV | 5 |
| 2023 | Cooperative Self-Training for Multi-Target Adaptive Semantic SegmentationabstractIn this work we address multi-target domain adaptation (MTDA) in semantic segmentation, which consists in adapting a single model from an annotated source dataset to multiple unannotated target datasets that differ in their underlying data distributions. To address MTDA, we propose a self-training strategy that employs pseudo-labels to induce cooperation among multiple domain-specific classifiers. We employ feature stylization as an efficient way to generate image views that forms an integral part of self-training. Additionally, to prevent the network from overfitting to noisy pseudo-labels, we devise a rectification strategy that leverages the predictions from different classifiers to estimate the quality of pseudo-labels. Our extensive experiments on numerous settings, based on four different semantic segmentation datasets, validates the effectiveness of the proposed self-training strategy and shows that our method outperforms state-of-the-art MTDA approaches. https://github.com/Mael-zys/CoaST. Yangsong Zhang 0002, Subhankar Roy, Hongtao Lu 0001, Elisa Ricci 0001, Stéphane Lathuilière |
WACV | 4 |
| 2023 | Self-training transformer for source-free domain adaptation
Guanglei Yang, Zhun Zhong, Mingli Ding, Nicu Sebe, Elisa Ricci 0001 |
Appl. Intell. | 5 |
| 2023 | Interactive Neural PaintingabstractIn the last few years, Neural Painting (NP) techniques became capable of producing extremely realistic artworks. This paper advances the state of the art in this emerging research domain by proposing the first approach for Interactive NP. Considering a setting where a user looks at a scene and tries to reproduce it on a painting, our objective is to develop a computational framework to assist the user’s creativity by suggesting the next strokes to paint, that can be possibly used to complete the artwork. To accomplish such a task, we propose I-Paint, a novel method based on a conditional transformer Variational AutoEncoder (VAE) architecture with a two-stage decoder. To evaluate the proposed approach and stimulate research in this area, we also introduce two novel datasets. Our experiments show that our approach provides good stroke suggestions and compares favorably to the state of the art. Elia Peruzzo, Willi Menapace, Vidit Goel, Federica Arrigoni, Hao Tang 0005, Xingqian Xu, Arman Chopikyan, Nikita Orlov, Humphrey Shi, Nicu Sebe, Elisa Ricci 0001 |
Comput. Vis. Image Underst. | 12 |
| 2023 | Compositional Semantic Mix for Domain Adaptation in Point Cloud SegmentationabstractDeep-learning models for 3D point cloud semantic segmentation exhibit limited generalization capabilities when trained and tested on data captured with different sensors or in varying environments due to domain shift. Domain adaptation methods can be employed to mitigate this domain shift, for instance, by simulating sensor noise, developing domain-agnostic generators, or training point cloud completion networks. Often, these methods are tailored for range view maps or necessitate multi-modal input. In contrast, domain adaptation in the image domain can be executed through sample mixing, which emphasizes input data manipulation rather than employing distinct adaptation modules. In this study, we introduce compositional semantic mixing for point cloud domain adaptation, representing the first unsupervised domain adaptation technique for point cloud segmentation based on semantic and geometric sample mixing. We present a two-branch symmetric network architecture capable of concurrently processing point clouds from a source domain (e.g. synthetic) and point clouds from a target domain (e.g. real-world). Each branch operates within one domain by integrating selected data fragments from the other domain and utilizing semantic information derived from source labels and target (pseudo) labels. Additionally, our method can leverage a limited number of human point-level annotations (semi-supervised) to further enhance performance. We assess our approach in both synthetic-to-real and real-to-real scenarios using LiDAR datasets and demonstrate that it significantly outperforms state-of-the-art methods in both unsupervised and semi-supervised settings. Cristiano Saltori, Fabio Galasso, Giuseppe Fiameni, Nicu Sebe, Fabio Poiesi, Elisa Ricci 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Uncertainty-Aware Contrastive Distillation for Incremental Semantic SegmentationabstractA fundamental and challenging problem in deep learning is catastrophic forgetting, i.e., the tendency of neural networks to fail to preserve the knowledge acquired from old tasks when learning new tasks. This problem has been widely investigated in the research community and several Incremental Learning (IL) approaches have been proposed in the past years. While earlier works in computer vision have mostly focused on image classification and object detection, more recently some IL approaches for semantic segmentation have been introduced. These previous works showed that, despite its simplicity, knowledge distillation can be effectively employed to alleviate catastrophic forgetting. In this paper, we follow this research direction and, inspired by recent literature on contrastive learning, we propose a novel distillation framework, Uncertainty-aware Contrastive Distillation (UCD). In a nutshell, UCDis operated by introducing a novel distillation loss that takes into account all the images in a mini-batch, enforcing similarity between features associated to all the pixels from the same classes, and pulling apart those corresponding to pixels from different classes. In order to mitigate catastrophic forgetting, we contrast features of the new model with features extracted by a frozen model learned at the previous incremental step. Our experimental results demonstrate the advantage of the proposed distillation technique, which can be used in synergy with previous IL approaches, and leads to state-of-art performance on three commonly adopted benchmarks for incremental semantic segmentation. Guanglei Yang, Enrico Fini, Dan Xu 0002, Paolo Rota, Mingli Ding, Moin Nabi, Xavier Alameda-Pineda, Elisa Ricci 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Continual Attentive Fusion for Incremental Learning in Semantic SegmentationabstractInternational audience Guanglei Yang, Enrico Fini, Dan Xu 0002, Paolo Rota, Mingli Ding, Hao Tang 0005, Xavier Alameda-Pineda, Elisa Ricci 0001 |
IEEE Trans. Multim. | 8 |
| 2022 | Cluster-level pseudo-labelling for source-free cross-domain facial expression recognition
Alessandro Conti, Paolo Rota, Yiming Wang 0002, Elisa Ricci 0001 |
BMVC | 4 |
| 2022 | Data Augmentation-free Unsupervised Learning for 3D Point Cloud Understanding
Guofeng Mei, Cristiano Saltori, Fabio Poiesi, Jian Zhang 0002, Elisa Ricci 0001, Nicu Sebe, Qiang Wu 0001 |
BMVC | 5 |
| 2022 | Self-Supervised Models are Continual LearnersabstractSelf-supervised models have been shown to produce comparable or better visual representations than their su-pervised counterparts when trained offline on unlabeled data at scale. However, their efficacy is catastrophically reduced in a Continual Learning (CL) scenario where data is presented to the model sequentially. In this paper, we show that self-supervised loss functions can be seamlessly converted into distillation mechanisms for CL by adding a predictor network that maps the current state of the repre-sentations to their past state. This enables us to devise a framework for Continual self-supervised visual representation Learning that (i) significantly improves the quality of the learned representations, (ii) is compatible with several state-of-the-art self-supervised objectives, and (iii) needs little to no hyperparameter tuning. We demonstrate the ef-fectiveness of our approach empirically by training six pop-ular self-supervised models in various CL settings. Code: github.com/DonkeyShot21/cassle. Enrico Fini, Victor G. T. da Costa, Xavier Alameda-Pineda, Elisa Ricci 0001, Karteek Alahari, Julien Mairal |
CVPR | 4 |
| 2022 | Playable Environments: Video Manipulation in Space and TimeabstractWe present Playable Environments-a new representation for interactive video generation and manipulation in space and time. With a single image at inference time, our novel framework allows the user to move objects in 3D while generating a video by providing a sequence of desired actions. The actions are learnt in an unsupervised manner. The camera can be controlled to get the desired viewpoint. Our method builds an environment state for each frame, which can be manipulated by our proposed action mod-ule and decoded back to the image space with volumetric rendering. To support diverse appearances of objects, we extend neural radiance fields with style-based modulation. Our method trains on a collection of various monocular videos requiring only the estimated camera parameters and 2D object locations. To set a challenging benchmark, we in-troduce two large scale video datasets with significant cam-era movements. As evidenced by our experiments, playable environments enable several creative applications not at-tainable by prior video synthesis works, including playable 3D video generation, stylization and manipulation11willi-menapace.github.io/playable-environments-website. Willi Menapace, Stéphane Lathuilière, Aliaksandr Siarohin, Christian Theobalt, Sergey Tulyakov, Vladislav Golyanik, Elisa Ricci 0001 |
CVPR | 7 |
| 2022 | Quantum Motion Segmentation
Federica Arrigoni, Willi Menapace, Marcel Seelbach Benkner, Elisa Ricci 0001, Vladislav Golyanik |
ECCV (29) | 4 |
| 2022 | Class-Incremental Novel Class Discovery
Subhankar Roy, Zhun Zhong, Nicu Sebe, Elisa Ricci 0001 |
ECCV (33) | 5 |
| 2022 | Uncertainty-Guided Source-Free Domain Adaptation
Subhankar Roy, Martin Trapp 0001, Andrea Pilzer, Juho Kannala, Nicu Sebe, Elisa Ricci 0001, Arno Solin |
ECCV (25) | 6 |
| 2022 | CoSMix: Compositional Semantic Mix for Domain Adaptation in 3D LiDAR Segmentation
Cristiano Saltori, Fabio Galasso, Giuseppe Fiameni, Nicu Sebe, Elisa Ricci 0001, Fabio Poiesi |
ECCV (33) | 5 |
| 2022 | GIPSO: Geometrically Informed Propagation for Online Adaptation in 3D LiDAR Segmentation
Cristiano Saltori, Evgeny Krivosheev, Stéphane Lathuilière, Nicu Sebe, Fabio Galasso, Giuseppe Fiameni, Elisa Ricci 0001, Fabio Poiesi |
ECCV (33) | 7 |
| 2022 | Multimodal Across Domains Gaze Target DetectionabstractThis paper addresses the gaze target detection problem in single images captured from the third-person perspective. We present a multimodal deep architecture to infer where a person in a scene is looking. This spatial model is trained on the head images of the person-of-interest, scene and depth maps representing rich context information. Our model, unlike several prior art, do not require supervision of the gaze angles, do not rely on head orientation information and/or location of the eyes of person-of-interest. Extensive experiments demonstrate the stronger performance of our method on multiple benchmark datasets. We also investigated several variations of our method by altering joint-learning of multimodal data. Some variations outperform a few prior art as well. First time in this paper, we inspect domain adaptation for gaze target detection, and we empower our multimodal network to effectively handle the domain gap across datasets. The code of the proposed method is available at https://github.com/francescotonini/multimodal-across-domains-gaze-target-detection. Francesco Tonini, Cigdem Beyan, Elisa Ricci 0001 |
ICMI | 3 |
| 2022 | Unsupervised Domain Adaptation for Video Transformers in Action RecognitionabstractOver the last few years, Unsupervised Domain Adaptation (UDA) techniques have acquired remarkable importance and popularity in computer vision. However, when compared to the extensive literature available for images, the field of videos is still relatively unexplored. On the other hand, the performance of a model in action recognition is heavily affected by domain shift. In this paper, we propose a simple and novel UDA approach for video action recognition. Our approach leverages recent advances on spatio-temporal transformers to build a robust source model that better generalises to the target domain. Furthermore, our architecture learns domain invariant features thanks to the introduction of a novel alignment loss term derived from the Information Bottleneck principle. We report results on two video action recognition benchmarks for UDA, showing state-of-the-art performance on HMDB ↔ UCF, as well as on Kinetics→NEC-Drone, which is more challenging. This demonstrates the effectiveness of our method in handling different levels of domain shift. The source code is available at https://github.com/vturrisi/UDAVT. Victor G. T. da Costa, Giacomo Zara, Paolo Rota, Thiago Oliveira-Santos, Nicu Sebe, Vittorio Murino, Elisa Ricci 0001 |
ICPR | 7 |
| 2022 | Multimodal Emotion Recognition with Modality-Pairwise Unsupervised Contrastive LossabstractEmotion recognition is involved in several real-world applications. With an increase in available modalities, automatic understanding of emotions is being performed more accurately. The success in Multimodal Emotion Recognition (MER), primarily relies on the supervised learning paradigm. However, data annotation is expensive, time-consuming, and as emotion expression and perception depends on several factors (e.g., age, gender, culture) obtaining labels with a high reliability is hard. Motivated by these, we focus on unsupervised feature learning for MER. We consider discrete emotions, and as modalities text, audio and vision are used. Our method, as being based on contrastive loss between pairwise modalities, is the first attempt in MER literature. Our end-to-end feature learning approach has several differences (and advantages) compared to existing MER methods: i) it is unsupervised, so the learning is lack of data labelling cost; ii) it does not require data spatial augmentation, modality alignment, large number of batch size or epochs; iii) it applies data fusion only at inference; and iv) it does not require backbones pre-trained on emotion recognition task. The experiments on benchmark datasets show that our method outperforms several baseline approaches and unsupervised learning methods applied in MER. Particularly, it even surpasses a few supervised MER state-of-the-art. Riccardo Franceschini, Enrico Fini, Cigdem Beyan, Alessandro Conti, Federica Arrigoni, Elisa Ricci 0001 |
ICPR | 6 |
| 2022 | Dual-Head Contrastive Domain Adaptation for Video Action RecognitionabstractUnsupervised domain adaptation (UDA) methods have become very popular in computer vision. However, while several techniques have been proposed for images, much less attention has been devoted to videos. This paper introduces a novel UDA approach for action recognition from videos, inspired by recent literature on contrastive learning. In particular, we propose a novel two-headed deep architecture that simultaneously adopts cross-entropy and contrastive losses from different network branches to robustly learn a target classifier. Moreover, this work introduces a novel large-scale UDA dataset, Mixamo→Kinetics, which, to the best of our knowledge, is the first dataset that considers the domain shift arising when transferring knowledge from synthetic to real video sequences. Our extensive experimental evaluation conducted on three publicly available benchmarks and on our new Mixamo→Kinetics dataset demonstrate the effectiveness of our approach, which outperforms the current state-of-the-art methods. Code is available at https://github.com/vturrisi/CO2A. Victor G. T. da Costa, Giacomo Zara, Paolo Rota, Thiago Oliveira-Santos, Nicu Sebe, Vittorio Murino, Elisa Ricci 0001 |
WACV | 7 |
| 2022 | Multi-frame Motion Segmentation by Combining Two-Frame ResultsabstractAbstract In this paper we consider the motion segmentation problem on sparse and unstructured datasets involving rigid motions, motivated by multibody structure from motion. In particular, we assume only two-frame correspondences as input without prior knowledge about trajectories. Inspired by the success of synchronization methods, we address this problem by introducing a two-stage approach: first, motion segmentation is addressed on image pairs independently; then, two-frame results are combined in a robust way to compute the final multi-frame segmentation. Our synthetic and real experiments demonstrate that the proposed approach is very effective in reducing the errors among two-frame results and it can cope with a large amount of mismatches. Moreover, our method can be profitably used to build a multibody structure from motion pipeline. Federica Arrigoni, Elisa Ricci 0001, Tomás Pajdla |
Int. J. Comput. Vis. | 2 |
| 2022 | solo-learn: A Library of Self-supervised Methods for Visual Representation LearningabstractThis paper presents solo-learn, a library of self-supervised methods for visual representation learning. Implemented in Python, using Pytorch and Pytorch lightning, the library fits both research and industry needs by featuring distributed training pipelines with mixed-precision, faster data loading via Nvidia DALI, online linear evaluation for better prototyping, and many additional training tricks. Our goal is to provide an easy-to-use library comprising a large amount of Self-supervised Learning (SSL) methods, that can be easily extended and fine-tuned by the community. solo-learn opens up avenues for exploiting large-budget SSL solutions on inexpensive smaller infrastructures and seeks to democratize SSL by making it accessible to all. The source code is available at https://github.com/vturrisi/solo-learn. Victor G. T. da Costa, Enrico Fini, Moin Nabi, Nicu Sebe, Elisa Ricci 0001 |
J. Mach. Learn. Res. | 5 |
| 2022 | Modeling the Background for Incremental and Weakly-Supervised Semantic SegmentationabstractDeep neural networks have enabled major progresses in semantic segmentation. However, even the most advanced neural architectures suffer from important limitations. First, they are vulnerable to catastrophic forgetting, i.e., they perform poorly when they are required to incrementally update their model as new classes are available. Second, they rely on large amount of pixel-level annotations to produce accurate segmentation maps. To tackle these issues, we introduce a novel incremental class learning approach for semantic segmentation taking into account a peculiar aspect of this task: since each training step provides annotation only for a subset of all possible classes, pixels of the background class exhibit a semantic shift. Therefore, we revisit the traditional distillation paradigm by designing novel loss terms which explicitly account for the background shift. Additionally, we introduce a novel strategy to initialize classifier's parameters at each step in order to prevent biased predictions toward the background class. Finally, we demonstrate that our approach can be extended to point- and scribble-based weakly supervised segmentation, modeling the partial annotations to create priors for unlabeled pixels. We demonstrate the effectiveness of our approach with an extensive evaluation on the Pascal-VOC, ADE20K, and Cityscapes datasets, significantly outperforming state-of-the-art methods. Fabio Cermelli, Massimiliano Mancini, Samuel Rota Bulò, Elisa Ricci 0001, Barbara Caputo |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Probabilistic Graph Attention Network With Conditional Kernels for Pixel-Wise PredictionabstractMulti-scale representations deeply learned via convolutional neural networks have shown tremendous importance for various pixel-level prediction problems. In this paper we present a novel approach that advances the state of the art on pixel-level prediction in a fundamental aspect, i.e. structured multi-scale features learning and fusion. In contrast to previous works directly considering multi-scale feature maps obtained from the inner layers of a primary CNN architecture, and simply fusing the features with weighted averaging or concatenation, we propose a probabilistic graph attention network structure based on a novel Attention-Gated Conditional Random Fields (AG-CRFs) model for learning and fusing multi-scale representations in a principled manner. In order to further improve the learning capacity of the network structure, we propose to exploit feature dependant conditional kernels within the deep probabilistic framework. Extensive experiments are conducted on four publicly available datasets (i.e. BSDS500, NYUD-V2, KITTI and Pascal-Context) and on three challenging pixel-wise prediction problems involving both discrete and continuous labels (i.e. monocular depth estimation, object contour prediction and semantic segmentation). Quantitative and qualitative results demonstrate the effectiveness of the proposed latent AG-CRF model and the overall probabilistic graph attention network with feature conditional kernels for structured feature learning and pixel-wise prediction. Dan Xu 0002, Xavier Alameda-Pineda, Wanli Ouyang, Elisa Ricci 0001, Xiaogang Wang 0001, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Introduction to the Special Section on Learning Representations, Similarity, and Associations in Dynamic Multimedia EnvironmentsabstractNo abstract available. Xun Yang 0001, Liang Zheng 0001, Elisa Ricci 0001, Meng Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Playable Video GenerationabstractThis paper introduces the unsupervised learning problem of playable video generation (PVG). In PVG, we aim at allowing a user to control the generated video by selecting a discrete action at every time step as when playing a video game. The difficulty of the task lies both in learning semantically consistent actions and in generating realistic videos conditioned on the user input. We propose a novel framework for PVG that is trained in a self-supervised manner on a large dataset of unlabelled videos. We employ an encoder-decoder architecture where the predicted action labels act as bottleneck. The network is constrained to learn a rich action space using, as main driving loss, a reconstruction loss on the generated video. We demonstrate the effectiveness of the proposed approach on several datasets with wide environment variety. Further details, code and examples are available on our project page willi-menapace.github.io/playable-video-generation-website. Willi Menapace, Stéphane Lathuilière, Sergey Tulyakov, Aliaksandr Siarohin, Elisa Ricci 0001 |
CVPR | 5 |
| 2021 | Curriculum Graph Co-Teaching for Multi-Target Domain AdaptationabstractIn this paper we address multi-target domain adaptation (MTDA), where given one labeled source dataset and multiple unlabeled target datasets that differ in data distributions, the task is to learn a robust predictor for all the target domains. We identify two key aspects that can help to alleviate multiple domain-shifts in the MTDA: feature aggregation and curriculum learning. To this end, we propose Curriculum Graph Co-Teaching (CGCT) that uses a dual classifier head, with one of them being a graph convolutional network (GCN) which aggregates features from similar samples across the domains. To prevent the classifiers from over-fitting on its own noisy pseudo-labels we develop a co-teaching strategy with the dual classifier head that is assisted by curriculum learning to obtain more reliable pseudo-labels. Furthermore, when the domain labels are available, we propose Domain-aware Curriculum Learning (DCL), a sequential adaptation strategy that first adapts on the easier target domains, followed by the harder ones. We experimentally demonstrate the effectiveness of our proposed frameworks on several benchmarks and advance the state-of-the-art in the MTDA by large margins (e.g. +5.6% on the DomainNet). Subhankar Roy, Evgeny Krivosheev, Zhun Zhong, Nicu Sebe, Elisa Ricci 0001 |
CVPR | 5 |
| 2021 | Neighborhood Contrastive Learning for Novel Class DiscoveryabstractIn this paper, we address Novel Class Discovery (NCD), the task of unveiling new classes in a set of unlabeled samples given a labeled dataset with known classes. We exploit the peculiarities of NCD to build a new framework, named Neighborhood Contrastive Learning (NCL), to learn discriminative representations that are important to clustering performance. Our contribution is twofold. First, we find that a feature extractor trained on the labeled set generates representations in which a generic query sample and its neighbors are likely to share the same class. We exploit this observation to retrieve and aggregate pseudo-positive pairs with contrastive learning, thus encouraging the model to learn more discriminative representations. Second, we notice that most of the instances are easily discriminated by the network, contributing less to the contrastive loss. To overcome this issue, we propose to generate hard negatives by mixing labeled and unlabeled samples in the feature space. We experimentally demonstrate that these two ingredients significantly contribute to clustering performance and lead our model to outperform state-of-the-art methods by a large margin (e.g., clustering accuracy +13% on CIFAR-100 and +8% on ImageNet). Zhun Zhong, Enrico Fini, Subhankar Roy, Zhiming Luo, Elisa Ricci 0001, Nicu Sebe |
CVPR | 5 |
| 2021 | Click to Move: Controlling Video Generation with Sparse MotionabstractThis paper introduces Click to Move (C2M), a novel framework for video generation where the user can control the motion of the synthesized video through mouse clicks specifying simple object trajectories of the key objects in the scene. Our model receives as input an initial frame, its corresponding segmentation map and the sparse motion vectors encoding the input provided by the user. It outputs a plausible video sequence starting from the given frame and with a motion that is consistent with user input. Notably, our proposed deep architecture incorporates a Graph Convolution Network (GCN) modelling the movements of all the objects in the scene in a holistic manner and effectively combining the sparse user motion information and image features. Experimental results show that C2M outperforms existing methods on two publicly available datasets, thus demonstrating the effectiveness of our GCN framework at modelling object interactions. The source code is publicly available at https://github.com/PierfrancescoArdino/C2M. Pierfrancesco Ardino, Marco De Nadai, Bruno Lepri, Elisa Ricci 0001, Stéphane Lathuilière |
ICCV | 4 |
| 2021 | Viewing Graph Solvability via Cycle ConsistencyabstractIn structure-from-motion the viewing graph is a graph where vertices correspond to cameras and edges represent fundamental matrices. We provide a new formulation and an algorithm for establishing whether a viewing graph is solvable, i.e. it uniquely determines a set of projective cameras. Known theoretical conditions either do not fully characterize the solvability of all viewing graphs, or are exceedingly hard to compute for they involve solving a system of polynomial equations with a large number of unknowns. The main result of this paper is a method for reducing the number of unknowns by exploiting the cycle consistency. We advance the understanding of the solvability by (i) finishing the classification of all previously undecided minimal graphs up to 9 nodes, (ii) extending the practical solvability testing up to minimal graphs with up to 90 nodes, and (iii) definitely answering an open research question by showing that the finite solvability is not equivalent to the solvability. Finally, we present an experiment on real data showing that unsolvable graphs are appearing in practical situations. Federica Arrigoni, Andrea Fusiello, Elisa Ricci 0001, Tomás Pajdla |
ICCV | 3 |
| 2021 | A Unified Objective for Novel Class DiscoveryabstractIn this paper, we study the problem of Novel Class Discovery (NCD). NCD aims at inferring novel object categories in an unlabeled set by leveraging from prior knowledge of a labeled set containing different, but related classes. Existing approaches tackle this problem by considering multiple objective functions, usually involving specialized loss terms for the labeled and the unlabeled samples respectively, and often requiring auxiliary regularization terms. In this paper we depart from this traditional scheme and introduce a UNified Objective function (UNO) for discovering novel classes, with the explicit purpose of favoring synergy between supervised and unsupervised learning. Using a multi-view self-labeling strategy, we generate pseudo-labels that can be treated homogeneously with ground truth labels. This leads to a single classification objective operating on both known and unknown classes. De-spite its simplicity, UNO outperforms the state of the art by a significant margin on several benchmarks (≈ +10% on CIFAR-100 and +8% on ImageNet). The project page is available at : https://ncd-uno.github.io. Enrico Fini, Enver Sangineto, Stéphane Lathuilière, Zhun Zhong, Moin Nabi, Elisa Ricci 0001 |
ICCV | 6 |
| 2021 | Are we Missing Confidence in Pseudo-LiDAR Methods for Monocular 3D Object Detection?abstractPseudo-LiDAR-based methods for monocular 3D object detection have received considerable attention in the community due to the performance gains exhibited on the KITTI3D benchmark, in particular on the commonly reported validation split. This generated a distorted impression about the superiority of Pseudo-LiDAR-based (PL-based) approaches over methods working with RGB images only. Our first contribution consists in rectifying this view by pointing out and showing experimentally that the validation results published by PL-based methods are substantially biased. The source of the bias resides in an overlap between the KITTI3D object detection validation set and the training/validation sets used to train depth predictors feeding PL-based methods. Surprisingly, the bias remains also after geographically removing the overlap. This leaves the test set as the only reliable set for comparison, where published PL-based methods do not excel. Our second contribution brings PL-based methods back up in the ranking with the design of a novel deep architecture which introduces a 3D confidence prediction module. We show that 3D confidence estimation techniques derived from RGB-only 3D detection approaches can be successfully integrated into our framework and, more importantly, that improved performance can be obtained with a newly designed 3D confidence measure, leading to state-of-the-art performance on the KITTI3D benchmark. Andrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Peter Kontschieder, Elisa Ricci 0001 |
ICCV | 5 |
| 2021 | Transformer-Based Attention Networks for Continuous Pixel-Wise PredictionabstractWhile convolutional neural networks have shown a tremendous impact on various computer vision tasks, they generally demonstrate limitations in explicitly modeling long-range dependencies due to the intrinsic locality of the convolution operation. Initially designed for natural language processing tasks, Transformers have emerged as alternative architectures with innate global self-attention mechanisms to capture long-range dependencies. In this paper, we propose TransDepth, an architecture that benefits from both convolutional neural networks and transformers. To avoid the network losing its ability to capture locallevel details due to the adoption of transformers, we propose a novel decoder that employs attention mechanisms based on gates. Notably, this is the first paper that applies transformers to pixel-wise prediction problems involving continuous labels (i.e., monocular depth prediction and surface normal estimation). Extensive experiments demonstrate that the proposed TransDepth achieves state-of-theart performance on three challenging datasets. Our code is available at: https://github.com/ygjwd12345/TransDepth. Guanglei Yang, Hao Tang 0005, Mingli Ding, Nicu Sebe, Elisa Ricci 0001 |
ICCV | 5 |
| 2021 | TriGAN: image-to-image translation for multi-source domain adaptationabstractAbstract Most domain adaptation methods consider the problem of transferring knowledge to the target domain from a single-source dataset. However, in practical applications, we typically have access to multiple sources. In this paper we propose the first approach for multi-source domain adaptation (MSDA) based on generative adversarial networks. Our method is inspired by the observation that the appearance of a given image depends on three factors: thedomain, thestyle(characterized in terms of low-level features variations) and thecontent. For this reason, we propose to project the source image features onto a space where only the dependence from the content is kept, and then re-project this invariant representation onto the pixel space using the target domain and style. In this way, new labeled images can be generated which are used to train a final target classifier. We test our approach using common MSDA benchmarks, showing that it outperforms state-of-the-art methods. Subhankar Roy, Aliaksandr Siarohin, Enver Sangineto, Nicu Sebe, Elisa Ricci 0001 |
Mach. Vis. Appl. | 5 |
| 2021 | MultiDIAL: Domain Alignment Layers for (Multisource) Unsupervised Domain AdaptationabstractOne of the main challenges for developing visual recognition systems working in the wild is to devise computational models immune from the domain shift problem, i.e., accurate when test data are drawn from a (slightly) different data distribution than training samples. In the last decade, several research efforts have been devoted to devise algorithmic solutions for this issue. Recent attempts to mitigate domain shift have resulted into deep learning models for domain adaptation which learn domain-invariant representations by introducing appropriate loss terms, by casting the problem within an adversarial learning framework or by embedding into deep network specific domain normalization layers. This paper describes a novel approach for unsupervised domain adaptation. Similarly to previous works we propose to align the learned representations by embedding them into appropriate network feature normalization layers. Opposite to previous works, our Domain Alignment Layers are designed not only to match the source and target feature distributions but also to automatically learn the degree of feature alignment required at different levels of the deep network. Differently from most previous deep domain adaptation methods, our approach is able to operate in a multi-source setting. Thorough experiments on four publicly available benchmarks confirm the effectiveness of our approach. Fabio Maria Carlucci, Lorenzo Porzi, Barbara Caputo, Elisa Ricci 0001, Samuel Rota Bulò |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Inferring Latent Domains for Unsupervised Deep Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) refers to the problem of learning a model in a target domain where labeled data are not available by leveraging information from annotated data in a source domain. Most deep UDA approaches operate in a single-source, single-target scenario, i.e., they assume that the source and the target samples arise from a single distribution. However, in practice most datasets can be regarded as mixtures of multiple domains. In these cases, exploiting traditional single-source, single-target methods for learning classification models may lead to poor results. Furthermore, it is often difficult to provide the domain labels for all data points, i.e. latent domains should be automatically discovered. This paper introduces a novel deep architecture which addresses the problem of UDA by automatically discovering latent domains in visual datasets and exploiting this information to learn robust target classifiers. Specifically, our architecture is based on two main components, i.e. a side branch that automatically computes the assignment of each sample to its latent domain and novel layers that exploit domain membership information to appropriately align the distribution of the CNN internal feature representations to a reference distribution. We evaluate our approach on publicly available benchmarks, showing that it outperforms state-of-the-art domain adaptation methods. Massimiliano Mancini, Lorenzo Porzi, Samuel Rota Bulò, Barbara Caputo, Elisa Ricci 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | SF-UDA3D: Source-Free Unsupervised Domain Adaptation for LiDAR-Based 3D Object Detectionabstract3D object detectors based only on LiDAR point clouds hold the state-of-the-art on modern street-view benchmarks. However, LiDAR-based detectors poorly generalize across domains due to domain shift. In the case of LiDAR, in fact, domain shift is not only due to changes in the environment and in the object appearances, as for visual data from RGB cameras, but is also related to the geometry of the point clouds (e.g., point density variations). This paper proposes SF-UDA3D, the first Source-Free Unsupervised Domain Adaptation (SF-UDA) framework to domain-adapt the state-of-the-art PointRCNN 3D detector to target domains for which we have no annotations (unsupervised), neither we hold images nor annotations of the source domain (source-free). SF-UDA3Dis novel on both aspects. Our approach is based on pseudo-annotations, reversible scale-transformations and motion coherency. SFUDA3Doutperforms both previous domain adaptation techniques based on features alignment and state-of-the-art 3D object detection methods which additionally use few-shot target annotations or target annotation statistics. This is demonstrated by extensive experiments on two large-scale datasets, i.e., KITTI and nuScenes. Cristiano Saltori, Stéphane Lathuilière, Nicu Sebe, Elisa Ricci 0001, Fabio Galasso |
3DV | 4 |
| 2020 | Modeling the Background for Incremental Learning in Semantic SegmentationabstractDespite their effectiveness in a wide range of tasks, deep architectures suffer from some important limitations. In particular, they are vulnerable to catastrophic forgetting, i.e. they perform poorly when they are required to update their model as new classes are available but the original training set is not retained. This paper addresses this problem in the context of semantic segmentation. Current strategies fail on this task because they do not consider a peculiar aspect of semantic segmentation: since each training step provides annotation only for a subset of all possible classes, pixels of the background class (i.e. pixels that do not belong to any other classes) exhibit a semantic distribution shift. In this work we revisit classical incremental learning methods, proposing a new distillation-based framework which explicitly accounts for this shift. Furthermore, we introduce a novel strategy to initialize classifier's parameters, thus preventing biased predictions toward the background class. We demonstrate the effectiveness of our approach with an extensive evaluation on the Pascal-VOC 2012 and ADE20K datasets, significantly outperforming state of the art incremental learning methods. Fabio Cermelli, Massimiliano Mancini, Samuel Rota Bulò, Elisa Ricci 0001, Barbara Caputo |
CVPR | 4 |
| 2020 | Online Depth Learning Against Forgetting in Monocular VideosabstractOnline depth learning is the problem of consistently adapting a depth estimation model to handle a continuously changing environment. This problem is challenging due to the network easily overfits on the current environment and forgets its past experiences. To address such problem, this paper presents a novel Learning to Prevent Forgetting (LPF) method for online mono-depth adaptation to new target domains in unsupervised manner. Instead of updating the universal parameters, LPF learns adapter modules to efficiently adjust the feature representation and distribution without losing the pre-learned knowledge in online condition. Specifically, to adapt temporal-continuous depth patterns in videos, we introduce a novel meta-learning approach to learn adapter modules by combining online adaptation process into the learning objective. To further avoid overfitting, we propose a novel temporal-consistent regularization to harmonize the gradient descent procedure at each online learning step. Extensive evaluations on real-world datasets demonstrate that the proposed method, with very limited parameters, significantly improves the estimation quality. Zhenyu Zhang 0005, Stéphane Lathuilière, Elisa Ricci 0001, Nicu Sebe, Yan Yan 0002, Jian Yang 0003 |
CVPR | 3 |
| 2020 | Online Continual Learning Under Extreme Memory Constraints
Enrico Fini, Stéphane Lathuilière, Enver Sangineto, Moin Nabi, Elisa Ricci 0001 |
ECCV (28) | 5 |
| 2020 | Towards Recognizing Unseen Categories in Unseen Domains
Massimiliano Mancini, Zeynep Akata, Elisa Ricci 0001, Barbara Caputo |
ECCV (23) | 3 |
| 2020 | Learning to Cluster Under Domain Shift
Willi Menapace, Stéphane Lathuilière, Elisa Ricci 0001 |
ECCV (28) | 3 |
| 2020 | Towards Generalization Across Depth for Monocular 3D Object Detection
Andrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Elisa Ricci 0001, Peter Kontschieder |
ECCV (22) | 4 |
| 2020 | Semantic-Guided Inpainting Network for Complex Urban Scenes ManipulationabstractManipulating images of complex scenes to reconstruct, insert and/or remove specific object instances is a challenging task. Complex scenes contain multiple semantics and objects, which are frequently cluttered or ambiguous, thus hampering the performance of inpainting models. Conventional techniques often rely on structural information such as object contours in multi-stage approaches that generate unreliable results and boundaries. In this work, we propose a novel deep learning model to alter a complex urban scene by removing a user-specified portion of the image and coherently inserting a new object (e.g. car or pedestrian) in that scene. Inspired by recent works on image inpainting, our proposed method leverages the semantic segmentation to model the content and structure of the image, and learn the best shape and location of the object to insert. To generate reliable results, we design a new decoder block that combines the semantic segmentation and generation task to guide better the generation of new objects and scenes, which have to be semantically consistent with the image. Our experiments, conducted on two large-scale datasets of urban scenes (Cityscapes and Indian Driving), show that our proposed approach successfully address the problem of semantically-guided inpainting of complex urban scene. Pierfrancesco Ardino, Elisa Ricci 0001, Bruno Lepri, Marco De Nadai |
ICPR | 3 |
| 2020 | Multi-Domain Image-to-Image Translation with Adaptive Inference GraphabstractIn this work, we address the problem of multidomain image-to-image translation with particular attention paid to computational cost. In particular, current state of the art models require a large and deep model in order to handle the visual diversity of multiple domains. In a context of limited computational resources, increasing the network size may not be possible. Therefore, we propose to increase the network capacity by using an adaptive graph structure. At inference time, the network estimates its own graph by selecting specific sub-networks. Sub-network selection is implemented using Gumbel-Softmax in order to allow end-to-end training. This approach leads to an adjustable increase in number of parameters while preserving an almost constant computational cost. Our evaluation on two publicly available datasets of facial and painting images shows that our adaptive strategy generates better images with fewer artifacts than literature methods. The-Phuc Nguyen, Stéphane Lathuilière, Elisa Ricci 0001 |
ICPR | 3 |
| 2020 | Motion-supervised Co-Part SegmentationabstractRecent co-part segmentation methods mostly operate in a supervised learning setting, which requires a large amount of annotated data for training. To overcome this limitation, we propose a self-supervised deep learning method for co-part segmentation. Differently from previous works, our approach develops the idea that motion information inferred from videos can be leveraged to discover meaningful object parts. To this end, our method relies on pairs of frames sampled from the same video. The network learns to predict part segments together with a representation of the motion between two frames, which permits reconstruction of the target image. Through extensive experimental evaluation on publicly available video sequences we demonstrate that our approach can produce improved segmentation maps with respect to previous self-supervised co-part segmentation approaches. Aliaksandr Siarohin, Subhankar Roy, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci 0001, Nicu Sebe |
ICPR | 5 |
| 2020 | Shape Consistent 2D Keypoint Estimation under Domain ShiftabstractRecent unsupervised domain adaptation methods based on deep architectures have shown remarkable performance not only in traditional classification tasks but also in more complex problems involving structured predictions (e.g. semantic segmentation, depth estimation). Following this trend, in this paper we present a novel deep adaptation framework for estimating keypoints under domain shift, i.e, when the training (source) and the test (target) images significantly differ in terms of visual appearance. Our method seamlessly combines three different components: feature alignment, adversarial training and self-supervision. Specifically, our deep architecture leverages from domain-specific distribution alignment layers to perform target adaptation at the feature level. Furthermore, a novel loss is proposed which combines an adversarial term for ensuring aligned predictions in the output space and a geometric consistency term which guarantees coherent predictions between a target sample and its perturbed version. Our extensive experimental evaluation conducted on three publicly available benchmarks shows that our approach outperforms state-of-the-art domain adaptation methods in the 2D keypoint prediction task. Levi O. Vasconcelos, Massimiliano Mancini, Davide Boscaini, Samuel Rota Bulò, Barbara Caputo, Elisa Ricci 0001 |
ICPR | 6 |
| 2020 | Coping with Pandemics: Opportunities and Challenges for AI Multimedia in the "New Normal"abstractTheworld iswelcoming the newnormal - the coronavirus pandemic has significantly changed the way people live, work, communicate and learn. Almost everyone now is wearing a face mask when they go in public. People are working from home, some taking care of children at the same time. Bars and restaurants are limited to carry-out and delivery only. Meetings and conferences go online. Schools are closed and educators are instead holding video conference classes regularly. All these become the new normal as our ways of life. The panel thus provides a valuable opportunity for people from a variety of backgrounds to exchange views on opportunities and challenges for AI multimedia in the current and post pandemics era. Jiaying Liu 0001, Wen-Huang Cheng, Klara Nahrstedt, Ramesh Jain 0001, Elisa Ricci 0001, Hyeran Byun |
ACM Multimedia | 5 |
| 2020 | Special Issue on Generating Realistic Visual Data of Human Behavior
Xavier Alameda-Pineda, Elisa Ricci 0001, Albert Ali Salah, Nicu Sebe, Shuicheng Yan |
Int. J. Comput. Vis. | 2 |
| 2020 | Boosting binary masks for multi-domain learning through affine transformations
Massimiliano Mancini, Elisa Ricci 0001, Barbara Caputo, Samuel Rota Bulò |
Mach. Vis. Appl. | 2 |
| 2020 | Progressive Fusion for Unsupervised Binocular Depth Estimation Using Cycled NetworksabstractRecent deep monocular depth estimation approaches based on supervised regression have achieved remarkable performance. However, they require costly ground truth annotations during training. To cope with this issue, in this paper we present a novel unsupervised deep learning approach for predicting depth maps. We introduce a new network architecture, named Progressive Fusion Network (PFN), that is specifically designed for binocular stereo depth estimation. This network is based on a multi-scale refinement strategy that combines the information provided by both stereo views. In addition, we propose to stack twice this network in order to form a cycle. This cycle approach can be interpreted as a form of data-augmentation since, at training time, the network learns both from the training set images (in the forward half-cycle) but also from the synthesized images (in the backward half-cycle). The architecture is jointly trained with adversarial learning. Extensive experiments on the publicly available datasets KITTI, Cityscapes and ApolloScape demonstrate the effectiveness of the proposed model which is competitive with other unsupervised deep learning methods for depth prediction. Andrea Pilzer, Stéphane Lathuilière, Dan Xu 0002, Mihai Marian Puscas, Elisa Ricci 0001, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Deep Learning for Classification and Localization of COVID-19 Markers in Point-of-Care Lung UltrasoundabstractDeep learning (DL) has proved successful in medical imaging and, in the wake of the recent COVID-19 pandemic, some works have started to investigate DL-based solutions for the assisted diagnosis of lung diseases. While existing works focus on CT scans, this paper studies the application of DL techniques for the analysis of lung ultrasonography (LUS) images. Specifically, we present a novel fully-annotated dataset of LUS images collected from several Italian hospitals, with labels indicating the degree of disease severity at a frame-level, video-level, and pixel-level (segmentation masks). Leveraging these data, we introduce several deep models that address relevant tasks for the automatic analysis of LUS images. In particular, we present a novel deep network, derived from Spatial Transformer Networks, which simultaneously predicts the disease severity score associated to a input frame and provides localization of pathological artefacts in a weakly-supervised way. Furthermore, we introduce a new method based on uninorms for effective frame score aggregation at a video-level. Finally, we benchmark state of the art deep models for estimating pixel-level segmentations of COVID-19 imaging biomarkers. Experiments on the proposed dataset demonstrate satisfactory results on all the considered tasks, paving the way to future research on DL for the assisted diagnosis of COVID-19 from LUS data. Subhankar Roy, Willi Menapace, Sebastiaan Oei, Ben Luijten, Enrico Fini, Cristiano Saltori, Iris A. M. Huijben, Nishith Chennakeshava, Federico Mento, Alessandro Sentelli, Emanuele Peschiera, Riccardo Trevisan, Giovanni Maschietto, Elena Torri, Riccardo Inchingolo, Andrea Smargiassi, Gino Soldati, Paolo Rota, Andrea Passerini, Ruud van Sloun, Elisa Ricci 0001, Libertario Demi |
IEEE Trans. Medical Imaging | 21 |
| 2020 | Learning How to Smile: Expression Video Generation With Conditional Adversarial Recurrent NetsabstractWhile several research studies have focused on analyzing human behavior and, in particular, emotional signals from visual data, the problem of synthesizing face video sequences with specific attributes (e.g. age, facial expressions) received much less attention. This paper proposes a novel deep generative model able to produce face videos from a given image of a neutral face and a label indicating a specific facial expression, e.g. spontaneous smile. Our framework consists of two main building blocks: an image generator and a frame sequence generator. The image generator is implemented as a deep neural model which combines generative adversarial networks and variational auto-encoders, while the sequence generator is a label-conditioned recurrent neural network. In the proposed framework, given as input a neural face and a label, the sequence generator outputs a set of hidden representations with smooth transitions corresponding to video frames. Then, the image generator is used to decode the hidden representations into the actual face images. To impose that the net generates videos consistent with the given label, a novel identity adversarial loss is proposed. Our experimental results demonstrate the effectiveness of the framework and the advantage of introducing an adversarial component into recurrent models for face video generation. Wei Wang 0108, Xavier Alameda-Pineda, Dan Xu 0002, Elisa Ricci 0001, Nicu Sebe |
IEEE Trans. Multim. | 4 |
| 2019 | Structured Domain Adaptation for 3D Keypoint EstimationabstractMotivated by recent advances in deep domain adaptation, this paper introduces a deep architecture for estimating 3D keypoints when the training (source) and the test (target) images greatly differ in terms of visual appearance (domain shift). Our approach operates by promoting domain distribution alignment in the feature space adopting batch normalization-based techniques. Furthermore, we propose to collect statistics about 3D keypoints positions of the source training data and to use this prior information to constrain predictions on the target domain introducing a loss derived from Multidimensional Scaling. We conduct an extensive experimental evaluation considering three publicly available benchmarks and show that our approach out-performs state-of-the-art domain adaptation methods for 3D keypoints predictions. Levi O. Vasconcelos, Massimiliano Mancini, Davide Boscaini, Barbara Caputo, Elisa Ricci 0001 |
3DV | 5 |
| 2019 | AdaGraph: Unifying Predictive and Continuous Domain Adaptation Through GraphsabstractThe ability to categorize is a cornerstone of visual intelligence, and a key functionality for artificial, autonomous visual machines. This problem will never be solved without algorithms able to adapt and generalize across visual domains. Within the context of domain adaptation and generalization, this paper focuses on the predictive domain adaptation scenario, namely the case where no target data are available and the system has to learn to generalize from annotated source images plus unlabeled samples with associated metadata from auxiliary domains. Our contribution is the first deep architecture that tackles predictive domain adaptation, able to leverage over the information brought by the auxiliary domains through a graph. Moreover, we present a simple yet effective strategy that allows us to take advantage of the incoming target data at test time, in a continuous domain adaptation scenario. Experiments on three benchmark databases support the value of our approach. Massimiliano Mancini, Samuel Rota Bulò, Barbara Caputo, Elisa Ricci 0001 |
CVPR | 4 |
| 2019 | Refine and Distill: Exploiting Cycle-Inconsistency and Knowledge Distillation for Unsupervised Monocular Depth EstimationabstractNowadays, the majority of state of the art monocular depth estimation techniques are based on supervised deep learning models. However, collecting RGB images with associated depth maps is a very time consuming procedure. Therefore, recent works have proposed deep architectures for addressing the monocular depth prediction task as a reconstruction problem, thus avoiding the need of collecting ground-truth depth. Following these works, we propose a novel self-supervised deep model for estimating depth maps. Our framework exploits two main strategies: refinement via cycle-inconsistency and distillation. Specifically, first a student network is trained to predict a disparity map such as to recover from a frame in a camera view the associated image in the opposite view. Then, a backward cycle network is applied to the generated image to re-synthesize back the input image, estimating the opposite disparity. A third network exploits the inconsistency between the original and the reconstructed input frame in order to output a refined depth map. Finally, knowledge distillation is exploited, such as to transfer information from the refinement network to the student. Our extensive experimental evaluation demonstrate the effectiveness of the proposed framework which outperforms state of the art unsupervised methods on the KITTI benchmark. Andrea Pilzer, Stéphane Lathuilière, Nicu Sebe, Elisa Ricci 0001 |
CVPR | 4 |
| 2019 | Unsupervised Domain Adaptation Using Feature-Whitening and Consensus LossabstractA classifier trained on a dataset seldom works on other datasets obtained under different conditions due to domain shift. This problem is commonly addressed by domain adaptation methods. In this work we introduce a novel deep learning framework which unifies different paradigms in unsupervised domain adaptation. Specifically, we propose domain alignment layers which implement feature whitening for the purpose of matching source and target feature distributions. Additionally, we leverage the unlabeled target data by proposing the Min-Entropy Consensus loss, which regularizes training while avoiding the adoption of many user-defined hyper-parameters. We report results on publicly available datasets, considering both digit classification and object recognition tasks. We show that, in most of our experiments, our approach improves upon previous methods, setting new state-of-the-art performances. Subhankar Roy, Aliaksandr Siarohin, Enver Sangineto, Samuel Rota Bulò, Nicu Sebe, Elisa Ricci 0001 |
CVPR | 6 |
| 2019 | Animating Arbitrary Objects via Deep Motion TransferabstractThis paper introduces a novel deep learning framework for image animation. Given an input image with a target object and a driving video sequence depicting a moving object, our framework generates a video in which the target object is animated according to the driving sequence. This is achieved through a deep architecture that decouples appearance and motion information. Our framework consists of three main modules: (i) a Keypoint Detector unsupervisely trained to extract object keypoints, (ii) a Dense Motion prediction network for generating dense heatmaps from sparse keypoints, in order to better encode motion information and (iii) a Motion Transfer Network, which uses the motion heatmaps and appearance information extracted from the input image to synthesize the output frames. We demonstrate the effectiveness of our method on several benchmark datasets, spanning a wide variety of object appearances, and show that our approach outperforms state-of-the-art image animation and video generation methods. Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci 0001, Nicu Sebe |
CVPR | 4 |
| 2019 | Budget-Aware Adapters for Multi-Domain LearningabstractMulti-Domain Learning (MDL) refers to the problem of learning a set of models derived from a common deep architecture, each one specialized to perform a task in a certain domain (e.g., photos, sketches, paintings). This paper tackles MDL with a particular interest in obtaining domain-specific models with an adjustable budget in terms of the number of network parameters and computational complexity. Our intuition is that, as in real applications the number of domains and tasks can be very large, an effective MDL approach should not only focus on accuracy but also on having as few parameters as possible. To implement this idea we derive specialized deep models for each domain by adapting a pre-trained architecture but, differently from other methods, we propose a novel strategy to automatically adjust the computational complexity of the network. To this aim, we introduce Budget-Aware Adapters that select the most relevant feature channels to better handle data from a novel domain. Some constraints on the number of active switches are imposed in order to obtain a network respecting the desired complexity budget. Experimentally, we show that our approach leads to recognition accuracy competitive with state-of-the-art approaches but with much lighter networks both in terms of storage and computation. Rodrigo Ferreira Berriel, Stéphane Lathuilière, Moin Nabi, Tassilo Klein, Thiago Oliveira-Santos, Nicu Sebe, Elisa Ricci 0001 |
ICCV | 7 |
| 2019 | Knowledge is Never Enough: Towards Web Aided Deep Open World RecognitionabstractWhile today's robots are able to perform sophisticated tasks, they can only act on objects they have been trained to recognize. This is a severe limitation: any robot will inevitably see new objects in unconstrained settings, and thus will always have visual knowledge gaps. However, standard visual modules are usually built on a limited set of classes and are based on the strong prior that an object must belong to one of those classes. Identifying whether an instance does not belong to the set of known categories (i.e. open set recognition), only partially tackles this problem, as a truly autonomous agent should be able not only to detect what it does not know, but also to extend dynamically its knowledge about the world. We contribute to this challenge with a deep learning architecture that can dynamically update its known classes in an end-to-end fashion. The proposed deep network, based on a deep extension of a non-parametric model, detects whether a perceived object belongs to the set of categories known by the system and learns it without the need to retrain the whole system from scratch. Annotated images about the new category can be provided by an `oracle' (i.e. human supervision), or by autonomous mining of the Web. Experiments on two different databases and on a robot platform demonstrate the promise of our approach. Massimiliano Mancini, Hakan Karaoguz, Elisa Ricci 0001, Patric Jensfelt, Barbara Caputo |
ICRA | 3 |
| 2019 | The RGB-D Triathlon: Towards Agile Visual Toolboxes for RobotsabstractDeep networks have brought significant advances in robot perception, enabling to improve the capabilities of robots in several visual tasks, ranging from object detection and recognition to pose estimation, semantic scene segmentation and many others. Still, most approaches typically address visual tasks in isolation, resulting in overspecialized models which achieve strong performances in specific applications but work poorly in other (often related) tasks. This is clearly sub-optimal for a robot which is often required to perform simultaneously multiple visual recognition tasks in order to properly act and interact with the environment. This problem is exacerbated by the limited computational and memory resources typically available onboard to a robotic platform. The problem of learning flexible models which can handle multiple tasks in a lightweight manner has recently gained attention in the computer vision community and benchmarks supporting this research have been proposed. In this work we study this problem in the robot vision context, proposing a new benchmark, the RGB-D Triathlon, and evaluating state of the art algorithms in this novel challenging scenario. We also define a new evaluation protocol, better suited to the robot vision setting. Results shed light on the strengths and weaknesses of existing approaches and on open issues, suggesting directions for future research. Fabio Cermelli, Massimiliano Mancini, Elisa Ricci 0001, Barbara Caputo |
IROS | 3 |
| 2019 | First Order Motion Model for Image AnimationabstractImage animation consists of generating a video sequence so that an object in a source image is animated according to the motion of a driving video. Our framework addresses this problem without using any annotation or prior information about the specific object to animate. Once trained on a set of videos depicting objects of the same category (e.g. faces, human bodies), our method can be applied to any object of this class. To achieve this, we decouple appearance and motion information using a self-supervised formulation. To support complex motions, we use a representation consisting of a set of learned keypoints along with their local affine transformations. A generator network models occlusions arising during target motions and combines the appearance extracted from the source image and the motion derived from the driving video. Our framework scores best on diverse benchmarks and on a variety of object categories. Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci 0001, Nicu Sebe |
NeurIPS | 4 |
| 2019 | Monocular Depth Estimation Using Multi-Scale Continuous CRFs as Sequential Deep NetworksabstractDepth cues have been proved very useful in various computer vision and robotic tasks. This paper addresses the problem of monocular depth estimation from a single still image. Inspired by the effectiveness of recent works on multi-scale convolutional neural networks (CNN), we propose a deep model which fuses complementary information derived from multiple CNN side outputs. Different from previous methods using concatenation or weighted average schemes, the integration is obtained by means of continuous Conditional Random Fields (CRFs). In particular, we propose two different variations, one based on a cascade of multiple CRFs, the other on a unified graphical model. By designing a novel CNN implementation of mean-field updates for continuous CRFs, we show that both proposed models can be regarded as sequential deep networks and that training can be performed end-to-end. Through an extensive experimental evaluation, we demonstrate the effectiveness of the proposed approach and establish new state of the art results for the monocular depth estimation task on three publicly available datasets, i.e., NYUD-V2, Make3D and KITTI. Dan Xu 0002, Elisa Ricci 0001, Wanli Ouyang, Xiaogang Wang 0001, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Increasing Image Memorability with Neural Style TransferabstractRecent works in computer vision and multimedia have shown that image memorability can be automatically inferred exploiting powerful deep-learning models. This article advances the state of the art in this area by addressing a novel and more challenging issue: “ Given an arbitrary input image, can we make it more memorable? ” To tackle this problem, we introduce an approach based on an editing-by-applying-filters paradigm: given an input image, we propose to automatically retrieve a set of “style seeds,” i.e., a set of style images that, applied to the input image through a neural style transfer algorithm, provide the highest increase in memorability. We show the effectiveness of the proposed approach with experiments on the publicly available LaMem dataset, performing both a quantitative evaluation and a user study. To demonstrate the flexibility of the proposed framework, we also analyze the impact of different implementation choices, such as using different state-of-the-art neural style transfer methods. Finally, we show several qualitative results to provide additional insights on the link between image style and memorability. Aliaksandr Siarohin, Gloria Zen, Cveta Majtanovic, Xavier Alameda-Pineda, Elisa Ricci 0001, Nicu Sebe |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2018 | Unsupervised Adversarial Depth Estimation Using Cycled Generative NetworksabstractWhile recent deep monocular depth estimation approaches based on supervised regression have achieved remarkable performance, costly ground truth annotations are required during training. To cope with this issue, in this paper we present a novel unsupervised deep learning approach for predicting depth maps and show that the depth estimation task can be effectively tackled within an adversarial learning framework. Specifically, we propose a deep generative network that learns to predict the correspondence field (i.e. the disparity map) between two image views in a calibrated stereo camera setting. The proposed architecture consists of two generative sub-networks jointly trained with adversarial learning for reconstructing the disparity map and organized in a cycle such as to provide mutual constraints and supervision to each other. Extensive experiments on the publicly available datasets KITTI and Cityscapes demonstrate the effectiveness of the proposed model and competitive results with state of the art methods. The code is available at https://github.com/andrea-pilzer/unsup-stereo-depthGAN. Andrea Pilzer, Dan Xu 0002, Mihai Marian Puscas, Elisa Ricci 0001, Nicu Sebe |
3DV | 4 |
| 2018 | Enhancing Perceptual Attributes with Bayesian Style Generation
Aliaksandr Siarohin, Gloria Zen, Nicu Sebe, Elisa Ricci 0001 |
ACCV (1) | 4 |
| 2018 | Structured Attention Guided Convolutional Neural Fields for Monocular Depth EstimationabstractRecent works have shown the benefit of integrating Conditional Random Fields (CRFs) models into deep architectures for improving pixel-level prediction tasks. Following this line of research, in this paper we introduce a novel approach for monocular depth estimation. Similarly to previous works, our method employs a continuous CRF to fuse multi-scale information derived from different layers of a front-end Convolutional Neural Network (CNN). Differently from past works, our approach benefits from a structured attention model which automatically regulates the amount of information transferred between corresponding features at different scales. Importantly, the proposed attention model is seamlessly integrated into the CRF, allowing end-to-end training of the entire architecture. Our extensive experimental evaluation demonstrates the effectiveness of the proposed method which is competitive with previous methods on the KITTI benchmark and outperforms the state of the art on the NYU Depth V2 dataset. Dan Xu 0002, Wei Wang 0108, Hao Tang 0005, Hong Liu 0008, Nicu Sebe, Elisa Ricci 0001 |
CVPR | 6 |
| 2018 | Every Smile Is Unique: Landmark-Guided Diverse Smile GenerationabstractEach smile is unique: one person surely smiles in different ways (e.g. closing/opening the eyes or mouth). Given one input image of a neutral face, can we generate multiple smile videos with distinctive characteristics? To tackle this one-to-many video generation problem, we propose a novel deep learning architecture named Conditional Multi-Mode Network (CMM-Net). To better encode the dynamics of facial expressions, CMM-Net explicitly exploits facial landmarks for generating smile sequences. Specifically, a variational auto-encoder is used to learn a facial landmark embedding. This single embedding is then exploited by a conditional recurrent network which generates a landmark embedding sequence conditioned on a specific expression (e.g. spontaneous smile). Next, the generated landmark embeddings are fed into a multi-mode recurrent landmark generator, producing a set of landmark sequences still associated to the given smile class but clearly distinct from each other. Finally, these landmark sequences are translated into face videos. Our experimental results demonstrate the effectiveness of our CMM-Net in generating realistic videos of multiple smile expressions. Wei Wang 0108, Xavier Alameda-Pineda, Dan Xu 0002, Pascal Fua, Elisa Ricci 0001, Nicu Sebe |
CVPR | 5 |
| 2018 | Boosting Domain Adaptation by Discovering Latent DomainsabstractCurrent Domain Adaptation (DA) methods based on deep architectures assume that the source samples arise from a single distribution. However, in practice most datasets can be regarded as mixtures of multiple domains. In these cases exploiting single-source DA methods for learning target classifiers may lead to sub-optimal, if not poor, results. In addition, in many applications it is difficult to manually provide the domain labels for all source data points, i.e. latent domains should be automatically discovered. This paper introduces a novel Convolutional Neural Network (CNN) architecture which (i) automatically discovers latent domains in visual datasets and (ii) exploits this information to learn robust target classifiers. Our approach is based on the introduction of two main components, which can be embedded into any existing CNN architecture: (i) a side branch that automatically computes the assignment of a source sample to a latent domain and (ii) novel layers that exploit domain membership information to appropriately align the distribution of the CNN internal feature representations to a reference distribution. We test our approach on publicly-available datasets, showing that it outperforms state-of-the-art multi-source DA methods by a large margin. Massimiliano Mancini, Lorenzo Porzi, Samuel Rota Bulò, Barbara Caputo, Elisa Ricci 0001 |
CVPR | 5 |
| 2018 | Best Sources Forward: Domain Generalization through Source-Specific NetsabstractA long standing problem in visual object categorization is the ability of algorithms to generalize across different testing conditions. The problem has been formalized as a covariate shift among the probability distributions generating the training data (source) and the test data (target) and several domain adaptation methods have been proposed to address this issue. While these approaches have considered the single source-single target scenario, it is plausible to have multiple sources and require adaptation to any possible target domain. This last scenario, named Domain Generalization (DG), is the focus of our work. Differently from previous DG methods which learn domain invariant representations from source data, we design a deep network with multiple domain-specific classifiers, each associated to a source domain. At test time we estimate the probabilities that a target sample belongs to each source domain and exploit them to optimally fuse the classifiers predictions. To further improve the generalization ability of our model, we also introduced a domain agnostic component supporting the final classifier. Experiments on two public benchmarks demonstrate the power of our approach. Massimiliano Mancini, Samuel Rota Bulò, Barbara Caputo, Elisa Ricci 0001 |
ICIP | 4 |
| 2018 | Kitting in the Wild through Online Domain AdaptationabstractTechnological developments call for increasing perception and action capabilities of robots. Among other skills, vision systems that can adapt to any possible change in the working conditions are needed. Since these conditions are unpredictable, we need benchmarks which allow to assess the generalization and robustness capabilities of our visual recognition algorithms. In this work we focus on robotic kitting in unconstrained scenarios. As a first contribution, we present a new visual dataset for the kitting task. Differently from standard object recognition datasets, we provide images of the same objects acquired under various conditions where camera, illumination and background are changed. This novel dataset allows for testing the robustness of robot visual recognition algorithms to a series of different domain shifts both in isolation and unified. Our second contribution is a novel online adaptation algorithm for deep models, based on batch-normalization layers, which allows to continuously adapt a model to the current working conditions. Differently from standard domain adaptation algorithms, it does not require any image from the target domain at training time. We benchmark the performance of the algorithm on the proposed dataset, showing its capability to fill the gap between the performances of a standard architecture and its counterpart adapted offline to the given target domain. Massimiliano Mancini, Hakan Karaoguz, Elisa Ricci 0001, Patric Jensfelt, Barbara Caputo |
IROS | 3 |
| 2018 | Joint Estimation of Human Pose and Conversational Groups from Social Scenes
Jagannadan Varadarajan, Subramanian Ramanathan, Samuel Rota Bulò, Narendra Ahuja, Oswald Lanz, Elisa Ricci 0001 |
Int. J. Comput. Vis. | 6 |
| 2018 | Cross-Paced Representation Learning With Partial Curricula for Sketch-Based Image RetrievalabstractIn this paper, we address the problem of learning robust cross-domain representations for sketch-based image retrieval (SBIR). While, most SBIR approaches focus on extracting low- and mid-level descriptors for direct feature matching, recent works have shown the benefit of learning coupled feature representations to describe data from two related sources. However, cross-domain representation learning methods are typically cast into non-convex minimization problems that are difficult to optimize, leading to unsatisfactory performance. Inspired by self-paced learning (SPL), a learning methodology designed to overcome convergence issues related to local optima by exploiting the samples in a meaningful order (i.e., easy to hard), we introduce the cross-paced partial curriculum learning (CPPCL) framework. Compared with existing SPL methods which only consider a single modality and cannot deal with prior knowledge, CPPCL is specifically designed to assess the learning pace by jointly handling data from dual sources and modality-specific prior information provided in the form of partial curricula. In addition, thanks to the learned dictionaries, we demonstrate that the proposed CPPCL embeds robust coupled representations for SBIR. Our approach is extensively evaluated on four publicly available datasets (i.e., CUFS, Flickr15K, QueenMary SBIR, and TU-Berlin Extension datasets), showing superior performance over competing SBIR methods. Dan Xu 0002, Xavier Alameda-Pineda, Jingkuan Song, Elisa Ricci 0001, Nicu Sebe |
IEEE Trans. Image Process. | 4 |
| 2017 | Viraliency: Pooling Local ViralityabstractIn our overly-connected world, the automatic recognition of virality - the quality of an image or video to be rapidly and widely spread in social networks - is of crucial importance, and has recently awaken the interest of the computer vision community. Concurrently, recent progress in deep learning architectures showed that global pooling strategies allow the extraction of activation maps, which highlight the parts of the image most likely to contain instances of a certain class. We extend this concept by introducing a pooling layer that learns the size of the support area to be averaged: the learned top-N average (LENA) pooling. We hypothesize that the latent concepts (feature maps) describing virality may require such a rich pooling strategy. We assess the effectiveness of the LENA layer by appending it on top of a convolutional siamese architecture and evaluate its performance on the task of predicting and localizing virality. We report experiments on two publicly available datasets annotated for virality and show that our method outperforms state-of-the-art approaches. Xavier Alameda-Pineda, Andrea Pilzer, Dan Xu 0002, Nicu Sebe, Elisa Ricci 0001 |
CVPR | 5 |
| 2017 | Learning Cross-Modal Deep Representations for Robust Pedestrian Detection
Dan Xu 0002, Wanli Ouyang, Elisa Ricci 0001, Xiaogang Wang 0001, Nicu Sebe |
CVPR | 3 |
| 2017 | Multi-scale Continuous CRFs as Sequential Deep Networks for Monocular Depth EstimationabstractThis paper addresses the problem of depth estimation from a single still image. Inspired by recent works on multi-scale convolutional neural networks (CNN), we propose a deep model which fuses complementary information derived from multiple CNN side outputs. Different from previous methods, the integration is obtained by means of continuous Conditional Random Fields (CRFs). In particular, we propose two different variations, one based on a cascade of multiple CRFs, the other on a unified graphical model. By designing a novel CNN implementation of mean-field updates for continuous CRFs, we show that both proposed models can be regarded as sequential deep networks and that training can be performed end-to-end. Through extensive experimental evaluation we demonstrate the effectiveness of the proposed approach and establish new state of the art results on publicly available datasets. Dan Xu 0002, Elisa Ricci 0001, Wanli Ouyang, Xiaogang Wang 0001, Nicu Sebe |
CVPR | 2 |
| 2017 | AutoDIAL: Automatic Domain Alignment LayersabstractClassifiers trained on given databases perform poorly when tested on data acquired in different settings. This is explained in domain adaptation through a shift among distributions of the source and target domains. Attempts to align them have traditionally resulted in works reducing the domain shift by introducing appropriate loss terms, measuring the discrepancies between source and target distributions, in the objective function. Here we take a different route, proposing to align the learned representations by embedding in any given network specific Domain Alignment Layers, designed to match the source and target feature distributions to a reference one. Opposite to previous works which define a priori in which layers adaptation should be performed, our method is able to automatically learn the degree of feature alignment required at different levels of the deep network. Thorough experiments on different public benchmarks, in the unsupervised setting, confirm the power of our approach. Fabio Maria Carlucci, Lorenzo Porzi, Barbara Caputo, Elisa Ricci 0001, Samuel Rota Bulò |
ICCV | 4 |
| 2017 | Depth-aware convolutional neural networks for accurate 3D pose estimation in RGB-D imagesabstractMost recent approaches to 3D pose estimation from RGB-D images address the problem in a two-stage pipeline. First, they learn a classifier-typically a random forest-to predict the position of each input pixel on the object surface. These estimates are then used to define an energy function that is minimized w.r.t. the object pose. In this paper, we focus on the first stage of the problem and propose a novel classifier based on a depth-aware Convolutional Neural Network. This classifier is able to learn a scale-adaptive regression model that yields very accurate pixel-level predictions, allowing to finally estimate the pose using a simple RANSAC-based scheme, with no need to optimize complex ad hoc energy functions. Our experiments on publicly available datasets show that our approach achieves remarkable improvements over state-of-the-art methods. Lorenzo Porzi, Adrián Peñate Sánchez, Elisa Ricci 0001, Francesc Moreno-Noguer |
IROS | 3 |
| 2017 | How to Make an Image More Memorable?: A Deep Style Transfer ApproachabstractRecent works have shown that it is possible to automatically predict intrinsic image properties like memorability. In this paper, we take a step forward addressing the question: "Can we make an image more memorable?". Methods for automatically increasing image memorability would have an impact in many application fields like education, gaming or advertising. Our work is inspired by the popular editing-by-applying-filters paradigm adopted in photo editing applications, like Instagram and Prisma. In this context, the problem of increasing image memorability maps to that of retrieving ``memorabilizing'' filters or style ``seeds''. Still, users generally have to go through most of the available filters before finding the desired solution, thus turning the editing process into a resource and time consuming task. In this work, we show that it is possible to automatically retrieve the best style seeds for a given image, thus remarkably reducing the number of human attempts needed to find a good match. Our approach leverages from recent advances in the field of image synthesis and adopts a deep architecture for generating a memorable picture from a given input image and a style seed. Importantly, to automatically select the best style a novel learning-based solution, also relying on deep models, is proposed. Our experimental evaluation, conducted on publicly available benchmarks, demonstrates the effectiveness of the proposed approach for generating memorable images through automatic style seed selection. Aliaksandr Siarohin, Gloria Zen, Cveta Majtanovic, Xavier Alameda-Pineda, Elisa Ricci 0001, Nicu Sebe |
ICMR | 5 |
| 2017 | Learning Deep Structured Multi-Scale Features using Attention-Gated CRFs for Contour PredictionabstractRecent works have shown that exploiting multi-scale representations deeply learned via convolutional neural networks (CNN) is of tremendous importance for accurate contour detection. This paper presents a novel approach for predicting contours which advances the state of the art in two fundamental aspects, i.e. multi-scale feature generation and fusion. Different from previous works directly considering multi-scale feature maps obtained from the inner layers of a primary CNN architecture, we introduce a hierarchical deep model which produces more rich and complementary representations. Furthermore, to refine and robustly fuse the representations learned at different scales, the novel Attention-Gated Conditional Random Fields (AG-CRFs) are proposed. The experiments ran on two publicly available datasets (BSDS500 and NYUDv2) demonstrate the effectiveness of the latent AG-CRF model and of the overall hierarchical framework. Dan Xu 0002, Wanli Ouyang, Xavier Alameda-Pineda, Elisa Ricci 0001, Xiaogang Wang 0001, Nicu Sebe |
NIPS | 4 |
| 2017 | Detecting anomalous events in videos by learning deep representations of appearance and motion
Dan Xu 0002, Yan Yan 0002, Elisa Ricci 0001, Nicu Sebe |
Comput. Vis. Image Underst. | 3 |
| 2017 | An automatic image-to-DEM alignment approach for annotating mountains pictures on a smartphone
Lorenzo Porzi, Samuel Rota Bulò, Oswald Lanz, Paolo Valigi, Elisa Ricci 0001 |
Mach. Vis. Appl. | 5 |
| 2016 | Recognizing Emotions from Abstract Paintings Using Non-Linear Matrix CompletionabstractAdvanced computer vision and machine learning techniques tried to automatically categorize the emotions elicited by abstract paintings with limited success. Since the annotation of the emotional content is highly resourceconsuming, datasets of abstract paintings are either constrained in size or partially annotated. Consequently, it is natural to address the targeted task within a transductive framework. Intuitively, the use of multi-label classification techniques is desirable so to synergically exploit the relations between multiple latent variables, such as emotional content, technique, author, etc. A very popular approach for transductive multi-label recognition under linear classification settings is matrix completion. In this study we introduce non-linear matrix completion (NLMC), thus extending classical linear matrix completion techniques to the non-linear case. Together with the theory grounding the model, we propose an efficient optimization solver. As shown by our extensive experimental validation on two publicly available datasets, NLMC outperforms state-of-the-art methods when recognizing emotions from abstract paintings. Xavier Alameda-Pineda, Elisa Ricci 0001, Yan Yan 0002, Nicu Sebe |
CVPR | 2 |
| 2016 | Self-Adaptive Matrix Completion for Heart Rate Estimation from Face Videos under Realistic ConditionsabstractRecent studies in computer vision have shown that, while practically invisible to a human observer, skin color changes due to blood flow can be captured on face videos and, surprisingly, be used to estimate the heart rate (HR). While considerable progress has been made in the last few years, still many issues remain open. In particular, state of-the-art approaches are not robust enough to operate in natural conditions (e.g. in case of spontaneous movements, facial expressions, or illumination changes). Opposite to previous approaches that estimate the HR by processing all the skin pixels inside a fixed region of interest, we introduce a strategy to dynamically select face regions useful for robust HR estimation. Our approach, inspired by recent advances on matrix completion theory, allows us to predict the HR while simultaneously discover the best regions of the face to be used for estimation. Thorough experimental evaluation conducted on public benchmarks suggests that the proposed approach significantly outperforms state-of the-art HR estimation methods in naturalistic conditions. Sergey Tulyakov, Xavier Alameda-Pineda, Elisa Ricci 0001, Lijun Yin 0001, Jeffrey F. Cohn, Nicu Sebe |
CVPR | 3 |
| 2016 | Multi-Paced Dictionary Learning for cross-domain retrieval and recognitionabstractSeveral applications benefit from learning coupled representations able to describe data from multiple sources. For instance, cross-domain dictionary learning methods demonstrated to be particularly effective. In this paper we introduce Multi-Paced Dictionary Learning (MPDL) and propose an instantiation of it under the framework of cross-domain dictionary learning. MPDL is inspired by previous works on self-paced learning, a framework able to enhance the accuracy of conventional learning models by presenting the training data in a meaningful order, i.e. easy samples are provided first. However, most of existing self-paced learning methods only consider a single modality, while MPDL is specifically designed to assess the learning pace when data from multiple sources are available. We present the model and propose an efficient algorithm to learn the dictionaries and codes. The approach is validated via experiments on two different tasks, namely cross-media retrieval and sketch-to-photo face recognition, using publicly available datasets. Dan Xu 0002, Jingkuan Song, Xavier Alameda-Pineda, Elisa Ricci 0001, Nicu Sebe |
ICPR | 4 |
| 2016 | Emerging Topics in Learning from Noisy and Missing DataabstractWhile vital for handling most multimedia and computer vision problems, collecting large scale fully annotated datasets is a resource-consuming, often unaffordable task. Indeed, on the one hand datasets need to be large and variate enough so that learning strategies can successfully exploit the variability inherently present in real data, but on the other hand they should be small enough so that they can be fully annotated at a reasonable cost. With the overwhelming success of (deep) learning methods, the traditional problem of balancing between dataset dimensions and resources needed for annotations became a full-fledged dilemma. In this context, methodological approaches able to deal with partially described data sets represent a one-of-a-kind opportunity to find the right balance between data variability and resource-consumption in annotation. These include methods able to deal with noisy, weak or partial annotations. In this tutorial we will present several recent methodologies addressing different visual tasks under the assumption of noisy, weakly annotated data sets. Xavier Alameda-Pineda, Timothy M. Hospedales, Elisa Ricci 0001, Nicu Sebe, Xiaogang Wang 0001 |
ACM Multimedia | 3 |
| 2016 | A Deeply-Supervised Deconvolutional Network for Horizon Line DetectionabstractAutomatic skyline detection from mountain pictures is an important task in many applications, such as web image retrieval, augmented reality and autonomous robot navigation. Recent works addressing the problem of Horizon Line Detection (HLD) demonstrated that learning-based boundary detection techniques are more accurate than traditional filtering methods. In this paper we introduce a novel approach for skyline detection, which adheres to a learning-based paradigm and exploits the representation power of deep architectures to improve the horizon line detection accuracy. Differently from previous works, we explore a novel deconvolutional architecture, which introduces intermediate levels of supervision to support the learning process. Our experiments, conducted on a publicly available dataset, confirm that the proposed method outperforms previous learning-based HLD techniques by reducing the number of spurious edge pixels. Lorenzo Porzi, Samuel Rota Bulò, Elisa Ricci 0001 |
ACM Multimedia | 3 |
| 2016 | Academic Coupled Dictionary Learning for Sketch-based Image RetrievalabstractIn the last few years, the query-by-visual-example paradigm gained popularity, specially for content based retrieval systems. As sketches represent a natural way of expressing a synthetic query, recent research efforts focused on developing algorithmic solutions to address the sketch-based image retrieval (SBIR) problem. Within this context, we propose a novel approach for SBIR that, unlike previous methods, is able to exploit the visual complexity inherently present in sketches and images. We introduce academic learning, a paradigm in which the sample learning order is constructed both from the data, as in self-paced learning, and from partial curricula. We propose an instantiation of this paradigm within the framework of coupled dictionary learning to address the SBIR task. We also present an efficient algorithm to learn the dictionaries and the codes, and to pace the learning combining the reconstruction error, the prior knowledge suggested by the partial curricula and the cross-domain code coherence. In order to evaluate the proposed approach, we report an extensive experimental validation showing that the proposed method outperforms the state-of-the-art in coupled dictionary learning and in SBIR on three different publicly available datasets. Dan Xu 0002, Xavier Alameda-Pineda, Jingkuan Song, Elisa Ricci 0001, Nicu Sebe |
ACM Multimedia | 4 |
| 2016 | SALSA: A Novel Dataset for Multimodal Group Behavior AnalysisabstractStudying free-standing conversational groups (FCGs) in unstructured social settings (e.g., cocktail party ) is gratifying due to the wealth of information available at the group (mining social networks) and individual (recognizing native behavioral and personality traits) levels. However, analyzing social scenes involving FCGs is also highly challenging due to the difficulty in extracting behavioral cues such as target locations, their speaking activity and head/body pose due to crowdedness and presence of extreme occlusions. To this end, we propose SALSA, a novel dataset facilitating multimodal and Synergetic sociAL Scene Analysis, and make two main contributions to research on automated social interaction analysis: (1) SALSA records social interactions among 18 participants in a natural, indoor environment for over 60 minutes, under the poster presentation and cocktail party contexts presenting difficulties in the form of low-resolution images, lighting variations, numerous occlusions, reverberations and interfering sound sources; (2) To alleviate these problems we facilitate multimodal analysis by recording the social interplay using four static surveillance cameras and sociometric badges worn by each participant, comprising the microphone, accelerometer, bluetooth and infrared sensors. In addition to raw data, we also provide annotations concerning individuals' personality as well as their position, head, body orientation and F-formation information over the entire event duration. Through extensive experiments with state-of-the-art approaches, we show (a) the limitations of current methods and (b) how the recorded multiple cues synergetically aid automatic analysis of social interactions. SALSA is available at http://tev.fbk.eu/salsa. Xavier Alameda-Pineda, Jacopo Staiano, Subramanian Ramanathan, Ligia Maria Batrinca, Elisa Ricci 0001, Bruno Lepri, Oswald Lanz, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2016 | A Multi-Task Learning Framework for Head Pose Estimation under Target MotionabstractRecently, head pose estimation (HPE) from low-resolution surveillance data has gained in importance. However, monocular and multi-view HPE approaches still work poorly under target motion, as facial appearance distorts owing to camera perspective and scale changes when a person moves around. To this end, we propose FEGA-MTL, a novel framework based on Multi-Task Learning (MTL) for classifying the head pose of a person who moves freely in an environment monitored by multiple, large field-of-view surveillance cameras. Upon partitioning the monitored scene into a dense uniform spatial grid, FEGA-MTL simultaneously clusters grid partitions into regions with similar facial appearance, while learning region-specific head pose classifiers. In the learning phase, guided by two graphs which a-priori model the similarity among (1) grid partitions based on camera geometry and (2) head pose classes, FEGA-MTL derives the optimal scene partitioning and associated pose classifiers. Upon determining the target's position using a person tracker at test time, the corresponding region-specific classifier is invoked for HPE. The FEGA-MTL framework naturally extends to a weakly supervised setting where the target's walking direction is employed as a proxy in lieu of head orientation. Experiments confirm that FEGA-MTL significantly outperforms competing single-task and multi-task learning methods in multi-view settings. Yan Yan 0002, Elisa Ricci 0001, Subramanian Ramanathan, Gaowen Liu, Oswald Lanz, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Learning Personalized Models for Facial Expression Analysis and Gesture RecognitionabstractFacial expression and gesture recognition algorithms are key enabling technologies for human-computer interaction (HCI) systems. State of the art approaches for automatic detection of body movements and analyzing emotions from facial features heavily rely on advanced machine learning algorithms. Most of these methods are designed for the average user, but the assumption “one-size-fits-all” ignores diversity in cultural background, gender, ethnicity, and personal behavior, and limits their applicability in real-world scenarios. A possible solution is to build personalized interfaces, which practically implies learning person-specific classifiers and usually collecting a significant amount of labeled samples for each novel user. As data annotation is a tedious and time-consuming process, in this paper we present a framework for personalizing classification models which does not require labeled target data. Personalization is achieved by devising a novel transfer learning approach. Specifically, we propose a regression framework which exploits auxiliary (source) annotated data to learn the relation between person-specific sample distributions and parameters of the corresponding classifiers. Then, when considering a new target user, the classification model is computed by simply feeding the associated (unlabeled) sample distribution into the learned regression function. We evaluate the proposed approach in different applications: pain recognition and action unit detection using visual data and gestures classification using inertial measurements, demonstrating the generality of our method with respect to different input data types and basic classifiers. We also show the advantages of our approach in terms of accuracy and computational time both with respect to user-independent approaches and to previous personalization techniques. Gloria Zen, Lorenzo Porzi, Enver Sangineto, Elisa Ricci 0001, Nicu Sebe |
IEEE Trans. Multim. | 4 |
| 2015 | Online Multitask Learning for Machine Translation Quality EstimationabstractJosé G. C. de Souza, Matteo Negri, Elisa Ricci, Marco Turchi. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. José Guilherme Camargo de Souza, Matteo Negri, Elisa Ricci 0001, Marco Turchi |
ACL (1) | 3 |
| 2015 | Learning Deep Representations of Appearance and Motion for Anomalous Event DetectionabstractWe present a novel unsupervised deep learning framework for anomalous event detection in complex video scenes.While most existing works merely use hand-crafted appearance and motion features, we propose Appearance and Motion DeepNet (AMDN) which utilizes deep neural networks to automatically learn feature representations.To exploit the complementary information of both appearance and motion patterns, we introduce a novel double fusion framework, combining both the benefits of traditional early fusion and late fusion strategies.Specifically, stacked denoising autoencoders are proposed to separately learn both appearance and motion features as well as a joint representation (early fusion).Based on the learned representations, multiple one-class SVM models are used to predict the anomaly scores of each input, which are then integrated with a late fusion strategy for final anomaly detection.We evaluate the proposed method on two publicly available video surveillance datasets, showing competitive performance with respect to state of the art approaches. Dan Xu 0002, Elisa Ricci 0001, Yan Yan 0002, Jingkuan Song, Nicu Sebe |
BMVC | 2 |
| 2015 | Uncovering Interactions and Interactors: Joint Estimation of Head, Body Orientation and F-Formations from Surveillance VideosabstractWe present a novel approach for jointly estimating targets' head, body orientations and conversational groups called F-formations from a distant social scene (e.g., a cocktail party captured by surveillance cameras). Differing from related works that have (i) coupled head and body pose learning by exploiting the limited range of orientations that the two can jointly take, or (ii) determined F-formations based on the mutual head (but not body) orientations of interactors, we present a unified framework to jointly infer both (i) and (ii). Apart from exploiting spatial and orientation relationships, we also integrate cues pertaining to temporal consistency and occlusions, which are beneficial while handling low-resolution data under surveillance settings. Efficacy of the joint inference framework reflects via increased head, body pose and F-formation estimation accuracy over the state-of-the-art, as confirmed by extensive experiments on two social datasets. Elisa Ricci 0001, Jagannadan Varadarajan, Subramanian Ramanathan, Samuel Rota Bulò, Narendra Ahuja, Oswald Lanz |
ICCV | 1 |
| 2015 | Inferring Painting Style with Multi-Task Dictionary Learning
Gaowen Liu, Yan Yan 0002, Elisa Ricci 0001, Yi Yang 0001, Yahong Han, Stefan Winkler 0001, Nicu Sebe |
IJCAI | 3 |
| 2015 | Analyzing Free-standing Conversational Groups: A Multimodal ApproachabstractDuring natural social gatherings, humans tend to organize themselves in the so-called free-standing conversational groups. In this context, robust head and body pose estimates can facilitate the higher-level description of the ongoing interplay. Importantly, visual information typically obtained with a distributed camera network might not suffice to achieve the robustness sought. In this line of thought, recent advances in wearable sensing technology open the door to multimodal and richer information flows. In this paper we propose to cast the head and body pose estimation problem into a matrix completion task. We introduce a framework able to fuse multimodal data emanating from a combination of distributed and wearable sensors, taking into account the temporal consistency, the head/body coupling and the noise inherent to the scenario. We report results on the novel and challenging SALSA dataset, containing visual, auditory and infrared recordings of 18 people interacting in a regular indoor environment. We demonstrate the soundness of the proposed method and the usability for higher-level tasks such as the detection of F-formations and the discovery of social attention attractors. Xavier Alameda-Pineda, Yan Yan 0002, Elisa Ricci 0001, Oswald Lanz, Nicu Sebe |
ACM Multimedia | 3 |
| 2015 | Predicting and Understanding Urban Perception with Convolutional Neural NetworksabstractCities' visual appearance plays a central role in shaping human perception and response to the surrounding urban environment. For example, the visual qualities of urban spaces affect the psychological states of their inhabitants and can induce negative social outcomes. Hence, it becomes critically important to understand people's perceptions and evaluations of urban spaces. Previous works have demonstrated that algorithms can be used to predict high level attributes of urban scenes (e.g. safety, attractiveness, uniqueness), accurately emulating human perception. In this paper we propose a novel approach for predicting the perceived safety of a scene from Google Street View Images. Opposite to previous works, we formulate the problem of learning to predict high level judgments as a ranking task and we employ a Convolutional Neural Network (CNN), significantly improving the accuracy of predictions over previous methods. Interestingly, the proposed CNN architecture relies on a novel pooling layer, which permits to automatically discover the most important areas of the images for predicting the concept of perceived safety. An extensive experimental evaluation, conducted on the publicly available Place Pulse dataset, demonstrates the advantages of the proposed approach over state-of-the-art methods. Lorenzo Porzi, Samuel Rota Bulò, Bruno Lepri, Elisa Ricci 0001 |
ACM Multimedia | 4 |
| 2015 | Jointly Estimating Interactions and Head, Body Pose of Interactors from Distant Social ScenesabstractWe present joint estimation of F-formations and head, body pose of interactors in a social scene captured by surveillance cameras. Unlike prior works that have focused on (a) discovering F-formations based on head pose and position cues, or (b) jointly learned head and body pose of individuals based on anatomic constraints, we exploit positional and pose cues characterizing interactors and interactions to jointly infer both (a) and (b). We show how the joint inference framework benefits both F-formation and head, body pose estimation accuracy via experiments on two social datasets. Subramanian Ramanathan, Jagannadan Varadarajan, Elisa Ricci 0001, Oswald Lanz, Stefan Winkler 0001 |
ACM Multimedia | 3 |
| 2015 | Egocentric Daily Activity Recognition via Multitask ClusteringabstractRecognizing human activities from videos is a fundamental research problem in computer vision. Recently, there has been a growing interest in analyzing human behavior from data collected with wearable cameras. First-person cameras continuously record several hours of their wearers' life. To cope with this vast amount of unlabeled and heterogeneous data, novel algorithmic solutions are required. In this paper, we propose a multitask clustering framework for activity of daily living analysis from visual data gathered from wearable cameras. Our intuition is that, even if the data are not annotated, it is possible to exploit the fact that the tasks of recognizing everyday activities of multiple individuals are related, since typically people perform the same actions in similar environments, e.g., people working in an office often read and write documents). In our framework, rather than clustering data from different users separately, we propose to look for clustering partitions which are coherent among related tasks. In particular, two novel multitask clustering algorithms, derived from a common optimization problem, are introduced. Our experimental evaluation, conducted both on synthetic data and on publicly available first-person vision data sets, shows that the proposed approach outperforms several single-task and multitask learning methods. Yan Yan 0002, Elisa Ricci 0001, Gaowen Liu, Nicu Sebe |
IEEE Trans. Image Process. | 2 |
| 2014 | Recognizing Daily Activities from First-Person Videos with Multi-task Clustering
Yan Yan 0002, Elisa Ricci 0001, Gaowen Liu, Nicu Sebe |
ACCV (4) | 2 |
| 2014 | Exploiting transfer learning for personalized view invariant gesture recognitionabstractA robust gesture recognition system is an essential component in many human-computer interaction applications. In particular, the widespread adoption of portable devices and the diffusion of autonomous systems with limited power and load capacity has increased the need of developing efficient recognition algorithms which operates on video streams recorded from low cost devices and which can cope with the challenging issue of point of view changes. A further challenge arises as different users tend to perform the same gesture with different styles and speeds. Thus a classifier trained with gestures data of certain set of users may work poorly when data from other users are being processed. However, as often a mobile device or a robot are intended to be used by a single or by a small group of people, it would be desirable to have a gesture recognition system designed specifically for these users. In this paper we introduce a novel approach to face the problems of view-invariance and user personalization in the context of gesture interaction systems. More specifically, we propose a domain adaptation framework based on a feature space augmentation approach operating on robust view-invariant Self Similarity Matrix descriptors. To prove the effectiveness of our method a dataset corresponding to 17 users performing 10 different gestures under 3 point of views is collected and an extensive experimental evaluation is performed. Gabriele Costante, Valerio Galieni, Yan Yan 0002, Mario Luca Fravolini, Elisa Ricci 0001, Paolo Valigi |
ICASSP | 5 |
| 2014 | It's all about habits: Exploiting multi-task clustering for activities of daily living analysisabstractMotivated by applications in areas such as patient monitoring, tele-rehabilitation and ambient assisted living, analyzing activities of daily living is an active research topic in computer vision and image processing. In this paper we address the problem of everyday activity recognition from unlabeled data proposing a novel multi-task clustering (MTC) approach. Our intuition is that, when analyzing activities of daily living, we can take advantage of the fact that people tend to perform the same actions in the same environment (e.g. people working in an office environment use to read and write documents). Thus, even if labels are not available, information about typical activities can be exploited in the learning process. Arguing that the tasks of recognizing activities of specific individuals are related, we resort on multi-task learning and rather than clustering the data of each individual separately, we also look for clustering results which are coherent among related tasks. Extensive experimental results show that our method outperforms several state-of-the-art approaches by up to 11% on the Rochester activities of daily living dataset. Yan Yan 0002, Elisa Ricci 0001, Negar Rostamzadeh, Nicu Sebe |
ICIP | 2 |
| 2014 | Unsupervised Domain Adaptation for Personalized Facial Emotion RecognitionabstractThe way in which human beings express emotions depends on their specific personality and cultural background. As a consequence, person independent facial expression classifiers usually fail to accurately recognize emotions which vary between different individuals. On the other hand, training a person-specific classifier for each new user is a time consuming activity which involves collecting hundreds of labeled samples. In this paper we present a personalization approach in which only unlabeled target-specific data are required. The method is based on our previous paper [20] in which a regression framework is proposed to learn the relation between the user's specific sample distribution and the parameters of her/his classifier. Once this relation is learned, a target classifier can be constructed using only the new user's sample distribution to transfer the personalized parameters. The novelty of this paper with respect to [20] is the introduction of a new method to represent the source sample distribution based on using only the Support Vectors of the source classifiers. Moreover, we present here a simplified regression framework which achieves the same or even slightly superior experimental results with respect to [20] but it is much easier to reproduce. Gloria Zen, Enver Sangineto, Elisa Ricci 0001, Nicu Sebe |
ICMI | 3 |
| 2014 | Clustered Multi-task Linear Discriminant Analysis for View Invariant Color-Depth Action RecognitionabstractThe widespread adoption of low-cost depth cameras has opened new opportunities to improve traditional action recognition systems. In this paper we focus on the specific problem of action recognition under view point changes and propose a novel approach for view-invariant action recognition operating jointly on visual data of color and depth camera channels. Our method is based on the unique combination of robust Self-Similarity Matrix (SSM) descriptors and multi-task learning. Indeed, multi-view action recognition is inherently a multi-task learning problem: images from a camera view can be modeled as visual data associated to the same task and it is reasonable to assume that the data of different tasks (camera views) are related to each other. In this work we propose a novel algorithm extending Multi-Task Linear Discriminant Analysis (MT-LDA) to enhance its flexibility by learning the dependencies between different views. Extensive experimental results on the publicly available ACT42dataset demonstrate the effectiveness of the proposed method. Yan Yan 0002, Elisa Ricci 0001, Gaowen Liu, Subramanian Ramanathan, Nicu Sebe |
ICPR | 2 |
| 2014 | Evaluating Multi-task Learning for Multi-view Head-Pose Classification in Interactive EnvironmentsabstractSocial attention behavior offers vital cues towards inferring one's personality traits from interactive settings such as round-table meetings and cocktail parties. Head orientation is typically employed as a proxy for determining the social attention direction when faces are captured at low-resolution. Recently, multi-task learning has been proposed to robustly compute head pose under perspective and scale-based facial appearance variations when multiple, distant and large field-of-view cameras are employed for visual analysis in smart-room applications. In this paper, we evaluate the effectiveness of an SVM-based MTL (SVM+MTL) framework with various facial descriptors (KL, HOG, LBP, etc.). The KL+HOG feature combination is found to produce the best classification performance, with SVM+MTL outperforming classical SVM irrespective of the feature used. Yan Yan 0002, Subramanian Ramanathan, Elisa Ricci 0001, Oswald Lanz, Nicu Sebe |
ICPR | 3 |
| 2014 | Simultaneous Ground Metric Learning and Matrix Factorization with Earth Mover's DistanceabstractNon-negative matrix factorization is widely used in pattern recognition as it has been proved to be an effective method for dimensionality reduction and clustering. We propose a novel approach for matrix factorization which is based on Earth Mover's Distance (EMD) as a measure of reconstruction error. Differently from previous works on EMD matrix decomposition, we consider a semi-supervised learning setting and we also propose to learn the ground distance parameters. While few previous works have addressed the problem of ground distance computation, these methods do not learn simultaneously the optimal metric and the reconstruction matrices. We demonstrate the effectiveness of the proposed approach both on synthetic data experiments and on a real world scenario, i.e. addressing the problem of complex video scene analysis in the context of video surveillance applications. Our experiments show that our method allows not only to achieve state-of-the-art performance on video segmentation, but also to learn the relationship among elementary activities which characterize the high level events in the video scene. Gloria Zen, Elisa Ricci 0001, Nicu Sebe |
ICPR | 2 |
| 2014 | Personalizing vision-based gestural interfaces for HRI with UAVs: a transfer learning approachabstractFollowing recent works on HRI for UAVs, we present a gesture recognition system which operates on the video stream recorded from a passive monocular camera installed on a quadcopter. While many challenges must be addressed for building a real-time vision-based gestural interface, in this paper we specifically focus on the problem of user personalization. Different users tend to perform the same gesture with different styles and speed. Thus, a system trained on visual sequences depicting some users may work poorly when data from other people are available. On the other hand, collecting and annotating many user-specific data is time consuming. To avoid these issues, in this paper we propose a personalized gestural interface. We introduce a novel transfer learning algorithm which, exploiting both data downloaded from the web and gestures collected from other users, permits to learn a set of person-specific classifiers. We integrate the proposed gesture recognition module into a HRI system with a flying quadrotor robot. In our system first the UAV localizes a person and individuates her identity. Then, when a user performs a specific gesture, the system recognizes it adopting the associated user-specific classifier and the quadcopter executes the corresponding task. Our experimental evaluation demonstrates that the proposed personalized gesture recognition solution is advantageous with respect to generic ones. Gabriele Costante, Enrico Bellocchio, Paolo Valigi, Elisa Ricci 0001 |
IROS | 4 |
| 2014 | We are not All Equal: Personalizing Models for Facial Expression Analysis with Transductive Parameter TransferabstractPrevious works on facial expression analysis have shown that person specific models are advantageous with respect to generic ones for recognizing facial expressions of new users added to the gallery set. This finding is not surprising, due to the often significant inter-individual variability: different persons have different morphological aspects and express their emotions in different ways. However, acquiring person-specific labeled data for learning models is a very time consuming process. In this work we propose a new transfer learning method to compute personalized models without labeled target data Our approach is based on learning multiple person-specific classifiers for a set of source subjects and then directly transfer knowledge about the parameters of these classifiers to the target individual. The transfer process is obtained by learning a regression function which maps the data distribution associated to each source subject to the corresponding classifier's parameters. We tested our approach on two different application domains, Action Units (AUs) detection and spontaneous pain recognition, using publicly available datasets and showing its advantages with respect to the state-of-the-art both in term of accuracy and computational cost. Enver Sangineto, Gloria Zen, Elisa Ricci 0001, Nicu Sebe |
ACM Multimedia | 3 |
| 2014 | Exploring Transfer Learning Approaches for Head Pose Classification from Multi-view Surveillance Images
Anoop Kolar Rajagopal, Subramanian Ramanathan, Elisa Ricci 0001, Radu L. Vieriu, Oswald Lanz, Kalpathi Ramakrishnan, Nicu Sebe |
Int. J. Comput. Vis. | 3 |
| 2014 | Multitask Linear Discriminant Analysis for View Invariant Action RecognitionabstractRobust action recognition under viewpoint changes has received considerable attention recently. To this end, self-similarity matrices (SSMs) have been found to be effective view-invariant action descriptors. To enhance the performance of SSM-based methods, we propose multitask linear discriminant analysis (LDA), a novel multitask learning framework for multiview action recognition that allows for the sharing of discriminative SSM features among different views (i.e., tasks). Inspired by the mathematical connection between multivariate linear regression and LDA, we model multitask multiclass LDA as a single optimization problem by choosing an appropriate class indicator matrix. In particular, we propose two variants of graph-guided multitask LDA: 1) where the graph weights specifying view dependencies are fixed a priori and 2) where graph weights are flexibly learnt from the training data. We evaluate the proposed methods extensively on multiview RGB and RGBD video data sets, and experimental results confirm that the proposed approaches compare favorably with the state-of-the-art. Yan Yan 0002, Elisa Ricci 0001, Subramanian Ramanathan, Gaowen Liu, Nicu Sebe |
IEEE Trans. Image Process. | 2 |
| 2013 | No Matter Where You Are: Flexible Graph-Guided Multi-task Learning for Multi-view Head Pose Classification under Target MotionabstractWe propose a novel Multi-Task Learning framework (FEGA-MTL) for classifying the head pose of a person who moves freely in an environment monitored by multiple, large field-of-view surveillance cameras. As the target (person) moves, distortions in facial appearance owing to camera perspective and scale severely impede performance of traditional head pose classification methods. FEGA-MTL operates on a dense uniform spatial grid and learns appearance relationships across partitions as well as partition-specific appearance variations for a given head pose to build region-specific classifiers. Guided by two graphs which a-priori model appearance similarity among (i) grid partitions based on camera geometry and (ii) head pose classes, the learner efficiently clusters appearance wise related grid partitions to derive the optimal partitioning. For pose classification, upon determining the target's position using a person tracker, the appropriate region specific classifier is invoked. Experiments confirm that FEGA-MTL achieves state-of-the-art classification with few training data. Yan Yan 0002, Elisa Ricci 0001, Subramanian Ramanathan, Oswald Lanz, Nicu Sebe |
ICCV | 2 |
| 2013 | Multi-task linear discriminant analysis for multi-view action recognitionabstractAction recognition is a central problem in many practical applications, such as video annotation, video surveillance and human-computer interaction. Most action recognition approaches are currently based on localized spatio-temporal features that can vary significantly when the viewpoint changes. Therefore, the performance rapidly drops when training and test data correspond to different cameras/viewpoints. Recently, Self-Similarity Matrix (SSM) features have been introduced to circumvent this problem. To improve the performance of current SSM-based methods, in this paper we propose a multi-task learning framework for multi-view action recognition where discriminative SSM features are shared among different views. Inspired by the mathematical connection between multivariate linear regression and Linear Discriminant Analysis (LDA), we propose a novel learning algorithm, where a single optimization framework is defined for multi-task multi-class LDA by choosing an appropriate class indicator matrix. Experimental results on the popular IXMAS dataset demonstrate that our approach achieves accurate performance and compares favorably with state-of-the-art methods. Yan Yan 0002, Gaowen Liu, Elisa Ricci 0001, Nicu Sebe |
ICIP | 3 |
| 2013 | A transfer learning approach for multi-cue semantic place recognitionabstractAs researchers are striving for developing robotic systems able to move into the `the wild', the interest towards novel learning paradigms for domain adaptation has increased. In the specific application of semantic place recognition from cameras, supervised learning algorithms are typically adopted. However, once learning has been performed, if the robot is moved to another location, the acquired knowledge may be not useful, as the novel scenario can be very different from the old one. The obvious solution would be to retrain the model updating the robot internal representation of the environment. Unfortunately this procedure involves a very time consuming data-labeling effort at the human side. To avoid these issues, in this paper we propose a novel transfer learning approach for place categorization from visual cues. With our method the robot is able to decide automatically if and how much its internal knowledge is useful in the novel scenario. Differently from previous approaches, we consider the situation where the old and the novel scenario may differ significantly (not only the visual room appearance changes but also different room categories are present). Importantly, our approach does not require labeling from a human operator. We also propose a strategy for improving the performance of the proposed method by fusing two complementary visual cues. Our extensive experimental evaluation demonstrates the advantages of our approach on several sequences from publicly available datasets. Gabriele Costante, Thomas A. Ciarfuglia, Paolo Valigi, Elisa Ricci 0001 |
IROS | 4 |
| 2013 | Automated defect detection in uniform and structured fabrics using Gabor filters and PCA
Lucia Bissi, Giuseppe Baruffa, Pisana Placidi, Elisa Ricci 0001, Andrea Scorzoni, Paolo Valigi |
J. Vis. Commun. Image Represent. | 4 |
| 2013 | A Prototype Learning Framework Using EMD: Application to Complex Scenes AnalysisabstractIn the last decades, many efforts have been devoted to develop methods for automatic scene understanding in the context of video surveillance applications. This paper presents a novel nonobject centric approach for complex scene analysis. Similarly to previous methods, we use low-level cues to individuate atomic activities and create clip histograms. Differently from recent works, the task of discovering high-level activity patterns is formulated as a convex prototype learning problem. This problem results in a simple linear program that can be solved efficiently with standard solvers. The main advantage of our approach is that, using as the objective function the Earth Mover's Distance (EMD), the similarity among elementary activities is taken into account in the learning phase. To improve scalability we also consider some variants of EMD adopting L1 as ground distance for 1D and 2D, linear and circular histograms. In these cases, only the similarity between neighboring atomic activities, corresponding to adjacent histogram bins, is taken into account. Therefore, we also propose an automatic strategy for sorting atomic activities. Experimental results on publicly available datasets show that our method compares favorably with state-of-the-art approaches, often outperforming them. Elisa Ricci 0001, Gloria Zen, Nicu Sebe, Stefano Messelodi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | An Adaptation Framework for Head-Pose Classification in Dynamic Multi-view Scenarios
Anoop Kolar Rajagopal, Subramanian Ramanathan, Radu L. Vieriu, Elisa Ricci 0001, Oswald Lanz, Kalpathi Ramakrishnan, Nicu Sebe |
ACCV (2) | 4 |
| 2012 | Exploiting Sparse Representations for Robust Analysis of Noisy Complex Video Scenes
Gloria Zen, Elisa Ricci 0001, Nicu Sebe |
ECCV (6) | 2 |
| 2012 | A Multiple Kernel Learning framework for detecting altered fingerprints
Michela Tiribuzi, Marco Pastorelli, Paolo Valigi, Elisa Ricci 0001 |
ICPR | 4 |
| 2012 | Enhanced semantic descriptors for functional scene categorization
Gloria Zen, Negar Rostamzadeh, Jacopo Staiano, Elisa Ricci 0001, Nicu Sebe |
ICPR | 4 |
| 2012 | A discriminative approach for appearance based loop closingabstractThe place recognition module is a fundamental component in SLAM systems, as incorrect loop closures may result in severe errors in trajectory estimation. In the case of appearance-based methods the bag-of-words approach is typically employed for recognizing locations. This paper introduces a novel algorithm for improving loop closures detection performance by adopting a set of visual words weights, learned offline accordingly to a discriminative criterion. The proposed weights learning approach, based on the large margin paradigm, can be used for generic similarity functions and relies on an efficient online leaning algorithm in the training phase. As the computed weights are usually very sparse, a gain in terms of computational cost at recognition time is also obtained. Our experiments, conducted on publicly available datasets, demonstrate that the discriminative weights lead to loop closures detection results that are more accurate than the traditional bag-of-words method and that our place recognition approach is competitive with state-of-the-art methods. Thomas A. Ciarfuglia, Gabriele Costante, Paolo Valigi, Elisa Ricci 0001 |
IROS | 4 |
| 2011 | Earth mover's prototypes: A convex learning approach for discovering activity patterns in dynamic scenesabstractWe present a novel approach for automatically discovering spatio-temporal patterns in complex dynamic scenes. Similarly to recent non-object centric methods, we use low level visual cues to detect atomic activities and then construct clip histograms. Differently from previous works, we formulate the task of discovering high level activity patterns as a prototype learning problem where the correlation among atomic activities is explicitly taken into account when grouping clip histograms. Interestingly at the core of our approach there is a convex optimization problem which allows us to efficiently extract patterns at multiple levels of detail. The effectiveness of our method is demonstrated on publicly available datasets. Gloria Zen, Elisa Ricci 0001 |
CVPR | 2 |
| 2010 | Learning Pedestrian Trajectories with KernelsabstractWe present a novel method for learning pedestrian trajectories which is able to describe complex motion patterns such as multiple crossing paths. This approach adopts Kernel Canonical Correlation Analysis (KCCA) to build a mapping between the physical location space and the trajectory patterns space. To model crossing paths we rely on a clustering algorithm based on Kernel K-means with a Dynamic Time Warping (DTW) kernel. We demonstrate the effectiveness of our method incorporating the learned motion model into a multi-person tracking algorithm and testing it on several video surveillance sequences. Elisa Ricci 0001, Francesco Tobia, Gloria Zen |
ICPR | 1 |
| 2010 | Tracking Multiple People with Illumination MapsabstractWe address the problem of multiple people tracking under non-homogenous and time-varying illumination conditions. We propose a unified framework for jointly estimating the position of the targets and their illumination conditions. For each target multiple templates are considered to model appearance variations due to lighting changes. The template choice is driven by an illumination map which describes the light conditions in different areas of the scene. This map is computed with a novel algorithm for efficient inference in a hierarchical Markov Random Field (MRF) and is updated online to adapt to slow lighting changes. Experimental results demonstrate the effectiveness of our approach. Gloria Zen, Oswald Lanz, Stefano Messelodi, Elisa Ricci 0001 |
ICPR | 4 |
| 2009 | Dynamic Partitioned Sampling For Tracking With Discriminative FeaturesabstractWe present a multi-cue fusion method for tracking with particle filters which relies on a novel hierarchical sampling strategy. Similarly to previous works, it tackles the problem of tracking in a relatively high-dimensional state space by dividing such a space into partitions, each one corresponding to a single cue, and sampling from them in a hierarchical manner. However, unlike other approaches, the order of partitions is not fixed a priori but changes dynamically depending on the reliability of each cue, i.e. more reliable cues are sampled first. We call this approach Dynamic Partitioned Sampling (DPS). The reliability of each cue is measured in terms of its ability to discriminate the object with respect to the background, where the background is not described by a fixed model or by random patches but is represented by a set of informative "background particles" which are tracked in order to be as similar as possible to the object. The effectiveness of this general framework is demonstrated on the specific problem of head tracking with three different cues: colour, edge and contours. Experimental results prove the robustness of our algorithm in several challenging video sequences. Stefan Duffner, Jean-Marc Odobez, Elisa Ricci 0001 |
BMVC | 3 |
| 2009 | Learning large margin likelihoods for realtime head pose tracking
Elisa Ricci 0001, Jean-Marc Odobez |
ICIP | 1 |
| 2008 | Recurrent Correlation Associative Memories: A Feature Space PerspectiveabstractIn this paper, we analyze a model of recurrent kernel associative memory (RKAM) recently proposed by Garcia and Moreno. We show that this model consists in a kernelization of the recurrent correlation associative memory (RCAM) of Chiueh and Goodman. In particular, using an exponential kernel, we obtain a generalization of the well-known exponential correlation associative memory (ECAM), while using a polynomial kernel, we obtain a generalization of higher order Hopfield networks with Hebbian weights. We show that the RKAM can outperform the aforementioned associative memory models, becoming equivalent to them when a dominance condition is fulfilled by the kernel matrix. To ascertain the dominance condition, we propose a statistical measure which can be easily computed from the probability distribution of the interpattern Hamming distance or directly estimated from the memory vectors. The RKAM can be used below saturation to realize associative memories with reduced dynamic range with respect to the ECAM and with reduced number of synaptic coefficients with respect to higher order Hopfield networks. Renzo Perfetti, Elisa Ricci 0001 |
IEEE Trans. Neural Networks | 2 |
| 2007 | Discriminative Sequence Labeling by Z-Score Optimization
Elisa Ricci 0001, Tijl De Bie, Nello Cristianini |
ECML | 1 |
| 2007 | Learning to Align: A Statistical Approach
Elisa Ricci 0001, Tijl De Bie, Nello Cristianini |
IDA | 1 |
| 2007 | Retinal Blood Vessel Segmentation Using Line Operators and Support Vector ClassificationabstractIn the framework of computer-aided diagnosis of eye diseases, retinal vessel segmentation based on line operators is proposed. A line detector, previously used in mammography, is applied to the green channel of the retinal image. It is based on the evaluation of the average grey level along lines of fixed length passing through the target pixel at different orientations. Two segmentation methods are considered. The first uses the basic line detector whose response is thresholded to obtain unsupervised pixel classification. As a further development, we employ two orthogonal line detectors along with the grey level of the target pixel to construct a feature vector for supervised classification using a support vector machine. The effectiveness of both methods is demonstrated through receiver operating characteristic analysis on two publicly available databases of color fundus images. Elisa Ricci 0001, Renzo Perfetti |
IEEE Trans. Medical Imaging | 1 |
| 2006 | Reduced complexity RBF classifiers with support vector centres and dynamic decay adjustment
Renzo Perfetti, Elisa Ricci 0001 |
Neurocomputing | 2 |
| 2006 | Improved pruning strategy for radial basis function networks with dynamic decay adjustment
Elisa Ricci 0001, Renzo Perfetti |
Neurocomputing | 1 |
| 2006 | SVM-based CDMA receiver with incremental active learning
Elisa Ricci 0001, Luca Rugini, Renzo Perfetti |
Neurocomputing | 1 |
| 2006 | Associative Memory Design Using Support Vector MachinesabstractThe relation existing between support vector machines (SVMs) and recurrent associative memories is investigated. The design of associative memories based on the generalized brain-state-in-a-box (GBSB) neural model is formulated as a set of independent classification tasks which can be efficiently solved by standard software packages for SVM learning. Some properties of the networks designed in this way are evidenced, like the fact that surprisingly they follow a generalized Hebb's law. The performance of the SVM approach is compared to existing methods with nonsymmetric connections, by some design examples. Daniele Casali, Giovanni Costantini, Renzo Perfetti, Elisa Ricci 0001 |
IEEE Trans. Neural Networks | 4 |
| 2006 | Analog neural network for support vector machine learningabstractAn analog neural network for support vector machine learning is proposed, based on a partially dual formulation of the quadratic programming problem. It results in a simpler circuit implementation with respect to existing neural solutions for the same application. The effectiveness of the proposed network is shown through some computer simulations concerning benchmark problems. Renzo Perfetti, Elisa Ricci 0001 |
IEEE Trans. Neural Networks | 2 |