EDBT 2026 Demo / reviewers in the wild / expert
Tiberio Uricchio
dblp:74/11533
· DBLP profile ↗
27ranked-venue papers
2as first author
8since 2021 · last 2024
0000-0003-1025-4541ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 1 first-author · 7 since 2021Computer networks · 5 · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Composed Image Retrieval using Contrastive Learning and Task-oriented CLIP-based FeaturesabstractGiven a query composed of a reference image and a relative caption, the Composed Image Retrieval goal is to retrieve images visually similar to the reference one that integrates the modifications expressed by the caption. Given that recent research has demonstrated the efficacy of large-scale vision and language pre-trained (VLP) models in various tasks, we rely on features from the OpenAI CLIP model to tackle the considered task. We initially perform a task-oriented fine-tuning of both CLIP encoders using the element-wise sum of visual and textual features. Then, in the second stage, we train a Combiner network that learns to combine the image-text features integrating the bimodal information and providing combined features used to perform the retrieval. We use contrastive learning in both stages of training. Starting from the bare CLIP features as a baseline, experimental results show that the task-oriented fine-tuning and the carefully crafted Combiner network are highly effective and outperform more complex state-of-the-art approaches on FashionIQ and CIRR, two popular and challenging datasets for composed image retrieval. Code and pre-trained models are available at https://github.com/ABaldrati/CLIP4Cir . Alberto Baldrati, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Meta-learning Advisor Networks for Long-tail and Noisy Labels in Social Image ClassificationabstractDeep neural networks (DNNs)for social image classification are prone to performance reduction and overfitting when trained on datasets plagued by noisy or imbalanced labels. Weight loss methods tend to ignore the influence of noisy or frequent category examples during the training, resulting in a reduction of final accuracy and, in the presence of extreme noise, even a failure of the learning process. A new advisor network is introduced to address both imbalance and noise problems, and is able to pilot learning of a main network by adjusting the visual features and the gradient with a meta-learning strategy. In a curriculum learning fashion, the impact of redundant data is reduced while recognizable noisy label images are downplayed or redirected.Meta Feature Re-Weighting (MFRW)andMeta Equalization Softmax (MES)methods are introduced to let the main network focus only on the information in an image deemed relevant by the advisor network and to adjust the training gradient to reduce the adverse effects of frequent or noisy categories. The proposed method is first tested on synthetic versions of CIFAR10 and CIFAR100, and then on the more realistic ImageNet-LT, Places-LT, and Clothing1M datasets, reporting state-of-the-art results. Simone Ricci, Tiberio Uricchio, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Effective conditioned and composed image retrieval combining CLIP-based featuresabstractConditioned and composed image retrieval extend CBIR systems by combining a query image with an additional text that expresses the intent of the user, describing additional requests w.r.t. the visual content of the query image. This type of search is interesting for e-commerce applications, e.g. to develop interactive multimodal searches and chat-bots. In this demo, we present an interactive system based on a combiner network, trained using contrastive learning, that combines visual and textual features obtained from the OpenAI CLIP network to address conditioned CBIR. The system can be used to improve e-shop search engines. For example, considering the fashion domain it lets users search for dresses, shirts and toptees using a candidate start image and expressing some visual differences w.r.t. its visual con-tent, e.g. asking to change color, pattern or shape. The pro-posed network obtains state-of-the-art performance on the FashionIQ dataset and on the more recent CIRR dataset, showing its applicability to the fashion domain for conditioned retrieval, and to more generic content considering the more general task of composed image retrieval. Alberto Baldrati, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
CVPR | 3 |
| 2022 | Memory NetworksabstractMemory Networks are models equipped with a storage component where information can generally be written and successively retrieved for any purpose. Simple forms of memory networks like the popular recurrent neural networks (RNN), LSTMs or GRUs, have limited storage capabilities and for specific tasks. In contrast, recent works, starting from Memory Augmented Neural Networks, overcome storage and computational limitations with the addition of a controller network with an external element-wise addressable memory. This tutorial aims at providing an overview of such memory-based techniques and their applications in multimedia. It will cover an explanation of the basic concepts behind recurrent neural networks and will then delve into the advanced details of memory augmented neural networks, their structure and how such models can be trained. We target a broad audience, from beginners to experienced researchers, offering an in-depth introduction to an important crop of literature which is starting to gain interest in the multimedia, computer vision and natural language processing communities. Federico Becattini, Tiberio Uricchio |
ACM Multimedia | 2 |
| 2022 | Effective triplet mining improves training of multi-scale pooled CNN for image retrieval
Federico Vaccaro, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
Mach. Vis. Appl. | 3 |
| 2021 | Fast Video Visual Quality and Resolution Improvement using SR-UNetabstractIn this paper, we address the problem of real-time video quality enhancement, considering both frame super-resolution and compression artifact-removal. The first operation increases the sampling resolution of video frames, the second removes visual artifacts such as blurriness, noise, aliasing, or blockiness introduced by lossy compression techniques, such as JPEG encoding for single-images, or H.264/H.265 for video data. We propose to use SR-UNet, a novel network architecture based on UNet, that has been specialized for fast visual quality improvement (i.e. capable of operating in less than 40ms, to be able to operate on videos at 25FPS). We show how this network can be used in a streaming context where the content is generated live, e.g. in video calls, and how it can be optimized when video to be streamed are prepared in advance. The network can be used as a final post processing, to optimize the visual appearance of a frame before showing it to the end-user in a video player. Thus, it can be applied without any change to existing video coding and transmission pipelines. Experiments carried on standard video datasets, also considering the H.265 compression, show that the proposed approach is able to either improve visual quality metrics given a fixed bandwidth budget, or video distortion given a fixed quality goal. Federico Vaccaro, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
ACM Multimedia | 3 |
| 2021 | Conditioned Image Retrieval for Fashion using Contrastive Learning and CLIP-based FeaturesabstractBuilding on the recent advances in multimodal zero-shot representation learning, in this paper we explore the use of features obtained from the recent CLIP model to perform conditioned image retrieval. Starting from a reference image and an additive textual description of what the user wants with respect to the reference image, we learn a Combiner network that is able to understand the image content, integrate the textual description and provide combined feature used to perform the conditioned image retrieval. Starting from the bare CLIP features and a simple baseline, we show that a carefully crafted Combiner network, based on such multimodal features, is extremely effective and outperforms more complex state of the art approaches on the popular FashionIQ dataset. Alberto Baldrati, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
MMAsia | 3 |
| 2021 | Am I Done? Predicting Action Progress in VideosabstractIn this article, we deal with the problem of predicting action progress in videos. We argue that this is an extremely important task, since it can be valuable for a wide range of interaction applications. To this end, we introduce a novel approach, named ProgressNet, capable of predicting when an action takes place in a video, where it is located within the frames, and how far it has progressed during its execution. To provide a general definition of action progress, we ground our work in the linguistics literature, borrowing terms and concepts to understand which actions can be the subject of progress estimation. As a result, we define a categorization of actions and their phases. Motivated by the recent success obtained from the interaction of Convolutional and Recurrent Neural Networks, our model is based on a combination of the Faster R-CNN framework, to make framewise predictions, and LSTM networks, to estimate action progress through time. After introducing two evaluation protocols for the task at hand, we demonstrate the capability of our model to effectively predict action progress on the UCF-101 and J-HMDB datasets. Federico Becattini, Tiberio Uricchio, Lorenzo Seidenari, Lamberto Ballan, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Learning Group Activities from Skeletons without Individual Action LabelsabstractTo understand human behavior we must not just recognize individual actions but model possibly complex group activity and interactions. Hierarchical models obtain the best results in group activity recognition but require fine grained individual action annotations at the actor level. In this paper we show that using only skeletal data we can train a state-of-the art end-to-end system using only group activity labels at the sequence level. Our experiments show that models trained without individual action supervision perform poorly. On the other hand we show that pseudo-labels can be computed from any pre-trained feature extractor with comparable final performance. Finally our carefully designed lean pose only architecture shows highly competitive results versus more complex multimodal approaches even in the self-supervised variant. Fabio Zappardino, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo |
ICPR | 2 |
| 2020 | Image Retrieval using Multi-scale CNN Features PoolingabstractIn this paper, we address the problem of image retrieval by learning images representation based on the activations of a Convolutional Neural Network. We present an end-to-end trainable network architecture that exploits a novel multi-scale local pooling based on NetVLAD and a triplet mining procedure based on samples difficulty to obtain an effective image representation. Extensive experiments show that our approach is able to reach state-of-the-art results on three standard datasets. Federico Vaccaro, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
ICMR | 3 |
| 2020 | Increasing Video Perceptual Quality with GANs and Semantic CodingabstractWe have seen a rise in video based user communication in the last year, unfortunately fueled by the spread of COVID-19 disease. Efficient low-latency delay of transmission of video is a challenging problem which must also deal with the segmented nature of network infrastructure not always allowing a high throughput. Lossy video compression is a basic requirement to enable such technology widely. While this may compromise the quality of the streamed video there are recent deep learning based solutions to restore quality of a lossy compressed video. Leonardo Galteri, Marco Bertini 0001, Lorenzo Seidenari, Tiberio Uricchio, Alberto Del Bimbo |
ACM Multimedia | 4 |
| 2020 | Learning Visual Elements of Images for Discovery of Brand PostsabstractOnline Social Network Sites have become a primary platform for brands and organizations to engage their audience by sharing image and video posts on their timelines. Different from traditional advertising, these posts are not restricted to the products or logo but include visual elements that express more in general the values and attributes of the brand, called brand associations. Since marketers are increasingly spending time in discovering and re-posting user generated posts that reflect the brand attributes, there is an increasing demand for such discovery systems. The goal of these systems is to assist brand experts in filtering through online collections of new user media to discover actionable posts, which match the brand value and have the potential to engage the consumers. Driven by this real-life application, we define and formulate a new task of content discovery for brands and propose a framework that learns to rank posts for brands from their historical timeline. We design a Personalized Content Discovery (PCD) framework to address the three challenges of high inter-brand similarity, sparsity of brand--post interactions, and diversification of timeline. To learn fine-grained brand representation and to generate explanations for the ranking, we automatically learn visual elements of posts from the timeline of brands and from a set of brand attributes in the domain of marketing. To test our framework we use two large-scale Instagram datasets that contain a total of more than 1.5 million image and video posts from the historical timeline of hundreds of brands from multiple verticals such as food and fashion. Extensive experiments indicate that our model can effectively learn fine-grained brand representations and outperform the closest state-of-the-art solutions. Francesco Gelli, Tiberio Uricchio, Xiangnan He 0001, Alberto Del Bimbo, Tat-Seng Chua |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Fast Video Quality Enhancement using GANsabstractVideo compression algorithms result in a reduction of image quality, because of their lossy approach to reduce the required bandwidth. This affects commercial streaming services such as Netflix, or Amazon Prime Video, but affects also video conferencing and video surveillance systems. In all these cases it is possible to improve the video quality, both for human view and for automatic video analysis, without changing the compression pipeline, through a post-processing that eliminates the visual artifacts created by the compression algorithms. Generative Adversarial Networks have obtained extremely high quality results in image enhancement tasks; however, to obtain such results large generators are usually employed, resulting in high computational costs and processing time. In this work we present an architecture that can be used to reduce the computational cost and that has been implemented on mobile devices. A possible application is to improve video conferencing, or live streaming. In these cases there is no original uncompressed video stream available. Therefore, we report results using no-reference video quality metric showing high naturalness and quality even for efficient networks. Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
ACM Multimedia | 4 |
| 2019 | Learning Subjective Attributes of Images from Auxiliary SourcesabstractRecent years have seen unprecedented research on using artificial intelligence to understand the subjective attributes of images and videos. These attributes are not objective properties of the content but are highly dependent on the perception of the viewers. Subjective attributes are extremely valuable in many applications where images are tailored to the needs of a large group, which consists of many individuals with inherently different ideas and preferences. For instance, marketing experts choose images to establish specific associations in the consumers' minds, while psychologists look for pictures with adequate emotions for therapy. Unfortunately, most of the existing frameworks either focus on objective attributes or rely on large scale datasets of annotated images, making them costly and unable to clearly measure multiple interpretations of a single input. Meanwhile, we can see that users or organizations often interact with images in a multitude of real-life applications, such as the sharing of photographs by brands on social media or the re-posting of image microblogs by users. We argue that these aggregated interactions can serve as auxiliary information to infer image interpretations. To this end, we propose a probabilistic learning framework capable of transferring such subjective information to the image-level labels based on a known aggregated distribution. We use our framework to rank images by subjective attributes from the domain knowledge of social media marketing and personality psychology. Extensive studies and visualizations show that using auxiliary information is a viable line of research for the multimedia community to perform subjective attributes prediction. Francesco Gelli, Tiberio Uricchio, Xiangnan He 0001, Alberto Del Bimbo, Tat-Seng Chua |
ACM Multimedia | 2 |
| 2018 | Beyond the Product: Discovering Image Posts for Brands in Social MediaabstractBrands and organizations are using social networks such as Instagram to share image or video posts regularly, in order to engage and maximize their presence to the users. Differently from the traditional advertising paradigm, these posts feature not only specific products, but also the value and philosophy of the brand, known as brand associations in marketing literature. In fact, marketers are spending considerable resources to generate their content in-house, and increasingly often, to discover and repost the content generated by users. However, to choose the right posts for a brand in social media remains an open problem. Driven by this real-life application, we define the new task of content discovery for brands, which aims to discover posts that match the marketing value and brand associations of a target brand. We identify two main challenges in this new task: high inter-brand similarity and brand-post sparsity; and propose a tailored content-based learning-to-rank system to discover content for a target brand. Specifically, our method learns fine-grained brand representation via explicit modeling of brand associations, which can be interpreted as visual words shared among brands. We collected a new large-scale Instagram dataset, consisting of more than 1.1 million image and video posts from the history of 927 brands of fourteen verticals such as food and fashion. Extensive experiments indicate that our model can effectively learn fine-grained brand representations and outperform the closest state-of-the-art solutions. Francesco Gelli, Tiberio Uricchio, Xiangnan He 0001, Alberto Del Bimbo, Tat-Seng Chua |
ACM Multimedia | 2 |
| 2017 | Deep Sentiment Features of Context and Faces for Affective Video AnalysisabstractGiven the huge quantity of hours of video available on video sharing platforms such as YouTube, Vimeo, etc. development of automatic tools that help users find videos that fit their interests has attracted the attention of both scientific and industrial communities. So far the majority of the works have addressed semantic analysis, to identify objects, scenes and events depicted in videos, but more recently affective analysis of videos has started to gain more attention. In this work we investigate the use of sentiment driven features to classify the induced sentiment of a video, i.e. the sentiment reaction of the user. Instead of using standard computer vision features such as CNN features or SIFT features trained to recognize objects and scenes, we exploit sentiment related features such as the ones provided by Deep-SentiBank, and features extracted from models that exploit deep networks trained on face expressions. We experiment on two recently introduced datasets: LIRIS-ACCEDE and MEDIAEVAL-2015, that provide sentiment annotations of a large set of short videos. We show that our approach not only outperforms the current state-of-the-art in terms of valence and arousal classification accuracy, but it also uses a smaller number of features, requiring thus less video processing. Claudio Baecchi, Tiberio Uricchio, Marco Bertini 0001, Alberto Del Bimbo |
ICMR | 2 |
| 2017 | Outdoor Object Recognition for Smart Audio GuidesabstractWe present a smart audio guide that adapts itself to the environment the user is navigating into. The system builds automatically a point of interest database exploiting Wikipedia and Google APIs as source. We rely on a computer vision system, to overcome the likely sensor limitations, and determine with high accuracy if the user is facing a certain landmark or if he is not facing any. Thanks to this the guide presents audio description at the most appropriate moment without any user intervention, using text-to-speech augmenting the experience. Claudio Baecchi, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo |
ACM Multimedia | 2 |
| 2017 | Automatic image annotation via label transfer in the semantic space
Tiberio Uricchio, Lamberto Ballan, Lorenzo Seidenari, Alberto Del Bimbo |
Pattern Recognit. | 1 |
| 2017 | Deep Artwork Detection and Retrieval for Automatic Context-Aware Audio GuidesabstractIn this article, we address the problem of creating a smart audio guide that adapts to the actions and interests of museum visitors. As an autonomous agent, our guide perceives the context and is able to interact with users in an appropriate fashion. To do so, it understands what the visitor is looking at, if the visitor is moving inside the museum hall, or if he or she is talking with a friend. The guide performs automatic recognition of artworks, and it provides configurable interface features to improve the user experience and the fruition of multimedia materials through semi-automatic interaction. Our smart audio guide is backed by a computer vision system capable of working in real time on a mobile device, coupled with audio and motion sensors. We propose the use of a compact Convolutional Neural Network (CNN) that performs object classification and localization. Using the same CNN features computed for these tasks, we perform also robust artwork recognition. To improve the recognition accuracy, we perform additional video processing using shape-based filtering, artwork tracking, and temporal filtering. The system has been deployed on an NVIDIA Jetson TK1 and a NVIDIA Shield Tablet K1 and tested in a real-world environment (Bargello Museum of Florence). Lorenzo Seidenari, Claudio Baecchi, Tiberio Uricchio, Andrea Ferracani, Marco Bertini 0001, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2016 | Do Textual Descriptions Help Action Recognition?abstractWe present a novel method to improve action recognition by leveraging a set of captioned videos. By learning linear projections to map videos and text onto a common space, our approach shows that improved results on unseen videos can be obtained. We also propose a novel structure preserving loss that further ameliorates the quality of the projections. We tested our method on the challenging, realistic, Hollywood2 action recognition dataset where a considerable gain in performance is obtained. We show that the gain is proportional to the number of training samples used to learn the projections. Matteo Bruni, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo |
ACM Multimedia | 2 |
| 2016 | A multimodal feature learning approach for sentiment analysis of social network multimedia
Claudio Baecchi, Tiberio Uricchio, Marco Bertini 0001, Alberto Del Bimbo |
Multim. Tools Appl. | 2 |
| 2015 | Image Popularity Prediction in Social Media Using Sentiment and Context FeaturesabstractImages in social networks share different destinies: some are going to become popular while others are going to be completely unnoticed. In this paper we propose to use visual sentiment features together with three novel context features to predict a concise popularity score of social images. Experiments on large scale datasets show the benefits of proposed features on the performance of image popularity prediction. Exploiting state-of-the-art sentiment features, we report a qualitative analysis of which sentiments seem to be related to good or poor popularity. To the best of our knowledge, this is the first work understanding specific visual sentiments that positively or negatively influence the eventual popularity of images. Francesco Gelli, Tiberio Uricchio, Marco Bertini 0001, Alberto Del Bimbo, Shih-Fu Chang |
ACM Multimedia | 2 |
| 2015 | Image Tag Assignment, Refinement and RetrievalabstractThis tutorial focuses on challenges and solutions for content-based image annotation and retrieval in the context of online image sharing and tagging. We present a unified review on three closely linked problems, i.e., tag assignment, tag refinement, and tag-based image retrieval. We introduce a taxonomy to structure the growing literature, understand the ingredients of the main works, clarify their connections and difference, and recognize their merits and limitations. Moreover, we present an open-source testbed, with training sets of varying sizes and three test datasets, to evaluate methods of varied learning complexity. A selected set of eleven representative works have been implemented and evaluated. During the tutorial we provide a practice session for hands on experience with the methods, software and datasets. For repeatable experiments all data and code are online at http://www.micc.unifi.it/tagsurvey Xirong Li 0001, Tiberio Uricchio, Lamberto Ballan, Marco Bertini 0001, Cees Snoek, Alberto Del Bimbo |
ACM Multimedia | 2 |
| 2015 | Data-driven approaches for social image and video tagging
Lamberto Ballan, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
Multim. Tools Appl. | 3 |
| 2014 | A Cross-media Model for Automatic Image AnnotationabstractAutomatic image annotation is still an important open problem in multimedia and computer vision. The success of media sharing websites has led to the availability of large collections of images tagged with human-provided labels. Many approaches previously proposed in the literature do not accurately capture the intricate dependencies between image content and annotations. We propose a learning procedure based on Kernel Canonical Correlation Analysis which finds a mapping between visual and textual words by projecting them into a latent meaning space. The learned mapping is then used to annotate new images using advanced nearest-neighbor voting methods. We evaluate our approach on three popular datasets, and show clear improvements over several approaches relying on more standard representations. Lamberto Ballan, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo |
ICMR | 2 |
| 2013 | An evaluation of nearest-neighbor methods for tag refinementabstractThe success of media sharing and social networks has led to the availability of extremely large quantities of images that are tagged by users. The need of methods to manage efficiently and effectively the combination of media and metadata poses significant challenges. In particular, automatic image annotation of social images has become an important research topic for the multimedia community. In this paper we propose and thoroughly evaluate the use of nearest-neighbor methods for tag refinement. Extensive and rigorous evaluation using two standard large-scale datasets shows that the performance of these methods is comparable with that of more complex and computationally intensive approaches and that, differently from these latter approaches, nearest-neighbor methods can be applied to `web-scale' data. Tiberio Uricchio, Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo |
ICME | 1 |
| 2012 | LIT: transcription, annotation, search and visualization tools for the Lexicon of the Italian Television
Thomas M. Alisi, Alberto Del Bimbo, Andrea Ferracani, Tiberio Uricchio, Ervin Hoxha, Besmir Bregasi |
Multim. Tools Appl. | 4 |