Marco Bertini 0001

dblp:70/1173-1 · DBLP profile ↗
← Back
10ranked-venue papers in the field
3as first author
3since 2021 · last 2025
0000-0002-1364-218XORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 6 (2 first)Other / Interdisciplinary · 3Database Systems & Data Management · 1 (1 first)
YearPublicationVenuePosition
2025 ComicsPAP: Understanding Comic Strips by Picking the Correct Panel
Emanuele Vivoli, Artemis Llabrés, Mohamed Ali Souibgui, Marco Bertini 0001, Ernest Valveny, Dimosthenis Karatzas
ICDAR (1)4
2021 Conditioned Image Retrieval for Fashion using Contrastive Learning and CLIP-based Features
abstract
Building on the recent advances in multimodal zero-shot representation learning, in this paper we explore the use of features obtained from the recent CLIP model to perform conditioned image retrieval. Starting from a reference image and an additive textual description of what the user wants with respect to the reference image, we learn a Combiner network that is able to understand the image content, integrate the textual description and provide combined feature used to perform the conditioned image retrieval. Starting from the bare CLIP features and a simple baseline, we show that a carefully crafted Combiner network, based on such multimodal features, is extremely effective and outperforms more complex state of the art approaches on the popular FashionIQ dataset.
Alberto Baldrati, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo
MMAsia2
2021 Language Based Image Quality Assessment
abstract
Evaluation of generative models, in the visual domain, is often performed providing anecdotal results to the reader. In the case of image enhancement, reference images are usually available. Nonetheless, using signal based metrics often leads to counterintuitive results: highly natural crisp images may obtain worse scores than blurry ones. On the other hand, blind reference image assessment may rank images reconstructed with GANs higher than the original undistorted images. To avoid time consuming human based image assessment, semantic computer vision tasks may be exploited instead [9, 25, 33]. In this paper we advocate the use of language generation tasks to evaluate the quality of restored images. We show experimentally that image captioning, used as a downstream task, may serve as a method to score image quality. Captioning scores are better aligned with human rankings with respect to signal based metrics or no-reference image quality metrics. We show insights on how the corruption, by artifacts, of local image structure may steer image captions in the wrong direction.
Lorenzo Seidenari, Leonardo Galteri, Pietro Bongini, Marco Bertini 0001, Alberto Del Bimbo
MMAsia4
2020 Image Retrieval using Multi-scale CNN Features Pooling
abstract
In this paper, we address the problem of image retrieval by learning images representation based on the activations of a Convolutional Neural Network. We present an end-to-end trainable network architecture that exploits a novel multi-scale local pooling based on NetVLAD and a triplet mining procedure based on samples difficulty to obtain an effective image representation. Extensive experiments show that our approach is able to reach state-of-the-art results on three standard datasets.
Federico Vaccaro, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo
ICMR2
2017 Deep Sentiment Features of Context and Faces for Affective Video Analysis
abstract
Given the huge quantity of hours of video available on video sharing platforms such as YouTube, Vimeo, etc. development of automatic tools that help users find videos that fit their interests has attracted the attention of both scientific and industrial communities. So far the majority of the works have addressed semantic analysis, to identify objects, scenes and events depicted in videos, but more recently affective analysis of videos has started to gain more attention. In this work we investigate the use of sentiment driven features to classify the induced sentiment of a video, i.e. the sentiment reaction of the user. Instead of using standard computer vision features such as CNN features or SIFT features trained to recognize objects and scenes, we exploit sentiment related features such as the ones provided by Deep-SentiBank, and features extracted from models that exploit deep networks trained on face expressions. We experiment on two recently introduced datasets: LIRIS-ACCEDE and MEDIAEVAL-2015, that provide sentiment annotations of a large set of short videos. We show that our approach not only outperforms the current state-of-the-art in terms of valence and arousal classification accuracy, but it also uses a smaller number of features, requiring thus less video processing.
Claudio Baecchi, Tiberio Uricchio, Marco Bertini 0001, Alberto Del Bimbo
ICMR3
2016 Item-Based Video Recommendation: An Hybrid Approach considering Human Factors
abstract
In this paper we propose a method for video recommendation in Social Networks based on crowdsourced and automatic video annotations of salient frames. We show how two human factors, users' self-expression in user profiles and perception of visual saliency in videos, can be exploited in order to stimulate annotations and to obtain an efficient representation of video content features. Results are assessed through experiments conducted on a prototype of social network for video sharing. Several baseline approaches are evaluated and we show how the proposed method improves over them.
Andrea Ferracani, Daniele Pezzatini, Marco Bertini 0001, Alberto Del Bimbo
ICMR3
2016 Web Video Popularity Prediction using Sentiment and Content Visual Features
abstract
Hundreds of hours of videos are uploaded every minute on YouTube and other video sharing sites: some will be viewed by millions of people and other will go unnoticed by all but the uploader. In this paper we propose to use visual sentiment and content features to predict the popularity of web videos. The proposed approach outperforms current state-of-the-art methods on two publicly available datasets.
Giulia Fontanini, Marco Bertini 0001, Alberto Del Bimbo
ICMR2
2011 A flexible environment for multimedia management and publishing
abstract
In this paper, we describe the IM3I system, which provides a flexible approach to managing and publishing collections of images and videos. The system is based on web services that allow automatic and manual annotation, retrieval, browsing and authoring of multimedia. Results of user evaluations, performed by professional archivists and archive managers on a real-world system deployment have confirmed that the system is easy to be used and delivers a complete set of functionalities.
Marco Bertini 0001, Alberto Del Bimbo, George Ioannidis, Alexandru Stan, Emile Bijk
ICMR1
2006 Using Knowledge Representation Languages for Video Annotation and Retrieval
Marco Bertini 0001, Gianpaolo D'Amico, Alberto Del Bimbo, Carlo Torniai
FQAS1
2003 Annotation and Retrieval of Structured Video Documents
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati
ECIR1