VLDB 2026 Research / reviewers in the wild / expert
Marco Bertini 0001
dblp:70/1173-1
· DBLP profile ↗
118ranked-venue papers
31as first author
32since 2021 · last 2026
0000-0002-1364-218XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 101 · 25 first-author · 25 since 2021Artificial intelligence and machine learning · 25 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 10 · 3 first-author · 3 since 2021Computer networks · 6 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Diffusion Models with One-Step Distillation for Image and Video Super-Resolution
Alessio Bugetti, Leonardo Galteri, Marco Bertini 0001 |
MMSys | 3 |
| 2026 | DocWaveDiff: A Predict-and-Refine approach for Document Image Enhancement with Wavelet U-Nets and Diffusion modelsabstractOCR and document layout analysis algorithms are essential components of AI-based document systems, yet they are typically trained on clean, degradation-free images. When applied to degraded documents such as blurred scans or pages spoiled by undesired handwritten text their performance drops significantly. To address this issue, we propose DocWaveDiff, a novel document restoration method based on a predict-and-refine diffusion framework incorporating wavelet U-Nets. Given a degraded image patch and optionally its prior features, our Early Predictor generates an initial restoration, which is then refined by a Denoiser Refiner that estimates the residual image. The combination of these outputs yields the final restored result. We evaluate DocWaveDiff on multiple public benchmarks and demonstrate its strong performance across various document degradation scenarios, including deblurring and handwriting removal. Our results confirm that integrating wavelet transforms into the predict-and-refine framework enhances restoration quality and supports more robust document understanding systems. Matteo Marulli, Marco Bertini 0001 |
WACV | 2 |
| 2026 | Multimodal-Conditioned Latent Diffusion Models for Fashion Image EditingabstractFashion illustration is a crucial medium for designers to convey their creative vision and transform design concepts into tangible representations that showcase the interplay between clothing and the human body. In the context of fashion design, computer vision techniques have the potential to enhance and streamline the design process. Departing from prior research primarily focused on virtual try-on, this article tackles the task of multimodal-conditioned fashion image editing. Our approach aims to generate human-centric fashion images guided by multimodal prompts, including text, human body poses, garment sketches, and fabric textures. To address this problem, we propose extending latent diffusion models to incorporate these multiple modalities and modifying the structure of the denoising network, taking multimodal prompts as input. To condition the proposed architecture on fabric textures, we employ textual inversion techniques and let diverse cross-attention layers of the denoising network attend to textual and texture information, thus incorporating different granularity conditioning details. Given the lack of datasets for the task, we extend two existing fashion datasets, Dress Code and VITON-HD, with multimodal annotations. Experimental evaluations demonstrate the effectiveness of our proposed approach in terms of realism and coherence concerning the provided multimodal inputs. Alberto Baldrati, Davide Morelli, Marcella Cornia, Marco Bertini 0001, Rita Cucchiara |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Text-Oriented Image Query Representation for Zero-Shot Composed Image RetrievalabstractZero-Shot Composed Image Retrieval (ZS-CIR) is the task of retrieving a target image based on a query that combines a reference image with a textual description specifying desired modifications in a zero-shot setting. Existing ZS-CIR models typically fuse visual and textual modalities into a single query representation, but often struggle to capture the fine-grained distinctions essential for accurate retrieval. In this paper, we present TEOZCIR, a transformer-based model that introduces a balanced semantic fusion module and an enhancement mechanism to more effectively integrate multimodal information. The model is built around two core components: the Text-Aware Query Combiner (TAQC) and the Query Enhancer Network (QENet). These components operate in tandem: TAQC dynamically adjusts the semantic contributions of the visual context based on the input text, generating a balanced query representation. This representation is then further refined by QENet, which enhances the fused features to better align with the target image. Throughout the entire process, the model maintains a lightweight architecture with significantly fewer trainable parameters compared to conventional training-based methods. Experiments carried out on three benchmark datasets CIRR, Fashion IQ, and CIRCO to demonstrate that TEOZCIR significantly improves ZS-CIR performance, setting a new bench-mark for multimodal retrieval. Pavan K. Rachabathuni, Andrea Ciamarra, Roberto Caldelli, Marco Bertini 0001 |
CBMI | 4 |
| 2025 | ComicsPAP: Understanding Comic Strips by Picking the Correct Panel
Emanuele Vivoli, Artemis Llabrés, Mohamed Ali Souibgui, Marco Bertini 0001, Ernest Valveny, Dimosthenis Karatzas |
ICDAR (1) | 4 |
| 2025 | Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality InversionabstractPre-trained multi-modal Vision-Language Models like CLIP are widely used off-the-shelf for a variety of applications. In this paper, we show that the common practice of individually exploiting the text or image encoders of these powerful multi-modal models is highly suboptimal for intra-modal tasks like image-to-image retrieval. We argue that this is inherently due to the CLIP-style inter-modal contrastive loss that does not enforce any intra-modal constraints, leading to what we call intra-modal misalignment. To demonstrate this, we leverage two optimization-based modality inversion techniques that map representations from their input modality to the complementary one without any need for auxiliary data or additional trained adapters. We empirically show that, in the intra-modal tasks of image-to-image and text-to-text retrieval, approaching these tasks inter-modally significantly improves performance with respect to intra-modal baselines on more than fifteen datasets. Additionally, we demonstrate that approaching a native inter-modal task (e.g. zero-shot image classification) intra-modally decreases performance, further validating our findings. Finally, we show that incorporating an intra-modal term in the pre-training objective or narrowing the modality gap between the text and image feature embedding spaces helps reduce the intra-modal misalignment. The code is publicly available at: https://github.com/miccunifi/Cross-the-Gap. Marco Mistretta, Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini 0001, Andrew D. Bagdanov |
ICLR | 4 |
| 2025 | Real-time GenAI Solutions for Video Streaming in Low-bandwidth SettingsabstractSurveillance systems such as bodycams and drones often operate under bandwidth constraints that limit video quality and degrade both human monitoring and AI-based analytics. Traditional compression techniques introduce artifacts that obscure critical details, especially in high-motion scenarios. We present a generative AI-powered video compression framework developed by Small Pixels, a spin-off of the University of Florence, designed to deliver Full HD video at significantly reduced bitrates. The system combines edge-side preprocessing for compression resilience with real-time receiver-side super-resolution, enabling up to 50% bandwidth savings while preserving perceptual quality and detection accuracy. Objective evaluations on EgoSeg and VisDrone datasets show +6.2 VMAF improvement and stable YOLOv11 detection performance with 30% less bitrate. Live trials in Singapore, within the Singapore Hatch-X Global Innovation Program, validated real-time operation with minimal latency, demonstrating clearer faces and motion in challenging conditions. The solution integrates seamlessly into existing infrastructures without hardware upgrades, offering a practical path to reliable, high-quality video streaming in bandwidth-limited environments. Claudio Baecchi, Matteo Bruni, Fabio Clabot, Marco Bertini 0001 |
ACM Multimedia | 4 |
| 2025 | Navigating social contexts: A transformer approach to relationship recognition
Lorenzo Berlincioni, Luca Cultrera, Marco Bertini 0001, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 3 |
| 2025 | iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image RetrievalabstractGiven a query consisting of a reference image and a relative caption, Composed Image Retrieval (CIR) aims to retrieve target images visually similar to the reference one while incorporating the changes specified in the relative caption. The reliance of supervised methods on labor-intensive manually labeled datasets hinders their broad applicability to CIR. In this work, we introduce a new task, Zero-Shot CIR (ZS-CIR), that addresses CIR without the need for a labeled training dataset. We propose an approach, named iSEARLE (improved zero-Shot composEd imAge Retrieval with textuaL invErsion), that involves mapping the visual information of the reference image into a pseudo-word token in the CLIP token embedding space and combining it with the relative caption. To foster research on ZS-CIR, we present an open-domain benchmarking dataset named CIRCO (Composed Image Retrieval on Common Objects in context), the first CIR dataset where each query is labeled with multiple ground truths and a semantic categorization. The experimental results illustrate that iSEARLE obtains state-of-the-art performance on three different CIR datasets - FashionIQ, CIRR, and the proposed CIRCO - and two additional evaluation settings, namely domain conversion and object composition. Lorenzo Agnolucci, Alberto Baldrati, Alberto Del Bimbo, Marco Bertini 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Improving Zero-Shot Generalization of Learned Prompts via Unsupervised Knowledge Distillation
Marco Mistretta, Alberto Baldrati, Marco Bertini 0001, Andrew D. Bagdanov |
ECCV (84) | 3 |
| 2024 | Context-aware chatbot using MLLMs for Cultural HeritageabstractMulti-modal Large Language Models (MLLMs) are currently an extremely active research topic for the multimedia and computer vision communities, and show a significant impact in visual analysis and text generation tasks. MLLM's are well-versed in integrated understanding, analysis of complex data from cross modalities (i.e. text-image) and text generation with chat abilities. Almost all MLLM's, focus on alignment of image features to textual features for downstream text generation tasks includes detailed image description, visual question answering, stories and poems generation, phrase grounding, etc.. However, when focusing on visual question answering, questions that are highly relevant to the context of an image may not be answered correctly with the existing MLLM's, contrary to questions that are related to visual aspects. Moreover, generating meta data (context) for an image using present day MLLM's is hard task due to hallucinating characteristic of underlying Large Language Models (LLM's), and adequate contextual information cannot be directly derived from an image based perspective. Pavan Kartheek Rachabatuni, Filippo Principi, Paolo Mazzanti, Marco Bertini 0001 |
MMSys | 4 |
| 2024 | CoMix: A Comprehensive Benchmark for Multi-Task Comic UnderstandingabstractThe comic domain is rapidly advancing with the development of single-page analysis and synthesis models. However, evaluation metrics and datasets lag behind, often limited to small-scale or single-style test sets. We introduce a novel benchmark, CoMix, designed to evaluate the multi-task capabilities of models in comic analysis. Unlike existing benchmarks that focus on isolated tasks such as object detection or text recognition, CoMix addresses a broader range of tasks including object detection, speaker identification, character re-identification, reading order, and multi-modal reasoning tasks like character naming and dialogue generation. Our benchmark comprises three existing datasets with expanded annotations to support multi-task evaluation. To mitigate the over-representation of manga-style data, we have incorporated a new dataset of carefully selected American comic-style books, thereby enriching the diversity of comic styles. CoMix is designed to assess pre-trained models in zero-shot and limited fine-tuning settings, probing their transfer capabilities across different comic styles and tasks. The validation split of the benchmark is publicly available for research purposes, and an evaluation server for the held-out test split is also provided. Comparative results between human performance and state-of-the-art models reveal a significant performance gap, highlighting substantial opportunities for advancements in comic understanding. The dataset, baseline models, and code are accessible at https://github.com/emanuelevivoli/CoMix-dataset. This initiative sets a new standard for comprehensive comic analysis, providing the community with a common benchmark for evaluation on a large and varied set. Emanuele Vivoli, Marco Bertini 0001, Dimosthenis Karatzas |
NeurIPS | 2 |
| 2024 | ARNIQA: Learning Distortion Manifold for Image Quality AssessmentabstractNo-Reference Image Quality Assessment (NR-IQA) aims to develop methods to measure image quality in alignment with human perception without the need for a high-quality reference image. In this work, we propose a self-supervised approach named ARNIQA (leArning distoRtion maNifold for Image Quality Assessment) for modeling the image distortion manifold to obtain quality representations in an intrinsic manner. First, we introduce an image degradation model that randomly composes ordered sequences of consecutively applied distortions. In this way, we can synthetically degrade images with a large variety of degradation patterns. Second, we propose to train our model by maximizing the similarity between the representations of patches of different images distorted equally, despite varying content. Therefore, images degraded in the same manner correspond to neighboring positions within the distortion manifold. Finally, we map the image representations to the quality scores with a simple linear regressor, thus without fine-tuning the encoder weights. The experiments show that our approach achieves state-of-the-art performance on several datasets. In addition, ARNIQA demonstrates improved data efficiency, generalization capabilities, and robustness compared to competing methods. The code and the model are publicly available at https://github.com/miccunifi/ARNIQA. Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini 0001, Alberto Del Bimbo |
WACV | 3 |
| 2024 | Reference-based Restoration of Digitized Analog VideotapesabstractAnalog magnetic tapes have been the main video data storage device for several decades. Videos stored on analog videotapes exhibit unique degradation patterns caused by tape aging and reader device malfunctioning that are different from those observed in film and digital video restoration tasks. In this work, we present a reference-based approach for the resToration of digitized Analog videotaPEs (TAPE). We leverage CLIP for zero-shot artifact detection to identify the cleanest frames of each video through textual prompts describing different artifacts. Then, we select the clean frames most similar to the input ones and employ them as references. We design a transformer-based Swin-UNet network that exploits both neighboring and reference frames via our Multi-Reference Spatial Feature Fusion (MRSFF) blocks. MRSFF blocks rely on cross-attention and attention pooling to take advantage of the most useful parts of each reference frame. To address the absence of ground truth in real-world videos, we create a synthetic dataset of videos exhibiting artifacts that closely resemble those commonly found in analog videotapes. Both quantitative and qualitative experiments show the effectiveness of our approach compared to other state-of-the-art methods. The code, the model, and the synthetic dataset are publicly available at https://github.com/miccunifi/TAPE. Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini 0001, Alberto Del Bimbo |
WACV | 3 |
| 2024 | Special issue on content-based image retrieval
Gianluigi Ciocca, Raimondo Schettini, Simone Santini, Marco Bertini 0001 |
Multim. Tools Appl. | 4 |
| 2024 | Perceptual Quality Improvement in Videoconferencing Using Keyframes-Based GANabstractIn the latest years, videoconferencing has taken a fundamental role in interpersonal relations, both for personal and business purposes. Lossy video compression algorithms are the enabling technology for videoconferencing, as they reduce the bandwidth required for real-time video streaming. However, lossy video compression decreases the perceived visual quality. Thus, many techniques for reducing compression artifacts and improving video visual quality have been proposed in recent years. In this work, we propose a novel GAN-based method for compression artifacts reduction in videoconferencing. Given that, in this context, the speaker is typically in front of the camera and remains the same for the entire duration of the transmission, we can maintain a set of reference keyframes of the person from the higher-quality I-frames that are transmitted within the video stream and exploit them to guide the visual quality improvement; a novel aspect of this approach is the update policy that maintains and updates a compact and effective set of reference keyframes. First, we extract multi-scale features from the compressed and reference frames. Then, our architecture combines these features in a progressive manner according to facial landmarks. This allows the restoration of the high-frequency details lost after the video compression. Experiments show that the proposed approach improves visual quality and generates photo-realistic results even with high compression rates. Code and pre-trained networks are publicly available at https://github.com/LorenzoAgnolucci/Keyframes-GAN. Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini 0001, Alberto Del Bimbo |
IEEE Trans. Multim. | 3 |
| 2024 | Composed Image Retrieval using Contrastive Learning and Task-oriented CLIP-based FeaturesabstractGiven a query composed of a reference image and a relative caption, the Composed Image Retrieval goal is to retrieve images visually similar to the reference one that integrates the modifications expressed by the caption. Given that recent research has demonstrated the efficacy of large-scale vision and language pre-trained (VLP) models in various tasks, we rely on features from the OpenAI CLIP model to tackle the considered task. We initially perform a task-oriented fine-tuning of both CLIP encoders using the element-wise sum of visual and textual features. Then, in the second stage, we train a Combiner network that learns to combine the image-text features integrating the bimodal information and providing combined features used to perform the retrieval. We use contrastive learning in both stages of training. Starting from the bare CLIP features as a baseline, experimental results show that the task-oriented fine-tuning and the carefully crafted Combiner network are highly effective and outperform more complex state-of-the-art approaches on FashionIQ and CIRR, two popular and challenging datasets for composed image retrieval. Code and pre-trained models are available at https://github.com/ABaldrati/CLIP4Cir . Alberto Baldrati, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Zero-Shot Composed Image Retrieval with Textual InversionabstractComposed Image Retrieval (CIR) aims to retrieve a target image based on a query composed of a reference image and a relative caption that describes the difference between the two images. The high effort and cost required for labeling datasets for CIR hamper the widespread usage of existing methods, as they rely on supervised learning. In this work, we propose a new task, Zero-Shot CIR (ZS-CIR), that aims to address CIR without requiring a labeled training dataset. Our approach, named zero-Shot composEd imAge Retrieval with textuaL invErsion (SEARLE), maps the visual features of the reference image into a pseudo-word token in CLIP token embedding space and integrates it with the relative caption. To support research on ZS-CIR, we introduce an open-domain benchmarking dataset named Composed Image Retrieval on Common Objects in context (CIRCO), which is the first dataset for CIR containing multiple ground truths for each query. The experiments show that SEARLE exhibits better performance than the baselines on the two main datasets for CIR tasks, FashionIQ and CIRR, and on the proposed CIRCO. The dataset, the code and the model are publicly available at https://github.com/miccunifi/SEARLE. Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini 0001, Alberto Del Bimbo |
ICCV | 3 |
| 2023 | Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image EditingabstractFashion illustration is used by designers to communicate their vision and to bring the design idea from conceptualization to realization, showing how clothes interact with the human body. In this context, computer vision can thus be used to improve the fashion design process. Differently from previous works that mainly focused on the virtual try-on of garments, we propose the task of multimodal-conditioned fashion image editing, guiding the generation of human-centric fashion images by following multimodal prompts, such as text, human body poses, and garment sketches. We tackle this problem by proposing a new architecture based on latent diffusion models, an approach that has not been used before in the fashion domain. Given the lack of existing datasets suitable for the task, we also extend two existing fashion datasets, namely Dress Code and VITON-HD, with multimodal annotations collected in a semi-automatic manner. Experimental results on these new datasets demonstrate the effectiveness of our proposal, both in terms of realism and coherence with the given multimodal inputs. Source code and collected multimodal annotations are publicly available at: https://github.com/aimagelab/multimodal-garment-designer. Alberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia, Marco Bertini 0001, Rita Cucchiara |
ICCV | 5 |
| 2023 | Zero-Shot Image Retrieval with Human FeedbackabstractComposed image retrieval extends traditional content-based image retrieval (CBIR) combining a query image with additional descriptive text to express user intent and specify supplementary requests related to the visual attributes of the query image. This approach holds significant potential for e-commerce applications, such as interactive multimodal searches and chatbots. In our demo, we present an interactive composed image retrieval system based on the SEARLE approach, which tackles this task in a zero-shot manner efficiently and effectively. The demo allows users to perform image retrieval iteratively refining the results using textual feedback. Lorenzo Agnolucci, Alberto Baldrati, Marco Bertini 0001, Alberto Del Bimbo |
ACM Multimedia | 3 |
| 2023 | LaDI-VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-OnabstractThe rapidly evolving fields of e-commerce and metaverse continue to seek innovative approaches to enhance the consumer experience. At the same time, recent advancements in the development of diffusion models have enabled generative networks to create remarkably realistic images. In this context, image-based virtual try-on, which consists in generating a novel image of a target model wearing a given in-shop garment, has yet to capitalize on the potential of these powerful generative solutions. This work introduces LaDI-VTON, the first Latent Diffusion textual Inversion-enhanced model for the Virtual Try-ON task. The proposed architecture relies on a latent diffusion model extended with a novel additional autoencoder module that exploits learnable skip connections to enhance the generation process preserving the model's characteristics. To effectively maintain the texture and details of the in-shop garment, we propose a textual inversion component that can map the visual features of the garment to the CLIP token embedding space and thus generate a set of pseudo-word token embeddings capable of conditioning the generation process. Experimental results on Dress Code and VITON-HD datasets demonstrate that our approach outperforms the competitors by a consistent margin, achieving a significant milestone for the task. Source code and trained models are publicly available at: https://github.com/miccunifi/ladi-vton. Davide Morelli, Alberto Baldrati, Giuseppe Cartella, Marcella Cornia, Marco Bertini 0001, Rita Cucchiara |
ACM Multimedia | 5 |
| 2023 | Optimization Techniques of Deep Learning Models for Visual Quality ImprovementabstractVideo restoration is a widely studied task in the field of computer vision and image processing. The primary objective of video restoration is to improve the visual quality of degraded videos caused by various factors, such as noise, blur, compression artifacts, and other distortions. In this study, the integration of post-training quantization techniques was investigated to optimize deep learning models for super-resolution inference. The results indicate that reducing the precision of weights and activations in these models substantially decreases the computational complexity and memory requirements without compromising performance, rendering them more practical and cost-effective for real-world applications, where real-time inference is often required. When TensorRT was integrated with PyTorch, the efficiency of the model was further improved taking advantage of the INT8 computational capabilities of recent NVIDIA GPUs. Lorenzo Palloni, Leonardo Galteri, Marco Bertini 0001 |
SoMeT | 3 |
| 2022 | Effective conditioned and composed image retrieval combining CLIP-based featuresabstractConditioned and composed image retrieval extend CBIR systems by combining a query image with an additional text that expresses the intent of the user, describing additional requests w.r.t. the visual content of the query image. This type of search is interesting for e-commerce applications, e.g. to develop interactive multimodal searches and chat-bots. In this demo, we present an interactive system based on a combiner network, trained using contrastive learning, that combines visual and textual features obtained from the OpenAI CLIP network to address conditioned CBIR. The system can be used to improve e-shop search engines. For example, considering the fashion domain it lets users search for dresses, shirts and toptees using a candidate start image and expressing some visual differences w.r.t. its visual con-tent, e.g. asking to change color, pattern or shape. The pro-posed network obtains state-of-the-art performance on the FashionIQ dataset and on the more recent CIRR dataset, showing its applicability to the fashion domain for conditioned retrieval, and to more generic content considering the more general task of composed image retrieval. Alberto Baldrati, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
CVPR | 2 |
| 2022 | Restoration of Analog Videos Using Swin-UNetabstractIn this paper we present a system to restore analog videos of historical archives. These videos often contain severe visual degradation due to the deterioration of their tape supports that require costly and slow manual interventions to recover the original content. The proposed system uses a multi-frame approach and is able to deal also with severe tape mistracking, which results in completely scrambled frames. Tests on real-world videos from a major historical video archive show the effectiveness of our demo system. Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini 0001, Alberto Del Bimbo |
ACM Multimedia | 3 |
| 2022 | Engaging Museum Visitors with Gamification of Body and Facial ExpressionsabstractIn this demo we present two applications designed for the cultural heritage domain that exploit gamification techniques in order to improve fruition and learning of museum artworks. The two applications encourage users to replicate the poses and facial expressions of characters from paintings or statues, to help museum visitors make connections with works of art. Both applications challenge the user to fulfill a task in a funny way and provide the user with a visual report of the his/her experience that can be shared on social media, improving the engagement of the museums, and providing information on the artworks replicated in the challenge. Maria Giovanna Donadio, Filippo Principi, Andrea Ferracani, Marco Bertini 0001, Alberto Del Bimbo |
ACM Multimedia | 4 |
| 2022 | Effective triplet mining improves training of multi-scale pooled CNN for image retrieval
Federico Vaccaro, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
Mach. Vis. Appl. | 2 |
| 2022 | LANBIQUE: LANguage-based Blind Image QUality EvaluationabstractImage quality assessment is often performed with deep networks that are fine-tuned to regress a human provided quality score of a given image. Usually, this approach may lack generalization capabilities and, while being highly precise on similar image distribution, it may yield lower correlation on unseen distortions. In particular, they show poor performances, whereas images corrupted by noise, blur, or compression have been restored by generative models. As a matter of fact, evaluation of these generative models is often performed providing anecdotal results to the reader. In the case of image enhancement and restoration, reference images are usually available. Nevertheless, using signal based metrics often leads to counterintuitive results: Highly natural crisp images may obtain worse scores than blurry ones. However, blind reference image assessment may rank images reconstructed with GANs higher than the original undistorted images. To avoid time-consuming human-based image assessment, semantic computer vision tasks may be exploited instead. In this article, we advocate the use of language generation tasks to evaluate the quality of restored images. We refer to our assessment approach as LANguage-based Blind Image QUality Evaluation (LANBIQUE). We show experimentally that image captioning, used as a downstream task, may serve as a method to score image quality, independently of the distortion process that affects the data. Captioning scores are better aligned with human rankings with respect to classic signal based or No-reference image quality metrics. We show insights on how the corruption, by artefacts, of local image structure may steer image captions in the wrong direction. Leonardo Galteri, Lorenzo Seidenari, Pietro Bongini, Marco Bertini 0001, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | Partially Fake it Till you Make It: Mixing Real and Fake Thermal Images for Improved Object DetectionabstractIn this paper we propose a novel data augmentation approach for visual content domains that have scarce training datasets, compositing synthetic 3D objects within real scenes. We show the performance of the proposed system in the context of object detection in thermal videos, a domain where i) training datasets are very limited compared to visible spectrum datasets and ii) creating full realistic synthetic scenes is extremely cumbersome and expensive due to the difficulty in modeling the thermal properties of the materials of the scene. We compare different augmentation strategies, including state of the art approaches obtained through RL techniques, the injection of simulated data and the employment of a generative model, and study how to best combine our proposed augmentation with these other techniques. Experimental results demonstrate the effectiveness of our approach, and our single-modality detector achieves state-of-the-art results on the FLIR ADAS dataset. Francesco Bongini, Lorenzo Berlincioni, Marco Bertini 0001, Alberto Del Bimbo |
ACM Multimedia | 3 |
| 2021 | Fast Video Visual Quality and Resolution Improvement using SR-UNetabstractIn this paper, we address the problem of real-time video quality enhancement, considering both frame super-resolution and compression artifact-removal. The first operation increases the sampling resolution of video frames, the second removes visual artifacts such as blurriness, noise, aliasing, or blockiness introduced by lossy compression techniques, such as JPEG encoding for single-images, or H.264/H.265 for video data. We propose to use SR-UNet, a novel network architecture based on UNet, that has been specialized for fast visual quality improvement (i.e. capable of operating in less than 40ms, to be able to operate on videos at 25FPS). We show how this network can be used in a streaming context where the content is generated live, e.g. in video calls, and how it can be optimized when video to be streamed are prepared in advance. The network can be used as a final post processing, to optimize the visual appearance of a frame before showing it to the end-user in a video player. Thus, it can be applied without any change to existing video coding and transmission pipelines. Experiments carried on standard video datasets, also considering the H.265 compression, show that the proposed approach is able to either improve visual quality metrics given a fixed bandwidth budget, or video distortion given a fixed quality goal. Federico Vaccaro, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
ACM Multimedia | 2 |
| 2021 | Conditioned Image Retrieval for Fashion using Contrastive Learning and CLIP-based FeaturesabstractBuilding on the recent advances in multimodal zero-shot representation learning, in this paper we explore the use of features obtained from the recent CLIP model to perform conditioned image retrieval. Starting from a reference image and an additive textual description of what the user wants with respect to the reference image, we learn a Combiner network that is able to understand the image content, integrate the textual description and provide combined feature used to perform the conditioned image retrieval. Starting from the bare CLIP features and a simple baseline, we show that a carefully crafted Combiner network, based on such multimodal features, is extremely effective and outperforms more complex state of the art approaches on the popular FashionIQ dataset. Alberto Baldrati, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
MMAsia | 2 |
| 2021 | Language Based Image Quality AssessmentabstractEvaluation of generative models, in the visual domain, is often performed providing anecdotal results to the reader. In the case of image enhancement, reference images are usually available. Nonetheless, using signal based metrics often leads to counterintuitive results: highly natural crisp images may obtain worse scores than blurry ones. On the other hand, blind reference image assessment may rank images reconstructed with GANs higher than the original undistorted images. To avoid time consuming human based image assessment, semantic computer vision tasks may be exploited instead [9, 25, 33]. In this paper we advocate the use of language generation tasks to evaluate the quality of restored images. We show experimentally that image captioning, used as a downstream task, may serve as a method to score image quality. Captioning scores are better aligned with human rankings with respect to signal based metrics or no-reference image quality metrics. We show insights on how the corruption, by artifacts, of local image structure may steer image captions in the wrong direction. Lorenzo Seidenari, Leonardo Galteri, Pietro Bongini, Marco Bertini 0001, Alberto Del Bimbo |
MMAsia | 4 |
| 2021 | Bottom-up and Layerwise Domain Adaptation for Pedestrian Detection in Thermal ImagesabstractPedestrian detection is a canonical problem for safety and security applications, and it remains a challenging problem due to the highly variable lighting conditions in which pedestrians must be detected. This article investigates several domain adaptation approaches to adapt RGB-trained detectors to the thermal domain. Building on our earlier work on domain adaptation for privacy-preserving pedestrian detection, we conducted an extensive experimental evaluation comparing top-down and bottom-up domain adaptation and also propose two new bottom-up domain adaptation strategies. For top-down domain adaptation, we leverage a detector pre-trained on RGB imagery and efficiently adapt it to perform pedestrian detection in the thermal domain. Our bottom-up domain adaptation approaches include two steps: first, training an adapter segment corresponding to initial layers of the RGB-trained detector adapts to the new input distribution; then, we reconnect the adapter segment to the original RGB-trained detector for final adaptation with a top-down loss. To the best of our knowledge, our bottom-up domain adaptation approaches outperform the best-performing single-modality pedestrian detection results on KAIST and outperform the state of the art on FLIR. My Kieu, Andrew D. Bagdanov, Marco Bertini 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Task-Conditioned Domain Adaptation for Pedestrian Detection in Thermal Imagery
My Kieu, Andrew D. Bagdanov, Marco Bertini 0001, Alberto Del Bimbo |
ECCV (22) | 3 |
| 2020 | Inner Eye Canthus Localization for Human Body Temperature ScreeningabstractIn this paper, we propose an automatic approach for localizing the inner eye canthus in thermal face images. We first coarsely detect 5 facial keypoints corresponding to the center of the eyes, the nosetip and the ears. Then we compute a sparse 2D-3D points correspondence using a 3D Morphable Face Model (3DMM). This correspondence is used to project the entire 3D face onto the image, and subsequently locate the inner eye canthus. Detecting this location allows to obtain the most precise body temperature measurement for a person using a thermal camera. We evaluated the approach on a thermal face dataset provided with manually annotated landmarks. However, such manual annotations are normally conceived to identify facial parts such as eyes, nose and mouth, and are not specifically tailored for localizing the eye canthus region. As additional contribution, we enrich the original dataset by using the annotated landmarks to deform and project the 3DMM onto the images. Then, by manually selecting a small region corresponding to the eye canthus, we enrich the dataset with additional annotations. By using the manual landmarks, we ensure the correctness of the 3DMM projection, which can be used as ground-truth for future evaluations. Moreover, we supply the dataset with the 3D head poses and per-point visibility masks for detecting self-occlusions. The data is publicly available at https://www.micc.unifi.it/resources/datasets/thermal-face/. Claudio Ferrari, Lorenzo Berlincioni, Marco Bertini 0001, Alberto Del Bimbo |
ICPR | 3 |
| 2020 | Robust pedestrian detection in thermal imagery using synthesized imagesabstractIn this paper we propose a method for improving pedestrian detection in the thermal domain using two stages: first, a generative data augmentation approach is used, then a domain adaptation method using generated data adapts an RGB pedestrian detector. Our model, based on the Least-Squares Generative Adversarial Network, is trained to synthesize realistic thermal versions of input RGB images which are then used to augment the limited amount of labeled thermal pedestrian images available for training. We apply our generative data augmentation strategy in order to adapt a pretrained YOLOv3 pedestrian detector to detection in the thermal-only domain. Experimental results demonstrate the effectiveness of our approach: using less than 50% of available real thermal training data, and relying on synthesized data generated by our model in the domain adaptation phase, our detector achieves state-of-the-art results on the KAIST Multispectral Pedestrian Detection Benchmark; even if more real thermal data is available adding GAN generated images to the training data results in improved performance, thus showing that these images act as an effective form of data augmentation. To the best of our knowledge, our detector achieves the best single-modality detection results on KAIST with respect to the state-of-the-art. My Kieu, Lorenzo Berlincioni, Leonardo Galteri, Marco Bertini 0001, Andrew D. Bagdanov, Alberto Del Bimbo |
ICPR | 4 |
| 2020 | A NoGAN approach for image and video restoration and compression artifact removalabstractLossy image and video compression algorithms introduce several different types of visual artifacts that reduce the visual quality of the compressed media, and the higher the compression rate the higher is the strength of these artifacts. In this work, we describe an approach for visual quality improvement of compressed images and videos to be performed at presentation time, as to obtain the benefits of fast data transfer and reduced data storage, while enjoying a visual quality that could be obtained only reducing the compression rate. To obtain this result we propose to use a deep neural network trained using the NoGAN approach, adapting the popular DeOldify architecture used for colorization. We show how the proposed method can be applied both to image and video compression artifact removal and restoration. Filippo Mameli, Marco Bertini 0001, Leonardo Galteri, Alberto Del Bimbo |
ICPR | 2 |
| 2020 | Image Retrieval using Multi-scale CNN Features PoolingabstractIn this paper, we address the problem of image retrieval by learning images representation based on the activations of a Convolutional Neural Network. We present an end-to-end trainable network architecture that exploits a novel multi-scale local pooling based on NetVLAD and a triplet mining procedure based on samples difficulty to obtain an effective image representation. Extensive experiments show that our approach is able to reach state-of-the-art results on three standard datasets. Federico Vaccaro, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
ICMR | 2 |
| 2020 | Increasing Video Perceptual Quality with GANs and Semantic CodingabstractWe have seen a rise in video based user communication in the last year, unfortunately fueled by the spread of COVID-19 disease. Efficient low-latency delay of transmission of video is a challenging problem which must also deal with the segmented nature of network infrastructure not always allowing a high throughput. Lossy video compression is a basic requirement to enable such technology widely. While this may compromise the quality of the streamed video there are recent deep learning based solutions to restore quality of a lossy compressed video. Leonardo Galteri, Marco Bertini 0001, Lorenzo Seidenari, Tiberio Uricchio, Alberto Del Bimbo |
ACM Multimedia | 2 |
| 2020 | Image and Video Restoration and Compression Artefact Removal Using a NoGAN ApproachabstractLossy image and video compression algorithms introduce several types of visual artefacts that reduce the visual quality of the compressed media. In this work, we report results obtained using the NoGAN training approach and adapting the popular DeOldify architecture used for colorization, for image and video compression artefact removal and restoration. Filippo Mameli, Marco Bertini 0001, Leonardo Galteri, Alberto Del Bimbo |
ACM Multimedia | 2 |
| 2019 | Towards Real-Time Image Enhancement GANs
Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo |
CAIP (1) | 3 |
| 2019 | Fast Video Quality Enhancement using GANsabstractVideo compression algorithms result in a reduction of image quality, because of their lossy approach to reduce the required bandwidth. This affects commercial streaming services such as Netflix, or Amazon Prime Video, but affects also video conferencing and video surveillance systems. In all these cases it is possible to improve the video quality, both for human view and for automatic video analysis, without changing the compression pipeline, through a post-processing that eliminates the visual artifacts created by the compression algorithms. Generative Adversarial Networks have obtained extremely high quality results in image enhancement tasks; however, to obtain such results large generators are usually employed, resulting in high computational costs and processing time. In this work we present an architecture that can be used to reduce the computational cost and that has been implemented on mobile devices. A possible application is to improve video conferencing, or live streaming. In these cases there is no original uncompressed video stream available. Therefore, we report results using no-reference video quality metric showing high naturalness and quality even for efficient networks. Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
ACM Multimedia | 3 |
| 2019 | DeepPhysio: Monitored Physiotherapeutic Exercise in the Comfort of your Own HomeabstractThis paper describes an action classification pipeline for detecting and evaluating correct execution of actions in video recorded by smartphone cameras; the use case is that of simplifying monitoring of how physiotherapeutic exercises are performed by patients in the comfort of their own home, reducing the need of physical presence of therapists. Our approach is based on applying DensePose to every frame of acquired video and subsequent sequence analysis by an LSTM network. We validate our proposed recognition approach on a subset of the NTU RGB+D dataset in order to determine the best classification pipeline for this application. We also describe a mobile, cross-platform application called DeepPhysio that is designed to allow at physiotherapy patients to obtain immediate feedback about the correctness of the physical exercises. Preliminary usability analysis shows that this type of application can be effective at monitoring physiotherapy exercises. Gianmarco Sanesi, Andrew D. Bagdanov, Marco Bertini 0001, Alberto Del Bimbo |
ACM Multimedia | 3 |
| 2019 | Deep Universal Generative Adversarial Compression Artifact RemovalabstractImage compression is a need that arises in many circumstances. Unfortunately, whenever a lossy compression algorithm is used, artifacts will manifest. Image artifacts, caused by compression tend to eliminate higher frequency details and, in certain cases, may add noise or small image structures. There are two main drawbacks of this phenomenon. First, images appear much less pleasant to the human eye. Second, computer vision algorithms, such as object detectors, may be hindered and their performance reduced. Removing such artifacts means recovering the original image from a perturbed version of it. This means that one ideally should invert the compression process through a complicated nonlinear image transformation. We propose an image transformation approach based on a feedforward fully convolutional residual network model. We show that this model can be optimized either traditionally, directly optimizing an image similarity loss (SSIM), or using a generative adversarial approach (GAN). Our GAN is able to produce images with more photorealistic details than SSIM-based networks. We describe a novel training procedure based on subpatches and devise a novel testing protocol to evaluate restored images quantitatively. We show that our approach can be used as a preprocessing step for different computer vision tasks in case images are degraded by compression to a point that state-of-the art algorithms fail. In this case, our GAN-based approach obtains better performance than MSE or SSIM trained networks. Different from previously proposed approaches, we are able to remove artifacts generated at any QF by inferring the image quality directly from data. Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo |
IEEE Trans. Multim. | 3 |
| 2018 | Video Compression for Object Detection AlgorithmsabstractVideo compression algorithms have been designed aiming at pleasing human viewers, and are driven by video quality metrics that are designed to account for the capabilities of the human visual system. However, thanks to the advances in computer vision systems more and more videos are going to be watched by algorithms, e.g. implementing video surveillance systems or performing automatic video tagging. This paper describes an adaptive video coding approach for computer vision-based systems. We show how to control the quality of video compression so that automatic object detectors can still process the resulting video, improving their detection performance, by preserving the elements of the scene that are more likely to contain meaningful content. Our approach is based on computation of saliency maps exploiting a fast objectness measure. The computational efficiency of this approach makes it usable in a real-time video coding pipeline. Experiments show that our technique outperforms standard H.265 in speed and coding efficiency, and can be applied to different types of video domains, from surveillance to web videos. Leonardo Galteri, Marco Bertini 0001, Lorenzo Seidenari, Alberto Del Bimbo |
ICPR | 2 |
| 2017 | Deep Generative Adversarial Compression Artifact RemovalabstractCompression artifacts arise in images whenever a lossy compression algorithm is applied. These artifacts eliminate details present in the original image, or add noise and small structures; because of these effects they make images less pleasant for the human eye, and may also lead to decreased performance of computer vision algorithms such as object detectors. To eliminate such artifacts, when decompressing an image, it is required to recover the original image from a disturbed version. To this end, we present a feed-forward fully convolutional residual network model trained using a generative adversarial framework. To provide a baseline, we show that our model can be also trained optimizing the Structural Similarity (SSIM), which is a better loss with respect to the simpler Mean Squared Error (MSE). Our GAN is able to produce images with more photorealistic details than MSE or SSIM based networks. Moreover we show that our approach can be used as a pre-processing step for object detection in case images are degraded by compression to a point that state-of-the art detectors fail. In this task, our GAN method obtains better performance than MSE or SSIM trained networks. Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo |
ICCV | 3 |
| 2017 | Deep Sentiment Features of Context and Faces for Affective Video AnalysisabstractGiven the huge quantity of hours of video available on video sharing platforms such as YouTube, Vimeo, etc. development of automatic tools that help users find videos that fit their interests has attracted the attention of both scientific and industrial communities. So far the majority of the works have addressed semantic analysis, to identify objects, scenes and events depicted in videos, but more recently affective analysis of videos has started to gain more attention. In this work we investigate the use of sentiment driven features to classify the induced sentiment of a video, i.e. the sentiment reaction of the user. Instead of using standard computer vision features such as CNN features or SIFT features trained to recognize objects and scenes, we exploit sentiment related features such as the ones provided by Deep-SentiBank, and features extracted from models that exploit deep networks trained on face expressions. We experiment on two recently introduced datasets: LIRIS-ACCEDE and MEDIAEVAL-2015, that provide sentiment annotations of a large set of short videos. We show that our approach not only outperforms the current state-of-the-art in terms of valence and arousal classification accuracy, but it also uses a smaller number of features, requiring thus less video processing. Claudio Baecchi, Tiberio Uricchio, Marco Bertini 0001, Alberto Del Bimbo |
ICMR | 3 |
| 2017 | Spatio-Temporal Closed-Loop Object DetectionabstractObject detection is one of the most important tasks of computer vision. It is usually performed by evaluating a subset of the possible locations of an image, that are more likely to contain the object of interest. Exhaustive approaches have now been superseded by object proposal methods. The interplay of detectors and proposal algorithms has not been fully analyzed and exploited up to now, although this is a very relevant problem for object detection in video sequences. We propose to connect, in a closed-loop, detectors and object proposal generator functions exploiting the ordered and continuous nature of video sequences. Different from tracking we only require a previous frame to improve both proposal and detection: no prediction based on local motion is performed, thus avoiding tracking errors. We obtain three to four points of improvement in mAP and a detection time that is lower than Faster Regions with CNN features (R-CNN), which is the fastest Convolutional Neural Network (CNN) based generic object detector known at the moment. Leonardo Galteri, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo |
IEEE Trans. Image Process. | 3 |
| 2017 | Compact Hash Codes for Efficient Visual Descriptors Retrieval in Large Scale DatabasesabstractIn this paper, we present an efficient method for visual descriptors retrieval based on compact hash codes computed using a multiple k-means assignment. The method has been applied to the problem of approximate nearest neighbor (ANN) search of local and global visual content descriptors, and it has been tested on different datasets: three large scale standard datasets of engineered features of up to one billion descriptors (BIGANN) and, supported by recent progress in convolutional neural networks (CNNs), on CIFAR-10, MNIST, INRIA Holidays, Oxford 5K, and Paris 6K datasets; also, the recent DEEP1B dataset, composed by one billion CNN-based features, has been used. Experimental results show that, despite its simplicity, the proposed method obtains a very high performance that makes it superior to more complex state-of-the-art methods. Simone Ercoli, Marco Bertini 0001, Alberto Del Bimbo |
IEEE Trans. Multim. | 2 |
| 2017 | Deep Artwork Detection and Retrieval for Automatic Context-Aware Audio GuidesabstractIn this article, we address the problem of creating a smart audio guide that adapts to the actions and interests of museum visitors. As an autonomous agent, our guide perceives the context and is able to interact with users in an appropriate fashion. To do so, it understands what the visitor is looking at, if the visitor is moving inside the museum hall, or if he or she is talking with a friend. The guide performs automatic recognition of artworks, and it provides configurable interface features to improve the user experience and the fruition of multimedia materials through semi-automatic interaction. Our smart audio guide is backed by a computer vision system capable of working in real time on a mobile device, coupled with audio and motion sensors. We propose the use of a compact Convolutional Neural Network (CNN) that performs object classification and localization. Using the same CNN features computed for these tasks, we perform also robust artwork recognition. To improve the recognition accuracy, we perform additional video processing using shape-based filtering, artwork tracking, and temporal filtering. The system has been deployed on an NVIDIA Jetson TK1 and a NVIDIA Shield Tablet K1 and tested in a real-world environment (Bargello Museum of Florence). Lorenzo Seidenari, Claudio Baecchi, Tiberio Uricchio, Andrea Ferracani, Marco Bertini 0001, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2016 | Bloom Filters and Compact Hash Codes for Efficient and Distributed Image RetrievalabstractThis paper presents a novel method for efficient image retrieval, based on a simple and effective hashing of CNN features and the use of an indexing structure based on Bloom filters. These filters are used as gatekeepers for the database of image features, allowing to avoid to perform a query if the query features are not stored in the database and speeding up the query process, without affecting retrieval performance. Thanks to the limited memory requirements the system is suitable for mobile applications and distributed databases, associating each filter to a distributed portion of the database (database shard), addressing large scale archives and allowing query parallelization. Experimental validation has been performed on three standard image retrieval datasets, outperforming state-of-the-art hashing methods in terms of precision, while the proposed indexing method obtains a 2x speedup. Andrea Salvi, Simone Ercoli, Marco Bertini 0001, Alberto Del Bimbo |
ISM | 3 |
| 2016 | Item-Based Video Recommendation: An Hybrid Approach considering Human FactorsabstractIn this paper we propose a method for video recommendation in Social Networks based on crowdsourced and automatic video annotations of salient frames. We show how two human factors, users' self-expression in user profiles and perception of visual saliency in videos, can be exploited in order to stimulate annotations and to obtain an efficient representation of video content features. Results are assessed through experiments conducted on a prototype of social network for video sharing. Several baseline approaches are evaluated and we show how the proposed method improves over them. Andrea Ferracani, Daniele Pezzatini, Marco Bertini 0001, Alberto Del Bimbo |
ICMR | 3 |
| 2016 | Web Video Popularity Prediction using Sentiment and Content Visual FeaturesabstractHundreds of hours of videos are uploaded every minute on YouTube and other video sharing sites: some will be viewed by millions of people and other will go unnoticed by all but the uploader. In this paper we propose to use visual sentiment and content features to predict the popularity of web videos. The proposed approach outperforms current state-of-the-art methods on two publicly available datasets. Giulia Fontanini, Marco Bertini 0001, Alberto Del Bimbo |
ICMR | 2 |
| 2016 | Real-time Wearable Computer Vision System for Improved Museum ExperienceabstractThe goal of this work is to implement a real-time computer vision system that can run on wearable devices to perform object classification and artwork recognition, to improve the experience of a museum visit through understanding the interests of users. Object classification helps to understand the context of the visit, e.g. differentiating when a visitor is talking with people, or just wandering through the museum, or if he is looking at an exhibit that interests him. Artwork recognition allows to provide automatically information of the observed item or to create a user profile based on what and how long a user has observed artworks. Giovanni Taverriti, Stefano Lombini, Lorenzo Seidenari, Marco Bertini 0001, Alberto Del Bimbo |
ACM Multimedia | 4 |
| 2016 | A multimodal feature learning approach for sentiment analysis of social network multimedia
Claudio Baecchi, Tiberio Uricchio, Marco Bertini 0001, Alberto Del Bimbo |
Multim. Tools Appl. | 3 |
| 2016 | Guest Editorial: Analysis and Retrieval of Events/Actions and Workflows in Video Streams
Anastasios Doulamis, Nikolaos D. Doulamis, Marco Bertini 0001, Jordi Gonzàlez 0001, Thomas B. Moeslund |
Multim. Tools Appl. | 3 |
| 2015 | Movie's Affect Communication Using Multisensory ModalitiesabstractThe goal of the system presented in this demo is to make possible for the visually and hearing impaired audience to live empathetic viewing experiences using their home theatre. In this work we suggest the incorporation of new emotion communication modalities into the standard television, to provide the targeted audience with sensations that they do not have the opportunity to enjoy because of their disability. Joël Dumoulin, Diana Affi, Elena Mugellini, Omar Abou Khaled, Marco Bertini 0001, Alberto Del Bimbo |
ACM Multimedia | 5 |
| 2015 | A System for Video Recommendation using Visual Saliency, Crowdsourced and Automatic AnnotationsabstractIn this paper we present a system for content-based video recommendation that exploits visual saliency to better represent video features and content\footnote{Demo video available at http://bit.ly/1FYloeQ}. Visual saliency is used to select relevant frames to be presented in a web-based interface to tag and annotate video frames in a social network; it is also employed to summarize video content to create a more effective video representation used in the recommender system. The system exploits automatic annotations from CNN-based classifiers on salient frames and user generated annotations. We evaluate several baseline approaches and show how the proposed method improves over them. Andrea Ferracani, Daniele Pezzatini, Marco Bertini 0001, Saverio Meucci, Alberto Del Bimbo |
ACM Multimedia | 3 |
| 2015 | Image Popularity Prediction in Social Media Using Sentiment and Context FeaturesabstractImages in social networks share different destinies: some are going to become popular while others are going to be completely unnoticed. In this paper we propose to use visual sentiment features together with three novel context features to predict a concise popularity score of social images. Experiments on large scale datasets show the benefits of proposed features on the performance of image popularity prediction. Exploiting state-of-the-art sentiment features, we report a qualitative analysis of which sentiments seem to be related to good or poor popularity. To the best of our knowledge, this is the first work understanding specific visual sentiments that positively or negatively influence the eventual popularity of images. Francesco Gelli, Tiberio Uricchio, Marco Bertini 0001, Alberto Del Bimbo, Shih-Fu Chang |
ACM Multimedia | 3 |
| 2015 | Image Tag Assignment, Refinement and RetrievalabstractThis tutorial focuses on challenges and solutions for content-based image annotation and retrieval in the context of online image sharing and tagging. We present a unified review on three closely linked problems, i.e., tag assignment, tag refinement, and tag-based image retrieval. We introduce a taxonomy to structure the growing literature, understand the ingredients of the main works, clarify their connections and difference, and recognize their merits and limitations. Moreover, we present an open-source testbed, with training sets of varying sizes and three test datasets, to evaluate methods of varied learning complexity. A selected set of eleven representative works have been implemented and evaluated. During the tutorial we provide a practice session for hands on experience with the methods, software and datasets. For repeatable experiments all data and code are online at http://www.micc.unifi.it/tagsurvey Xirong Li 0001, Tiberio Uricchio, Lamberto Ballan, Marco Bertini 0001, Cees Snoek, Alberto Del Bimbo |
ACM Multimedia | 4 |
| 2015 | A data-driven approach for tag refinement and localization in web videos
Lamberto Ballan, Marco Bertini 0001, Giuseppe Serra 0001, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 2 |
| 2015 | Data-driven approaches for social image and video tagging
Lamberto Ballan, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
Multim. Tools Appl. | 2 |
| 2014 | Loki+Lire: a framework to create web-based multimedia search enginesabstractIn this paper we present Loki+Lire, a framework for the creation of web-based interfaces for search, annotation and presentation of multimedia data. The framework provides tools to ingest, transcode, present, annotate and index different types of media such as images, videos, audio files and textual documents. The front-end is compliant with the latest HTML5 standards, while the back-end allows system administrators to create processing pipelines that can be adapted for different tasks and purposes. Giuseppe Becchi, Marco Bertini 0001, Lorenzo Cioni, Alberto Del Bimbo, Andrea Ferracani, Daniele Pezzatini, Mathias Lux |
ACM Multimedia | 2 |
| 2013 | An evaluation of nearest-neighbor methods for tag refinementabstractThe success of media sharing and social networks has led to the availability of extremely large quantities of images that are tagged by users. The need of methods to manage efficiently and effectively the combination of media and metadata poses significant challenges. In particular, automatic image annotation of social images has become an important research topic for the multimedia community. In this paper we propose and thoroughly evaluate the use of nearest-neighbor methods for tag refinement. Extensive and rigorous evaluation using two standard large-scale datasets shows that the performance of these methods is comparable with that of more complex and computationally intensive approaches and that, differently from these latter approaches, nearest-neighbor methods can be applied to `web-scale' data. Tiberio Uricchio, Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo |
ICME | 3 |
| 2013 | A novel framework for collaborative video recommendation, interest discovery and friendship suggestion based on semantic profilingabstractTwo important challenges for social networks are the creation of targeted and personalized content for their users, selecting the most interesting material from the huge amount of user-generated content, and keeping user engagement , e.g. through creation and curation of users' profiles. In this demo we show a system for video commenting, sharing and interest discovery that combines recommendation algorithms, clustering techniques, tools for video tagging and evaluation of semantic resources relatedness. Combining these tools and techniques it becomes possible to provide personalized multimedia services and to improve and propagate interests and inter-personal connections through the network. Marco Bertini 0001, Alberto Del Bimbo, Andrea Ferracani, Francesco Gelli, Daniele Maddaluno, Daniele Pezzatini |
ACM Multimedia | 1 |
| 2013 | euTV: a system for media monitoring and publishingabstractIn this paper, we describe the euTV system, which provides a flexible approach to collect, manage, annotate and publish collections of images, videos and textual documents. The system is based on a Service Oriented Architecture that allows to combine and orchestrate a large set of web services for automatic and manual annotation, retrieval, browsing, ingestion and authoring of multimedia sources. euTV tools have been used to create several publicly available vertical applications, addressing different use cases. Positive results of user evaluations have shown that the system can be effectively used to create different types of applications. Marco Bertini 0001, Alberto Del Bimbo, George Ioannidis, Emile Bijk, Isabel Trancoso, Hugo Meinedo |
ACM Multimedia | 1 |
| 2013 | 4th ACM/IEEE ARTEMIS 2013 international workshop on analysis and retrieval of tracked events and motion in imagery streamsabstractIn this paper, we give a short summary of the papers proposed in ACMARTEMIS 2013 which is held in Barcelona Spain in conjunction with ACM Multimedia. The workshop handles the areas of features analysis both at low and high level for efficient events detection, retrieval of multimedia events and objects and video synchronization issues and also events and behavior recognition from visual data. All papers were classified into three session of a single track workshop. The first session named "Video Features and Scene Analysis" includes articles that handle low level and high level visual analysis appropriate for event detection. The second session entitled "Retrieval of Multimedia Objects/Events" applies schemes for media data retrieval and video synchronization. Finally the third session "Analysis of Visual Events" describes algorithms for detecting actions, behaviors and events in complex visual scenes. Anastasios Doulamis, Nikolaos D. Doulamis, Marco Bertini 0001, Jordi Gonzàlez 0001, Thomas B. Moeslund |
ACM Multimedia | 3 |
| 2013 | Interactive multi-user video retrieval systems
Marco Bertini 0001, Alberto Del Bimbo, Andrea Ferracani, Lea Landucci, Daniele Pezzatini |
Multim. Tools Appl. | 1 |
| 2012 | Combining generative and discriminative models for classifying social images from 101 object categories
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Andrea M. Serain, Giuseppe Serra 0001, Benito F. Zaccone |
ICPR | 2 |
| 2012 | Social and automatic annotation of videos for semantic profiling and content discoveryabstractThis demo presents a system based on social relationships, social knowledge and automatic video and textual content analysis for the discovery of videos in social networks. The system, developed as a web application, allows users to annotate, manually and automatically, and comment video frames and scenes enriching their content with tags, references to Facebook users and pages and Wikipedia resources. These annotations are used to semantically model the profile of each user extracting and expanding his interests and folksonomy, as well as resources of interest in his social graph. The automatically generated profile page is used to suggest to users new resources, Facebook friends and videos whose content is related to their interests and allows profile curation. A screencast showing an example of these functionalities is publicly available at: http://vimeo.com/miccunifi/facetube Marco Bertini 0001, Alberto Del Bimbo, Andrea Ferracani, Daniele Pezzatini |
ACM Multimedia | 1 |
| 2012 | Multi-scale and real-time non-parametric approach for anomaly detection and localization
Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari |
Comput. Vis. Image Underst. | 1 |
| 2012 | Multimedia and semantic technologies for future computing environments
Seungmin Rho, Marco Bertini 0001, Gamhewage Chaminda de Silva, Stephan Kopf |
Multim. Tools Appl. | 2 |
| 2012 | Effective Codebooks for Human Action Representation and Classification in Unconstrained VideosabstractRecognition and classification of human actions for annotation of unconstrained video sequences has proven to be challenging because of the variations in the environment, appearance of actors, modalities in which the same action is performed by different persons, speed and duration, and points of view from which the event is observed. This variability reflects in the difficulty of defining effective descriptors and deriving appropriate and effective codebooks for action categorization. In this paper, we propose a novel and effective solution to classify human actions in unconstrained videos. It improves on previous contributions through the definition of a novel local descriptor that uses image gradient and optic flow to respectively model the appearance and motion of human actions at interest point regions. In the formation of the codebook, we employ radius-based clustering with soft assignment in order to create a rich vocabulary that may account for the high variability of human actions. We show that our solution scores very good performance with no need of parameter tuning. We also show that a strong reduction of computation time can be obtained by applying codebook size reduction with Deep Belief Networks with little loss of accuracy. Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari, Giuseppe Serra 0001 |
IEEE Trans. Multim. | 2 |
| 2011 | A web system for ontology-based multimedia annotation, browsing and searchabstractIn this paper we present a complete system for semantic and syn tactic annotation, browsing and search of multimedia data, that is based on a service oriented architecture, with web-based interfaces developed following the Rich Internet Application paradigm. The system has been designed to be: i) flexible and extendable, allowing users to select only the services they need or to add their own tools to the multimedia processing pipelines; ii) distributed, with services that can be executed in a cloud computing infrastructure and accessed through web applications; Hi) user-friendly, with interfaces that have a uniform interface on every platform and that have an interaction level similar to that of desktop applications. Extensive user trials in real-world setup, performed by archive and broadcaster professionals, have shown the efficacy and usability of the proposed solution. Marco Bertini 0001, Giuseppe Becchi, Alberto Del Bimbo, Andrea Ferracani, Daniele Pezzatini |
ICME | 1 |
| 2011 | Adaptive Video Compression for Video Surveillance ApplicationsabstractThis article describes an approach to adaptive video coding for video surveillance applications. Using a combination of low-level features with low computational cost, we show how it is possible to control the quality of video compression so that semantically meaningful elements of the scene are encoded with higher fidelity, while background elements are allocated fewer bits in the transmitted representation. Our approach is based on adaptive smoothing of individual video frames so that image features highly correlated to semantically interesting objects are preserved. Using only low-level image features on individual frames, this adaptive smoothing can be seamlessly inserted into a video coding pipeline as a pre-processing state. Experiments show that our technique is efficient, outperforms standard H.264 encoding at comparable bit rates, and preserves features critical for downstream detection and recognition. Andrew D. Bagdanov, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari |
ISM | 2 |
| 2011 | A flexible environment for multimedia management and publishingabstractIn this paper, we describe the IM3I system, which provides a flexible approach to managing and publishing collections of images and videos. The system is based on web services that allow automatic and manual annotation, retrieval, browsing and authoring of multimedia. Results of user evaluations, performed by professional archivists and archive managers on a real-world system deployment have confirmed that the system is easy to be used and delivers a complete set of functionalities. Marco Bertini 0001, Alberto Del Bimbo, George Ioannidis, Alexandru Stan, Emile Bijk |
ICMR | 1 |
| 2011 | Enriching and localizing semantic tags in internet videosabstractTagging of multimedia content is becoming more and more widespread as web 2.0 sites, like Flickr and Facebook for images, YouTube and Vimeo for videos, have popularized tagging functionalities among their users. These user-generated tags are used to retrieve multimedia content, and to ease browsing and exploration of media collections, e.g.~using tag clouds. However, not all media are equally tagged by users: using the current browsers is easy to tag a single photo, and even tagging a part of a photo, like a face, has become common in sites like Flickr and Facebook; on the other hand tagging a video sequence is more complicated and time consuming, so that users just tag the overall content of a video. In this paper we present a system for automatic video annotation that increases the number of tags originally provided by users, and localizes them temporally, associating tags to shots. This approach exploits collective knowledge embedded in tags and Wikipedia, and visual similarity of keyframes and images uploaded to social sites like YouTube and Flickr. Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Giuseppe Serra 0001 |
ACM Multimedia | 2 |
| 2011 | Event detection and recognition for semantic annotation of video
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari, Giuseppe Serra 0001 |
Multim. Tools Appl. | 2 |
| 2010 | Sirio, orione and pan: an integrated web system for ontology-based video search and annotationabstractIn this technical demonstration we show an integrated web system for video search and annotation based on ontologies. The system is composed by three components: the Orione ontology-based search engine, the Sirio\footnote{Sirio was the hound of Orione. It was a dog so swift that no prey could escape it.} search interface, and the Pan web-based video annotation tool. The system is currently being developed within the EU IM3I project. The goal of the system is to provide an integrated environment for video annotation and retrieval of videos, for both technical and non-technical users. In fact, the search engine has different interfaces that permit different query modalities: free-text, natural language, graphical composition of concepts using Boolean and temporal relations and query by visual example. In addition, the ontology structure is exploited to encode semantic relations between concepts permitting, for example, to expand queries to synonyms and concept specializations. The annotation tool can be used to create ground-truth annotations to train automatic annotations systems, or to complement the results of automatic annotation, e.g. adding geolocalized information. Marco Bertini 0001, Gianpaolo D'Amico, Andrea Ferracani, Marco Meoni, Giuseppe Serra 0001 |
ACM Multimedia | 1 |
| 2010 | Web-based semantic browsing of video collections using multimedia ontologiesabstractIn this technical demonstration we present a novel web-based tool that allows a user friendly semantic browsing of video collections, based on ontologies, concepts, concept relations and concept clouds. The system is developed as a Rich Internet Application (RIA) to achieve a fast responsiveness and ease of use that can not be obtained by other web application paradigms, and uses streaming to access and inspect the videos. Users can also use the tool to browse the content of social and media sharing sites like YouTube, Flickr and Twitter, accessing these external resources through the ontologies used in the system. The tool has won the second prize in the Adobe YouGC contest, in the RIA category. Marco Bertini 0001, Gianpaolo D'Amico, Andrea Ferracani, Marco Meoni, Giuseppe Serra 0001 |
ACM Multimedia | 1 |
| 2010 | Non-parametric anomaly detection exploiting space-time featuresabstractIn this paper a real-time anomaly detection system for video streams is proposed. Spatio-temporal features are exploited to capture scene dynamic statistics together with appearance. Anomaly detection is performed in a non-parametric fashion, evaluating directly local descriptor statistics. A method to update scene statistics, to cope with scene changes that typically happen in real world settings, is also provided. The proposed method is tested on publicly available datasets. Lorenzo Seidenari, Marco Bertini 0001 |
ACM Multimedia | 2 |
| 2010 | Video event classification using string kernels
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Giuseppe Serra 0001 |
Multim. Tools Appl. | 2 |
| 2010 | Semantic annotation of soccer videos by visual instance clustering and spatial/temporal reasoning in ontologies
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Giuseppe Serra 0001 |
Multim. Tools Appl. | 2 |
| 2009 | Recognizing human actions by fusing spatio-temporal appearance and motion descriptorsabstractIn this paper we propose a new method for human action categorization by using an effective combination of a new 3D gradient descriptor with an optic flow descriptor, to represent spatio-temporal interest points. These points are used to represent video sequences using a bag of spatio-temporal visual words, following the successful results achieved in object and scene classification. We extensively test our approach on the standard KTH and Weizmann actions datasets, showing its validity and good performance. Experimental results outperform state-of-the-art methods, without requiring fine parameter tuning. Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari, Giuseppe Serra 0001 |
ICIP | 2 |
| 2009 | Deep networks for audio event classification in soccer videosabstractIn this work is presented a novel approach for the classification of audio concepts in broadcast soccer videos using deep belief network (DBN), a probabilistic neural network with several hidden layers. Comparison with support vector machine (SVM) classifiers has been carried on, showing that our preliminary results are promisingly comparable to the state-of-the-art. Lamberto Ballan, Alessio Bazzica, Marco Bertini 0001, Alberto Del Bimbo, Giuseppe Serra 0001 |
ICME | 3 |
| 2009 | Arneb: a rich internet application for ground truth annotation of videosabstractIn this technical demonstration we show the current version of Arneb, a web-based system for manual annotation of videos, developed within the EU VidiVideo project. This tool has been developed with the aim of creating ground truth annotations, that can be used for training and evaluating automatic video annotation systems. Annotations can be exported to MPEG-7 and OWL ontologies. The system has been developed according to the Rich Internet Application paradigm, allowing collaborative web-based annotation. Thomas M. Alisi, Marco Bertini 0001, Gianpaolo D'Amico, Alberto Del Bimbo, Andrea Ferracani, Federico Pernici, Giuseppe Serra 0001 |
ACM Multimedia | 2 |
| 2009 | Sirio: an ontology-based web search engine for videosabstractIn this technical demonstration we show a web video search engine based on ontologies, the Sirio system, that has been developed within the EU VidiVideo project. The goal of the system is to provide a search engine for videos for both technical and non-technical users. In fact, the system has different interfaces that permit different query modalities: free-text, natural language, graphical composition of concepts using boolean and temporal relations and query by visual example. In addition, the ontology structure is exploited to encode semantic relations between concepts permitting, for example, to expand queries to synonyms and concept specializations. Thomas M. Alisi, Marco Bertini 0001, Gianpaolo D'Amico, Alberto Del Bimbo, Andrea Ferracani, Federico Pernici, Giuseppe Serra 0001 |
ACM Multimedia | 2 |
| 2008 | Automatic trademark detection and recognition in sport videosabstractIn this paper we describe a system for automatic detection and recognition of trademarks in sports videos. We propose a compact representation of trademarks based on SIFT feature points and a matching algorithm to robustly detect and retrieve trademarks in a variety of different sports video types. Trademark localization is performed through robust clustering of matched feature points in the video frame. A supervised machine learning approach is used to automatically adapt the similarity threshold used to assess the trademark matches. Experimental results are provided, along with an analysis of the precision and recall. Results show that our proposed technique is efficient and effectively detects and classifies trademarks. Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Arjun Jain |
ICME | 2 |
| 2008 | A system for automatic detection and recognition of advertising trademarks in sports videosabstractIn this technical demonstration we show the current version of our trademark detection and recognition system that has been developed in collaboration with a sport marketing firm1 with the aim of evaluating the visibility of advertising trademarks in broadcast sporting events. We propose a semi-automatic system for detecting and retrieving trademark appearances in sports videos. A human annotator supervises the results of the automatic annotation through an interface that shows the time and the position of the detected trademarks; due to this fact the aim of the system is to provide a good recall figure, so that the supervisor can safely skip the parts of the video that have been marked as not containing a trademark, thus speeding up his work. Lamberto Ballan, Marco Bertini 0001, Arjun Jain |
ACM Multimedia | 2 |
| 2007 | Content-Based Retrieval of 3-D Objects Using Spin Image SignaturesabstractRetrieval by content of 3-D models is becoming more and more important due to the advancements in 3-D hardware and software technologies for acquisition, authoring and display of 3-D objects, their ever-increasing availability at affordable costs, and the establishment of open standards for 3-D data interchange. In this paper, we present a new method, referred to as Spin Image Signatures, that develops on the original spin images approach, with adaptations to support effective retrieval by content. According to the method proposed, a set of spin images is derived for each model, to obtain a view-independent description of its 3-D shape and a signature is evaluated for each spin image in the set. Clustering is hence performed on the set of Spin Image Signatures to obtain a compact representation. Experimental results are presented, showing the effectiveness of the Spin Image Signatures method for retrieval, also in comparison with other methods, and its sensitivity to model deformations. Jürgen Assfalg, Marco Bertini 0001, Alberto Del Bimbo, Pietro Pala |
IEEE Trans. Multim. | 2 |
| 2006 | Using Knowledge Representation Languages for Video Annotation and Retrieval
Marco Bertini 0001, Gianpaolo D'Amico, Alberto Del Bimbo, Carlo Torniai |
FQAS | 1 |
| 2006 | Matching Faces with Textual Cues in Soccer VideosabstractIn soccer videos, most significant actions are usually followed by close-up shots of players that take part in the action itself. Automatically annotating the identity of the players present in these shots would be considerably valuable for indexing and retrieval applications. Due to high variations in pose and illumination across shots however, current face recognition methods are not suitable for this task. We show how the inherent multiple media structure of soccer videos can be exploited to understand the players' identity without relying on direct face recognition. The proposed method is based on a combination of interest point detector to "read" textual cues that allow to label a player with its name, such as the number depicted on its jersey, or the superimposed text caption showing its name. Players not identified by this process are then assigned to one of the labeled faces by means of a face similarity measure, again based on the appearance of local salient patches. We present results obtained from soccer videos taken from various recent games between national teams Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati |
ICME | 1 |
| 2006 | Automatic detection of player's identity in soccer videos using faces and text cuesabstractIn soccer videos, most significant actions are usually followed by close--up shots of players that take part in the action itself. Automatically annotating the identity of the players present in these shots would be considerably valuable for indexing and retrieval applications. Due to high variations in pose and illumination across shots however, current face recognition methods are not suitable for this task. We show how the inherent multiple media structure of soccer videos can be exploited to understand the players' identity without relying on direct face recognition. The proposed method is based on a combination of interest point detector to "read" textual cues that allow to label a player with its name, such as the number depicted on its jersey, or the superimposed text caption showing its name. Players not identified by this process are then assigned to one of the labeled faces by means of a face similarity measure, again based on the appearance of local salient patches. We present results obtained from soccer videos taken from various recent games between national teams. Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati |
ACM Multimedia | 1 |
| 2006 | Automatic annotation and semantic retrieval of video sequences using multimedia ontologiesabstractEffective usage of multimedia digital libraries has to deal with the problem of building efficient content annotation and retrieval tools. MOM (Multimedia Ontology Manager) is a complete system that allows the creation of multimedia ontologies, supports automatic annotation and creation of extended text (and audio) commentaries of video sequences, and permits complex queries by reasoning on the ontology. Marco Bertini 0001, Alberto Del Bimbo, Carlo Torniai |
ACM Multimedia | 1 |
| 2006 | MOM: multimedia ontology manager. A framework for automatic annotation and semantic retrieval of video sequencesabstractEffective usage of multimedia digital libraries has to deal with the problem of building efficient content annotation and retrieval tools. MOM (Multimedia Ontology Manager) is a complete system that allows the creation of multimedia ontologies, supports automatic annotation and creation of extended text (and audio) commentaries of video sequences, and permits complex queries by reasoning on the ontology. Marco Bertini 0001, Alberto Del Bimbo, Carlo Torniai, Rita Cucchiara, Costantino Grana |
ACM Multimedia | 1 |
| 2006 | PEANO: pictorial enriched annotation of videoabstractIn this DEMO, we present a tool set for video digital library management that allows i) structural annotation of edited videos in MPEG-7 by automatically extracting shots and clips; ii) automatic semantic annotation based on perceptual similarity against a taxonomy enriched with pictorial concepts iii) video clip access and hierarchical summarization with stand-alone and web interface iv) access to clips from mobile platform in GPRS-UMTS video-streaming. The tools can be applied in different domain-specific Video Digital Libraries. The main novelty is the possibility to enrich the annotation with pictorial concepts that are added to a textual taxonomy in order to make the automatic annotation process more fast and often effective. The resulting multimedia ontology is described in the MPEG-7 framework. The PEANO (Perceptual Annotation of Video) tool has been tested over video art , sport (Soccer, Olimpic Games 2006, Formula 1) and news clips. Costantino Grana, Roberto Vezzani, Daniele Bulgarelli, Giovanni Gualdi, Rita Cucchiara, Marco Bertini 0001, Carlo Torniai, Alberto Del Bimbo |
ACM Multimedia | 6 |
| 2006 | Parallel and distributed simulation of wireless vehicular ad hoc networksabstractIn this paper, we present a novel solution for the fine-grained parallel and distributed simulation of Vehicular ad hoc networks' (VANETs) services and applications, based on vehicular traffic mobility and wireless IEEE 802.11-standard models. A Mobile Wireless Vehicular Environment Simulation (MoVES) framework supports dynamic partition of geographic areas and dynamic entity mapping of the modeled scenarios over parallel and distributed simulation platforms. This could be a viable solution to enhance simulation performances and to exploit low-cost commercial-off-the-shelf (COTS) simulation architectures. Testbed performance evaluation for realistic modeled scenarios has shown the effectiveness of our approach in terms of simulation efficiency (speedup), and simulation accuracy, also providing guidelines for future enhancements of the framework performances, and model characterization of inter-vehicular communication. Luciano Bononi, Marco Di Felice, Marco Bertini 0001, Emidio Croci |
MSWiM | 3 |
| 2006 | Semantic adaptation of sport videos with user-centred performance analysisabstractIn semantic video adaptation measures of performance must consider the impact of the errors in the automatic annotation over the adaptation in relationship with the preferences and expectations of the user. In this paper, we define two new performance measures Viewing Quality Loss and Bit-rate Cost Increase,that are obtained from classical peak signal-to-noise ration (PSNR) and bitrate, and relate the results of semantic adaptation to the errors in the annotation of events and objects and the user's preferences and expectations. We present and discuss results obtained with a system that performs automatic annotation of soccer sport video highlights and applies different coding strategies to different parts of the video according to their relative importance for the end user. With reference to this framework, we analyze how highlights' statistics and the errors of the annotation engine influence the performance of semantic adaptation and reflect into the quality of the video displayed at the user's client and the increase of transmission costs. Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001 |
IEEE Trans. Multim. | 1 |
| 2005 | Domain Knowledge Extension with Pictorially Enriched Ontologies
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Carlo Torniai |
CAIP | 1 |
| 2005 | Automatic Annotation of Sport Video Content
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati |
CIARP | 1 |
| 2005 | Video Annotation with Pictorially Enriched OntologiesabstractVideo annotation is typically performed by classifying video elements according to some pre-defined ontology of the video content domain. Ontologies are defined by establishing relationships between linguistic terms, that specify domain concepts at different abstraction levels. However, although linguistic terms are appropriate to distinguish event and object categories, they are inadequate when they must describe specific patterns of events or video entities. Instead, in these cases, pattern specifications are better expressed through visual prototypes that capture the essence of the event or entity. Pictorially enriched ontologies, that include visual concepts together with linguistic keywords, are therefore needed to support video annotation up to the level of detail of pattern specification. This paper presents pictorially enriched ontologies and provide a solution for their implementation in the soccer video domain. The pictorially enriched ontology is used both to directly assign multimedia objects to concepts, providing a more meaningful definition than the linguistics terms, and to extend the initial knowledge of the domain, adding subclasses of highlights or new highlight classes that were not defined in the linguistic ontology. Automatic annotation of soccer clips up to the pattern specification level using a pictorially enriched ontology is discussed. Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Carlo Torniai |
ICME | 1 |
| 2005 | Automatic video annotation using ontologies extended with visual informationabstractClassifying video elements according to some pre-defined ontology of the video content domain is a typical way to perform video annotation. Ontologies are defined by establishing relationships between linguistic terms that specify domain concepts at different abstraction levels. However, although linguistic terms are appropriate to distinguish event and object categories, they are inadequate when they must describe specific patterns of events or video entities. Instead, in these cases, pattern specifications can be better expressed through visual prototypes that capture the essence of the event or entity. Therefore pictorially enriched ontologies, that include both visual and linguistic concepts, can be useful to support video annotation up to the level of detail of pattern specification.This paper presents pictorially enriched ontologies and discusses a solution for their implementation for the soccer video domain. An unsupervised clustering method is proposed in order to create the enriched ontologies by defining visual prototypes representing specific patterns of highlights and adding them as visual concepts to the ontology.An algorithm that uses pictorially enriched ontologies to perform automatic soccer video annotation is proposed and results for typical highlights are presented. Annotation is performed associating occurrences of events, or entities, to higher level concepts by checking their proximity to visual concepts that are hierarchically linked to higher level semantics. Marco Bertini 0001, Alberto Del Bimbo, Carlo Torniai |
ACM Multimedia | 1 |
| 2005 | Common Visual Cues for Sports Highlights Modeling
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati |
Multim. Tools Appl. | 1 |
| 2005 | An Integrated Framework for Semantic Annotation and Adaptation
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001 |
Multim. Tools Appl. | 1 |
| 2004 | Common visual cues for sports highlights detectionabstractAutomatic annotation of semantic events allows effective retrieval of video content. We present automatic annotation of sports highlights for some of the principal sports types. They are obtained by detecting and tracking a limited number of visual cues common to each sport. Highlights are represented as atomic entities at the semantic level. They have a limited temporal extension and can be modeled as the spatio-temporal concatenation of specific events. Visual cues encode position and speed information coming from the camera and from the objects/athletes that are present in the scene, and are estimated automatically from the video stream. Algorithms for model checking and for visual cue estimation are discussed. as well as applications of the representation to different sports domains Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati |
ICME | 1 |
| 2004 | Content-based video adaptation with user's preferencesabstractWe present an integrated system that has been designed to support automatic semantic extraction of highlights in sports video and automatic video adaptation according to user's preferences. To analyze the user's satisfaction, we propose a new performance measure that explicitly takes into account the user's preferences and considers the number and type of errors produced by the annotation engine and the way in which these errors affect the compressed video quality and bandwidth allocation. We provide experimental results with application to soccer and swimming. Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001 |
ICME | 1 |
| 2004 | Automatic annotation of video streamsabstractBroadcasters are demonstrating interest in systems that ease the process of annotating huge amount of live and archived video materials. Exploitation of such assets is considered a key method for the improvement of production quality and sport videos (one of the most marketable assets). In Europe, soccer is one of the most relevant sport types. This paper deals with detection and recognition of soccer highlights, using an approach based on temporal logic models. Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati |
MMSP | 1 |
| 2004 | Highlights modeling and detection in sports videos
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati |
Pattern Anal. Appl. | 1 |
| 2003 | Annotation and Retrieval of Structured Video Documents
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati |
ECIR | 1 |
| 2003 | Automatic extraction and annotation of soccer video highlightsabstractBroadcasters are demonstrating interest in systems that ease the process of annotation the huge amount of live and archived video materials. Exploitation of such assets is considered a key method for the improvement of production quality, and sport videos are one of the most marketable assets. In particular, in Europe, soccer is one of the most relevant sport types. This paper deals with detection and recognition of soccer highlights, using an approach based on temporal logic models. Jürgen Assfalg, Marco Bertini 0001, Carlo Colombo, Alberto Del Bimbo, Walter Nunziati |
ICIP (2) | 2 |
| 2003 | Object and event detection for semantic annotation and transcodingabstractVideo annotation provides a suitable way to describe, organize, and index stored videos. On the other hand, transcoding aims at adapting content to the user/client capabilities and requirements. Both cues are now mandatory, given the tremendous demand of multimedia access from remote clients, in particular nowadays that new terminals with limited resources (PDAs, HCCs, Smart phones) have access to the network. In this paper we propose a unified framework to define event-based and object-based semantic extraction from video to provide both semantic video annotation for video stored and semantic on-line transcoding from live cameras. Two case studies (highlights' extraction from soccer videos for the annotation and people behavior detection in domotic application for transcoding) and corresponding experimental results are reported. Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001 |
ICME | 1 |
| 2003 | Semantic annotation for live and posterity logging of video documents
Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati |
VCIP | 1 |
| 2003 | Semantic annotation of soccer videos: automatic highlights identification
Jürgen Assfalg, Marco Bertini 0001, Carlo Colombo, Alberto Del Bimbo, Walter Nunziati |
Comput. Vis. Image Underst. | 2 |
| 2002 | Soccer highlights detection and recognition using HMMsabstractIn this paper we report on our experience in the detection and recognition of soccer highlights in videos using hidden Markov models. A first approach relies on camera motion only, whereas a second one also includes information regarding the location of players on the playing field. While the former approach requires less information, the latter has proven to be more precise. Our experimental evaluation yields interesting results. Jürgen Assfalg, Marco Bertini 0001, Alberto Del Bimbo, Walter Nunziati, Pietro Pala |
ICME (1) | 2 |
| 2002 | Semantic Annotation and Indexing of News and Sports Videos
Jürgen Assfalg, Marco Bertini 0001, Carlo Colombo, Alberto Del Bimbo, Walter Nunziati |
SOFSEM | 2 |
| 2002 | Indexing for reuse of TV news shots
Marco Bertini 0001, Alberto Del Bimbo, Pietro Pala |
Pattern Recognit. | 1 |
| 2001 | Automatic Caption Localization in Videos Using Salient PointsabstractBroadcasters are demonstrating interest in building digital archives of their assets for reuse of archive materials for TV programs, on-line availability, and archiving. This requires tools for video indexing and retrieval by content exploiting high-level video information such as that contained in super-imposed text captions. In this paper we present a method to automatically detect and localize captions in digital video using temporal and spatial local properties of salient points in video frames. Results of experiments on both high-resolutionDV sequences and standard VHS videos are presented and discussed. 1. Marco Bertini 0001, Carlo Colombo, Alberto Del Bimbo |
ICME | 1 |
| 2001 | Classification Of Rawmaterial Sports Videos For Broadcasting Using Color And Edge FeaturesabstractThe authors discuss the method to classify raw material sports videos for broadcasting. Because the raw material sports videos sometimes do not get edited, one cannot use the knowledge on edited videos. The authors use the color and edge features and evaluate whether one can classify the sports videos with those features. Also introduced is the "player" and "audience" class - apart from each sport class - to improve the classification results. Masayuki Mukunoki, Marco Bertini 0001, Jürgen Assfalg, Alberto Del Bimbo |
ICME | 2 |
| 2001 | Content-based indexing and retrieval of TV news
Marco Bertini 0001, Alberto Del Bimbo, Pietro Pala |
Pattern Recognit. Lett. | 1 |